Volta Logo

Volta

Platform Engineer

Posted 8 Hours Ago
Be an Early Applicant
Hybrid
Palo Alto, CA, USA
225K-290K Annually
Mid level
Hybrid
Palo Alto, CA, USA
225K-290K Annually
Mid level
Build and operate production platform software for large-scale, Kubernetes-native GPU infrastructure. Design Kubernetes operators and controllers, develop APIs and services, manage compute, storage, networking, and confidential-computing capabilities, and improve observability and reliability. Collaborate with product, operations, and security teams, own services in production, participate in on-call and incident response, and maintain strong engineering, testing, versioning, and CI/CD practices.
The summary above was generated by AI
About Volta

Volta is the category-defining, fully vertically integrated AI infrastructure platform – from capital to clusters to software, under a founder-led enterprise. Our mission is The Utility of Compute™: AI infrastructure as dependable and available as electricity, for every organization that needs it. Launched with a $10B strategic partnership with one of the leading frontier AI labs, a Series A led by Andreessen Horowitz, and a $5B AI Infrastructure Fund, Volta is building the infrastructure layer of the AI era from the ground up. We are 100+ people across London, Palo Alto, and New York, with rapid growth expectations to hundreds.

About The Role

Volta builds and operates large scale GPU compute infrastructure for AI workloads. Our platform is Kubernetes-native, spans multiple regions, and delivers virtual machines, storage, and networking through a fully automated infrastructure stack built on custom Kubernetes operators.

We are building out several platform engineering teams that together own the full stack, from managed bare metal and IaaS through to higher-order platform services. Each team owns a different part of that stack: compute, networking, storage, the control plane and API layer, confidential computing, and the customer-facing surface. The area you work in depends on the team you join and your prior expertise, so no single engineer is expected to cover all of it.

Platform Engineers work at the intersection of infrastructure and software development. Across every team, you will translate three key inputs into durable platform capabilities: product roadmap requirements from the product team, operational learnings from the bring-up teams, and security guidance from the security engineering team. The output of this role is production platform code, not configuration, not runbooks.

 
What You Will Be Doing

Common across every platform engineering team:

  • Design and implement Kubernetes operators and controllers that manage the lifecycle of platform resources.

  • Work closely with the product team to turn roadmap requirements into the platform capabilities that support them.

  • Collaborate with the bring-up teams to identify operational pain points and turn them into scalable platform features.

  • Integrate security guidance from the security engineering team into platform-level controls, and remediate findings at the platform layer.

  • Treat observability as a platform concern: instrument services, define meaningful metrics, and build tooling that gives the team visibility into platform health.

  • Own the services you build in production, including participation in an on-call rotation, incident response, and the follow-up work that closes structural gaps rather than only the immediate issue.

  • Hold to clean interface and versioning practice on anything other teams or customers depend on, including disciplined handling of breaking changes.

  • Participate in code review, technical design discussions, and cross-team collaboration in an Agile (Kanban or Scrum) environment.

Depending on your team and background, you will go deep in some of the following:

  • Control plane and APIs: improve and extend the API layer between user-facing services and the underlying platform, with disciplined versioning and backward compatibility.

  • Compute: build and operate the lifecycle of virtualized and bare metal compute resources.

  • Storage: provisioning workflows, attachment reliability, performance tuning, and failure handling.

  • Confidential computing: build and extend confidential computing capabilities across the stack, from secure bare metal and confidential VMs to Confidential Containers (CoCo).

  • Customer-facing services: the platform surfaces customers interact with directly, including the APIs and interfaces through which they consume capacity, working alongside product and UX.

 
What You Bring
  • 3+ years of software engineering experience, with a meaningful portion spent on infrastructure or platform systems.

  • Strong backend or systems programming experience in a production environment. Our working languages are Python, Go, and Rust; we welcome strong engineers from other compiled or object-oriented languages (for example C++, C#, or Java) who are ready to work across our stack as it evolves.

  • Solid understanding of Kubernetes internals: the control loop model, CRDs, controllers and operators, and reliable reconciliation logic.

  • Comfortable working close to the infrastructure layer: Linux, networking fundamentals, and distributed systems behavior.

  • Experience designing, building, and versioning production-grade APIs or service interfaces that other teams depend on, including disciplined handling of breaking changes and backward compatibility.

  • Experience operating what you build: debugging production systems, and taking part in on-call or incident response.

  • Strong engineering fundamentals: clean code, testing, version control, code review, and CI/CD practices.

Nice to Have (But Not Essential)

None of these are required. Several map to specific teams, so strength in one or more helps us match you to the right one:

  • Fluency with AI-assisted development: agentic CLI tools, IDE assistants, and orchestrating multiple coding agents through MCP, skills, or APIs to amplify delivery.

  • Depth in Go or Rust beyond working proficiency.

  • Familiarity with confidential computing technologies: TEEs, AMD SEV, Intel TDX, or Confidential Containers (CoCo).

  • Experience integrating security requirements into platform or infrastructure systems.

  • Familiarity with high-performance networking: overlay protocols, BGP, RDMA, or packet-processing frameworks.

  • Hands-on experience with distributed storage systems (Ceph or similar) at an engineering level.

  • Background building Kubernetes operators using frameworks such as Kopf, controller-runtime, or similar.

  • Experience with observability tooling: Prometheus, Grafana, OpenTelemetry, or structured logging in distributed systems.

  • Experience building SaaS or PaaS layers on top of an IaaS platform.

  • Exposure to serverless or inference serving infrastructure.

  • Exposure to GPU infrastructure or HPC environments.

  • Experience working distributed across time zones with counterparts in other regions.

    REQ-0

What We Offer

At Volta, we believe people do their best work when they feel supported, trusted and able to grow. We're building a company where you can make an impact, keep a healthy balance between work and life, and build a career you're proud of.
As a global team, we do our best to provide great benefits wherever you're based. While some benefits vary by country due to local regulations, we believe looking after our people is simply the right thing to do.

  • Competitive salary based on the work you do here, not your previous salary

  • Equity in Volta, giving you the opportunity to share in the company's long-term success

  • Retirement/pension contributions

  • Comprehensive health, wellbeing and insurance benefits

  • Generous number of vacation days each year

Additional Information

Background Checks

All offers of employment at Volta are conditional on the satisfactory completion of pre-employment screening, which includes confirmation of your right to work, verification of your employment history and a criminal record check, where this is permitted by local law. Screening is carried out by Zinc, an accredited third-party provider, after an offer is made and all information is handled confidentially and in accordance with applicable data protection law.

Equal Opportunity

Volta is an equal opportunity employer. We are committed to building a diverse and inclusive team and make employment decisions based on skills, qualifications, experience and business needs. We do not discriminate on the basis of race, colour, religion, sex, sexual orientation, gender identity or expression, national origin, age, disability, veteran status or any other legally protected characteristic.

Accessibility

Volta is committed to providing an accessible recruitment experience for all candidates. If you require accommodations or adjustments at any stage of the application or interview process, please contact us at [email protected]. We will work with you to identify reasonable accommodations that enable you to participate fully in the hiring process.

Candidate Privacy Notice

By applying, you consent to the processing of your personal data for recruitment purposes in accordance with applicable data protection laws, including the UK GDPR, EU GDPR and relevant US state privacy regulations. Your data will be shared only with those involved in the hiring process and will not be used for unrelated purposes. For details, see our Recruitment & Candidate Privacy Notice.

Note to Recruitment Agencies

Volta does not accept unsolicited CVs or candidate profiles from recruitment agencies. Any unsolicited submissions, including those sent directly to hiring managers or employees, will be treated as the property of Volta. No agency fees will be payable unless a valid, signed recruitment agreement is in place, and the agency has been specifically engaged for the relevant vacancy.

Similar Jobs

2 Hours Ago
Remote or Hybrid
USA
150K-170K Annually
Mid level
150K-170K Annually
Mid level
eCommerce • Fintech • Food • Mobile • Social Impact
Build, operate, and evolve AWS cloud infrastructure supporting a high-growth financial and hospitality technology platform. Responsibilities include infrastructure architecture, deployment pipelines, reliability engineering, incident response, observability, security, governance, capacity planning, migrations, and on-call support. The role partners with application engineers, implements infrastructure changes, and improves operational standards and scalability.
Top Skills: AlbAlertingAWSCloudFormationCloudwatchContainersEc2EcsEksElasticacheFargateIamInfrastructure As CodeLogsMetricsNlbRuby on RailsRdsRedisTerraformTracingValkeyVpc
5 Days Ago
In-Office
San Mateo, CA, USA
142K-213K Annually
Senior level
142K-213K Annually
Senior level
Aerospace • Artificial Intelligence • Machine Learning • Robotics • Software
Designs and operates scalable platform infrastructure across cloud, on-premise, and hybrid environments. Builds infrastructure-as-code frameworks, CI/CD pipelines, developer tooling, automation, observability, and secure deployment patterns. Partners with engineering, security, and infrastructure teams to improve reliability, resilience, compliance, and developer productivity across mission-critical systems.
Top Skills: AnsibleAWSAzureBashCentralized LoggingCi/CdDevsecopsDistributed SystemsDistributed TracingFedrampGoInfrastructure As CodeKubernetesLinuxMonitoringNistPowershellPythonTerraform
5 Days Ago
Hybrid
2 Locations
Senior level
Senior level
Financial Services
Leads development of secure, scalable cloud platforms optimized for AI and machine learning workloads. Designs and operates Kubernetes and containerized infrastructure, builds CI/CD and infrastructure-as-code automation, and optimizes cloud performance and costs. Partners with AI teams to translate compute needs into platform requirements, promotes observability and operational stability, and guides adoption of secure, responsible AI-assisted software development practices across engineering teams.
Top Skills: C#Ci/CdCloud ComputingDockerGoGrafanaHigh Performance ComputingHybrid CloudIaasInfrastructure As CodeJavaKubernetesLinuxMachine LearningMicroservicesMl InferenceMl TrainingMlflowMlopsNoSQLNvidia BcmNvidia DcgmNvidia Dynamo InferencePaasPrivate CloudPrometheusPublic CloudPythonRay.IoSaaSSlurmSQLTransformer ArchitectureVllm

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account