Volta Logo

Volta

Senior Network Engineer (3x Openings)

Posted 8 Hours Ago
Be an Early Applicant
Hybrid
Palo Alto, CA, USA
225K-295K Annually
Senior level
Hybrid
Palo Alto, CA, USA
225K-295K Annually
Senior level
Designs, operates, and automates large-scale AI data center network fabrics using leaf-spine Ethernet, BGP, EVPN, and VXLAN. Develops production Python or Go, builds configuration and validation tooling, exposes network capabilities through APIs, and improves observability. Troubleshoots congestion, packet loss, and control-plane issues; supports cluster bring-up, evaluates vendor designs, participates in on-call response, and collaborates on technical design and code reviews.
The summary above was generated by AI
About Volta

Volta is the category-defining, fully vertically integrated AI infrastructure platform – from capital to clusters to software, under a founder-led enterprise. Our mission is The Utility of Compute™: AI infrastructure as dependable and available as electricity, for every organization that needs it. Launched with a $10B strategic partnership with one of the leading frontier AI labs, a Series A led by Andreessen Horowitz, and a $5B AI Infrastructure Fund, Volta is building the infrastructure layer of the AI era from the ground up. We are 100+ people across London, Palo Alto, and New York, with rapid growth expectations to hundreds.

About The Role

Volta builds and operates large scale GPU compute infrastructure for AI workloads. The network is not a layer underneath our platform, it is part of it. Fabric design, overlay and multi-tenancy, edge connectivity, and the software that programs and observes all of it sit in one platform engineering team, deliberately.

This is a network role with a software expectation attached. You will need real depth in fabric and routing, because a 20,000 GPU training cluster punishes anyone who understands the network only from documentation. You will also need to write production code, because at this scale anything that requires a human at a CLI does not happen reliably. Configuration is generated from a model, validated in CI, and applied by tooling. If your current work is mostly logging into devices and making changes by hand, this role will be a significant shift.

You will work alongside platform engineers, not adjacent to them: same repositories, same review standards, same definition of done.

 
What You Will Be Doing

Common across the team:

  • Design, build, and operate the network fabric across our sites: leaf-spine Ethernet, BGP underlay, EVPN and VXLAN overlay, and multi-tenant isolation.

  • Build and extend the software that manages the fabric: configuration generation from a source of truth, validation pipelines, drift detection, and the tooling that makes change safe at scale.

  • Contribute production Python or Go to the platform codebase, including the integration between the physical fabric and the IaaS control plane.

  • Own how network capability is exposed upward: the APIs and abstractions through which tenants receive isolated, performant networking.

  • Instrument the fabric: streaming telemetry, topology-aware metrics, and tooling that makes a fabric of this size understandable.

  • Work with the bring-up teams during cluster deployment: fabric build, validation, acceptance testing, and turning the pain points you find into platform features rather than tribal knowledge.

  • Debug the hard problems: congestion and packet loss under collective communication load, and the class of failure where controller state and hardware state disagree.

  • Take part in on-call, incident response, and the follow-up work that closes structural gaps rather than only the immediate issue.

  • Evaluate designs and configurations proposed by OEMs and partners, and challenge them where they do not fit our requirements.

  • Participate in code review, technical design discussions, and cross-team collaboration in an Agile (Kanban or Scrum) environment.

Depending on your background, you will go deeper in one of these areas:

  • GPU fabric: RoCE v2 and InfiniBand for training traffic, congestion control tuning, rail-optimized topology, and the performance validation that proves a cluster is fit for workloads.

  • SDN and overlay integration: the software boundary between the fabric and the platform, controller and API-driven fabric programming, and multi-tenant network provisioning.

  • Edge connectivity: external connectivity, BGP peering, transit and IX relationships, and our public autonomous system.

  • Fabric observability: telemetry pipelines, fabric health tooling, and making network state queryable rather than inspectable.

 
What You Bring
  • 4+ years in data center or cloud network engineering, in production environments where downtime had real consequences.

  • Ethernet fabric depth: leaf-spine design, BGP including unnumbered BGP, ECMP, and the day-two realities of operating it.

  • EVPN and VXLAN in production, including what happens to an overlay under multi-tenant load.

  • Production software development, not scripting. Python or Go in a shared repository, under normal review, testing, and CI standards. Tools only you can run are not what we mean.

  • Network automation in practice: configuration as code, a source of truth system such as NetBox or Nautobot, and declarative or idempotent workflows.

  • Solid Linux fundamentals and comfort at the command line as an engineer, not only as an operator.

  • Familiarity with Kubernetes networking and how workload networking interacts with the underlying fabric.

  • Multi-vendor capability. Able to work across major OEM platforms and not dependent on one vendor's CLI.

  • Willing to be on site during cluster bring-up when it matters.

  • Clear written communication. Designs, decisions, and failure analysis need to be readable by people who were not in the room.

Nice to Have (But Not Essential)

None of these are required. Several map to specific areas of the team's scope, so strength in one or more helps us place you well:

  • Fluency with AI-assisted development: agentic CLI tools, IDE assistants, and orchestrating multiple coding agents through MCP, skills, or APIs to amplify delivery.

  • RoCE v2 at scale: PFC and ECN tuning, DCQCN, and how it behaves differently from InfiniBand under training load.

  • InfiniBand production experience: fat tree topology, UFM, fabric partitioning, adaptive routing, and SHARP.

  • NVIDIA Spectrum-X, including NetQ and Cumulus, or SONiC and whitebox platforms.

  • gNMI, OpenConfig, or NETCONF and YANG for configuration and telemetry.

  • IPv6 at production scale: dual stack design, v6 BGP peering, and addressing architecture.

  • ASN operations: public autonomous systems, transit and IX peering, RPKI and IRR hygiene, and DDoS posture.

  • OVN and OVS, SR-IOV, DPDK, or BlueField DPU based networking.

  • Network simulation or emulation with containerlab, NVIDIA Air, or equivalent.

  • Familiarity with NVLink and NVSwitch topologies and NCCL behavior.

  • Depth in Go or Rust beyond working proficiency.

  • Open source contributions to networking or infrastructure projects.

  • Experience working distributed across time zones with counterparts in other regions.

    REQ-59

What We Offer

At Volta, we believe people do their best work when they feel supported, trusted and able to grow. We're building a company where you can make an impact, keep a healthy balance between work and life, and build a career you're proud of.
As a global team, we do our best to provide great benefits wherever you're based. While some benefits vary by country due to local regulations, we believe looking after our people is simply the right thing to do.

  • Competitive salary based on the work you do here, not your previous salary

  • Equity in Volta, giving you the opportunity to share in the company's long-term success

  • Retirement/pension contributions

  • Comprehensive health, wellbeing and insurance benefits

  • Generous number of vacation days each year

Additional Information

Background Checks

All offers of employment at Volta are conditional on the satisfactory completion of pre-employment screening, which includes confirmation of your right to work, verification of your employment history and a criminal record check, where this is permitted by local law. Screening is carried out by Zinc, an accredited third-party provider, after an offer is made and all information is handled confidentially and in accordance with applicable data protection law.

Equal Opportunity

Volta is an equal opportunity employer. We are committed to building a diverse and inclusive team and make employment decisions based on skills, qualifications, experience and business needs. We do not discriminate on the basis of race, colour, religion, sex, sexual orientation, gender identity or expression, national origin, age, disability, veteran status or any other legally protected characteristic.

Accessibility

Volta is committed to providing an accessible recruitment experience for all candidates. If you require accommodations or adjustments at any stage of the application or interview process, please contact us at [email protected]. We will work with you to identify reasonable accommodations that enable you to participate fully in the hiring process.

Candidate Privacy Notice

By applying, you consent to the processing of your personal data for recruitment purposes in accordance with applicable data protection laws, including the UK GDPR, EU GDPR and relevant US state privacy regulations. Your data will be shared only with those involved in the hiring process and will not be used for unrelated purposes. For details, see our Recruitment & Candidate Privacy Notice.

Note to Recruitment Agencies

Volta does not accept unsolicited CVs or candidate profiles from recruitment agencies. Any unsolicited submissions, including those sent directly to hiring managers or employees, will be treated as the property of Volta. No agency fees will be payable unless a valid, signed recruitment agreement is in place, and the agency has been specifically engaged for the relevant vacancy.

Similar Jobs

A Minute Ago
Hybrid
Mountain View, CA, USA
176K-308K Annually
Senior level
176K-308K Annually
Senior level
Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Build and scale Moveworks’ generative AI conversation engine for enterprise use. Responsibilities include designing scalable APIs, optimizing low-latency multilingual dialog systems, developing infrastructure for model customization, implementing logging and tracing frameworks, improving observability, and collaborating with ML, application engineering, product, and support teams. The role also champions engineering best practices, system robustness, performance optimization, and rapid innovation.
Top Skills: Api DesignGenerative AiLarge Language Models (Llms)LoggingMetricsMicrosoft TeamsSlackTracing
A Minute Ago
Hybrid
Santa Clara, CA, USA
221K-387K Annually
Expert/Leader
221K-387K Annually
Expert/Leader
Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Own the product strategy and roadmap for securing agentic AI across internal systems and customer offerings. Translate AI security threats, red-team findings, and research into prioritized solutions. Lead architectural reviews covering agentic AI frameworks, authentication protocols, and LLM vulnerabilities. Partner with engineering, product, documentation, training, professional services, and enterprise customers to embed security into development lifecycles, define success metrics, and drive adoption of secure-by-default products.
Top Skills: Agent-To-Agent (A2A) ProtocolsEu Ai ActFedrampGdprLangchainLanggraphModel Context Protocol (Mcp)Oauth 2.0Openid Connect (Oidc)
A Minute Ago
Hybrid
Santa Clara, CA, USA
191K-334K Annually
Senior level
191K-334K Annually
Senior level
Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Designs and operates cloud-native reliability, release, testing, and developer productivity platforms. Builds Kubernetes infrastructure, CI/CD and GitOps workflows, automated validation, observability, progressive delivery, resilience testing, and self-service engineering tools. Resolves complex platform and networking issues, improves release confidence and operational efficiency, partners across engineering teams, participates in architecture decisions, and mentors engineers. Requires extensive SRE, DevOps, platform, software, or infrastructure engineering experience and strong programming skills.
Top Skills: AnsibleArgo CdArgo WorkflowsAws EksAzure AksCi/CdCypressDockerFluxGateway ApiGitlab Ci/CdGitopsGoGoogle Cloud GkeHelmIngressIstioJavaJunitKubernetesKustomizeLinkerdOpentelemetryPlaywrightPrometheusPytestPythonRest AssuredRubySeleniumTerraformTestng

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account