Design, build, and operate large-scale AI compute infrastructure for training, fine-tuning, evaluation, and inference. Manage Kubernetes clusters, workload scheduling, accelerator enablement, networking, storage, monitoring, and capacity. Troubleshoot complex application, cloud, cluster, and hardware issues while improving reliability, performance, scalability, and developer productivity. Partner with AI teams to automate provisioning, upgrades, monitoring, and maintenance.
As an engineer on the AI Compute Infra team, you will design, build, and operate large-scale infrastructure for AI training, fine-tuning, evaluation, and inference. You will work across Kubernetes clusters, accelerator enablement, workload scheduling, high-performance networking, storage, and capacity management, partnering with AI researchers and engineers to improve reliability, performance, scalability, and developer productivity.
Responsibilities:
Necessary Skills and Experience:
"Preferred" Skills and Experience:
In Return:
You will join a driven group committed to developing world-class AI compute infrastructure. We provide a cooperative setting where your ideas can come to life. Your efforts will directly impact the success of our AI projects, guaranteeing smooth operations and outstanding results. Join us and help build the future of AI compute infrastructure!
Additional Information
Please note that a relocation package (including visa sponsorship support) is available for this role, for candidates who require it.
We are an equal opportunity employer and value diversity at our company. We do not discriminate on the basis of race, religion, color, national origin, sex, gender, gender expression, sexual orientation, age, marital status, veteran status, or disability status. We will ensure that individuals with disabilities are provided reasonable accommodation to participate in the job application or interview process, to perform essential job functions, and to receive other benefits and privileges of employment. Please contact us to request accommodation.
Salary Range:
$209,100-$282,900 per year
We value people as individuals and our dedication is to reward people competitively and equitably for the work they do and the skills and experience they bring to Arm. Salary is only one component of Arm's offering. The total reward package will be shared with candidates during the recruitment and selection process.
Accommodations at Arm
At Arm, we want to build extraordinary teams. If you need an adjustment or an accommodation during the recruitment process, please email [email protected] . To note, by sending us the requested information, you consent to its use by Arm to arrange for appropriate accommodations. All accommodation or adjustment requests will be treated with confidentiality, and information concerning these requests will only be disclosed as necessary to provide the accommodation. Although this is not an exhaustive list, examples of support include breaks between interviews, having documents read aloud, or office accessibility. Please email us about anything we can do to accommodate you during the recruitment process.
Hybrid Working at Arm
Arm's approach to hybrid working is designed to create a working environment that supports both high performance and personal wellbeing. We believe in bringing people together face to face to enable us to work at pace, whilst recognizing the value of flexibility. Within that framework, we empower groups/teams to determine their own hybrid working patterns, depending on the work and the team's needs. Details of what this means for each role will be shared upon application. In some cases, the flexibility we can offer is limited by local legal, regulatory, tax, or other considerations, and where this is the case, we will collaborate with you to find the best solution. Please talk to us to find out more about what this could look like for you.
Equal Opportunities at Arm
Arm is an equal opportunity employer, committed to providing an environment of mutual respect where equal opportunities are available to all applicants and colleagues. We are a diverse organization of dedicated and innovative individuals, and don't discriminate on the basis of race, color, religion, sex, sexual orientation, gender identity, national origin, disability, or status as a protected veteran.
Responsibilities:
- Build and operate Kubernetes clusters while improving workload scheduling, topology-aware placement, capacity use, and recovery.
- Enable new CPU and GPU systems by integrating and validating drivers, networking, storage, monitoring, and health checks.
- Investigate performance and reliability issues across applications, cloud infrastructure, clusters, and hardware, then turn findings into lasting improvements.
- Partner with AI teams to understand their workloads and automate cluster provisioning, upgrades, monitoring, and maintenance around their needs.
Necessary Skills and Experience:
- 5+ years of experience building or operating cloud, compute, HPC, or distributed infrastructure in a production environment.
- Programming experience in Go, Python, or another systems language, with an interest in developing reliable infrastructure software.
- Practical knowledge of Kubernetes, containers, Linux, networking, and storage.
- Experience supporting GPU, accelerator, or distributed machine-learning workloads.
- An ability to troubleshoot complex systems and communicate clearly with engineers from different technical backgrounds.
"Preferred" Skills and Experience:
- Familiarity with Kubernetes scheduling, operators, quotas, or resource management.
- Experience with NVIDIA technologies such as CUDA, NVLink, NVSwitch, NCCL, EFA, or DCGM.
- Knowledge of AWS EKS, Terraform, Argo CD, Helm, Prometheus, or Grafana.
- Familiarity with frameworks such as PyTorch, Ray, vLLM, SGLang, or TensorRT-LLM, or experience qualifying accelerators and tuning distributed workloads.
In Return:
You will join a driven group committed to developing world-class AI compute infrastructure. We provide a cooperative setting where your ideas can come to life. Your efforts will directly impact the success of our AI projects, guaranteeing smooth operations and outstanding results. Join us and help build the future of AI compute infrastructure!
Additional Information
Please note that a relocation package (including visa sponsorship support) is available for this role, for candidates who require it.
We are an equal opportunity employer and value diversity at our company. We do not discriminate on the basis of race, religion, color, national origin, sex, gender, gender expression, sexual orientation, age, marital status, veteran status, or disability status. We will ensure that individuals with disabilities are provided reasonable accommodation to participate in the job application or interview process, to perform essential job functions, and to receive other benefits and privileges of employment. Please contact us to request accommodation.
Salary Range:
$209,100-$282,900 per year
We value people as individuals and our dedication is to reward people competitively and equitably for the work they do and the skills and experience they bring to Arm. Salary is only one component of Arm's offering. The total reward package will be shared with candidates during the recruitment and selection process.
Accommodations at Arm
At Arm, we want to build extraordinary teams. If you need an adjustment or an accommodation during the recruitment process, please email [email protected] . To note, by sending us the requested information, you consent to its use by Arm to arrange for appropriate accommodations. All accommodation or adjustment requests will be treated with confidentiality, and information concerning these requests will only be disclosed as necessary to provide the accommodation. Although this is not an exhaustive list, examples of support include breaks between interviews, having documents read aloud, or office accessibility. Please email us about anything we can do to accommodate you during the recruitment process.
Hybrid Working at Arm
Arm's approach to hybrid working is designed to create a working environment that supports both high performance and personal wellbeing. We believe in bringing people together face to face to enable us to work at pace, whilst recognizing the value of flexibility. Within that framework, we empower groups/teams to determine their own hybrid working patterns, depending on the work and the team's needs. Details of what this means for each role will be shared upon application. In some cases, the flexibility we can offer is limited by local legal, regulatory, tax, or other considerations, and where this is the case, we will collaborate with you to find the best solution. Please talk to us to find out more about what this could look like for you.
Equal Opportunities at Arm
Arm is an equal opportunity employer, committed to providing an environment of mutual respect where equal opportunities are available to all applicants and colleagues. We are a diverse organization of dedicated and innovative individuals, and don't discriminate on the basis of race, color, religion, sex, sexual orientation, gender identity, national origin, disability, or status as a protected veteran.
Arm San Jose, California, USA Office
150 Rose Orchard Way, San Jose, CA, United States, 95134
Similar Jobs at Arm
Artificial Intelligence • Internet of Things • Semiconductor
Leads project management across AI and Developer Platforms, establishing scalable operating models, planning frameworks, governance, risk management, and delivery standards. Manages and develops a distributed Project Management team, oversees portfolio and execution rhythms, aligns cross-functional stakeholders, resolves trade-offs, and provides leadership visibility into progress, dependencies, risks, resources, and outcomes. Uses data, automation, dashboards, and AI to improve delivery efficiency and organizational decision-making.
Top Skills:
Ai PlatformsDashboardsDelivery Data And MetricsDeveloper ToolingSoftware InfrastructureWorkflow Automation
Artificial Intelligence • Internet of Things • Semiconductor
Build and operate secure, scalable AI compute platform services, including backend APIs, asynchronous workflows, Kubernetes controllers, and multi-tenant workload capabilities. The role focuses on distributed training and inference, platform reliability, observability, developer tooling, and infrastructure simplification. Responsibilities include collaborating across engineering teams, improving scalability and developer productivity, and providing technical leadership for production AI platform capabilities.
Top Skills:
APIsAsynchronous ProcessingContainersGitopsGoGrafanaKafkaKubernetesKubernetes OperatorsOpentelemetryPostgresPrometheusPythonPyTorchRayRelational DatabasesRpcVllm
Artificial Intelligence • Internet of Things • Semiconductor
Design and build secure, scalable platform services for distributed AI training and inference. Develop backend APIs, Kubernetes controllers, orchestration systems, asynchronous workflows, multi-tenant access controls, and developer tools. Improve reliability, observability, scalability, and engineering productivity while partnering with infrastructure and AI teams. The role also provides technical leadership, mentoring, and cross-functional collaboration throughout platform development and production delivery.
Top Skills:
ContainersGitopsGoGrafanaKafkaKubernetesOpentelemetryPostgresPrometheusPythonPyTorchRayRelational DatabasesRpcVllm
What you need to know about the San Francisco Tech Scene
San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.
Key Facts About San Francisco Tech
- Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
- Major Tech Employers: Google, Apple, Salesforce, Meta
- Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
- Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
- Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
- Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

