Periodic Labs Logo

Periodic Labs

ML Systems Engineer

Reposted One Month Ago
Be an Early Applicant
In-Office
Menlo Park, CA, USA
300K-400K Annually
Expert/Leader
In-Office
Menlo Park, CA, USA
300K-400K Annually
Expert/Leader
The ML Systems Engineer will design and manage efficient training and inference systems, optimize hardware utilization, and collaborate with researchers on RL loop integration, enhancing scientific discovery.
The summary above was generated by AI
About Periodic Labs

We're an AI and physical sciences company building state-of-the-art models to accelerate breakthroughs across materials, energy, and beyond. Backed by world-class investors and growing rapidly, we operate at the pace the frontier requires. Our team brings deep expertise, genuine ownership, and an insatiable drive to push the boundaries of what's scientifically possible.

About the Role

You’ll work alongside some of the world’s leading ML systems engineers, including leaders behind Megatron-LM, SGLang, Liger Kernel, TorchRec, CleanRL, TorchRL, and JAX-MD.

We’re looking for exceptional ML Systems Engineers to build the agentic infrastructure powering our large-scale training, inference, and reinforcement learning. You’ll own critical pieces of the ML systems stack to maximize performance, scalability, reliability, and productivity for both engineers and AI agents.

What You'll Do
  • Build and optimize large-scale training and reinforcement learning infrastructure while ensuring its correctness

  • Develop high-performance inference and serving systems

  • Design distributed runtimes and scheduling systems for complex ML workloads

  • Build secure and large-scale sandboxing and execution environments

  • Optimize memory, GPU kernels and communication for maximum throughput and end-to-end efficiency

  • Improve scalability, reliability, and efficiency across the ML systems stack

What We're Looking For
  • Strong systems programming and performance engineering skills

  • Experience building high-performance ML infrastructure at scale

  • Ability to own complex technical problems end-to-end

  • Strong coding ability and engineering judgment, including the ability to work effectively with AI agents to design, implement, test, and debug complex systems

  • High ownership, fast execution, and a passion for pushing the frontier of AI systems and accelerating scientific discovery


You should have deep expertise in at least one of the following:
  • Training: Strong experience building, debugging and optimizing large-scale training systems with Megatron-LM. Familiarity with TorchTitan, FSDP, veRL, Slime, or other distributed training systems is a plus.

  • Distributed Runtime: Strong experience with Ray. Familiarity with Monarch or other distributed execution frameworks is a plus.

  • Inference: Strong experience with SGLang. Familiarity with vLLM, TensorRT-LLM, or production LLM serving systems is a plus.

  • Sandboxing: Strong experience with secure execution environments, containers, virtualization, or code sandboxing.

  • GPU Kernels: Strong experience with CUDA, Triton, CUTLASS, CuTe, or custom GPU kernel development.

  • GPU Communication: Strong experience with NCCL, NVLink, InfiniBand, RDMA, GPUDirect RDMA, or large-scale communication optimization.

Mechanics

Minimum education: Bachelor’s degree or similar experience

Location: Menlo Park, CA (Soon: San Francisco, too)

Compensation: $250,000-$350,000 base + equity

Visa sponsorship: Yes, we sponsor visas.

We’re building a team of the world’s best — the scientists, engineers, and problem-solvers who don’t just follow the frontier, they define it. If you’re driven to bring AI to life in the physical world and make discoveries that have never been made before, you belong here.

Similar Jobs

3 Days Ago
Hybrid
Mountain View, CA, USA
Senior level
Senior level
Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Build and operate reliable production systems for machine learning models, agentic workflows, and self-learning capabilities. Responsibilities include ML lifecycle automation, continuous delivery, safe deployments, feedback loops, observability, SLOs, distributed workload optimization, incident response, platform automation, and cloud infrastructure standards. The role provides staff-level technical leadership across ML, data, product, and platform engineering teams.
Top Skills: C++Ci/CdCloud InfrastructureContainersDistributed SystemsGoGpu ComputingInfrastructure As CodeJavaKubernetesMachine Learning OperationsObservabilityPythonRustSlis/Slos
9 Days Ago
In-Office or Remote
7 Locations
277K-415K Annually
Expert/Leader
277K-415K Annually
Expert/Leader
Blockchain • eCommerce • Fintech • Payments • Software • Financial Services • Cryptocurrency
Build and operate production machine learning systems for ranking, retrieval, recommendations, search, propensity, churn, LTV, and next-best-action decisioning. Design reliable signal contracts covering freshness, provenance, confidence, eligibility, and calibration. Lead feature pipelines, model serving, experimentation, monitoring, and feedback loops while evaluating fairness, risk, compliance, trust, and long-term customer impact. Collaborate across product, growth, data, platform, modeling, risk, and compliance teams.
Top Skills: Ai AgentsBatch PipelinesData LakehousesData WarehousesEmbeddingsEvent StreamsExperimentation SystemsFeature StoresJavaKotlinKubernetesLarge Language ModelsLightgbmModel-Serving InfrastructureObservability ToolingPythonPyTorchRecommendation SystemsSemantic SearchSQLTensorFlowWorkflow OrchestrationXgboost
10 Days Ago
Remote or Hybrid
7 Locations
277K-415K Annually
Expert/Leader
277K-415K Annually
Expert/Leader
Blockchain • Fintech • Mobile • Payments • Software • Financial Services
Build and operate production machine learning systems for ranking, retrieval, recommendations, search, propensity, churn, lifecycle intelligence, and next-best-action decisioning. Design reliable signal contracts with freshness, provenance, confidence, and calibration guarantees. Lead experimentation, monitoring, feedback loops, and impact evaluation focused on fairness, trust, risk, compliance, and long-term engagement. Partner across product, growth, data, platform, modeling, risk, and compliance teams.
Top Skills: JavaKotlinKubernetesLightgbmPythonPyTorchSQLTensorFlowXgboost

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account