Periodic Labs Logo

Periodic Labs

Distributed Training Engineer

Reposted 2 Days Ago
Be an Early Applicant
In-Office or Remote
Hiring Remotely in Menlo Park, CA, USA
Mid level
In-Office or Remote
Hiring Remotely in Menlo Park, CA, USA
Mid level
Optimize and develop large-scale distributed LLM training systems, support reinforcement learning workflows, and contribute to open-source frameworks.
The summary above was generated by AI

About Periodic Labs

We are an AI + physical sciences lab building state of the art models to make novel scientific discoveries. We are well funded and growing rapidly. Team members are owners who identity and solve problems without boundaries or bureaucracy. We eagerly learn new tools and new science to push forward our mission.

About the role

You will optimize, operate and develop large-scale distributed LLM training systems that power AI scientific research. You will work closely with researchers to bring up, debug, and maintain mid-training and reinforcement learning workflows. You will build tools and directly support frontier-scale experiments to make Periodic Labs the world’s best AI + science lab for physicists, computational materials scientists, AI researchers, and engineers. You will contribute open-source large scale LLM training frameworks.

You might thrive in this role if you have experience with:

  • Training on clusters with ≥5,000 GPUs

  • 5D parallel LLM training

  • Distributed training frameworks such as Megatron-LM, FSDP, DeepSpeed, TorchTitan

  • Optimizing training throughput for large scale Mixture-of-Expert models

Similar Jobs

10 Days Ago
In-Office or Remote
7 Locations
Senior level
Senior level
Agency • Artificial Intelligence • Blockchain • Web3
Design, orchestrate, and optimize large-scale LLM pre-training across 1,000+ GPUs. Implement 3D parallelism, manage GPU clusters (SLURM/Kubernetes), optimize InfiniBand/RDMA networking and memory, and automate checkpointing and failure recovery for long training runs.
Top Skills: 3D ParallelismC++CudaDeepspeedGpuInfinibandKubernetesMegatron-LmPythonPyTorchRdmaSlurm
21 Days Ago
In-Office or Remote
San Francisco, CA, USA
Mid level
Mid level
Artificial Intelligence • Software
The Research Engineer will advance decentralized AI by optimizing distributed training, developing open-source tools, publishing research, and enhancing platform capabilities.
Top Skills: Ai/MlCi/CdDeepspeedMosaicmlPytorch DistributedRay
2 Hours Ago
Remote or Hybrid
United States
61K-92K Annually
Junior
61K-92K Annually
Junior
Artificial Intelligence • Automotive • Greentech • Information Technology • Machine Learning • Software • Cybersecurity
Execute end-to-end digital advertising across Search, Display, Social, and Video for a 50+ account portfolio. Monitor KPIs, perform keyword research, create ads, troubleshoot and optimize campaigns, produce monthly reports, consult with clients to retain and grow budgets, and maintain required platform certifications. Handle client communications, track actions for audit, and perform limited travel (5%).
Top Skills: CSSFtpGoogle AdsGoogle AnalyticsHTMLHTTPMicrosoft AdvertisingSalesforceSeo

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account