Design, orchestrate, and optimize large-scale LLM pre-training across 1,000+ GPUs. Implement 3D parallelism, manage GPU clusters (SLURM/Kubernetes), optimize InfiniBand/RDMA networking and memory, and automate checkpointing and failure recovery for long training runs.
We are seeking a highly skilled LLM Pre-training & Distributed Systems Engineer. This role is essential for orchestrating large-scale machine learning training runs and optimizing distributed infrastructure. The ideal candidate will have a deep understanding of GPU clusters and extensive experience in system engineering to ensure efficient and reliable training processes.
Responsibilities:
- Orchestrate distributed training runs across 1,000+ GPUs using PyTorch, DeepSpeed, or Megatron-LM.
- Optimize networking (InfiniBand/RDMA) and memory management to prevent out-of-memory errors.
- Automate checkpointing and failure recovery during month-long training runs.
Required Skills:
- Deep expertise in 3D parallelism (Data, Tensor, Pipeline).
- Experience managing SLURM or Kubernetes-based GPU clusters.
- Strong systems engineering background (C++, CUDA, Python).
Similar Jobs
Artificial Intelligence • Big Data • Cloud • Information Technology • Software • Big Data Analytics • Automation
Deliver Dynatrace consulting services including application monitoring, performance analysis, troubleshooting, dashboard and report development, integrations, customer training, adoption support, and project documentation. Serve as a trusted digital transformation advisor and product expert while supporting sales opportunities. The role requires enterprise Java or .NET experience, web programming, application and database technology knowledge, and strong communication, analytical, and problem-solving skills.
Top Skills:
.NetAjaxAWSCitrixDb2DockerDynatraceGoogle Cloud PlatformJ2EeJavaJavaScriptKubernetesMicroservicesAzureMicrosoft Sql ServerOracle
Artificial Intelligence • Big Data • Cloud • Information Technology • Software • Big Data Analytics • Automation
Delivers Dynatrace consulting services, including platform deployment, application monitoring, performance troubleshooting, dashboard development, customer training, platform adoption, deployment health, documentation, and project reporting. The consultant advises clients on implementation strategies and performance improvements, supports sales opportunities, and serves as a trusted product expert in digital transformation.
Top Skills:
.NetAjaxAWSCitrixDb2DockerDynatrace Unified PlatformGoogle Cloud PlatformJ2EeJavaJavaScriptKubernetesMicroservicesAzureMicrosoft Sql ServerObject-Oriented ProgrammingOracle
Blockchain • eCommerce • Fintech • Payments • Software • Financial Services • Cryptocurrency
Leads ambiguous, high-impact revenue strategy initiatives from analysis through executive recommendation. Responsibilities include sales forecasting, pipeline health measurement, capacity planning, organizational sizing, commercial deal economics, partnership modeling, and building scalable operating processes. The role partners cross-functionally with Sales, Finance, Partnerships, Services, Data, and Systems, influencing outcomes without direct authority.
Top Skills:
ExcelGoogle SheetsSQL
What you need to know about the San Francisco Tech Scene
San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.
Key Facts About San Francisco Tech
- Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
- Major Tech Employers: Google, Apple, Salesforce, Meta
- Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
- Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
- Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
- Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine


