Dyna Robotics Logo

Dyna Robotics

Research Engineer

Reposted One Month Ago
Be an Early Applicant
In-Office
Redwood City, CA, USA
220K-320K Annually
Senior level
In-Office
Redwood City, CA, USA
220K-320K Annually
Senior level
The role involves designing and maintaining large-scale ML infrastructure, optimizing distributed training systems, and enhancing computing performance for model training.
The summary above was generated by AI

Dyna Robotics builds general-purpose robots powered by a proprietary embodied AI foundation model with top-in-industry generalization and real-world performance. Already deployed with customers across multiple industries, our robots do commercial-grade work in the physical world. Our team comes from Google DeepMind, Meta, and Cruise, and we're backed by CRV, First Round, and other leading investors.

The Role

As a Research Engineer, you will architect and build the systems that turn our multi-cloud GPU fleet into a training engine our researchers love. Your charter is singular and broad: own training infrastructure end-to-end so that every GPU is busy, every run is reproducible, and every researcher's next experiment is one command away.

What You’ll Do
  • Scale Distributed Training: Architect and own the infrastructure for large-scale GPU clusters. You’ll implement sharding, activation checkpointing, and memory optimization (ZeRO, FSDP) to enable the training of massive multimodal models.

  • Optimize Researcher Ergonomics: Build a research codebase and job scheduling system (Kubernetes/SLURM) that prioritizes fast iteration, automated retries, and seamless failure recovery.

  • High-Performance Data Handling: Design high-throughput pipelines to ingest and transform terabytes of multimodal robot data (video, proprioception, 3D signals), ensuring dataloaders never starve the GPUs.

  • Production Inference: Build low-latency inference pipelines for real-time robot control. You’ll apply quantization, distillation, and model compilation (TensorRT, Triton) to move models from the lab to the physical world.

  • Deep Systems Profiling: Dive into the weeds of GPU utilization, I/O bottlenecks, and memory fragmentation to squeeze every bit of performance out of our expanding compute fleet.

What You’ll Bring
  • 7+ Years of Engineering: With a track record of leading technical projects in high-performance computing (HPC) or ML infrastructure.

  • ML Systems Mastery: Deep experience with PyTorch and distributed training frameworks (DeepSpeed, Accelerate). You understand the nuances of mixed precision and gradient accumulation.

  • Infrastructure Expertise: Hands-on experience managing cloud GPU environments (GCP/AWS) and container orchestration (Kubernetes).

  • Low-Level Intuition: A fundamental understanding of distributed systems, including race conditions, memory management, and NCCL/inter-node communication.

  • Ownership Mindset: You don't just "deploy" code; you design, build, and operate systems end-to-end to unblock fast-moving research.

Bonus Points For
  • Experience with Robotics Data Formats (MCAP, Protobuf) or multimodal models (VLAs).

  • Deep ML systems experience: custom kernels (Triton), compilers, or runtime optimization.

  • Experience as a founding or early-stage infrastructure hire.

At Dyna Robotics, we build technology for the real world, which requires a team as diverse as the environments our robots inhabit. We are an equal opportunity employer committed to technical rigor and mutual respect.

Don’t let a checklist stop you. Data shows that underrepresented groups often only apply if they meet 100% of the criteria. We value problem-solving and grit over keyword matching. If you’re passionate about robotics, no matter your discipline, we want to hear from you, even if you don't check every box.

Similar Jobs

17 Days Ago
Hybrid
Santa Clara, CA, USA
203K-354K Annually
Senior level
203K-354K Annually
Senior level
Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Lead research in agent learning and recursive self-improvement for enterprise AI agents. Design post-training methods, agent harnesses, training environments, evaluations, and distributed pipelines across language and multimodal systems. Analyze failures, develop training signals, run rigorous experiments, and transition validated improvements into reliable production systems. Collaborate with research, engineering, infrastructure, security, and product teams while communicating results through publications, patents, reports, and open-source work.
Top Skills: DeepspeedFsdpKubernetesMcpMegatronOpenrlhfPythonPyTorchPytorch DistributedRaySglangSlurmTrlVerlVllm
18 Days Ago
Hybrid
Santa Clara, CA, USA
232K-405K Annually
Senior level
232K-405K Annually
Senior level
Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Lead high-impact research in agent learning and recursive self-improvement. Design post-training methods, agent harnesses, training environments, evaluations, and distributed pipelines for reliable enterprise agents. Investigate planning, reasoning, memory, tool use, multimodal capabilities, and long-horizon execution. Analyze failures, run rigorous experiments, build improvement loops, and transition validated research into production systems. Provide technical leadership across teams and communicate results through reports, publications, patents, benchmarks, and open-source contributions.
Top Skills: Distributed InferenceDistributed TrainingDpoGpu InfrastructureGrpoKnowledge RetrievalLarge Language Models (Llms)Model Context Protocol (Mcp)Multimodal ModelsPythonPyTorchReinforcement LearningReward ModelingSupervised Fine-Tuning (Sft)
25 Days Ago
Hybrid
San Francisco, CA, USA
136K-265K Annually
Junior
136K-265K Annually
Junior
Cloud • Healthtech • Social Impact • Software • Biotech
Build datasets, evaluations, and scalable data infrastructure to improve frontier AI models for scientific and biological applications. Analyze model failure modes, run experiments, translate scientific expertise into rigorous evaluation criteria, and collaborate with scientists and AI labs. The role requires working at the intersection of software engineering, biology, and frontier AI in an in-person, fast-paced San Francisco environment.
Top Skills: AIData InfrastructureData PipelinesFrontier ModelsLarge Language Models (Llms)

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account