Build and operate production machine learning training systems on Amazon SageMaker using AWS Trainium. The role focuses on optimizing PyTorch code for Trainium hardware, understanding NeuronCore compilation, diagnosing device-level issues, tuning distributed training for cost and throughput, and delivering production pipelines. The engineer will collaborate directly with client and internal teams in a forward-deployed engineering environment.
Robots & Pencils is an AWS Partner building production AI systems for enterprise clients who need real engineering, not proofs of concept that never ship. We work forward-deployed, embedded directly with client teams, solving the problems that are too new or too specialized for a typical vendor relationship to handle.
The Role
We're looking for an engineer who can operate and train models on Amazon SageMaker running on AWS Trainium, AWS's custom silicon built specifically for large-scale model training. This isn't a role where you call an API and wait. You'll be walking up the stack: understanding what a training request actually looks like at the Trainium hardware and compiler level, then carrying that understanding all the way up through PyTorch training code and into a production SageMaker pipeline.
PyTorch is the backbone of this work. If you know the framework deeply and you're comfortable reasoning about how your code actually behaves on custom accelerator hardware rather than treating it as a black box, this role is built around that skill set specifically.
What You'll Do
- Train and operate models on Amazon SageMaker with AWS Trainium as the underlying compute
- Write and optimize PyTorch training code with a real understanding of how it compiles and executes on Trainium (NeuronCore architecture, compiler behavior, memory and throughput tradeoffs)
- Diagnose training run issues that show up specifically because of the hardware, not just the model, distinguishing a data or code problem from a compiler or device-level one
- Translate a request for "a Trainium job" into an actual working, cost-aware training pipeline, end to end
- Tune distributed training runs for throughput and cost on SageMaker's training infrastructure
- Work directly with client and internal engineering teams to scope and deliver real production training workloads, not experiments that stay in a notebook
What You'll Bring
- Strong, hands-on PyTorch experience, ideally including distributed or multi-device training
- Production experience with Amazon SageMaker for training and/or inference
- Comfort working close to the hardware layer: you understand device-specific compilation and can debug issues that are actually about the accelerator, not just the model
- AWS Trainium or Inferentia (Neuron SDK) experience is a strong plus; if you don't have it yet but have deep PyTorch and a track record of picking up new hardware targets fast, we want to talk to you
- Solid Python fundamentals and comfort operating in a client-facing, production engineering environment
Similar Jobs
Healthtech • Social Impact • Telehealth
Guide older adults and their caregivers through the mental healthcare intake process. Responsibilities include providing empathetic support, coordinating care, matching patients with therapists, managing multiple patients through different stages, and communicating clearly during sensitive conversations. The role requires organization, attention to detail, comfort with healthcare and CRM software, and a commitment to improving mental health access for older adults.
Top Skills:
Crm SystemsEhr PlatformsGoogle WorkspaceHealthieHubspot
Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Leads complex enterprise AI deployments and high-stakes go-live engagements, owns deployment playbooks and shared tooling, mentors engineers, establishes quality standards, captures reusable deployment knowledge, and partners with product and engineering teams. The role requires hands-on software development, customer-facing communication, AI and GenAI integration, prompt engineering, and production delivery experience.
Top Skills:
Ai AgentsGenerative AiJavaJavaScriptNow AssistPrompt EngineeringPython
Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Leads global system integrator partnerships through joint business planning, executive alignment, co-selling, partner enablement, governance, and strategic solution development. Drives sourced and influenced revenue, pipeline expansion, deal collaboration, and multi-year growth plans across ServiceNow, partners, field sales, and customer outcome teams. Represents the partnership internally and externally while overseeing global-regional teams and operational execution.
Top Skills:
AIAutomationCloud ComputingSaaSServicenow
What you need to know about the San Francisco Tech Scene
San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.
Key Facts About San Francisco Tech
- Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
- Major Tech Employers: Google, Apple, Salesforce, Meta
- Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
- Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
- Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
- Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

.png)
