Design, deploy, and operate a high-performance SaaS manufacturing platform on cloud infrastructure. Ensure scalability, reliability, automation, monitoring, and KPIs. Integrate AI into dev/ops workflows, lead impactful projects, and improve uptime, latency, and efficiency. Work cross-functionally in a fast-paced startup while meeting U.S. government data access citizenship requirements.
Instrumental builds the manufacturing acceleration platform behind the world’s most complex electronics. We capture digital exhaust and engineering context from assembly lines—images, test logs, BOM data, performance, repair cycles—and our AI engines identify insights that are difficult or impossible for human engineers to find. We accelerate the companies building the AI era by improving manufacturing yield, throughput, and ramp. NVIDIA, Meta, L3Harris, and their manufacturing partners rely on Instrumental to accelerate new product introduction and production.
The Instrumental platform collects, intelligently transforms, and contextually presents manufacturing data to technical end-users, enabling them to optimize their manufacturing process in real-time. Our core technology is proprietary ML algorithms, packaged in an accessible, user-centric user interface – we believe we must have both the best technology and the best access to that technology to win.
Requirements:
- 5 or more years of DevOps or SRE experience deploying and operating commercial SaaS platforms on public cloud infrastructure, AWS preferred.
- Expert knowledge with Linux, shell, containerization, Kubernetes, IaC (terraform preferred), monitoring, logging, and APM tools.
- Proven ability to take initiative and drive impactful projects to completion efficiently and independently.
- Comfort with ambiguity, pace, and frequent pivots inherent in a startup environment, with a track record of creating clarity for teams.
- Experience introducing and integrating AI tools/processes into development and operation workflows.
- Demonstrated skill in setting, iterating on, and measuring KPIs to ensure ongoing performance, reliability and efficiency.
- Network/application security and compliance experience is a plus.
Who You Are:
- Dead serious about performance, scalability, and reliability (PSR): You care deeply about how systems behave in the real world and sweat the details around latency, uptime, and scale.
- Systems engineering & infrastructure expertise: You’ve spent real time building and running distributed systems and know your way around cloud infrastructure, networks, and operating systems.
- Automation, automation, automation: If something is repetitive or error-prone, your first instinct is to automate it and make it disappear.
- Operating in ambiguity & high-growth environments: You’re comfortable making good calls without perfect information and adapting as the system and company grow fast.
- Dependable, trustworthy: People trust you to own problems, show up when things are broken, and follow through.
This position requires access to items and data that are developed under U.S. government contracts and subject to dissemination controls that limit access to U.S. citizens only.
We’re a growing team that works collaboratively, is supportive of each other, and is highly energized by the opportunity for a large impact. We actively work to promote an inclusive environment, valuing passion and the ability to learn. You’re encouraged to apply even if your experience doesn’t precisely match the job description!
The following is a representative annual base salary range for this position within the Bay Area: $175-229k. We consider candidates at multiple levels for this role. Job level and salary opportunities are evaluated through our interview process – we review the experience, knowledge, skills, and abilities of each applicant.
Instrumental is proud to offer a highly-rated variety of benefits, including health, vision, dental, commuter plans, and parental leave.
At Instrumental, protecting company and customer information is a shared responsibility. Employees are expected to comply with company engineering, security, access control, and privacy policies, and promptly report suspected security incidents or policy violations.
Instrumental Palo Alto, California, USA Office
909 Alma Street, Palo Alto, CA, United States, 94301
Similar Jobs
Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Lead architecture and implementation of reliability improvements across CrowdStrike's cloud-native platform. Build shared libraries and services, drive observability and SLO practices, perform performance and cost optimization, run resilience engineering and chaos experiments, automate infrastructure-as-code, mentor engineers, and embed with product teams to deliver scalable, highly reliable distributed systems at organizational scale.
Top Skills:
AIAlertingAWSCassandraElasticsearchGCPGoInfrastructure-As-CodeJavaKafkaKotlinKubernetesNode.jsObservability (TracingOciOpensearchProfilingProtobufPythonScalaSlos)
Blockchain • Fintech • Payments • Financial Services • Cryptocurrency • Web3
Design, build, and operate secure, scalable Kubernetes platforms and infrastructure as code (Terraform). Develop backend services and automation (Go, Python, JS/TS), improve CI/CD and observability, run on-call and incident response, define SLIs/SLOs and disaster recovery, embed security and compliance, mentor team members, and partner with product and engineering to raise reliability, performance, and cost-efficiency across hybrid and public-cloud environments.
Top Skills:
Ci/CdGitopsGoJavaScriptKubernetesPythonTerraformTypescript
Cloud • Software
Design, deploy, and operate large-scale, multi-region cloud-native services to improve reliability, performance, and security. Partner with application teams to build automation, run SLO-driven incident response and on-call rotations, leverage Kubernetes and CNCF tooling, and implement scalable operations, chaos and scale testing, and infrastructure-as-code for a resilient SaaS platform.
Top Skills:
ArgocdAWSGoKubernetesLinux/UnixOpentelemetryPrometheusPythonService Mesh
What you need to know about the San Francisco Tech Scene
San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.
Key Facts About San Francisco Tech
- Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
- Major Tech Employers: Google, Apple, Salesforce, Meta
- Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
- Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
- Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
- Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine



