Rivian and Volkswagen Group Technologies, LLC Logo

Rivian and Volkswagen Group Technologies, LLC

Staff Site Reliability Engineer-Production Operations

Posted 11 Days Ago
Be an Early Applicant
Hybrid
Palo Alto, CA, USA
186K-256K Annually
Senior level
Hybrid
Palo Alto, CA, USA
186K-256K Annually
Senior level
Lead and operate ProdOps as senior SRE: coordinate and communicate during high-severity incidents, run blameless post-incident reviews, prioritize and drive systemic action items, verify fixes, build automation and observability (including AI-agent workflows), and grow a lean team while mentoring engineers and setting technical standards.
The summary above was generated by AI
About Us

Rivian and Volkswagen Group Technologies is a joint venture between two industry leaders with a clear vision for automotive’s next chapter. From operating systems to zonal controllers to cloud and connectivity solutions, we’re addressing the challenges of electric vehicles through technology that will set the standards for software-defined vehicles around the world.

The road to the future is uncharted. By combining our expertise across connectivity, AI, security and more, we’ll map a new way forward. Working together, we’ll create a future that’s more connected, more intelligent, more sustainable for everyone.

Role Summary

We are looking for an SRE Lead to serve as the senior technical leader and player-coach for ProdOps. This is a hybrid role: you will set the technical direction of the team and lead from the front during incidents, while also growing and managing a small group of exceptional engineers as the function scales.

As the calm center during a crisis, you will maintain a high-level mental model of the entire production ecosystem, freeing engineers to focus strictly on debugging and mitigation. You will own the reliability feedback loop end to end, which includes running blameless post-incident reviews, coordinating major incidents across Cloud systems, vehicle development pipelines, and Product Security, driving the systemic action-item backlog with TPMs and development teams, and verifying that fixes hold in production.

This role suits a senior systems engineer who still writes code, thinks in terms of systems and failure modes rather than single root causes, and wants to build a lean, automation-first reliability practice rather than staff a support queue.

 
ResponsibilitiesLead incident coordination and communication

Drive incident progress and coordinate cross-functional response as the central nervous system during a crisis. Own executive, customer, and parent-company communications, providing production expertise and communication leadership so engineers can concentrate on the technical problem. Maintain an accurate, high-level model of the full production ecosystem spanning Cloud, vehicle development pipelines, and Product Security.

Facilitate modern, blameless post-incident learning

Run post-incident reviews using Learning From Incidents (LFI) principles and HOWIE-style reporting. Move the organization away from the search for a single root cause and toward understanding how tooling, context, and multiple latent conditions combined to produce failure. Surface weaknesses in observability, process, testing, and tooling that impaired our ability to detect, mitigate, and recover.

Own and prioritize systemic action items

Manage the backlog of action items generated by reviews. While ProdOps does not write the fixes itself, you will prioritize, track, and drive these items to closure in partnership with TPMs and development teams, keeping leadership focused on customer impact.

Measure and verify efficacy

Once the fixes ship, measure and verify that they actually prevent recurrence. Drive accountability for outcomes and feed the results back into the Novel Incident Rate.

Build the automation and observability backbone

Design and build the systems, tooling, and AI-agent workflows that automate incident triage and administrative toil. Improve observability, including instrumentation, alerting, dashboards, and SLOs, so incidents are detected faster and understood more deeply. Write and review code where it multiplies the team's impact.

Lead and grow the team

Set technical standards and operating rhythm for ProdOps. Mentor and develop engineers, and as the team scales, take on hiring and people management while preserving the minimal-headcount, maximum-automation philosophy.

 
Qualifications

Minimum Qualifications:

  • 6+ years of experience in SRE, production/platform engineering, or systems engineering for large-scale distributed systems, including senior technical leadership or lead responsibilities.

  • Proven incident command experience: you have coordinated high-severity, cross-functional incidents and led communication with executives and external stakeholders under pressure.

  • Strong systems engineering fundamentals, including distributed systems, networking, cloud infrastructure, and an instinct for how complex systems fail.

  • Hands-on coding ability (e.g., Python, Go, or similar) sufficient to build automation, tooling, and integrations. This is not a code-free management role.

  • Deep observability expertise: instrumentation, metrics, logging, tracing, alerting, dashboards, and SLO/SLI design (Datadog or comparable platforms).

  • Fluency in modern reliability and post-incident practice: blameless reviews, Learning From Incidents (LFI), HOWIE, and systemic (non-single-root-cause) analysis.

  • A demonstrated bias toward eliminating toil through automation, and enthusiasm for using AI agents as force multipliers.

  • Excellent written and verbal communication; able to translate technical detail for executive and cross-company audiences.

  • Experience mentoring engineers, with the judgment and appetite to grow into formal people management.

Preferred Qualifications

  • Experience across both cloud services and hardware/vehicle or embedded development pipelines.

  • Exposure to Product Security or working closely with security teams during incidents.

  • Track record building an SRE or reliability function from an early stage.

  • Experience deploying LLM- or agent-based automation into production operational workflows.

Total Rewards

We build the exceptional — and we believe the people doing that work should be rewarded accordingly. In addition to a competitive base salary, full-time positions may be is eligible to participate in our annual company performance bonus program.

Payments are discretionary and not guaranteed; actual amounts depend on company results and the terms of the plan in effect, and require active employment at the time of payout. This role is also eligible for equity in the form of Restricted Stock Units (RSUs), subject to board approval and the terms of our equity incentive plans, including applicable vesting requirements.

In addition to our compensation programs, we invest in our people with a comprehensive benefits package designed to support the health, wellbeing, and financial future for full-time employees — including health coverage, retirement savings, time off, and family planning programs. Offerings vary by country. Learn more about our global benefit programs.

External candidates can apply for this role through the Rivian and Volkswagen Group Technologies careers site (https://rivianvw.tech/#careers). If you are a current employee, please apply through our internal job board.

Equal Opportunity

Rivian and Volkswagen Group Technologies is committed to creating a diverse environment and is proud to be an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, national origin, ancestry, sex, sexual orientation, gender, gender expression, gender identity, genetic information or characteristics, physical or mental disability, marital/domestic partner status, age, military/veteran status, medical condition, or any other characteristic protected by law. We are also committed to ensuring compliance with all applicable fair employment practice laws regarding citizenship and immigration status.

Rivian and Volkswagen Group Technologies is committed to ensuring that our hiring process is accessible for persons with disabilities. If you have a disability or limitation, such as those covered by the Americans with Disabilities Act, that requires accommodations to assist you in the search and application process, please email us at [email protected].

Candidate Data Privacy

Rivian and Volkswagen Group Technologies” may collect, use and disclose your personal information or personal data (within the meaning of the applicable data protection laws) when you apply for employment and/or participate in our recruitment processes (“Candidate Personal Data”). This data includes contact, demographic, communications, educational, professional, employment, social media/website, network/device, recruiting system usage/interaction, security and preference information. Rivian and Volkswagen Group Technologies may use your Candidate Personal Data for the purposes of (i) tracking interactions with our recruiting system; (ii) carrying out, analyzing and improving our application and recruitment process, including assessing you and your application and conducting employment, background and reference checks; (iii) establishing an employment relationship or entering into an employment contract with you; (iv) complying with our legal, regulatory and corporate governance obligations; (v) record keeping; (vi) ensuring network and information security and preventing fraud; and (vii) as otherwise required or permitted by applicable law.

Rivian and Volkswagen Group Technologies may share your Candidate Personal Data with (i) internal personnel who have a need to know such information in order to perform their duties, including individuals on our People Team, Finance, Legal, and the team(s) with the position(s) for which you are applying; (ii) Rivian and Volkswagen Group Technologies affiliates; and (iii) Rivian and Volkswagen Group Technologies’ service providers, including providers of background checks, staffing services, and cloud services.

Rivian and Volkswagen Group Technologies may transfer or store internationally your Candidate Personal Data, including to or in the United States, Canada, and the European Union and in the cloud, and this data may be subject to the laws and accessible to the courts, law enforcement and national security authorities of such jurisdictions.

If you provide a mobile telephone number as part of your application or during the recruitment process, Rivian and Volkswagen Group Technologies may use that number to contact you via SMS text message for recruitment-related purposes, including scheduling, logistics, and status updates. Message and data rates may apply. You may opt out of SMS communications at any time by replying STOP to any text message you receive from us. Consent to receive SMS messages is not a condition of applying for or being considered for employment.

Please see our Candidate Data Privacy Notice (English) and Candidate Data Privacy Notice (Serbian) for more information.

--

Please note this job posting represents an open, active vacancy. Additionally, we are not currently accepting applications from third party application services.

Similar Jobs

40 Minutes Ago
Hybrid
107K-141K Annually
Junior
107K-141K Annually
Junior
Cloud • Information Technology • Internet of Things • Machine Learning • Software • Cybersecurity • Infrastructure as a Service (IaaS)
Support RAN optimization and automation for 4G/5G/6G networks: design generative AI tools and automated workflows, launch carriers and sites, perform verification and drive testing, track KPIs, troubleshoot RF issues, and deliver engineering projects on time.
Top Skills: 4G5G6GCoreGenerative AiPower BIPythonRanRf OptimizationSmall CellsTableauTransport
2 Hours Ago
Remote or Hybrid
7 Locations
305K-457K Annually
Senior level
305K-457K Annually
Senior level
eCommerce • Fintech • Hardware • Payments • Software • Financial Services
Lead Square-wide Trust strategy across Risk and Identity, owning high-stakes Trust bets while coaching senior PMs. Drive product for Identity Verification, Access, Disputes, and Controls; optimize seller experience, risk loss, and ecosystem health. Operate AI-native, engage regulators and partners, and coordinate Product, Engineering, Compliance, Legal, and Operations to deliver Trust and Safety initiatives.
Top Skills: JumioMl/AiOnfidoPersonaVeriff
2 Hours Ago
Remote or Hybrid
7 Locations
164K-297K Annually
Senior level
164K-297K Annually
Senior level
eCommerce • Fintech • Hardware • Payments • Software • Financial Services
Lead end-to-end delivery of high-priority, cross-functional Revenue initiatives including product and partnership launches. Develop integrated program plans, establish governance and operating cadences, drive cross-functional alignment, manage risks and dependencies, and create repeatable launch frameworks and executive communications to ensure successful market readiness and post-launch stabilization.

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account