GridCARE Logo

GridCARE

Senior Site Reliability Engineer

Posted 8 Days Ago
Be an Early Applicant
Hybrid
Redwood City, CA, USA
180K-230K Annually
Senior level
Hybrid
Redwood City, CA, USA
180K-230K Annually
Senior level
Own the reliability, scalability, and observability of production systems. Design and operate AWS infrastructure with Terraform and Kubernetes, build monitoring and SLOs, automate deployments and capacity management, ensure database and data pipeline reliability, partner on architecture reviews, and drive infrastructure security, compliance, and incident response.
The summary above was generated by AI

About Us

GridCARE is a leading venture-backed startup solving the most critical constraint in AI’s growth trajectory: immediate access to power. As demand for computing skyrockets, access to energy has become the defining bottleneck in the AI infrastructure race. While leading tech companies invest billions in speculative, long-term solutions that may take decades to arrive, GridCARE’s pioneering physics-based generative AI platform unlocks gigawatts of hidden capacity in today’s electric grid — enabling hyperscalers, data center developers, and utilities to power AI infrastructure years sooner than conventional approaches and without costly upgrades.

Founded at Stanford’s Doerr School of Sustainability and backed by leading climate-tech and deep-tech investors, GridCARE has assembled a world-class team spanning power systems, AI, and infrastructure.

At GridCARE, you will:

⚡ Work at the intersection of AI, energy, and infrastructure — the foundation of the next industrial revolution.

🤝 Partner with hyperscalers, developers, and utilities on high-impact, real-world deployments.

🌎 Help shape a more abundant, efficient, and resilient energy future for the digital era.

🚀 Join a company defining a new category — capacity acceleration for AI.

💰 Receive competitive compensation, equity, and benefits in a fast-growth, mission-driven environment.

Learn more about GridCARE:

  • TechCrunch: GridCARE thinks more than 100 GW of data-center capacity is hiding in the grid

  • Utility Dive: Portland General Electric invests in AI-powered flexibility to speed data-center connection

  • Data Center Dynamics: From Years to Months — Creating an AI Fast Lane for Data Centers

  • GridCARE Raises $64 Million Series A to Create a New Category: Power Acceleration

Job Description

We're looking for a Senior SRE to own the reliability, scalability, and observability of our production systems. You'll work closely with platform and data engineering to keep high-throughput, data-intensive services running at the availability our customers (utilities, data center operators) require.

Responsibilities
  • Design and operate infrastructure on AWS using Terraform and Kubernetes

  • Build monitoring, alerting, and observability (Prometheus, Grafana, Datadog, or similar) with meaningful SLOs/SLIs

  • Automate away toil — deployment pipelines, capacity management, self-healing systems

  • Partner with engineering on architecture reviews to catch reliability and scalability risks before they ship

  • Manage database and data pipeline reliability for large-scale, real-time grid data processing

  • Drive security and compliance best practices across infrastructure

Qualifications

Required

  • 5+ years in SRE, DevOps, or infrastructure engineering roles

  • Deep experience with Kubernetes, Terraform/IaC, and cloud platforms (AWS Preferred)

  • Strong scripting/programming ability (Python, Bash)

  • Observability Experience (Prometheus, Grafana, Datadog)

  • Track record of running on-call for production systems and leading incident response

  • Experience with CI/CD pipelines (Github Actions) and infrastructure automation

  • Experience with Gitops concepts and tooling (ArgoCD/Flux)

  • Solid understanding of networking, distributed systems, and database reliability

  • Comfortable operating in a fast-moving startup environment with ambiguity

Preferred

  • Experience with data-intensive or real-time processing systems

  • Background in energy, climate tech, or critical infrastructure

  • Experience scaling infrastructure through hypergrowth

  • On-Prem Kubernetes Deployment Experience

  • Windows Server Administration Experience

What We Offer
  • Competitive salary, performance bonus, and equity.

  • Comprehensive health, dental, and vision coverage.

  • Lunch provided three days a week in office.

  • Hybrid schedule for local employees: 3 days in office for collaboration, 2 days remote for focused work.

  • Access to leading academic, industry, and government partners in the AI-energy ecosystem.

  • A mission-driven team focused on shaping the future of the energy transition.

Salary Range

$180,000-$230,000 Total

Join us in tackling one of the most important infrastructure challenges of our time — enabling the energy foundation for the age of AI.

HQ

GridCARE Redwood, California, USA Office

Redwood, California, United States

Similar Jobs

4 Days Ago
Easy Apply
Hybrid
San Francisco, CA, USA
Easy Apply
186K-232K Annually
Senior level
186K-232K Annually
Senior level
Artificial Intelligence • Big Data • Healthtech • Biotech • Pharmaceutical
Build and operate reliable cloud infrastructure, developer platforms, CI/CD systems, observability, and production workloads. Support applications, data systems, ML pipelines, and AI workloads across development, staging, and production. Establish SLOs, monitoring, incident response, automation, infrastructure-as-code practices, and operational standards. Collaborate with Product Engineering, Data Engineering, Data Science, and Security while mentoring engineers and participating in support rotations.
Top Skills: AWSAzureCi/CdDockerGCPGitInfrastructure As CodeKubernetesOpentofuPythonSnowflakeTerraformTerragruntVercelVirtual Networking
6 Days Ago
Remote or Hybrid
United States
Senior level
Senior level
Fintech • Software
The Senior Site Reliability Engineer ensures SaaS platforms remain reliable, performant, secure, and scalable. Responsibilities include building cloud infrastructure, implementing monitoring and alerting, automating operational runbooks and deployments, managing Infrastructure as Code, applying AI-powered observability and remediation, supporting Kubernetes and cloud networking, and leading incident triage and root-cause analysis during 24/7 on-call rotations.
Top Skills: AIAiopsAksAnsibleAppdynamicsAWSAzureAzure DevopsBashC# .NetCi/CdCloud NetworkingCloudopsCosmos DbDatadogDynatraceEksFirewallsHarnessIdera Sql Diagnostic ManagerInfrastructure As CodeJavaJenkinsKubernetesLinuxLoad BalancingNew RelicPowershellPythonRedgate Sql MonitorSolarwinds Database Performance AnalyzerSQLTerraformWindows
7 Days Ago
Easy Apply
Remote or Hybrid
USA
Easy Apply
180K-220K Annually
Senior level
180K-220K Annually
Senior level
Healthtech • Information Technology • Software • Telehealth
Develop, monitor, and maintain distributed production systems and AWS-based microservices infrastructure. Build automation, tooling, and repeatable processes that improve uptime, scalability, security, and operational efficiency. Support product engineering teams with performance, scaling, incident diagnosis, and production debugging. Analyze and tune systems, code, and networking while participating in on-call operations and blameless post-mortems.
Top Skills: AWSDnsDockerGCPGenaiHttp/HttpsKubernetesLoad BalancersNtpReverse ProxiesTcp/IpTlsWeb Application Firewalls

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account