Plenful Logo

Plenful

Senior Site Reliability Engineer

Posted 13 Days Ago
Be an Early Applicant
Hybrid
San Francisco, CA, USA
Senior level
Hybrid
San Francisco, CA, USA
Senior level
Own production reliability and performance for a healthcare AI platform: define SLIs/SLOs, run incident response and postmortems, build observability and automation, optimize scalability across serverless and container services, and partner with engineering, security, and compliance to improve operational maturity.
The summary above was generated by AI
About Plenful

Plenful is on a mission to transform healthcare operations from the inside out. Fresh off our $50M Series B and backed by Notable Capital, Bessemer Venture Partners, TQ Ventures, Susa/Kivu Ventures, and other leading investors, we’re building the category-defining AI workflow automation platform that healthcare teams rely on to operate smarter, faster, and more efficiently. Our technology empowers healthcare operators across hospital and health systems, pharmacies and payors to eliminate manual work, reduce administrative burden, and improve compliance, all while unlocking critical revenue to fund programs for their in-need patient populations.


Built by healthcare operators for healthcare operators, Plenful is driven by a deep understanding of the challenges facing today’s care teams. We’re passionate about equipping healthcare workers with world-class tools that deliver real, measurable impact, and we’re proud to serve 90+ leading health systems across the country. If you’re excited to help shape the future of healthcare, we’d love to meet you. Apply now to join our growing team.

About the Role

Plenful is hiring a Senior Site Reliability Engineer (SRE) to keep our production systems reliable, performant, and scalable as we grow.

This role is centered on operating real systems at scale — not just building infrastructure, but understanding deeply how it behaves under load, fails in production, and recovers. You'll define reliability standards, own production health, and build the feedback loops that make our systems more resilient over time.

You'll work closely with backend, data, and ML engineers to keep the platform highly available, measurable, and continuously improving — from incident response and performance debugging to SLO design and system-level optimization. This role is hybrid.

What You’ll Do

Reliability Engineering & System Ownership

  • Define and implement SLIs, SLOs, and error budgets across core services.

  • Own production system health: uptime, latency, and availability targets.

  • Improve system resilience through proactive reliability work.

  • Find and mitigate single points of failure across distributed systems.

Production Operations & Incident Response

  • Take part in and improve on-call rotations and incident response.

  • Lead incident triage, mitigation, and resolution in real time.

  • Run blameless postmortems and follow through on action items.

  • Build tooling and automation to cut MTTR (Mean Time to Recovery).

Observability & System Insight

  • Design and evolve observability across metrics, logs, and distributed tracing (OpenTelemetry), using tools like Datadog, CloudWatch, Grafana, and Sentry.

  • Improve signal quality to cut noise and alert fatigue.

  • Build dashboards and alerts that reflect real system health and user impact.

  • Use observability data to drive performance and reliability improvements.

Performance & Scalability

  • Analyze system performance under load and find bottlenecks.

  • Optimize latency, throughput, and resource use across serverless (AWS Lambda), containerized services (ECS), and data systems (Aurora Postgres, ClickHouse).

  • Partner with engineering teams to improve system efficiency and scaling behavior.

Automation & Reliability Tooling

  • Build automation that eliminates repetitive operational work.

  • Improve deployment safety through reliability checks and safeguards.

  • Contribute to CI/CD pipelines (GitHub Actions) with a focus on stability.

  • Build tools for incident response, debugging, and capacity planning.

Security, Compliance & Operational Maturity

  • Partner with security and compliance to keep systems meeting operational standards.

  • Support audit readiness and reliability-related compliance requirements (Vanta).

  • Integrate monitoring and alerting into security and SIEM workflows.

  • Help mature operational practices across engineering.

You'll know it's working when SLOs and error budgets are clear and enforced, incidents are rare and shrinking over time, engineers trust their signals about system health, alerts are actionable instead of noisy, systems scale predictably under load, postmortems drive real improvement, and reliability is a shared responsibility, not a reactive function.

You May Be a Fit If
  • You've spent 5+ years in Site Reliability Engineering, SRE-adjacent roles, or production infrastructure.

  • You've operated and debugged distributed systems in production.

  • You have hands-on experience with observability tooling (Datadog, Grafana, OpenTelemetry, or similar), incident response and on-call practices, and performance and reliability debugging.

  • You've defined and worked with SLOs, SLIs, and error budgets.

  • You're familiar with AWS environments, serverless and container-based architectures, and Postgres or similar relational databases.

  • You can write code or scripts (Python, Bash, etc.) for automation and tooling.

  • You think in systems and reason clearly about failure modes.

  • Bonus points for experience in high-growth or high-scale environments, background in regulated industries like healthcare or fintech, experience with ClickHouse or analytical systems at scale, familiarity with chaos engineering or load testing, and exposure to ML infrastructure or data platforms.

Why You'll Love Working Here
  • 🚀 Mission-Driven, World-Class Team — Join an exceptional group of professionals aligned around a meaningful mission and committed to making an impact

  • 📈 Opportunities for Growth — Strengthen your expertise through collaboration with experienced, high-performing leaders across the organization

  • 🏢 Flexible Hybrid Work Environment — We're remote-first, with meaningful office presence in San Francisco and New York. R&D roles follow a hybrid model, with two days per week in our San Francisco office

Benefits & Perks
  • 🏥 Healthcare Coverage — Full medical, dental, and vision insurance for you and participation for your family

  • 💰 401(k) with Company Match — Plenful matches 50% of your first 3% contributed

  • 📊 Equity — Every full-time employee shares in our success

  • 🌴 Unlimited PTO — Take the time you need, when you need it

  • 🍽️ Daily Lunch Stipend — $100/week to cover your midday meals

  • 💪 Wellness Stipend — $100/month to support your health and well-being

  • 🚇 Commuter Benefits — $100/month for SF and NYC-based employees

  • 👶 Parental Leave — Paid leave to support growing families

HQ

Plenful San Francisco, California, USA Office

San Francisco, CA, United States

Similar Jobs

5 Days Ago
Remote or Hybrid
San Francisco, CA, USA
147K-278K Annually
Senior level
147K-278K Annually
Senior level
Cloud • Software
Design, deploy, and operate large-scale, multi-region cloud-native services to improve reliability, performance, and security. Partner with application teams to build automation, run SLO-driven incident response and on-call rotations, leverage Kubernetes and CNCF tooling, and implement scalable operations, chaos and scale testing, and infrastructure-as-code for a resilient SaaS platform.
Top Skills: ArgocdAWSGoKubernetesLinux/UnixOpentelemetryPrometheusPythonService Mesh
14 Days Ago
Remote or Hybrid
United States
Senior level
Senior level
Fintech • Software
Lead SRE efforts for DFIN SaaS: ensure availability, performance, scalability, and automation. Implement monitoring, CI/CD, IaC, container orchestration, AI-enhanced observability, incident response, RCA, and runbook automation while collaborating across engineering teams.
Top Skills: .NetAiopsAksAnsibleAppdynamicsAWSAzureAzure DevopsBashC#Ci/CdCloud Ai ServicesContainersCosmosDatadogDynatraceEksFirewallHarnessIdera Sql Diagnostic ManagerInfrastructure As Code (Iac)JavaJenkinsKubernetesLinuxLoad BalancingNew RelicPowershellPythonRedgate Sql MonitorSolarwinds Database Performance AnalyzerSQLTerraformWindows
21 Days Ago
Easy Apply
Hybrid
Easy Apply
170K-190K Annually
Senior level
170K-190K Annually
Senior level
AdTech • Big Data • Cloud • Marketing Tech • Software • Analytics
Lead SRE efforts to improve security, reliability, cost efficiency, and observability. Build automation, CI/CD, agentic AI platforms (MCPs), and tooling for capacity planning, incident response, and self-service. Evangelize SecDevOps and zero-trust designs across product and platform teams.
Top Skills: Ai Agentic FrameworksArgocdAWSBashCi/CdDevsecopsDockerEksGoGrafanaKubernetesLinuxLokiMcpNew RelicPrometheusPythonSamTerraformZero Trust

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account