Sieve

Reliability Engineer

Reposted 13 Days Ago

Be an Early Applicant

In-Office

San Francisco, CA, USA

150K-300K Annually

Mid level

In-Office

San Francisco, CA, USA

150K-300K Annually

Mid level

Sieve is seeking a Founding Reliability Engineer to build and maintain infrastructure for petabyte-scale video workloads, focusing on reliability and security. Responsibilities include incident response, cloud security, and observability systems management.

The summary above was generated by AI

About Us

Sieve is the only AI research lab exclusively focused on video data. We combine exabyte-scale video infrastructure, novel video understanding techniques, and dozens of data sources to develop datasets that push the frontier of video modeling. Video makes up 80% of internet traffic and has become the enabling digital medium powering creativity, communication, gaming, AR/VR, and robotics. Sieve exists to solve the biggest bottleneck in the growth of these applications: high-quality training data.

We’ve partnered with top AI labs and did $XXM last quarter alone, as a team of just 12 people. We also raised our Series A earlier this year from Tier 1 firms such as Matrix Partners, Swift Ventures, Y Combinator, and AI Grant.

About the Role

We process petabytes of video across thousands of nodes and multiple cloud environments. As we scale, reliability, observability, and security become existential.

We’re hiring our first engineer fully dedicated to the infrastructure foundation of Sieve. This is a high-ownership role for someone who thinks deeply about:

throughput and system stability
monitoring and incident response
security and least-privilege design
reducing operational burden for the entire engineering team

You’ll work directly with our CTO and our founding engineers to build the core tooling that powers all of engineering.

This role is for someone who spends their time thinking deeply about reliability, throughput, observability, and security. You’re the kind of engineer who is always anticipating failure modes, eliminating operational risk, and designing systems that don’t break.

If something goes down, you take it personally, and you thrive in that level of responsibility.

What You’ll Do

Work with engineering to design and validate the infrastructure powering PB-scale workloads
Build and maintain Terraform-managed multi-cloud deployments
Improve cloud and data security (SSO, IAM, least privilege, auditability)
Own incident response and harden systems against failure
Develop CI/CD systems that minimize user error and maximize safety
Build monitoring + alerting platforms (Prometheus, OpenTelemetry, VictoriaMetrics)
Wrap internal reliability tooling with simple UIs for engineers

Requirements

3+ years building internal infrastructure at scale
Experience on-call for Sev 0 / Sev 1 production incidents (L3 preferred)
Strong cloud experience (GCP, AWS, Oracle, Cloudflare, etc.)
Deep Infrastructure-as-Code experience (Terraform preferred)
Familiarity with Argo, Helm, Kustomize, or similar deployment tools
Experience operating observability systems (Prometheus, OTel, VictoriaMetrics)
Backend fundamentals in Python, Go, Rust, or C++
Strong networking + security intuition, including SSO implementation
High ownership mindset over critical systems
In-person at our SF HQ

Bonus

Experience building lightweight internal tooling (APIs, dashboards, Svelte)
Familiarity with object storage systems (“buckets”)
Active GitHub or portfolio projects

Benefits

401k + Full Health Insurance
Breakfast, Lunch, and Dinner covered and your choice of snacks
Ubers covered home

San Francisco, CA, United States

Similar Jobs

Sprinter Health

Site Reliability Engineer

Yesterday

Remote or Hybrid

160K-235K Annually

Senior level

160K-235K Annually

Senior level

Artificial Intelligence • Healthtech • Logistics • Social Impact • Software • Telehealth

The Senior Site Reliability Engineer will enhance the reliability and security of infrastructure for in-home healthcare services, using cloud technology and automation to improve systems and processes.

Top Skills: AWSBashGCPPythonTerraformTypescript

Zscaler

Site Reliability Engineer

5 Days Ago

Easy Apply

Remote or Hybrid

San Jose, CA, USA

Easy Apply

193K-275K Annually

Expert/Leader

193K-275K Annually

Expert/Leader

Cloud • Information Technology • Security • Software • Cybersecurity

The Principal Site Reliability Engineer leads infrastructure projects, mentors junior engineers, ensures system reliability, and oversees networking services and scalable solutions with a focus on CI/CD and IaC/CaC.

Top Skills: AnsibleCi/CdEnterprise LinuxFreebsdGitGoHashicorp VaultKubernetesLdapLinux HypervisorsOidcPythonTerraform

Okta

Reliability Engineer

11 Hours Ago

In-Office

San Francisco, CA, USA

160K-220K Annually

Senior level

160K-220K Annually

Senior level

Cloud

The role involves designing, optimizing, and maintaining PostgreSQL and MySQL databases, ensuring high availability, reliability, and performance for mission-critical systems, while automating operational tasks and responding to incidents.

Top Skills: AnsibleAWSDatadogGCPGoGrafanaKubernetesMySQLPostgresPrometheusPythonTerraform

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
Major Tech Employers: Google, Apple, Salesforce, Meta
Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sieve

Reliability Engineer

Sieve San Francisco, California, USA Office

Similar Jobs

Site Reliability Engineer

Site Reliability Engineer

Reliability Engineer

What you need to know about the San Francisco Tech Scene

Key Facts About San Francisco Tech