Skydio Logo

Skydio

Staff Site Reliability Engineer

Posted One Month Ago
In-Office or Remote
Hiring Remotely in San Mateo, CA, USA
240K-300K Annually
Senior level
In-Office or Remote
Hiring Remotely in San Mateo, CA, USA
240K-300K Annually
Senior level
Build, operate, and scale production cloud infrastructure and Kubernetes/EKS clusters on AWS. Own Terraform-defined infra, CI/CD, observability, networking, and reliability. Troubleshoot production incidents, participate in on-call rotations, automate operational work with Python/Go, and expand multi-region deployments.
The summary above was generated by AI

Skydio is the leading US drone company and the world leader in autonomous flight, the key technology for the future of drones and aerial mobility. The Skydio team combines deep expertise in artificial intelligence, best-in-class hardware and software product development, operational excellence, and customer obsession to empower a broader, more diverse audience of drone users, from utility inspectors to first responders, soldiers in battlefield scenarios, and beyond.

About the team:

Skydio’s Cloud infrastructure team is here to ensure that the Skydio Cloud platform is always available to our customers when they need it: whether that’s performing a routine bridge inspection or it’s aiding in rescue operations during a natural disaster. With tens of thousands of drones operating all over the world, we are constantly improving the continuous delivery and deployment of our infrastructure. Our technology helps save lives. You’ll play a critical role in keeping the infrastructure behind it reliable, scalable, and available when it matters most.

About the role:

We are looking for a hands-on Staff Site Reliability Engineer to build, operate, and scale the cloud infrastructure that powers our products. This role is focused on owning production infrastructure, including Kubernetes, AWS, infrastructure as code, CI/CD, observability, networking, and reliability.

You don't need to be an expert in every area, but you should have strong Kubernetes and cloud fundamentals with meaningful depth in at least one infrastructure domain.

How you'll make an impact:

  • Build, operate, and troubleshoot production Kubernetes/EKS clusters.

  • Perform Kubernetes upgrades, node rollouts, and cluster maintenance.

  • Build and manage AWS infrastructure including VPCs, networking, subnets, load balancers, IAM, EKS, databases, and storage.

  • Define and maintain infrastructure using Terraform.

  • Build and operate CI/CD and deployment infrastructure.

  • Troubleshoot production issues across Kubernetes, AWS, Linux, networking, and databases.

  • Build monitoring, alerting, and observability for critical infrastructure.

  • Participate in on-call rotations and respond to production incidents.

  • Identify and solve infrastructure scaling and reliability problems.

  • Automate operational work using Python, Go, or similar languages.

  • Help expand infrastructure across new regions and deployment environments.

What makes you a good fit:

  • 8+ years of experience as a Site Reliability Engineer, Platform Engineer, DevOps, Production Engineer or equivalent infrastructure role.

  • Strong hands-on experience operating Kubernetes, not simply deploying applications to existing clusters.

  • Experience managing Kubernetes/EKS upgrades and production clusters.

  • Strong AWS fundamentals, including VPCs, public/private subnets, networking, load balancers, EKS, IAM, and databases.

  • Production experience with Terraform or similar infrastructure-as-code tooling.

  • Experience owning or maintaining CI/CD and deployment systems such as Argo CD, Spinnaker, GitHub Actions, GitLab CI/CD, or Jenkins.

  • Experience diagnosing production infrastructure and networking problems.

  • Experience solving meaningful scaling or reliability challenges.

  • Obtaining FAA Part 107 certification within the first 60 days of employment is strongly encouraged for all Skydio employees and required for certain positions.

  • This position requires access to export-controlled technical data, restricted government information, and/or information systems subject to U.S. government security and access-control requirements. Employment in this role is contingent upon verification of U.S. person status and the ability to access controlled or restricted information as required for the position.

Bonus points:

  • Helm and GitOps experience.

  • Datadog or similar observability tooling.

  • PostgreSQL/database operations experience.

  • Multi-region infrastructure experience.

  • On-premises or disconnected deployment experience.

  • Streaming or high-throughput distributed systems experience.

Compensation:

At Skydio, our compensation packages for regular, full-time employees include competitive base salaries, equity in the form of stock options, and comprehensive benefits packages. Compensation will vary based on factors, including skill level, proficiencies, transferable knowledge, and experience. Relocation assistance may also be provided for eligible roles. The annual base salary range for this position is $200,000 - $285,000*. Fundamentally, we believe that equity is the key to long-term financial growth, and we ensure all regular, full-time employees have the opportunity to significantly benefit from the company's success. Regular, full-time employees are eligible to enroll in the Company’s group health insurance plans. Regular, full-time employees are eligible to receive the following benefits: Paid vacation time, sick leave, holiday pay and 401K savings plan. This position and all associated benefits are subject to applicable federal, state, and local laws, as well as the Company’s policies and eligibility criteria.

*Compensation for certain positions may vary based on the position's location.

#LI-WA1

At Skydio we believe that diversity drives innovation. We have created a multidisciplinary environment that embraces the power of diverse perspectives to create elegant solutions for complex problems. We are committed to growing our network of people, programs, and resources to nurture an inclusive culture.

Qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, protected veteran status, or other characteristics protected by federal, state or local anti-discrimination laws.

For positions located in the United States of America, Skydio, Inc. uses E-Verify to confirm employment eligibility. To learn more about E-Verify, including your rights and responsibilities, please visit https://www.e-verify.gov/

HQ

Skydio San Mateo, California, USA Office

San Mateo, CA, United States

Skydio Hayward, California, USA Office

27317 Industrial Blvd, Hayward, United States, 94545

Similar Jobs

14 Days Ago
Easy Apply
Remote
United States
Easy Apply
Senior level
Senior level
Cloud • Security • Software • Cybersecurity • Automation
Build and operate reliable, scalable production infrastructure for GitLab’s user-facing services. Responsibilities include developing infrastructure automation and tooling, managing Kubernetes deployments, maintaining infrastructure as code, supporting CI/CD and GitOps, participating in on-call and incident response, improving observability and SLOs, troubleshooting production systems, and documenting operational practices. The role spans Intermediate through Senior Staff levels and requires strong software engineering, cloud, reliability, and asynchronous collaboration skills.
Top Skills: AlertingAWSCi/CdGCPGitopsGoInfrastructure As CodeKubernetesLoggingMetricsRubySlisSlosTerraform
4 Days Ago
Remote
United States
250K-325K Annually
Senior level
250K-325K Annually
Senior level
Artificial Intelligence • Cloud • Machine Learning • Software • Database • App development • Generative AI
Lead reliability engineering for Replit’s large-scale infrastructure by designing observability, defining SLOs and SLIs, leading incident response, automating operations, optimizing Kubernetes and GCP deployments, debugging distributed systems, and mentoring engineers. Build internal tools and integrations in Python or Go, maintain infrastructure as code and CI/CD pipelines, improve system performance and resilience, and establish reliability, security, and operational best practices across the engineering organization.
Top Skills: Ci/CdDatadogDockerGoGoogle Cloud Platform (Gcp)GrafanaKubernetesOpentelemetryPrometheusPulumiPythonTerraform
8 Days Ago
Remote
USA
177K-240K Annually
Expert/Leader
177K-240K Annually
Expert/Leader
Information Technology • Security • Software • Cybersecurity
Lead reliability strategy across eight engineering groups by defining SLIs, SLOs, and error budgets; strengthening incident response, alerting, postmortems, change safety, and failure testing; and coaching teams to own reliability. The role remains hands-on through production investigations, tooling, dashboards, and reference implementations. It also leads AI adoption in incident management and observability while partnering with architecture, platform, and product teams to improve distributed-system resilience.
Top Skills: Ai ToolsAWSDatadogDynamoDBElasticsearchGoInfrastructure As CodeKafkaKubernetesObservability ToolingRedisTypescript

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account