Airlock Digital Logo

Airlock Digital

Senior Site Reliability Engineer

Sorry, this job was removed at 05:09 p.m. (PST) on Tuesday, Apr 21, 2026
Be an Early Applicant
In-Office
San Francisco, CA, USA
In-Office
San Francisco, CA, USA

Similar Jobs

18 Hours Ago
Remote or Hybrid
United States
Senior level
Senior level
Fintech • Software
The Senior Site Reliability Engineer ensures SaaS platforms remain reliable, performant, secure, and scalable. Responsibilities include building cloud infrastructure, implementing monitoring and alerting, automating operational runbooks and deployments, managing Infrastructure as Code, applying AI-powered observability and remediation, supporting Kubernetes and cloud networking, and leading incident triage and root-cause analysis during 24/7 on-call rotations.
Top Skills: AIAiopsAksAnsibleAppdynamicsAWSAzureAzure DevopsBashC# .NetCi/CdCloud NetworkingCloudopsCosmos DbDatadogDynatraceEksFirewallsHarnessIdera Sql Diagnostic ManagerInfrastructure As CodeJavaJenkinsKubernetesLinuxLoad BalancingNew RelicPowershellPythonRedgate Sql MonitorSolarwinds Database Performance AnalyzerSQLTerraformWindows
Yesterday
Easy Apply
Remote or Hybrid
USA
Easy Apply
180K-220K Annually
Senior level
180K-220K Annually
Senior level
Healthtech • Information Technology • Software • Telehealth
Develop, monitor, and maintain distributed production systems and AWS-based microservices infrastructure. Build automation, tooling, and repeatable processes that improve uptime, scalability, security, and operational efficiency. Support product engineering teams with performance, scaling, incident diagnosis, and production debugging. Analyze and tune systems, code, and networking while participating in on-call operations and blameless post-mortems.
Top Skills: AWSDnsDockerGCPGenaiHttp/HttpsKubernetesLoad BalancersNtpReverse ProxiesTcp/IpTlsWeb Application Firewalls
Yesterday
In-Office or Remote
4 Locations
168K-334K Annually
Senior level
168K-334K Annually
Senior level
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
Operate and improve large-scale Kubernetes and GPU clusters across public and private clouds. Build automation, observability, capacity-management, and reliability systems; define SLOs and SLIs; support production launches; lead incident triage and root-cause analysis; conduct blameless postmortems; and participate in on-call support for AI workloads.
Top Skills: AirflowAnsibleArgo WorkflowsAWSAws Step FunctionsCadenceChefCudaElk StackGoGoogle Cloud PlatformGrafanaKubernetesKubevirtLightstepLinuxAzureNcclNvidia Dgx CloudNvidia DynamoOpentelemetryOracle Cloud InfrastructurePrometheusPuppetPythonPyTorchSglangSplunkTcp/IpTemporalTensorrt-LlmTerraformVllm
Location: San Francisco, California - RemoteWho Are We?  About Airlock Digital: 

Airlock Digital is a global leader in application control and allowlisting. We seek to empower every organization to run only what they trust and operate free from malware and ransomware.  

With rapid growth across Australia, North America, and EMEA. We are committed to our core values, respect, determination, and integrity. We support a diverse and expanding global customer base. At Airlock, we pride ourselves on being a team of humble, collaborative, and driven professionals who support one another and share a passion for cybersecurity.


What We are Looking For: 

The Senior Site Reliability Engineer (SSRE) is responsible for ensuring the reliability, scalability, performance and efficiency of our systems, applications and services. Working closely with cross-functional teams such as development, operations, and infrastructure to proactively identify, troubleshoot and resolve issues to ensure optimal performance and uptime.

Key Responsibilities:
  • Design, implement, and maintain highly available, scalable, and fault-tolerant systems and services.
  • Introduce best practices into Airlock Digital around observability, SLO’s and reliability.
  • Continuously monitor the performance, availability and security of Airlock Digital systems and services and proactively identify and resolve issues.
  • Identify areas for improvement across the organization and drive engineering-wide technical change in the field of site reliability.
  • Collaborate with cross-functional teams to implement and maintain deployment pipelines, monitoring tools and automated testing frameworks.
  • Develop and maintain document of systems, processes and procedures to ensure knowledge transfer and continuity.
  • Lead incident response, root cause analysis and post-mortem activities to identify and address underlying issues.
  • Work with Software Developers to design and implement scalable and resilient applications services and infrastructure.
  • Participate in on-call rotation to ensure 24/7 support for critical systems and services.
Required Skills & Qualifications: 
  • 5+ years of hands-on experience in Site Reliability and Observability Engineering, DevOps or Infrastructure Engineering, debugging, diagnosing and resolving high-severity incidents.
  • Commercial experience in in at least one programming language such as Python, or Go.
  • Solid experience with automation tools such as Ansible and containerization tools like Docker and Podman.
  • Deep understanding of distributed systems, networking, operating systems, and cloud computing.
  • Strong troubleshooting and problem-solving skills, and experience in incident response, root cause analysis, and post-mortem activities.
  • Systematic problem-solving approach, coupled with effective communication skills and a sense of ownership and drive.

What We Offer: 

We don’t think money is everything, but we know it is an important part of your decision to apply for a role. This position has a salary range of USD $148,000 - $185,000. Additional factors considered in extending an offer include responsibilities of the job, education, location, experience, knowledge, skills, abilities, and internal equity, alignment with market data, or applicable laws.  

At Airlock Digital we offer a wide range of benefits to our eligible team members. Our benefit programs vary by location and can include Medical, dental, and vision insurance – 401K Plan with 4% Company Match – Life and Disability Programs – Paid Parental Leave - Paid time off and Paid Holidays – Volunteer and Birthday Time off – Home Office Allowance 

Our Commitment: 

We believe in supporting our team members both personally and professionally. Named one of the USA’s Greatest Places to Work in 2024 and 2025, we value flexibility, trust, and a work environment that empowers our team to do their best work. 

We will be assessing applications as they come in, so we encourage you to send your resume through to us as soon as possible. All official job offers from our company are extended directly by our recruitment team and will be sent through an official BambooHR document for your review and signature. Please be aware that we do not ask for any personal information in the process of extending offers of employment, such as financial details or social security numbers. Upon acceptance of any offer, we will request such information as part of the onboarding process prior to or on your first day of employment, and only after completing a background check through an authorized third-party vendor. If you receive any communication asking for personal details outside of these processes, please contact us immediately to verify the authenticity of the request. Your security is important to us, and we are committed to a safe and transparent hiring experience. No contact from recruitment agencies, thank you 

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account