Ziff Davis Logo

Ziff Davis

Site Reliability Engineer

Posted 8 Days Ago
Remote
Hiring Remotely in United States
90K-100K Annually
Mid level
Remote
Hiring Remotely in United States
90K-100K Annually
Mid level
Build, maintain, and operate Ookla’s globally distributed infrastructure platform at massive scale. Responsibilities include managing cloud instances, containers, serverless applications, databases, streaming systems, and big-data tooling; supporting 24/7 production operations and on-call rotations; implementing security programs; improving deployment pipelines, monitoring, observability, and reliability; and guiding software and data engineering teams on operational best practices and troubleshooting.
The summary above was generated by AI

Site Reliability Engineer 

The Opportunity

 

We are looking for a highly capable engineer to join our Platform and Site Reliability engineering team. You will be responsible for building, maintaining and operating the infrastructure platform on which all Ookla services are built. In this role, you will build, maintain, and support a massive-scale dynamic infrastructure that is relied on by  hundreds of millions of users around the world. You will obsess over systems performance, scalability, reliability, observability, and security. Most importantly, you will help deliver critical application functionality and help make the internet experience better for our users to help us achieve our goal of better connectivity for all.

 

We are committed to providing you a flexible work environment where individuality, fun, and talent are all valued equally. If you consider yourself innovative, adept at collaboration, and you care deeply about the work you do, we want to talk!

 

Key Responsibilities

  • Maintaining a distributed, global ecosystem of thousands of cloud instances, containerized workflows, serverless applications, Linux servers, and associated infrastructure supporting billions of requests daily.

  • Maintaining transactional database infrastructure using MySQL, PostgreSQL, and managed services such as RDS/Aurora.

  • Supporting the use of NoSQL data storage engines such as DynamoDb and MongoDB.

  • Building and supporting data stream processing with Kinesis or Kafka.

  • Supporting data engineering and big data toolchains such as Spark.

  •  Supporting production systems in a 24x7x365 environment, including on-call responsibilities.

  • Providing architectural and operational support to software engineers in a wide variety of focus areas.

  • Support software and data engineering teams and guiding operational best practices.

  • Implementation and oversight of security programs including vulnerability remediation, patch management, IDS/IPS, penetration testing, and interfacing with our corporate InfoSec team.

  • Supporting the development to production code deploy pipeline for a range of production applications.

  • Providing the tooling and guidance for the software and data engineering team to implement our monitoring and observability best practices.

  • Assisting development teams with troubleshooting.

 

Job Qualifications

We are looking for the right person, not the exact list of requirements. If you believe your life experience has prepared you for similar challenges, we’d like to hear from you.

  • 4+ Years Systems/Platform engineering experience

  • Experience building globally-distributed systems

  • Strong understanding of security best practices 

  • Infrastructure as Code: Terraform, Cloudformation

  • Branching and Merge based Source Code Configuration Management: Git, Github

  • Configuration management systems such as Chef or Ansible

  • Container-based architectures including Docker, Kubernetes

  • Proficiency in one or more high level programming languages such as Typescript, Go, Python, PHP, Ruby, Java, etc.

  • Experience with AWS and other Cloud infrastructure platforms

  • Comfort writing SQL queries and analyzing query performance

  • Comfortable learning and working with new technologies in an ever-changing environment

  • Strong verbal and written communication skills 

  • Strong time management skills and a self-driven work ethic

  

About 

Ookla, an Accenture company, is a global leader in connectivity intelligence that brings together the trusted expertise of Speedtest®, Downdetector®, Ekahau®, and RootMetrics® to deliver unmatched network and connectivity insights. By combining multi-source data with industry-leading expertise, we transform network performance metrics into strategic, actionable insights.

 Our solutions empower service providers, enterprises, and governments with the critical data and insights needed to optimize networks, enhance digital experiences, and help close the digital divide. At the same time, we amplify the real-world experiences of individuals and businesses that rely on connectivity to work, learn, and communicate. From measuring and analyzing connectivity to driving industry innovation, Ookla helps the world stay connected.

 About Accenture

Accenture helps the world’s leading enterprises reinvent by building their digital core and unleashing the power of AI to create value at speed for organizations across industries. Our strategy is to be the reinvention partner of choice for our clients and lead in the safe, widespread adoption of AI, and to be the most client focused, AI-enabled, great place to work in the world. We bring together the talent of our approximately 799,000 people with proprietary assets and platforms, deep process and industry expertise, and leading ecosystem relationships to deliver end-to-end solutions and measurable outcomes at scale. Through our Reinvention Services, we offer broad expertise across Cybersecurity, Digital Core, Finance, Industry and Enterprise, Song, Supply Chain and Engineering, and Talent, with advanced capabilities in AI and Data, Industry and Process, and Technology. We serve approximately 9,000 clients and generated approximately $70 billion in FY25 revenue. Visit us at accenture.com.

 Compensation Range 

Ookla provides a range for the base pay. Factors that may be used to determine your actual pay may include your specific job related knowledge, skills, experience, and geographic location. The salary compensation for this role is $90,000 - $100,000. Individual pay within the compensation range for this business unit specific role is determined based on a variety of factors including experience, scope of the role, capabilities to perform the role, education and training, as well as business and company performance.


Similar Jobs

3 Days Ago
In-Office or Remote
73K-130K Annually
Mid level
73K-130K Annually
Mid level
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Architect, operate, and maintain resilient cloud infrastructure across Azure and AWS commercial and government environments. Support Kubernetes platforms, IaC, observability, monitoring, deployments, platform services, performance testing, and incident response. Define reliability metrics, participate in 24/7 on-call rotations, perform root cause analysis, and automate operational processes and remediation. The role requires U.S. citizenship and eligibility to obtain a Confidential, Secret, or Top Secret clearance.
Top Skills: ArgocdAWSAzureAzure MonitorDynatraceEncryptionFluxGitGitlabGitopsGrafanaHelmIaasIamKubernetesOwaspPaasPkiPrometheusPulumiRestful ApisSplunkTerraformVisual Studio Code
9 Days Ago
In-Office or Remote
135K-231K Annually
Expert/Leader
135K-231K Annually
Expert/Leader
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Leads AI-assisted site reliability engineering across Azure and AWS. Designs observability, incident response, automation, resiliency testing, disaster recovery, chaos engineering, and recovery-validation capabilities. Establishes OpenTelemetry, SLI, SLO, error-budget, and reliability-scorecard standards; improves alert quality and operational insights; creates human-in-the-loop mitigation workflows; and mentors engineers while driving cross-functional reliability improvements.
Top Skills: AnsibleAWSAzureDatadogGrafanaHelmKubernetesLlmsOpentelemetryPrometheusPulumiRagTerraform
21 Days Ago
Remote or Hybrid
USA
140K-215K Annually
Senior level
140K-215K Annually
Senior level
Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Senior SRE owning availability, automation, and observability for CI/CD platform services. Build and operate infrastructure, run on-call, lead incident response, mentor engineers, drive design/capacity planning, integrate AI-assisted workflows, and improve cross-team reliability.
Top Skills: Active DirectoryAnsibleApache AirflowSparkAWSAzureBashBazelBitbucketCassandraChefDatadogDnsFirewall RulesGCPGitGithub ActionsGitlabGitlab CiGoGrafanaHoneycombHumio/LogscaleJenkinsKafkaKubernetesLoad BalancersMongoDBMySQLNasNew RelicNfsObject StorageOpensearchOraclePostgresPowershellPrometheusPulsarPuppetPythonRabbitMQRedis/ValkeyRedpandaRoutingSaltSanSplunkTerraformVarnishVipsWindows Server

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account