AgileEngine Logo

AgileEngine

DevOps / Site Reliability Engineer ID70127

Reposted 3 Hours Ago
Be an Early Applicant
In-Office
San Francisco, CA, USA
Senior level
In-Office
San Francisco, CA, USA
Senior level
Maintain operational resilience across Azure, AWS, and GCP in a 24x7 environment. Build Terraform-based security baselines, optimize CI/CD pipelines, monitor cloud security posture with Wiz, and lead major-incident response as Incident Commander. Drive remediation through closure, communicate incident updates to technical and executive audiences, and create incident-management playbooks, runbooks, and escalation procedures. Mentor SREs and develop automated runbooks for compliant, secure platforms.
The summary above was generated by AI
AgileEngine is an Inc. 5000 company that creates award-winning software for Fortune 500 brands and trailblazing startups across 17+ industries. We rank among the leaders in areas like application development and AI/ML, and our people-first culture has earned us multiple Best Place to Work awards.

WHY JOIN US
If you're looking for a place to grow, make an impact, and work with people who care, we'd love to meet you!

ABOUT THE ROLE
We are looking for a DevOps / Site Reliability Engineer to maintain operational resilience across Azure, AWS, and GCP in a 24x7 environment. This role blends platform engineering with incident command, using Terraform, CI/CD pipelines, and CSPM tools like Wiz. You will lead major-incident calls, own remediation follow-through, and build the playbooks that guide response.

WHAT YOU WILL DO
- Scale and maintain the ability to drive operational stability across multi-cloud environments (Azure, AWS, GCP).
- Engineer unified security policies and configuration baselines using IaC (Terraform) to prevent misconfigurations.
- Design, maintain, and optimize enterprise CI/CD pipelines to support continuous ASPM ingestion and deployment.
- Act on continuous monitoring alerts, utilizing Cloud Security Posture Management (CSPM) tools like Wiz to secure workloads.
- Serve as Incident Commander on major and critical incidents — running the bridge, directing technical workstreams, making time-critical decisions, and coordinating cross-functional responders under pressure.
- Own the post-incident loop — track remediation items to closure, hold owning teams accountable to timelines, and drive systemic fixes and preventative actions across groups.
- Draft and send clear, accurate, audience-appropriate incident notifications and status updates to technical teams, management, and stakeholders throughout the incident lifecycle.
- Develop, maintain, and socialize divisional / group-level incident-management playbooks, runbooks, and escalation procedures that standardize response and reduce time-to-resolution.

MUST HAVES
- You must be authorized to work for ANY employer in the US (e.g., Green card holders, TN visa holders, GC EAD, H4 EAD, U4U with EAD), as we are unable to sponsor or take over employment visa sponsorship at this time;
- 5+ years of experience.
- In-depth architectural expertise in multi-cloud defense, federated IAM, and zero-trust principles.
- Strong practical experience with Kubernetes, Terraform, CI/CD orchestration, and Python/Go scripting.
- Senior-level, hands-on incident-command experience driving major/critical incident calls to resolution in a 24x7 production environment.
- Proven track record of remediation follow-up — coordinating with teams and holding owners accountable until issues are fully closed.
- Demonstrated skill drafting and issuing incident notification communications to both technical and executive audiences.
- Direct experience authoring divisional/group incident-management playbooks and escalation procedures.
- Fully autonomous.
- Drives the architecture of complex automated runbooks and mentors Middle-level SREs.
- Extensive experience deploying and tuning APIs from modern CNAPP/CSPM platforms, ideally Wiz.
- Prior experience building platforms subject to strict financial compliance standards (PCI-DSS, SOC2).
- Upper-intermediate English level.

NICE TO HAVES
- PagerDuty — hands-on experience with on-call scheduling, alert routing, and incident orchestration.
- ServiceNow — familiarity with incident, problem, and change management workflows and reporting.

PERKS AND BENEFITS
- Growth without limits: build your skills through mentorship, internal TechTalks, challenging projects, and a dedicated annual learning budget
- Competitive compensation: get recognition that reflects your skills and impact, with regular performance and compensation reviews
- Flexibility: work 100% remotely with flexible hours that support focus, autonomy, and a healthy work rhythm
- Meaningful, modern projects: build impactful products using modern technologies alongside global teams and leading brands
- Collaborative culture: join a supportive environment with zero micromanagement where ideas are welcomed and contributions are recognized
- Well-being & support: access local well-being programs and people-focused support tailored to your location

Similar Jobs

8 Minutes Ago
Hybrid
15-20 Hourly
Junior
15-20 Hourly
Junior
eCommerce • Fashion • Retail • Sales • Wearables • Design
Provides customer service and sales support in a luxury retail store. Responsibilities include welcoming clients, operating POS and cash wrap, suggesting products, processing shipments and transfers, maintaining stock levels, replenishing the sales floor, executing visual merchandising updates, supporting social media engagement, and following housekeeping and loss-prevention procedures. Requires flexible availability, physical ability to handle stockroom and sales-floor tasks, and strong communication and organizational skills.
Top Skills: InternetIpadLaptopMobile PosPosWalkie-Talkie
8 Minutes Ago
In-Office
99K-162K Annually
Senior level
99K-162K Annually
Senior level
Aerospace • Information Technology • Software • Cybersecurity • Design • Defense • Manufacturing
Develop and execute ASIC, FPGA, and SoC design verification activities using simulation, SystemVerilog, UVM, coverage analysis, debugging, requirements traceability, reusable verification infrastructure, and UVCs. Support integration and system-level investigations, contribute to coverage closure, communicate risks and dependencies, and progressively own verification scope of moderate complexity. Level II engineers develop independence, while Level III engineers independently manage assigned verification scope.
Top Skills: AsicFpgaOopRtlSimulationSocSystemverilogUvcUvm
8 Minutes Ago
In-Office
99K-162K Annually
Mid level
99K-162K Annually
Mid level
Aerospace • Information Technology • Software • Cybersecurity • Design • Defense • Manufacturing
Designs and supports ASIC, FPGA, and mixed digital systems for aerospace and defense programs. Responsibilities include requirements development, architecture support, HDL/RTL development, simulation, debugging, synthesis, FPGA implementation, timing analysis, verification collaboration, integration, troubleshooting, and technical status communication. The role may involve owning functional blocks, supporting subsystem efforts, and contributing to cross-functional technical leadership depending on experience level.
Top Skills: AsicClock-Domain Crossing (Cdc)FpgaHdlRtlSocStatic Timing AnalysisVlsi

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account