Improve AWS production infrastructure reliability, observability, performance, and operational maturity. Build Terraform infrastructure, enhance CI/CD, automate operational work, manage incident response and on-call operations, lead postmortems, improve application resilience, support capacity planning and database reliability, and collaborate on security hardening and compliance. Mentor engineers and promote reliability practices across the organization.
Senior Site Reliability Engineer
Improve the reliability, performance, and operational maturity of a platform that supports the future of educational fundraising.
CONTRACT-TO-HIRE REMOTE - UNITED STATES SEATTLE / WEST COAST PREFERRED
About the role
Our client is looking for a hands-on Senior Site Reliability Engineer to improve the reliability, performance, and operational maturity of our platform. With our migration to AWS complete, this role will focus on strengthening our production environment: improving observability, automating infrastructure and operational work, enhancing incident response, and partnering with product engineers to build resilient systems. This is a high-impact role for someone who understands both infrastructure and application development. You will work across the stack, contribute code when appropriate, and help ensure our systems remain secure, scalable, and dependable as we grow.
What you'll do
• Operate, maintain, and improve our production infrastructure in AWS.
• Build and maintain infrastructure as code using Terraform.
• Improve monitoring, alerting, dashboards, and service-level indicators using New Relic or comparable observability platforms.
• Reduce alert noise and build systems that identify problems before customers are affected.
• Participate in the 24/7 on-call rotation and help coordinate the response to production incidents.
• Lead blameless postmortems and ensure corrective actions result in durable improvements.
• Partner with product engineers to diagnose performance and reliability issues throughout the application stack.
• Improve application resilience through appropriate use of timeouts, retries, queuing, backpressure, and idempotency.
• Improve CI/CD pipelines and deployment practices using platforms such as GitHub Actions, GitLab CI, or CircleCI.
• Automate repetitive operational work and reduce engineering toil.
• Create and maintain runbooks, system diagrams, troubleshooting guides, and production documentation.
• Support capacity planning, performance testing, database reliability, and production-readiness reviews.
• Collaborate with Security and Engineering teams on infrastructure hardening, access controls, logging, and compliance-related operational practices.
• Mentor engineers and promote effective reliability practices across the Engineering organization
What we're looking for
• 10+ years of overall software engineering, infrastructure, or systems experience, including at least 5 years in an SRE, Platform Engineering, DevOps, or production operations role.
• Previous professional software development experience and the ability to read, debug, and contribute to application code. • Strong, hands-on experience operating production workloads in AWS. • Experience building and maintaining infrastructure with Terraform or a similar infrastructure-as-code tool.
• Strong observability skills using New Relic, Datadog, or another modern monitoring platform. • Experience with incident response, on-call operations, postmortems, and production troubleshooting.
• Experience building or maintaining CI/CD pipelines.
• Working knowledge of networking, Linux, distributed systems, and relational databases.
• Strong judgment when balancing immediate operational needs with long-term maintainability.
• Clear communication skills and the ability to collaborate effectively across engineering disciplines.
• A track record of using automation to improve reliability and create leverage for other engineers.
Bonus points
• Experience with Ruby or Ruby on Rails.
• Strong PostgreSQL administration or performance-tuning experience.
• Experience operating enterprise SaaS products at scale.
• Familiarity with SLOs, SLIs, error budgets, capacity modeling, and load testing.
• Experience with payments, fintech, or other highly regulated systems.
• Experience supporting SOC 2 or similar security and compliance programs. Role details
• Contract-to-hire.
• Remote within the United States.
• Seattle-area or West Coast candidates are preferred to support occasional in-person collaboration, but exceptional candidates elsewhere should also be considered.
• Participation in a shared on-call rotation is required.
Similar Jobs
Cloud • Information Technology • Security • Software • Cybersecurity
Leads complex, cross-functional Customer Success programs from strategy through execution and operational adoption. Establishes governance, charters, metrics, milestones, accountability, and risk management across Professional Services, Technical Success, and Support. Uses data-driven analysis to assess program health, guide course corrections, and provide executive-ready reporting. The role also supports AI-driven solutions and enterprise software implementation in a remote Customer Success PMO environment.
Top Skills:
Ai/MlEnterprise Software PlatformsSaaS
Cloud • Information Technology • Security • Software • Cybersecurity
Leads Data Security product growth across the AMS region by driving sales strategy, supporting technical and commercial opportunities, enabling field teams, and translating customer insights into product strategy and roadmap improvements. The role covers competitive analysis, pricing, pipeline and churn analysis, customer escalations, deployment support, and data-driven business updates. It focuses on AI-powered data security products, including DLP, DSPM, and AI security solutions.
Top Skills:
Ai Agent SecurityAi/MlAispmAWSAzureCasbCloud DlpDnsDspmEmail DlpEndpoint DlpFirewallsGCPGenerative AiHttp/SInline DlpProxiesSAMLScimSecure BrowserSsl/TlsSspmSwgTcpUdpVpnZtna
Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Leads employee communications programs, events, giveaway campaigns, and cultural initiatives. Responsibilities include developing content, managing event logistics, coordinating vendors and global distribution, maintaining the employee events calendar, and triaging communications. The role requires independently managing complex workstreams, writing for varied audiences, and partnering across internal and external teams to deliver large-scale employee experiences.
What you need to know about the San Francisco Tech Scene
San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.
Key Facts About San Francisco Tech
- Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
- Major Tech Employers: Google, Apple, Salesforce, Meta
- Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
- Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
- Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
- Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine


