Banyan Software Logo

Banyan Software

Senior SRE

Posted 18 Days Ago
Remote
Hiring Remotely in United States
120K-170K Annually
Senior level
Remote
Hiring Remotely in United States
120K-170K Annually
Senior level
Operate and maintain highly available, secure, containerized SaaS applications across AWS and Azure. Responsibilities include rotating 24x7 on-call coverage, observability, incident and security response, disaster recovery, Terraform infrastructure-as-code, CI/CD automation, cloud integration, performance optimization, and developing AI agents to automate SRE and DevSecOps workflows.
The summary above was generated by AI

Banyan Software is the best permanent home for software businesses that serve specialized industries, their employees, and their customers. With a buy-grow-and-hold-for-life approach and a permanent capital base, Banyan acquires and grows companies worldwide, honoring founder legacies and helping portfolio companies modernize through shared AI expertise and operational discipline. Founded in 2016, Banyan operates more than 120 portfolio companies across North and South America, Europe, and APAC, and has appeared on the Inc. 5000 list for six consecutive years. The Banyan Software Foundation, endowed with $100 million in Banyan stock, leverages technology to build a greener and more equitable world.


Job Title: Senior SRE (Site Reliability Engineer) – Modernized Application Operations

Remote: US/Canada 

Overview

We are seeking a highly experienced and hands-on SRE to own the operational excellence of the modernized SaaS applications produced by the Banyan AI Factory. This is not a role focused on building the factory itself; instead, you will run the reliability of the modernized applications the factory delivers to our Operating Companies (OpCos).

You will join a team that provides 24x7 coverage with rotating on-call responsibilities, serving as Tier 1 Site Reliability Engineering (SRE) for our OpCos’ distributed applications. Day to day this will include: automated deployments, cloud service integration, application performance and availability monitoring/observability, and security incident response across our two target clouds — Amazon Web Services (AWS) and Microsoft Azure. The ideal candidate has a track record of keeping secure, highly available production systems running at scale.

Key Responsibilities

  • 24x7 Operations & On-Call: Operate as part of a team providing round-the-clock coverage of OpCo containerized applications, participating in a rotating on-call schedule to ensure continuous availability and rapid response.
  • Tier 1 SRE & Operations: Serve as Tier 1 SRE for the modernized applications, managing day-to-day cloud integrations across our two target clouds — AWS and Azure — to keep production systems healthy, performant, and secure.
  • Performance & Availability Monitoring/Observability: Implement and maintain robust application observability tooling (monitoring, logging, tracing) to track performance and availability, proactively detect degradation, and drive down mean-time-to-detect and mean-time-to-resolve.
  • Disaster Recovery and Service Restoration: Develop, maintain, test, and execute disaster recovery and business continuity procedures. Ensure the timely recovery and restoration of services following geographic disruptions, cyber incidents, infrastructure failures, or other disaster events.
  • Security Incident Response: Respond to security incidents and operational events affecting OpCo SaaS platforms, executing established runbooks, coordinating remediation
  • Automation & Infrastructure-as-Code : Use Infrastructure-as-Code (Terraform) and CI/CD pipelines (e.g., GitHub Actions, GitLab CI) to manage, deploy, and automate the operational environments of modernized applications, reducing toil and improving consistency.
  • AI Agents & DevSecOps Scale: Build scale in our DevSecOps practice by designing, building, and operating AI agents that automate SRE tasks and incident response, reducing toil and accelerating detection, triage, and remediation.
  • Hands-on Problem Solving: Serve as a technical escalation point for operational challenges, applying strong analytical skills to resolve infrastructure, network, and automation issues across distributed, multi-tenant SaaS environments while navigating technical ambiguity.

Required Qualifications & Experience

  • Experience: 5–7 years of progressive experience in Software Engineering, and/or Site Reliability Engineering, with a focus on operating distributed systems.
  • Automation Coding Experience: Deep expertise in Python, Javascript, or Go. Building automation and integrations between tools. This may be with AI assistance, but you must have a deep understanding of the code and scripting principals such as: authentication, parallelization, triggering, APIs, data transformation, etc.
  • Containerization: Deep expertise in container technologies (Docker/Kubernetes) supporting highly scalable and resilient distributed systems.
  • Infrastructure-as-Code with Terraform: Have experience working with modules at scale. This is a requirement for the role.
  • Cloud Native Services: hands-on experience operating production workloads on Amazon Web Services (AWS) (e.g., EC2, Lambda, EKS, S3, RDS) and / or Microsoft Azure (e.g., Container Apps, AKS, Container Storage).
  • CI/CD & Automation: Deep history of hands-on work with CI/CD platforms (GitHub Actions, GitLab CI) and embedding DevSecOps practices directly into operational workflows.
  • Operations, Monitoring & Observability: Experience with application level logging, troubleshooting, and tracing tools, with a proven track record operating highly available production systems.
  • AI-Fluent Engineering: Experience with AI-assisted engineering tools such as Claude Code or similar
  • Application Performance Management (APM): Familiarity with APM tooling and practices (e.g., Datadog, New Relic, Dynatrace, or similar) to instrument, profile, and optimize application performance in production.
  • Incident & Security Response: Demonstrated experience participating in on-call rotations, responding to production and security incidents, and executing disaster recovery procedures.
  • Communication & Collaboration: Exceptional communication, presentation, and collaboration skills, with a proven ability to coordinate across teams.
  • Education: Bachelor’s degree in Computer Science or a related technical field.

Preferred Skills (A Plus)

    Familiarity with advanced cloud security tools like Wiz, Prisma Cloud, and Checkov.


The expected base salary for this position is approximately USD $145,000 - $170,000 for US-based candidates and CAD $120,000 -  $145,000 for Canada-based candidates, excluding annual bonus and equity (when applicable). Salary is based on a number of factors, including market conditions, location, job-related skills and experience, and may vary accordingly.



Diversity, Equity, Inclusion & Equal Employment Opportunity at Banyan: Banyan affirms that inequality is detrimental to our Global Teams, associates, our Operating Companies, and the communities we serve. As a collective, our goal is to impact lasting change through our actions. Together, we unite for equality and equity. Banyan is committed to equal employment opportunities regardless of any protected characteristic, including race, color, genetic information, creed, national origin, religion, sex, affectional or sexual orientation, gender identity or expression, lawful alien status, ancestry, age, marital status, or protected veteran status and will not discriminate against anyone on the basis of a disability. We support an inclusive workplace where associates excel based on personal merit, qualifications, experience, ability, and job performance.

Please Note: Banyan Software does not accept unsolicited resumes or applications submitted via email, LinkedIn, or other direct channels. All candidates must apply through our official Careers site to be considered for employment. Applications submitted outside of our applicant tracking system will not be reviewed.

Recruitment Notice
Banyan Software may use artificial intelligence (AI) tools to assist in screening and/or assessing applicants during the recruitment process. All hiring decisions are made by our team. Personal information submitted through your application will be collected and used for recruitment purposes in accordance with applicable privacy laws. Contact us at any time with questions about our process or to request accommodation.

Beware of Recruitment Scams

We have been made aware of individuals fraudulently posing as members of our Talent Acquisition team and extending fake job offers. These scams may involve requests for personal information or payment for equipment. 

Protect yourself by following these steps:

  • Verify that all communications from our recruiting team come from an @banyansoftware.com email address.
  • Remember, employers will never request payment or banking information during the hiring process.
  • If you receive a suspicious message, do not respond — instead, forward it to [email protected] and/or report it to the platform where you received it.

Your safety and security are important to us. Thank you for staying vigilant.

Similar Jobs

Yesterday
Remote
136K-201K Annually
Senior level
136K-201K Annually
Senior level
Insurance • Cybersecurity
Build and operate AWS infrastructure, automation, observability, CI/CD pipelines, and internal developer platform components. Improve system reliability through resilience testing, automated recovery, capacity planning, and SLOs. Own projects from planning through rollout, mentor engineers, establish infrastructure standards, collaborate with software and security teams, and participate in a low-volume on-call rotation.
Top Skills: AWSCi/CdDistributed TracingEcsGithub ActionsGoInfrastructure As CodeKafkaKubernetesMicroservicesPythonSlosSystem MetricsTerraform
3 Days Ago
Remote
United States
Senior level
Senior level
Other
Lead platform infrastructure reliability initiatives for a microservices-based cloud SaaS solution. Design scalable, automated, observable infrastructure; define platform roadmaps; guide engineering teams through design reviews; reduce operational toil; build monitoring and alerting; and participate in on-call incident response. The role uses AWS, Kubernetes on EKS, Istio, infrastructure-as-code, GitLab, Argo CD, and Datadog.
Top Skills: Amazon EksArgo CdAWSAws CdkDatadogEc2GitlabIstioKubernetesRdsS3Terraform
18 Days Ago
Remote
123K-167K Annually
Senior level
123K-167K Annually
Senior level
Software
Build and operate scalable, resilient, distributed services across AWS and Azure. The Senior Site Reliability Developer will automate infrastructure, develop platforms and frameworks, establish monitoring and incident response practices, define runbooks, reduce operational toil, lead delivery projects, mentor engineers, and participate in technical interviews and on-call rotations. The role requires expertise in cloud infrastructure, infrastructure as code, CI/CD, Kubernetes, observability, telemetry, and SRE practices including SLOs, SLIs, and error budgets.
Top Skills: Amazon CloudwatchAmazon EcrAmazon EcsAmazon EksAmazon KinesisAmazon RdsAmazon RedshiftAmazon SqsAnsibleArtifactoryAuth0AWSAzureAzure Container AppsAzure DevopsAzure MonitorCloudamqpDockerDropwizardElasticsearchGitGithub ActionsGitlabHibernateJavaJavaScriptJenkinsKubernetesMongoDBMySQLObserve Inc.OpentelemetryPrometheusPythonRabbitMQReactRedshift SpectrumReduxSpring BootTemporalTerraformTypescript

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account