OpusClip

Founding DevOps Engineer (Infrastructure & Reliability)

Posted 11 Days Ago

In-Office

Palo Alto, CA

180K-280K Annually

Expert/Leader

In-Office

Palo Alto, CA

180K-280K Annually

Expert/Leader

In this role, you will manage the reliability and scalability of the Opus platform, stabilize processing clusters, and enhance incident response.

The summary above was generated by AI

🎨 OpusClip is the world's No.1 AI video agent, built for authenticity on social media.

We envision a world where everyone can authentically share their story through video, with no expertise needed. Within just 18 months of our launch, over 10 million creators and businesses have used OpusClip to enhance their social presence.

We have raised $50 million in total funding and are fortunate to have some of the most supportive investors, including SoftBank Vision Fund, DCM Ventures, Millennium New Horizons, Fellows Fund, AI Grant, Jason Lemkin (SaaStr), Samsung Next, GTMfund, Alumni Ventures, and many more.

Check out our latest coverage by Business Insider featuring our product and funding milestones, and our recognition as one of The Information's 50 Most Promising Startups in 2024.

Headquartered in Palo Alto, we are a team of 100 passionate and experienced AI enthusiasts and video experts, driven by our core values:

Be a Champion Team
Prioritize Ruthlessly
Ship fast, Quality Follows
Obsess over customers

Be a part of this exciting journey with us!

The Mission:

The Mission we are on is to find a hands-on founding DevOps Engineer to own the reliability and scalability of the Agent Opus platform - a leading agentic video generation product. This person will stabilize our processing clusters, design the next phase of our agentic platform, and serve as the technical bridge between our cloud infrastructure and our 15M+ users.

You will engineer the infrastructure strategy that underpins our trust and reliability in the market. You will help setup on-call rotation and own the full incident lifecycle from minimizing Time-to-Detect to tracking post mortem actions ensuring that our high-velocity growth never compromises our performance.

Key Responsibilities:

Infrastructure Architecture & Cluster Operations

Architect Dedicated Environments: Lead the design and implementation of high-throughput, isolated processing environments and clusters for compliance needs.
Scale Production: Drive general improvements in our Temporal clusters and production Kubernetes environments.
Technical Execution: Be hands-on with the stack to optimize resource allocation, reduce latency, and enforce isolation strategies for critical accounts.

Monitoring, Alerting & Detectability

Beat the Customer to the Alert: Overhaul our Datadog observability suite to aggressively reduce Time-to-Detect (TTD). You ensure we identify latency spikes and stalled projects before users do.
Threshold Tuning: tune alert thresholds to eliminate noise and focus on "symptom-based" alerts that reflect the actual user experience.
External SLO Ownership: Define and report on Service Level Objectives (SLOs), acting as the internal guarantor that we are meeting the targets we sold.

Incident Command & "Extreme Ownership"

First Responder & Mitigation: Serve as the first line of defense during outages. You will own immediate mitigation, including cluster debugging and manual scaling intervention if Horizontal Pod Autoscalers (HPA) fail or lag.
Drive Recovery Metrics: You are accountable for shortening Time-to-Mitigation (TTM) and Time-to-Recover (TTR). Your priority is to stop the bleeding first, then fix the wound.
Root Cause Analysis: Lead the post-mortem process to determine Time-to-Root Cause and implement systemic fixes. You will translate these technical findings into clear updates for Customer Experience (CX) and Leadership.
Accountability: Work collaboratively with Engineering Owners to track improvements against the reliability roadmap. You are responsible for flagging risks early and resetting expectations on platform performance when necessary.
Cross-Functional Bridge: Serve as the primary technical voice to the Customer Experience (CX), Marketing, Sales, and Leadership teams. You will translate technical constraints and roadmaps into clear updates for stakeholders.

Qualifications:

Production K8s & Temporal: Expert-level ability to debug Kubernetes internals (HPA logic, node scaling) and operate stateful workflow engines (Temporal) at scale.
Incident Command: Proven track record as a primary first responder, demonstrating the ability to aggressively reduce Time-to-Mitigation (TTM) and Time-to-Recover (TTR).
Observability Architecture: Experience architecting Datadog SLOs and tuning alerts to distinguish system noise from actual user pain.
Automation: Strong proficiency in Python or Bash to automate manual recovery and operational tasks.
⭐️ Bonus: Experience scaling GPU/video rendering workloads or thriving in early-stage, high-velocity startups.

Tech Stack:

Orchestration & Compute: Kubernetes (GKE), Google Cloud, Docker, Horizontal Pod Autoscaling (HPA).
Workflow Engine: Temporal
Observability: Datadog (APM, Custom Metrics, Alerting).
Infrastructure as Code: Terraform or similar IaC tools.
Scripting & Backend: Python (primary), Bash.
Data & Storage: Redis, Postgres.

What We Offer!

Competitive Base Salary + Equity + Bonus
- Bonus Range: $15,000 - $70,000
Comprehensive medical, dental, and vision plans.
Global team offsite.
A challenging but exciting work environment with a strong culture of innovation and entrepreneurship.
Opportunities for advancement and career growth.

Location (Onsite):

US (Onsite) - Palo Alto, CA

EEO

OpusClip is proud to be an equal opportunity employer. We do not discriminate in hiring or any employment decision based on race, color, religion, national origin, age, sex (including pregnancy, childbirth, or related medical conditions), marital status, ancestry, physical or mental disability, genetic information, veteran status, gender identity or expression, sexual orientation, or other applicable legally protected characteristics. OpusClip considers qualified applicants with criminal histories, consistent with applicable federal, state and local law. Opus Clip is also committed to providing reasonable accommodations for qualified individuals with disabilities and disabled veterans in our job application procedures.

Top Skills

Bash

Datadog

Docker

GCP

Kubernetes

Postgres

Python

Redis

Temporal

Terraform

530 University Ave, Palo Alto, California, United States, 94301

Similar Jobs

Elekta

Senior Devops Engineer

2 Days Ago

In-Office

San Jose, CA, USA

145K-165K Annually

Senior level

145K-165K Annually

Senior level

Healthtech • Biotech

The Senior DevOps Engineer will build and maintain DevOps capabilities, enhance infrastructure, and facilitate collaboration across teams to ensure efficient software delivery.

Top Skills: Active DirectoryAzureAzure Automation DscDockerGitJenkinsPuppetSvnVmware Vcenter

Zoom

Devops Engineer

6 Days Ago

In-Office

San Jose, CA, USA

99K-229K Annually

Senior level

99K-229K Annually

Senior level

Artificial Intelligence • Information Technology • Software

The Senior DevOps Engineer will enhance the reliability and performance of SaaS platforms, manage incidents, and mentor teams while implementing automation and disaster recovery plans.

Top Skills: AutomationDevOpsDistributed SystemsIncident ManagementMonitoringObservabilitySaaSSre

Archer Aviation

Senior DevOps

8 Days Ago

Easy Apply

In-Office

San Jose, CA, USA

Easy Apply

133K-200K Annually

Senior level

133K-200K Annually

Senior level

Aerospace

Senior DevOps Engineer responsible for infrastructure strategy, focusing on automation, CI/CD, configuration management, and operational deployment of large language models in cloud and on-premise environments.

Top Skills: AnsibleAWSAzureBashDatadogDockerGCPKubernetesLinuxLitellmOpenrouterPythonTerraformWindows

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
Major Tech Employers: Google, Apple, Salesforce, Meta
Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

OpusClip

Founding DevOps Engineer (Infrastructure & Reliability)

Top Skills

OpusClip Palo Alto, California, USA Office

Similar Jobs

Senior Devops Engineer

Devops Engineer

Senior DevOps

What you need to know about the San Francisco Tech Scene

Key Facts About San Francisco Tech