Braintrust Logo

Braintrust

Cloud Infrastructure Engineer

Reposted 23 Days Ago
In-Office
San Francisco, CA, USA
Senior level
In-Office
San Francisco, CA, USA
Senior level
The Cloud Infrastructure Engineer will develop and maintain infrastructure using Terraform and Kubernetes, support multi-cloud deployments, and improve CI/CD processes.
The summary above was generated by AI
About the company

Braintrust is the agent observability platform. By actively applying intelligence to agent traces and automatically surfacing the most critical patterns, Braintrust gives teams the visibility to understand how agents behave in production and the tools to improve them.
Teams at Notion, Stripe, Box, OpenAI, and Cloudflare use Braintrust to trace their agents, find the issues in their observability data, and run evals that tell them how to improve.

About the role

We’re looking for a Cloud Infrastructure Engineer to help us build reliable, scalable infrastructure and give developers a world-class platform to ship code with speed and confidence. You’ll lead efforts across Terraform, Kubernetes, CI/CD, observability, and support, and play a key role in how we scale Braintrust both internally and for customers self-hosting our platform.

This is a high-impact role where you’ll contribute across our internal AWS environment and help customers deploy our stack in AWS, Azure, and GCP.

What you’ll do
  • Build and maintain Terraform modules for both internal infrastructure and customer deployments

  • Work directly with customers in Slack to support self-hosting and troubleshoot infrastructure issues. Build tools to make it easier for them to support themselves.

  • Own and improve our CI/CD pipeline: reduce build times, improve failure visibility, and enable safer, faster releases

  • Centralize and scale observability - including logs, metrics, dashboards, and alerts

  • Partner with engineering teams to build and evolve a secure, developer-friendly infrastructure platform

  • Support multi-cloud deployment patterns (AWS primarily, with Azure and GCP support for enterprise customers)

  • Implement tools and automation to improve deployment, rollback, and infrastructure reliability

Ideal candidate credentials
  • 5+ years of experience in DevOps, SRE, or Infrastructure Engineering roles

  • Deep experience with Terraform and at least one major cloud provider (AWS strongly preferred)

  • Strong Kubernetes skills: deploying, debugging, and scaling real workloads

  • Proficient in scripting or programming (Python, Typescript, or Go)

  • Experience supporting production systems and responding to incidents

  • Comfortable working directly with customers in a support or deployment context

  • Bonus: experience with multi-cloud environments or self-hosted enterprise software

Benefits include
  • Medical, dental, and vision insurance

  • Daily lunch, snacks, and beverages

  • Flexible time off

  • Competitive salary and equity

  • Wifi & cellphone stipend

Equal opportunity

Braintrust is an equal opportunity employer. All applicants will be considered for employment without attention to race, color, religion, sex, sexual orientation, gender identity, national origin, veteran or disability status.

HQ

Braintrust San Francisco, California, USA Office

San Francisco, CA, United States

Similar Jobs

8 Days Ago
Hybrid
San Francisco, CA, USA
Senior level
Senior level
Financial Services
Lead cloud networking architecture across AWS/Azure/GCP, define reusable network standards, manage technical risk and remediation, drive compliance and incident response, establish policy-as-code guardrails, mentor engineers, and coordinate cross-functional solution designs emphasizing security, resilience, and performance.
Top Skills: AnsibleApi-Driven Network AutomationAWSAzureBgpCloudFormationDnsEnterprise AiFirewallsGCPInfrastructure-As-CodeKubernetesLoad BalancingNatPolicy-As-CodeProxiesSd-WanService MeshTerraformVpnZero Trust
3 Days Ago
In-Office
94K-141K Annually
Mid level
94K-141K Annually
Mid level
Fintech • Payments • Financial Services
Own and manage a hybrid on-premises/Azure environment, including networks, servers, security, disaster recovery, and virtualization. Provide Tier 3 support, lead infrastructure projects and vendor implementations, document systems, participate in budgeting and on-call rotation, and drive modernization and security initiatives.
Top Skills: AzureDisaster RecoveryFirewallHybrid CloudMonitoringNetworkingServer InfrastructureVirtualization
8 Days Ago
Remote or Hybrid
United States
150K-170K Annually
Senior level
150K-170K Annually
Senior level
Information Technology • Database • Consulting
Lead design and delivery of agentic AI systems and LLMOps on AWS to generate, validate, and deploy infrastructure-as-code (Terraform). Build multi-agent applications, RAG pipelines, and AIOps for cloud operations; integrate AI into CI/CD and ITSM. Set technical direction, establish guardrails/policy-as-code, curate reusable Terraform modules, and mentor engineering teams.
Top Skills: AiopsAmazon Bedrock AgentsArtifactoryAWSCi/CdGitItsmJenkinsLangchainLlmopsRagSonarqubeTerraform

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account