Braintrust Logo

Braintrust

Platform Support Engineer

Posted 5 Days Ago
Hybrid
San Francisco, CA, USA
Mid level
Hybrid
San Francisco, CA, USA
Mid level
Provides customer-facing technical support for hybrid and self-hosted deployments across AWS, Azure, and GCP. Responsibilities include debugging Kubernetes, Terraform, networking, IAM, TLS, performance, database, and reliability issues; leading incident response; contributing backend and infrastructure fixes; building diagnostics and health checks; maintaining documentation; and participating in on-call support.
The summary above was generated by AI
About the company

Braintrust is the agent observability platform. By actively applying intelligence to agent traces and automatically surfacing the most critical patterns, Braintrust gives teams the visibility to understand how agents behave in production and the tools to improve them.
Teams at Notion, Stripe, Box, OpenAI, and Cloudflare use Braintrust to trace their agents, find the issues in their observability data, and run evals that tell them how to improve.

About the role

Our largest customers don't just use Braintrust — they run it. They deploy our stack inside their own AWS, Azure, and GCP accounts, behind their own VPCs, under their own compliance requirements, at their own scale. When a hybrid deployment stalls, when ingest backs up, when a query that was fast last week isn't, they come to us. Platform Support is the team that owns that. We're the technical front line for infrastructure, performance, and reliability.

We're hiring Platform Support Engineers at both mid and senior levels to join a small, high-ownership team. You'll work shoulder to shoulder with our Cloud Infrastructure and Engineering teams, and alongside our Developer Support Engineers, who own the SDK and API side of the customer experience. If you like hard infrastructure problems, and you like them more when a real customer is on the other end, this is the role.

What you'll do
  • Own customer-facing support for hybrid and self-hosted Braintrust deployments across AWS, Azure, and GCP — from first install through steady-state operation.

  • Debug real infrastructure problems: Kubernetes workloads, Terraform state, networking and VPC configuration, IAM and permissions, TLS, and cloud-provider quirks.

  • Diagnose performance and reliability issues in the backend — ingest throughput, query latency, database and object-store behavior — using logs, metrics, and traces to get to cause rather than symptom.

  • Lead incident response for customer-impacting issues: triage, communicate clearly while it's still on fire, and drive it to resolution.

  • Ship fixes. Submit PRs to our backend services, Terraform modules, and deployment tooling rather than handing every problem to Engineering.

  • Build the tooling that makes the next one easier — diagnostics, health checks, preflight validation, and self-service paths that let customers unblock themselves.

  • Write and maintain the runbooks and deployment documentation that turn one hard-won answer into a permanent one.

  • Feed patterns back to Engineering and Product, so the recurring failure modes stop recurring.

  • Participate in an on-call rotation for critical customer issues.

What we're looking for
  • Experience in a customer-facing technical role — Support Engineering, SRE, DevOps, Solutions Architecture, or Infrastructure Engineering — or backend/infra engineering experience with real appetite for customer work.

  • Strong Kubernetes fundamentals: you can deploy, debug, and scale actual workloads, and read a failing pod's story from its events and logs.

  • Hands-on Terraform, and depth in at least one major cloud (AWS strongly preferred).

  • Comfort in a backend codebase — Python, TypeScript, or Go — enough to reproduce a bug, trace it to its source, and fix it.

  • Fluency with observability tooling, and the instinct to reach for data before opinion.

  • Clear, calm, direct communication under pressure, especially when the customer is technical, blocked, and losing time.

  • Ownership. You take a problem personally and follow it until the customer is running again.

Bonus points for
  • Supporting self-hosted or on-prem enterprise software, especially in regulated environments.

  • Multi-cloud experience, particularly Azure or GCP alongside AWS.

  • Database and data-infrastructure depth — Postgres, ClickHouse, or similar analytical stores.

  • Experience with observability, ML infrastructure, or developer platforms.

  • Familiarity with LLM APIs and how teams are building and evaluating agents in production.

  • Having built support or diagnostic tooling that measurably reduced ticket volume.

Why join Braintrust
  • Work on genuinely hard infrastructure problems, at the scale and pace of the teams building the best AI products in the world.

  • Join a team early enough to shape how it operates — its standards, its tooling, and its bar.

  • Sit close to both the customer and the code, with the mandate to fix things in either direction.

Benefits include
  • Medical, dental, and vision insurance

  • Daily lunch, snacks, and beverages

  • Flexible time off

  • Competitive salary and equity

  • Wifi & cellphone stipend

Equal opportunity

Braintrust is an equal opportunity employer. All applicants will be considered for employment without attention to race, color, religion, sex, sexual orientation, gender identity, national origin, veteran or disability status.

HQ

Braintrust San Francisco, California, USA Office

1 Main St, San Francisco, CA, United States, 94105

Similar Jobs

3 Days Ago
In-Office or Remote
5 Locations
108K-173K Annually
Senior level
108K-173K Annually
Senior level
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
Provide Tier 1 support for complex cloud platforms, troubleshoot distributed software and customer issues, investigate root causes, improve operational workflows, create runbooks and documentation, build support tooling, and coordinate with engineering, SRE, and other internal teams. The role supports production systems through an on-call rotation and requires expertise across cloud infrastructure, networking, storage, Kubernetes, Linux, and DevOps tooling, with GPU, MLOps, HPC, or SLURM experience preferred.
Top Skills: AWSAzureBlob StorageBlock StorageDatabasesDevOpsDistributed Training SystemsFile StorageGoogle Cloud Platform (Gcp)Gpu WorkloadsHigh-Performance Computing (Hpc)InfrastructureKubernetesLinuxMachine Learning InfrastructureMlopsNetworkingOracle Cloud Infrastructure (Oci)SlurmStorage
5 Days Ago
Hybrid
San Francisco, CA, USA
Mid level
Mid level
Artificial Intelligence • Software • Database • Analytics
Provide customer-facing support for hybrid and self-hosted deployments across AWS, Azure, and GCP. Debug Kubernetes, Terraform, networking, IAM, TLS, backend performance, databases, and object stores. Lead incident response, submit infrastructure and backend fixes, build diagnostic tooling, maintain runbooks, and participate in on-call support. Collaborate with Cloud Infrastructure, Engineering, Developer Support, and Product teams to resolve recurring customer issues.
Top Skills: AWSAzureClickhouseGCPGoIamKubernetesLlm ApisObservability ToolingPostgresPythonTerraformTlsTypescriptVpc
An Hour Ago
In-Office or Remote
San Francisco, CA, USA
200K-260K Annually
Expert/Leader
200K-260K Annually
Expert/Leader
Blockchain • Fintech • Payments • Financial Services • Cryptocurrency • Web3
Lead regional growth strategy for USDC by building partnerships with exchanges and OTC desks, delivering regionally compliant product initiatives, and creating AI-driven, automated growth platforms. Own experimentation frameworks, dashboards, and metrics to scale liquidity, adoption, and measurable share shift versus competing stablecoins through cross-functional execution and data-driven prioritization.
Top Skills: Agent-Driven SystemsAIAnalyticsAutomation PlatformsBlockchainDashboardsExperimentation FrameworksStablecoins

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account