Together AI Logo

Together AI

Senior Software Engineer, Observability

Reposted 4 Days Ago
Be an Early Applicant
In-Office
San Francisco, CA, USA
200K-280K Annually
Senior level
In-Office
San Francisco, CA, USA
200K-280K Annually
Senior level
Design and implement a scalable observability platform (metrics, logs, traces) using Prometheus, Grafana, ClickHouse/ClickStack, and OpenTelemetry. Build telemetry pipelines, automated monitoring, SLIs/SLOs, alerting, anomaly detection, and runbooks. Develop IaC and custom tools with Go, Python, Terraform, Ansible, and Helm. Collaborate on distributed tracing, lead incident response, and define observability best practices.
The summary above was generated by AI

About the Role

Together AI is building the AI Acceleration Cloud, an end-to-end platform for the full generative AI lifecycle, combining the fastest LLM inference engine with state-of-the-art AI cloud infrastructure.

The AI Infrastructure team at Together AI is at the forefront of building and scaling the foundational systems that power our generative AI platform. The storage and observability team is crucial for designing, implementing, and maintaining robust distributed storage solutions, ensuring seamless data access and management. They are also responsible for developing comprehensive observability platforms, providing critical insights into system performance and GPU utilization, and proactively identifying and resolving issues.

Responsibilities

  • Design and implement a scalable observability platform (metrics, logs, traces) using tools like Prometheus, Grafana, ClickHouse, ClickStack, and OpenTelemetry, including telemetry data pipelines and log aggregation workflows.
  • Develop automated monitoring, alerting, and anomaly detection systems, including SLIs/SLOs, runbooks, and predictive analytics for critical services.
  • Build and deploy custom observability tools and infrastructure-as-code using Go, Python, Terraform, Ansible, and Helm.
  • Collaborate with engineering teams to enhance distributed tracing and application monitoring, and lead incident response with post-mortem analysis.
  • Define observability best practices.

Requirements

  • Expertise in observability platforms (Prometheus, Grafana, ClickStack, OpenTelemetry) and cloud-native monitoring services (AWS, GCP, Azure).
  • Strong programming skills in Go, Python, or similar languages, with proficiency in infrastructure-as-code tools (Terraform, Ansible, Helm).
  • Experience designing, operating, and scaling large-scale distributed systems and pipelines for high-volume data ingestion and real-time querying.
  • Deep understanding of containerization (Docker) and orchestration (Kubernetes).
  • Knowledge of microservices architecture, service mesh technologies, CI/CD pipelines, and GitOps workflows.
  • Expertise in managing databases (PostgreSQL, MongoDB, Redis) and time-series databases with high-cardinality data.

Preferred

  • Experience monitoring AI/ML infrastructure, GPU clusters, and custom metrics for model performance and training pipelines.
  • Background in high-frequency, low-latency systems monitoring, chaos engineering, and reliability testing.
  • Contributions to open-source observability projects.
  • Familiarity with security monitoring and compliance frameworks.
Compensation

We offer competitive compensation, startup equity, health insurance, and other benefits, as well as flexibility in terms of remote work. The US base salary range for this full-time position is: $200,000 - $280,000 + equity + benefits. Our salary ranges are determined by location, level and role. Individual compensation will be determined by experience, skills, and job-related knowledge.

Equal Opportunity

Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.

Please see our privacy policy at https://www.together.ai/privacy  


Together AI San Francisco, California, USA Office

584 Castro St, #2050, San Francisco, California , United States, 94114

Similar Jobs

2 Days Ago
In-Office
Mountain View, CA, USA
174K-299K Annually
Senior level
174K-299K Annually
Senior level
eCommerce
Design, build, and maintain observability platforms and tooling (metrics, logs, traces) for large-scale distributed systems. Collaborate with engineering and SRE teams to define SLOs/KPIs, evaluate observability tools, troubleshoot incidents, drive standardization, and mentor engineers while improving monitoring, alerting, and telemetry practices.
Top Skills: AppdynamicsAWSAzureCi/CdDockerDynatraceElastic StackGoGoogle Cloud PlatformGrafanaInfrastructure As CodeJaegerJavaKubernetesOpentelemetryPrometheusPythonRubyZipkin
6 Days Ago
In-Office
Santa Clara, CA, USA
200K-322K Annually
Senior level
200K-322K Annually
Senior level
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
Lead design, build, and operate AIOps and observability platforms (metrics, logs, traces, alerts, dashboards). Define roadmap and standards, collaborate with teams, mentor engineers, and implement ML/LLM-based anomaly detection, forecasting, root-cause analysis, and agentic debugging across large-scale cloud, on-prem, and bare-metal environments.
Top Skills: AlertmanagerBare MetalBigpandaC#ClickhouseDatadogDockerGenerative AiGoGrafanaJavaKafkaKubernetesLlmsLokiMachine LearningMicroservicesNatsNomadOpentelemetryPagerdutyPrometheusPythonVectorVictoriametrics
11 Days Ago
In-Office
San Francisco, CA, USA
173K-286K Annually
Senior level
173K-286K Annually
Senior level
Cloud • Software
Build and maintain high-volume event log pipelines and distributed observability services. Create tooling, automation, and auto-remediation, collaborate across platform teams, participate in on-call rotations, and enable teams to make data-driven performance and reliability decisions.
Top Skills: AstraAWSElasticsearchGoJavaKibanaLogstashPrometheusPython

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account