Together AI Logo

Together AI

Senior Software Engineer, Observability

Reposted One Month Ago
Be an Early Applicant
In-Office
San Francisco, CA, USA
200K-280K Annually
Senior level
In-Office
San Francisco, CA, USA
200K-280K Annually
Senior level
Design and implement a scalable observability platform (metrics, logs, traces) using Prometheus, Grafana, ClickHouse/ClickStack, and OpenTelemetry. Build telemetry pipelines, automated monitoring, SLIs/SLOs, alerting, anomaly detection, and runbooks. Develop IaC and custom tools with Go, Python, Terraform, Ansible, and Helm. Collaborate on distributed tracing, lead incident response, and define observability best practices.
The summary above was generated by AI

About the Role

Together AI is building the AI Acceleration Cloud, an end-to-end platform for the full generative AI lifecycle, combining the fastest LLM inference engine with state-of-the-art AI cloud infrastructure.

The AI Infrastructure team at Together AI is at the forefront of building and scaling the foundational systems that power our generative AI platform. The storage and observability team is crucial for designing, implementing, and maintaining robust distributed storage solutions, ensuring seamless data access and management. They are also responsible for developing comprehensive observability platforms, providing critical insights into system performance and GPU utilization, and proactively identifying and resolving issues.

Responsibilities

  • Design and implement a scalable observability platform (metrics, logs, traces) using tools like Prometheus, Grafana, ClickHouse, ClickStack, and OpenTelemetry, including telemetry data pipelines and log aggregation workflows.
  • Develop automated monitoring, alerting, and anomaly detection systems, including SLIs/SLOs, runbooks, and predictive analytics for critical services.
  • Build and deploy custom observability tools and infrastructure-as-code using Go, Python, Terraform, Ansible, and Helm.
  • Collaborate with engineering teams to enhance distributed tracing and application monitoring, and lead incident response with post-mortem analysis.
  • Define observability best practices.

Requirements

  • Expertise in observability platforms (Prometheus, Grafana, ClickStack, OpenTelemetry) and cloud-native monitoring services (AWS, GCP, Azure).
  • Strong programming skills in Go, Python, or similar languages, with proficiency in infrastructure-as-code tools (Terraform, Ansible, Helm).
  • Experience designing, operating, and scaling large-scale distributed systems and pipelines for high-volume data ingestion and real-time querying.
  • Deep understanding of containerization (Docker) and orchestration (Kubernetes).
  • Knowledge of microservices architecture, service mesh technologies, CI/CD pipelines, and GitOps workflows.
  • Expertise in managing databases (PostgreSQL, MongoDB, Redis) and time-series databases with high-cardinality data.

Preferred

  • Experience monitoring AI/ML infrastructure, GPU clusters, and custom metrics for model performance and training pipelines.
  • Background in high-frequency, low-latency systems monitoring, chaos engineering, and reliability testing.
  • Contributions to open-source observability projects.
  • Familiarity with security monitoring and compliance frameworks.
About Together AI

Together AI, the AI Native Cloud, is purpose-built for AI engineers. AI application developers get high-performance inference that scales reliably, fine-tuning and reinforcement learning for creating frontier-level specialized models, and pre-training at massive scale for fully custom intelligence, all around a marketplace of leading open models that teams can run, adapt, and own. Trusted by Cursor, Decagon, ElevenLabs, Salesforce, and Zoom, Together serves 400+ trillion tokens a month.

Compensation

We offer competitive compensation, startup equity, health insurance, and other benefits, as well as flexibility in terms of remote work. The US base salary range for this full-time position is: $200,000 - $280,000 + equity + benefits. Our salary ranges are determined by location, level and role. Individual compensation will be determined by experience, skills, and job-related knowledge.

Equal Opportunity

Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.

Please see our privacy policy at https://www.together.ai/privacy  


Together AI San Francisco, California, USA Office

584 Castro St, #2050, San Francisco, California , United States, 94114

Similar Jobs

24 Days Ago
In-Office
Sunnyvale, CA, USA
153K-204K Annually
Senior level
153K-204K Annually
Senior level
Cloud • Information Technology • Machine Learning
Design, develop, and maintain network datapath monitoring and observability infrastructure for GPU cloud services. Build real-time network telemetry and analytics pipelines across host networking, smart NICs, and overlay/underlay networks. Collaborate with DevOps, production, and data platform teams; troubleshoot Kubernetes and cloud networking issues; participate in on-call support, code reviews, architecture decisions, and continuous infrastructure improvements.
Top Skills: BgpCloud InfrastructureGoKernel NetworkingKubernetesKubernetes ControllersKubernetes OperatorsNetwork TelemetryNetwork VirtualizationObservabilityOverlay NetworksPythonSmart NicsSoftware-Defined Networking (Sdn)Tcp/IpUnderlay Networks
11 Days Ago
In-Office or Remote
Santa Clara, CA, USA
152K-242K Annually
Senior level
152K-242K Annually
Senior level
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
Design, implement, and operate large-scale observability and telemetry platforms focused on performance, reliability, monitoring, logging, and alerting. Manage services throughout their lifecycle, including system design, deployment, capacity planning, automation, incident response, and postmortems. Build infrastructure tools and frameworks for private and public cloud systems, improve availability and latency, and participate in on-call support for production environments.
Top Skills: Cloud ComputingContainersContinuous DeliveryContinuous DeploymentDistributed SystemsDockerGoGrafanaInfrastructure AutomationKubernetesLinuxNetworkingObservabilityOpenstackOpentelemetryPerlPrometheusPythonRubyTelemetry
One Month Ago
In-Office
Sunnyvale, CA, USA
182K-242K Annually
Senior level
182K-242K Annually
Senior level
Cloud • Information Technology • Machine Learning
Senior Software Engineer focused on AI infrastructure performance insights and observability. Responsible for designing and building monitoring, instrumentation, metrics, and tooling to measure and optimize compute and platform performance while collaborating with infrastructure and ML teams.

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account