xAI Logo

xAI

Software Engineer, Observability

Reposted One Month Ago
In-Office
Palo Alto, CA, USA
180K-440K Annually
Mid level
In-Office
Palo Alto, CA, USA
180K-440K Annually
Mid level
Build and operate the observability platform: design scalable metrics/logs/tracing infrastructure, implement high-performance telemetry pipelines, develop APIs/query engines/UIs, enforce instrumentation and alerting best practices, and own system reliability and performance.
The summary above was generated by AI

SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates.

ABOUT THE ROLE:

The Observability team builds and operates the core infrastructure that enables engineers to monitor, debug, and optimize the performance and reliability of their systems. We handle telemetry at massive scale — billions of time series and petabytes of logs — with strict performance and availability requirements.

You will be part of the small, high-impact team responsible for building and maintaining SpaceXAI’s observability platform. You’ll own critical systems that power metrics, logs, tracing, and alerting enabling engineering teams to operate services at scale, identify issues before they impact users, and drive systemic reliability improvements.

RESPONSIBILITIES:
  • Design and implement scalable observability infrastructure for metrics, logging, and tracing.
  • Build high-performance telemetry pipelines that handle massive ingestion volumes.
  • Develop APIs, query engines, and UIs that allow engineers to get real-time insights into their services.
  • Define and enforce best practices for instrumentation, alerting, and reliability across the company.
  • Partner with infrastructure and product teams to deeply integrate observability into our internal platforms.
  • Own the reliability, scalability, and performance of the observability stack end-to-end.
BASIC QUALIFICATIONS:
  • Production-level proficiency in Go, Rust, Scala, or a similar languages
  • Deep understanding of distributed systems and telemetry architecture.
  • Experience building and operating infrastructure at scale.
  • Familiarity with observability stacks such as Prometheus, Grafana, OpenTelemetry, VictoriaMetrics, or ClickHouse.
  • Experience with Kafka, Redis, or large-scale time series databases.
  • Experience operating observability pipelines in Kubernetes or similar orchestration environments.
COMPENSATION AND BENEFITS:

$180,000 - $440,000 USD

Base salary is just one part of our total rewards package at SpaceXAI, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long-term disability insurance, life insurance, and various other discounts and perks.

SpaceXAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice.

HQ

xAI Palo Alto, California, USA Office

1450 Page Mill Road, Palo Alto, CA, United States

xAI San Francisco, California, USA Office

3180 18th St., San Francisco, CA, United States

Similar Jobs

26 Days Ago
In-Office
Sunnyvale, CA, USA
153K-204K Annually
Senior level
153K-204K Annually
Senior level
Cloud • Information Technology • Machine Learning
Design, develop, and maintain network datapath monitoring and observability infrastructure for GPU cloud services. Build real-time network telemetry and analytics pipelines across host networking, smart NICs, and overlay/underlay networks. Collaborate with DevOps, production, and data platform teams; troubleshoot Kubernetes and cloud networking issues; participate in on-call support, code reviews, architecture decisions, and continuous infrastructure improvements.
Top Skills: BgpCloud InfrastructureGoKernel NetworkingKubernetesKubernetes ControllersKubernetes OperatorsNetwork TelemetryNetwork VirtualizationObservabilityOverlay NetworksPythonSmart NicsSoftware-Defined Networking (Sdn)Tcp/IpUnderlay Networks
One Month Ago
Easy Apply
Remote or Hybrid
United States
Easy Apply
106K-209K Annually
Mid level
106K-209K Annually
Mid level
Big Data • Cloud • Software • Database
Develop and improve MongoDB’s distributed database networking and observability components. Responsibilities include writing production-level C++ code, enhancing performance, availability, scalability, and resource efficiency, integrating observability frameworks, leading medium-sized projects, collaborating across the product lifecycle, and participating in code reviews and technical design. The role requires strong software architecture fundamentals and an interest in networking, observability, and computer systems internals.
Top Skills: C++MongoDBOpentelemetry
13 Days Ago
In-Office or Remote
Santa Clara, CA, USA
152K-242K Annually
Senior level
152K-242K Annually
Senior level
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
Design, implement, and operate large-scale observability and telemetry platforms focused on performance, reliability, monitoring, logging, and alerting. Manage services throughout their lifecycle, including system design, deployment, capacity planning, automation, incident response, and postmortems. Build infrastructure tools and frameworks for private and public cloud systems, improve availability and latency, and participate in on-call support for production environments.
Top Skills: Cloud ComputingContainersContinuous DeliveryContinuous DeploymentDistributed SystemsDockerGoGrafanaInfrastructure AutomationKubernetesLinuxNetworkingObservabilityOpenstackOpentelemetryPerlPrometheusPythonRubyTelemetry

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account