xAI Logo

xAI

Software Engineer - Network (C++)

Reposted One Month Ago
In-Office
Palo Alto, CA, USA
180K-440K Annually
Junior
In-Office
Palo Alto, CA, USA
180K-440K Annually
Junior
Design, implement, and operate core networking software for a large-scale GPU datacenter fabric. Develop routing and traffic-engineering algorithms, real-time switch software, prototypes and experiments, deployment and CI tooling, monitoring, and testing to maximize performance and reliability of AI training infrastructure.
The summary above was generated by AI

SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates.

ABOUT THE ROLE:

At SpaceXAI, we design, build, and operate Colossus from the ground up. This includes the massive GPU clusters, high-speed interconnect fabric, and the software that makes it all work at unprecedented scale. Colossus powers Grok and our frontier AI models with a custom, high-performance datacenter network that delivers ultra-low latency and massive bandwidth across hundreds of thousands of GPUs.

As a Software Engineer on the Colossus Networking team, you will develop the core networking software that maximizes the performance and reliability of our datacenter fabric. Your work will directly impact training efficiency, model convergence, and the speed at which we can push the frontier of AI.

Our engineers own the full lifecycle of their software — from design and implementation to deployment, monitoring, and iteration based on real-world performance at scale. You will solve hard problems in distributed systems, high-performance networking, and real-time control of one of the largest AI supercomputers on Earth.

RESPONSIBILITIES:
  • Develop routing and traffic-engineering algorithms for the Colossus high-performance datacenter network.
  • Develop highly reliable, real-time software designed to run on the switches that form the backbone of our low-latency, high-bandwidth AI training fabric.
  • Participate in and lead architecture, design, and code reviews.
  • Develop prototypes and run experiments to validate key design decisions at both small and full-cluster scale.
  • Build tools for software development, deployment, data analysis, visualization, and testing across virtualized environments, hardware-in-the-loop setups, and live production clusters.
  • Deploy reliable software updates through continuous integration and release systems with rigorous testing and monitoring.
BASIC QUALIFICATIONS:
  • Bachelor’s degree in computer science, engineering, math, or a related technical discipline; OR 2+ years of professional software development experience in lieu of a degree.
  • Strong development experience in C or C++.
PREFERRED SKILLS AND EXPERIENCE:
  • Strong professional experience writing high-performance C/C++ in production environments.
  • Experience developing, debugging, and deploying software that runs at scale in real-world systems.
  • Deep knowledge of networking protocols (UDP, TCP/IP, RDMA, etc.), distributed systems, and large-scale datacenter fabrics.
  • Background in real-time systems, high-performance computing, low-latency networking, or resource-constrained environments.
  • Creative problem-solving ability with exceptional analytical skills and strong engineering fundamentals.
  • Excellent written and verbal communication skills.
  • Ability to thrive in a fast-paced, dynamic environment with evolving requirements.
  • Experience with security considerations in large-scale distributed systems.
COMPENSATION AND BENEFITS

$180,000 - $440,000 USD

Base salary is just one part of our total rewards package at SpaceXAI, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long-term disability insurance, life insurance, and various other discounts and perks.

SpaceXAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice.

HQ

xAI Palo Alto, California, USA Office

1450 Page Mill Road, Palo Alto, CA, United States

xAI San Francisco, California, USA Office

3180 18th St., San Francisco, CA, United States

Similar Jobs

13 Minutes Ago
Hybrid
160K-195K Annually
Senior level
160K-195K Annually
Senior level
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Leads the strategy, architecture, and deployment of AI-enabled business intelligence infrastructure. Designs governed semantic layers, data models, and BI systems using dbt and enterprise data warehouses. Manages and coaches a technical engineering team, partners with executives and cross-functional stakeholders, and ensures reliable, high-performing data pipelines, warehouses, and reporting tools. The role also advances AI-driven analytics, data discovery, and self-service reporting across the organization.
Top Skills: Amazon RedshiftBigQueryData WarehousesDbtLookerNatural Language ProcessingPostgresPredictive ModelsSemantic LayersSnowflakeSQL
14 Minutes Ago
Remote or Hybrid
5 Locations
99K-232K Annually
Senior level
99K-232K Annually
Senior level
Artificial Intelligence • Professional Services • Business Intelligence • Consulting • Cybersecurity • Generative AI
Lead end-to-end NetSuite implementations as a functional consultant and team manager. Manage finance transformation initiatives, NetSuite order-to-cash, purchase-to-pay, and account-to-report workstreams, Advanced Revenue Management and SuiteBilling implementations, customizations, integrations, and testing. Lead onshore and offshore teams, address client needs, communicate complex concepts, and apply accounting and finance expertise to deliver effective technology solutions.
Top Skills: Erp IntegrationsNetSuiteNetsuite Advanced Revenue ManagementNetsuite Multibook AccountingOracleSuitebilling
15 Minutes Ago
Remote or Hybrid
5 Locations
77K-202K Annually
Senior level
77K-202K Annually
Senior level
Artificial Intelligence • Professional Services • Business Intelligence • Consulting • Cybersecurity • Generative AI
Implement and optimize Oracle NetSuite finance solutions, serving as a functional lead across full lifecycle implementations. Responsibilities include leading onshore and offshore teams, configuring order-to-cash, purchase-to-pay, account-to-report, Advanced Revenue Management, and SuiteBilling modules, designing customizations, conducting implementation testing, applying SuiteSuccess methodology, and advising clients on finance transformation and accounting processes.
Top Skills: Netsuite Advanced Revenue ManagementNetsuite SuitebillingNetsuite SuitesuccessOracle Netsuite

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account