Together AI Logo

Together AI

Senior Network Engineer

Reposted One Month Ago
Be an Early Applicant
In-Office
San Francisco, CA, USA
190K-270K Annually
Senior level
In-Office
San Francisco, CA, USA
190K-270K Annually
Senior level
Design, deploy, and maintain large-scale hybrid data center networks; troubleshoot and analyze network issues; develop automation and tooling; evaluate and recommend hardware/software; manage vendor relationships; ensure reliability, scalability, and compliance; lead complex network projects and contribute to architecture and roadmaps.
The summary above was generated by AI
About the Role

Together AI is looking for a Senior Network Engineer to design, deploy, and operate the global network infrastructure supporting our production services and high-performance AI compute environments.

This is a hands-on engineering role for someone with deep networking expertise who can also troubleshoot across Linux, Kubernetes, automation, and application boundaries. You will work on large-scale, multi-vendor data center networks and help ensure they remain highly available, reliable, scalable, and performant.

The ideal candidate has strong networking fundamentals, experience operating complex networks at scale, and a structured, evidence-based approach to troubleshooting. You should be comfortable owning problems from initial investigation through root cause and resolution, including situations where the issue may extend beyond the network itself.

Requirements

  • 8+ years of professional experience designing, building, and supporting large-scale production data center, cloud, service-provider, or high-performance computing networks (excluding enterprise networks).
  • Deep understanding of TCP/IP and strong experience with technologies such as BGP, OSPF, VXLAN, EVPN, ECMP, and QoS.
  • Experience designing and supporting multi-tenant network environments using technologies such as VRFs, VLANs, overlays, and policy-based segmentation.
  • Hands-on experience deploying and troubleshooting network platforms from vendors such as Arista, Cisco, Juniper, and NVIDIA.
  • Strong troubleshooting skills using tools such as Wireshark, tcpdump, MTR, curl, nmap, and standard Linux networking utilities.
  • Ability to diagnose connectivity, latency, packet-loss, routing, and performance issues across the network, host, and application layers.
  • Experience developing or maintaining network automation using Python, Ansible, or similar tools.
  • Experience working through a Git-based software development lifecycle, including branching, code review, validation, linting, testing, CI/CD, deployment, and rollback.
  • Working knowledge of Kubernetes networking, including pods, services, CNIs, and basic connectivity troubleshooting.
  • Foundational knowledge of RDMA networking and technologies such as RoCE or InfiniBand.
  • Experience with cloud networking in AWS, GCP, or Azure.
  • Strong Linux administration and troubleshooting skills.

Responsibilities

  • Design, deploy, operate, and maintain global, multi-vendor, multi-protocol networks supporting high-performance AI compute infrastructure.
  • Troubleshoot complex network and application-connectivity issues, identify root causes, and drive problems through resolution.
  • Analyze telemetry, packet captures, logs, and performance data to identify network degradation, congestion, packet loss, and capacity constraints.
  • Participate in architecture and design reviews to ensure solutions meet requirements for performance, availability, scalability, security, and operational supportability.
  • Develop and maintain automation, validation, and operational tooling that improves network reliability and reduces manual effort.
  • Evaluate network hardware, software, optics, and emerging technologies for use in production environments.
  • Establish standards and operational best practices for network design, deployment, monitoring, change management, and incident response.
  • Lead projects addressing complex technical challenges and contribute directly to the network engineering roadmap.
  • Partner with infrastructure, systems, security, and application teams to troubleshoot issues that cross traditional ownership boundaries.

Preferred

  • Hands-on experience deploying or operating RoCE and/or InfiniBand fabrics.
  • Experience supporting GPU clusters, HPC environments, distributed storage, or other high-bandwidth and latency-sensitive workloads.
  • Understanding of AI training and inference traffic patterns and the demands they place on network infrastructure.
  • Experience operating networks spanning thousands of devices, multiple data centers, and multiple geographic regions.
  • Familiarity with AI-assisted engineering tools and the ability to validate, test, and safely deploy AI-generated automation or code.
About Together AI

Together AI, the AI Native Cloud, is purpose-built for AI engineers. AI application developers get high-performance inference that scales reliably, fine-tuning and reinforcement learning for creating frontier-level specialized models, and pre-training at massive scale for fully custom intelligence, all around a marketplace of leading open models that teams can run, adapt, and own. Trusted by Cursor, Decagon, ElevenLabs, Salesforce, and Zoom, Together serves 400+ trillion tokens a month.

Compensation

We offer competitive compensation, startup equity, health insurance and other competitive benefits. The US base salary range for this full-time position is: $190,000 - $280,000 + equity + benefits. Our salary ranges are determined by location, level and role. Individual compensation will be determined by experience, skills, and job-related knowledge.

Equal Opportunity

Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.

Please see our privacy policy at https://www.together.ai/privacy  

Together AI San Francisco, California, USA Office

584 Castro St, #2050, San Francisco, California , United States, 94114

Similar Jobs

7 Days Ago
In-Office
159K-215K Annually
Senior level
159K-215K Annually
Senior level
Aerospace • Information Technology • Software • Cybersecurity • Design • Defense • Manufacturing
Designs, implements, maintains, and optimizes secure LAN, WAN, wireless, and cloud network infrastructure. Leads deployment of routers, switches, firewalls, VPNs, and load balancers; manages network security; develops resilience, disaster recovery, and business continuity plans; and maintains technical documentation. Provides technical leadership and mentorship while collaborating cross-functionally to identify and manage network risks and opportunities. The onsite role requires U.S. citizenship, security-clearance eligibility, and Secret clearance after starting.
Top Skills: AnsibleAWSAzureAzure ExpressrouteCcnaCiscoCloud NetworkingDellIat Level 2Intrusion Detection And Prevention SystemsIpsecJnciaJuniperLanNetwork AutomationNetwork MonitoringPalo Alto Networks FirewallsPalo Alto Networks PanoramaPcnsaPythonQuality Of Service (Qos)Routing ProtocolsSolarwindsSsl VpnSwitching ProtocolsVpnWanWireless NetworkingWireshark
13 Days Ago
In-Office
San Mateo, CA, USA
140K-211K Annually
Senior level
140K-211K Annually
Senior level
Aerospace • Artificial Intelligence • Machine Learning • Robotics • Software
Design, deploy, and maintain resilient hybrid network infrastructure across data centers, remote sites, and public clouds. Manage routing, switching, WLAN, firewalls, WAN, VPN, secure connectivity, network automation, capacity planning, monitoring, troubleshooting, upgrades, patching, and vulnerability remediation. Collaborate with cloud, platform, security, and infrastructure teams, engage vendors, document operations, mentor teams, participate in on-call rotations, and support classified and unclassified environments.
Top Skills: 802.1XAnsibleArista EosAWSAzureAzure Virtual WanCatoCcnpCloud WanExpressrouteFirewallsFortiosHttpsJncisJuniperJunosLinuxNacNetwork AutomationNx-OsPalo AltoPan-OsPanoramaPrismaPythonRadiusRoutingSaseSftpSwitchingTerraformTransit GatewayVcp-NvVMwareVpnWanWindowsWlanZscalerZtna
13 Days Ago
Easy Apply
Remote or Hybrid
USA
Easy Apply
225K-250K Annually
Senior level
225K-250K Annually
Senior level
Fintech • Information Technology • Software • Financial Services
Own end-to-end network reliability across data centers, AWS, and GCP, including routing, switching, DNS, load balancing, service mesh, VPNs, circuits, and cross-connects. Automate network configuration with Ansible, Terraform, GitOps, CI validation, and IPAM. Engineer failover, observability, security hardening, client connectivity, and Kubernetes ingress while participating in an interrupt-contained on-call rotation.
Top Skills: AnsibleAWSBgpCisco Ios-XeDnsElkEnvoyGCPGitopsGrafanaHaproxyInfrahubIpamKubernetesNautobotNetboxNginxPalo AltoPanoramaPrometheusTerraformTlsVpn

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account