TensorWave Logo

TensorWave

Senior Manager, Cluster Engineering & Deployment

Posted Yesterday
Be an Early Applicant
Remote
Hiring Remotely in USA
Senior level
Remote
Hiring Remotely in USA
Senior level
Leads large-scale cluster deployment from rack delivery through production acceptance, including network bring-up, cabling verification, GPU-fabric integration, validation, burn-in, and RCCL performance testing. Manages deployment engineers across concurrent builds, improves deployment velocity, establishes operational playbooks and acceptance gates, coordinates with data center teams and vendors, and drives defect resolution across network, cabling, and hardware partners.
The summary above was generated by AI

About TensorWave

Our mission is simple: deliver seamless, secure, reliable, and resilient AI compute at scale. We've built a versatile cloud platform that eliminates infrastructure barriers, empowering builders to focus on innovation instead of fighting their stack. Because breakthrough AI should move at the speed of ideas, not infrastructure.

About the Role

The Senior Manager, Cluster Engineering & Deployment owns and runs the machine that turns delivered racks into accepted clusters: network bring-up, fabric cabling verification against port maps, GPU node integration with the fabric, cluster-level validation and burn-in (including RCCL/collective performance), and the acceptance gate into production. This is one of the most schedule-critical roles in the pillar cluster revenue starts when this team says a cluster is ready.

What You’ll Do

  • Own the cluster deployment playbook and drive its evolution: staged bring-up, automated config push, link/optics validation, cabling verification against L1 port maps, and fault triage during deployment windows.

  • Lead deployment engineering across concurrent cluster builds, through team leads and on-site engineers; coordinate daily with Data Center Integration field teams and cabling vendors.

  • Drive deployment velocity engineering: cut bring-up time per cluster through tooling, pre-staging, and defect-source elimination, and set the targets the team is measured against.

  • Own defect feedback loops to Network Engineering (design), Layer One (cabling quality), and vendors (hardware/optics RMA patterns), holding those partners accountable to resolution.

  • Define spares, test equipment, and deployment tooling requirements per site, and standardize them across sites.

Who You Are

Required Qualifications

  • 10+ years across network deployment, cluster/HPC bring-up, or large-scale infrastructure delivery, including managing engineers in a field/deployment setting.

  • Hands-on fabric bring-up experience at scale (hundreds of switches / thousands of links per deployment).

  • Strong operational rigor: building and enforcing playbooks, gates, metrics, and blameless defect loops.

  • Team leadership with schedule accountability across multiple concurrent builds or sites.

Preferred Qualifications

  • GPU cluster validation experience (NCCL/RCCL benchmarking).

  • Automation skills (Python, Ansible) applied to deployment.

  • Optics/link-layer debugging depth.

  • Experience with acceptance testing as a commercial gate (revenue-linked).

  • Own cluster validation end to end: bandwidth/latency baselines, collective (RCCL) performance tests, burn-in criteria, and go/no-go acceptance gates and raise the bar on each as the fleet scales.

What We Offer

  • Stock Options

  • 100% paid Medical, Dental, and Vision insurance for Employees

  • Company Health Savings Account Contributions

  • 100% paid Short Term and Long Term Disability Insurance for Employees

  • Life and Voluntary Supplemental Insurance Options

  • Other Insurance Options, such as Pet & Legal Insurance

  • Various Supplementary Health Benefits, such as discounted Virtual Healthcare Appointments and Serious Illness Support

  • Flexible Spending Account

  • 401(k)

  • Employee Assistance Program

  • Flexible PTO

  • Paid Holidays

  • Parental Leave

  • Other In-Office Perks

Equal Employment Opportunity

TensorWave is an Equal Opportunity Employer. We celebrate diversity and are committed to creating an inclusive environment for all employees. We do not discriminate on the basis of any protected status under applicable law.

Reasonable Accommodations

TensorWave provides reasonable accommodations in accordance with applicable laws. If you require accommodation during the hiring process, please contact [email protected].

Employment Eligibility

All offers of employment are contingent upon verification of identity and authorization to work in the United States, as required by law.

Background Checks

Where permitted by law, employment may be contingent upon the successful completion of a job-related background check.

Data Privacy Notice

By submitting an application, you acknowledge that TensorWave may collect, use, and retain your personal information for recruiting and employment-related purposes in accordance with applicable data privacy laws.

Similar Jobs

31 Minutes Ago
Remote or Hybrid
IL, USA
111K-150K Annually
Senior level
111K-150K Annually
Senior level
Artificial Intelligence • Cloud • Information Technology • Sales • Security • Software • Cybersecurity
Drives Rapid7 solution adoption and joint revenue growth through CDW by delivering technical demos, proof-of-concept evaluations, training, enablement plans, and solution guidance. Partners with account managers, sales engineers, sellers, and customers to support opportunities, explain technical value, resolve complex questions, and strengthen channel relationships. Requires consistent travel to CDW headquarters in Chicago and partner offices.
Top Skills: Application SecurityAWSCloud SecurityDevOpsGoogle Cloud Platform (Gcp)Incident ResponseAzureRapid7Security AutomationVulnerability Management
An Hour Ago
Remote or Hybrid
2 Locations
164K-297K Annually
Expert/Leader
164K-297K Annually
Expert/Leader
Fintech • Payments • Software • Financial Services
Lead commercialization of emerging AI products by securing enterprise customers and strategic partners, structuring pilots, and converting them into scaled relationships. Own engagements from prospecting and negotiation through launch and optimization, while partnering with Product and Engineering to shape product strategy, pricing, implementation, and go-to-market models.
Top Skills: AICloud InfrastructureData InfrastructureDeveloper Platforms
An Hour Ago
Remote or Hybrid
2 Locations
240K-359K Annually
Expert/Leader
240K-359K Annually
Expert/Leader
Fintech • Payments • Software • Financial Services
Own commercialization of Block’s emerging AI products from early pilots through scaled enterprise adoption. Develop market strategy, positioning, pricing, partnerships, and routes to market; personally secure lighthouse enterprise customers and strategic partners; design pilots that convert to production relationships; establish repeatable GTM, sales, implementation, and expansion processes; and build the commercial organization as product-market fit develops. Partner closely with Product and Engineering while shaping an emerging AI business.
Top Skills: AICloud InfrastructureData InfrastructureDeveloper Platforms

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account