SambaNova Systems Logo

SambaNova Systems

Cloud Platform Architect

Reposted One Month Ago
In-Office
San Jose, CA, USA
245K-325K Annually
Senior level
In-Office
San Jose, CA, USA
245K-325K Annually
Senior level
As a Senior Cloud SRE, you'll ensure the reliability and performance of our AI inferencing service, handle incident management, and optimize cloud infrastructure and resource utilization.
The summary above was generated by AI

SambaNova is a leader in next-generation AI infrastructure, delivering a full-stack inference platform for customers worldwide. At the core of SambaNova's technology is the RDU (Reconfigurable Dataflow Unit) — a chip built on a dataflow architecture rather than the traditional GPU model. Its decode performance is especially strong for agentic workloads like multi-turn agents, code generation, and long-running applications. RDUs are packaged into SambaRack, rack-scale hardware that lets customers deploy state-of-the-art models with better performance, greater energy efficiency, and faster time to value.

About the team

The Cloud Platform team owns the production inferencing service that serves SambaNova's models to customers on RDU accelerators, including capacity planning, deployment, monitoring, and incident response across regions in the United States, Asia, Europe, and Latin America.

About the role

The Cloud Operations team is seeking an experienced engineering leader to scale the platform our internal and external customers use to access SambaNova RDUs.

Responsibilities

In this role, you'll architecting our next-generation system from the ground up, running the Kubernetes infrastructure that powers some of the most advanced AI workloads in the industry, and bridging multi-cloud and on-prem environments in ways no generic SaaS company can offer. Your work will directly impact the productivity of every engineer at SambaNova and by extension, the speed at which we ship the future of AI computing.

  • Architect, build, and maintain our next-generation internal developer platform, automating and streamlining our cloud and on-prem infrastructure
  • Design, write, and manage Terraform modules to provision and manage resources across AWS, GCP, and Azure, ensuring consistency and reproducibility
  • Build and manage highly available, secure, and performant Kubernetes clusters that serve as the primary runtime for our diverse AI workloads
  • Design and implement robust networking solutions (VPCs, load balancers, firewalls, service meshes) that seamlessly connect our multi-cloud and hybrid environments
  • Collaborate with AI and software engineering teams to understand their needs, provide golden paths to production, and build internal tools that accelerate their development cycles
  • Implement best practices for observability (monitoring, logging, tracing) to ensure system reliability and performance, and participate in on-call rotation
Required Qualifications
  • 7+ years of experience in DevOps, Site Reliability Engineering (SRE), or Cloud Infrastructure roles
  • Proficiency in at least one programming language (e.g., Python, Go, Rust)
  • Expertise with Kubernetes (EKS, GKE, or self-managed) in production environments - pods, operators, CRDs, CNIs, etc.
  • Expertise with Infrastructure as Code with the ability to manage complex, multi-cloud environments
  • Strong proficiency with at least one major cloud provider (AWS, GCP, or Azure), with a solid understanding of the core services (compute, storage, networking, IAM)
  • Networking fundamentals (TCP/IP, DNS, HTTP, load balancing) and security best practices in the cloud
Preferred Qualifications
  • Experience in a hybrid environment bridging cloud and on-premise/data center infrastructure
  • Experience managing infrastructure for data-intensive or ML/AI workloads
  • Knowledge of building and maintaining CI/CD pipelines (e.g., GitLab CI, Jenkins, ArgoCD)
  • Experience with service mesh technologies (e.g., Istio, Linkerd)
  • Contributions to open-source projects or a public portfolio of code (GitHub)

Base Salary Range:

Base Pay Range
$245,000$325,000 USD

Submission Guidelines
Please note that in order to be considered an applicant for any position at SambaNova Systems, you must submit an application form for each position for which you believe you are qualified. 

EEO Policy
SambaNova Systems is an Equal Opportunity/Affirmative Action Employer. All qualified applicants will receive consideration for employment without regard basis of age (40 and over), color, disability, gender identity, genetic information, marital status, military or veteran status, national origin/ancestry, race, religion, creed, sex (including pregnancy, childbirth, breastfeeding), sexual orientation, and any other applicable status protected by federal, state, or local laws.

Benefits Summary for US-Based, Full-Time Employment Positions
SambaNova offers a competitive total rewards package, including the base salary, plus equity and benefits. We cover 95% premium coverage for employee medical insurance, and 77% premium coverage for dependents and offer a Health Savings Account (HSA) with employer contribution. We also offer Dental, Vision, Short/Long term Disability, Basic Life, Voluntary Life, and AD&D insurance plans in addition to Flexible Spending Account (FSA) options like Health Care, Limited Purpose, and Dependent Care. Our library of well-being benefits available to you and your dependents includes a full subscription to Headspace, Gympass+ membership with access to physical gyms, One Medical membership, counseling services with an Employee Assistance Program, and much more.

HQ

SambaNova Systems Palo Alto, California, USA Office

Our Palo Alto office is in a tech complex known for incubating research facilities and borders the Bay Trail along the Don Edwards Wildlife Refuge. Only a 5-minute walk, our employees often fly out of PAO for lunch with colleagues or enjoy happy hour at the nearby Palo Alto Country Club.

Similar Jobs

One Month Ago
In-Office
Senior level
Senior level
Other • Transportation
Leads cloud, infrastructure, platform engineering, and DevSecOps architecture across AWS and Azure. Defines reference architectures, security guardrails, CI/CD pipelines, observability standards, infrastructure automation, and SRE practices. Builds and manages a team of architects and engineers, mentors staff, manages performance, influences enterprise stakeholders, and evaluates vendors and technologies.
Top Skills: ArmAWSAzureAzure DevopsBicepCi/CdCloudFormationDastDatadogDevsecopsDockerDynatraceGitlabHarnessInfrastructure As CodeKubernetesNew RelicSastSreTerraform
One Month Ago
In-Office
121K-219K Annually
Senior level
121K-219K Annually
Senior level
Information Technology • Security • Software
The Principal Cloud DevOps Engineer is responsible for deploying and maintaining IFIaaS applications in hybrid cloud environments, ensuring reliability, managing CI/CD pipelines, and providing on-call support.
Top Skills: AnsibleJenkinsAzureOctopus DeployTerraform
5 Minutes Ago
Hybrid
Senior level
Senior level
Financial Services
Leads engineering and SRE work for firm-wide identity platforms, designing reliable services, APIs, automation, and infrastructure across enterprise and public-cloud environments. Owns production reliability, observability, incident response, resiliency, disaster recovery, CI/CD, and infrastructure-as-code. Establishes service-level objectives, improves scalability and security, mentors engineers, and leads responsible adoption of AI-assisted development and incident-management workflows. Requires software engineering, IAM, cloud, networking, Kubernetes, and high-availability systems expertise.
Top Skills: AWSCi/CdCloudwatchGoGCPGrafanaJavaKubernetesMicrosoft EntraMtlsPkiPrometheusPythonSpring BootTempoTerraformX.509

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account