SambaNova Systems Logo

SambaNova Systems

Principal Cloud Platform Engineer

Reposted One Month Ago
In-Office
San Jose, CA, USA
210K-280K Annually
Mid level
In-Office
San Jose, CA, USA
210K-280K Annually
Mid level
Own and operate a global AI inferencing platform: ensure uptime, low-latency performance, scalability and cost efficiency. Build monitoring, alerting, CI/CD, IaC, autoscaling, capacity planning, incident response, and participate in a shared on-call rotation to maintain 24/7 service reliability.
The summary above was generated by AI

SambaNova is a leader in next-generation AI infrastructure, delivering a full-stack inference platform for customers worldwide. At the core of SambaNova's technology is the RDU (Reconfigurable Dataflow Unit) — a chip built on a dataflow architecture rather than the traditional GPU model. Its decode performance is especially strong for agentic workloads like multi-turn agents, code generation, and long-running applications. RDUs are packaged into SambaRack, rack-scale hardware that lets customers deploy state-of-the-art models with better performance, greater energy efficiency, and faster time to value.

About the team

The Cloud Platform team owns the production inferencing service that serves SambaNova's models to customers on RDU accelerators, including capacity planning, deployment, monitoring, and incident response across regions in the United States, Asia, Europe, and Latin America.

About the role

As a Principal Cloud Platform Engineer, you will be specializing in our AI Inferencing Service and will be the guardian of its reliability, performance, and scalability. You will bridge the gap between software development and operations, applying an engineering mindset to solve operational challenges. Your primary focus will be ensuring our inference endpoints have exceptional uptime, low-latency response times, and efficient resource utilization, directly impacting the experience of our customers and the success of our AI products. This role includes participating in a shared on-call rotation to maintain 24/7 service reliability. 

Responsibilities 

Some of your responsibilities will include:

  • Shared ownership of the production inferencing service across regions, covering availability, latency, performance, change management, and capacity planning
  • Standing-up and automating AI infrastructure in new regions
  • Participating in a shared primary/secondary on-call rotation, and leading incident response
  • Building monitoring, alerting, and dashboards in Prometheus, Grafana, and Datadog for service health, model latency and throughput, and accelerator utilization
  • Finding and eliminating performance bottlenecks
  • Designing auto-scaling policies that handle variable inference loads
  • Managing cloud and on-prem infrastructure as code in Terraform and Ansible
  • Building CI/CD pipelines that safely deploy new model versions and service updates
  • Forecasting infrastructure needs against the product roadmap and usage trends, and working with finance to manage cloud spend
  • Defining and reporting on SLOs and SLIs for the inferencing platform, using that data to prioritize reliability work
Required Qualifications
  • B.S. in Computer Science, Computer Engineering, or related field
  • 5+ years of experience in a Site Reliability Engineering, DevOps
  • Experience supporting a large-scale, customer-facing service in a public cloud environment (AWS, GCP, Azure)
  • Strong programming and scripting skills in languages like Python, Go, Rust, or Java
  • Proven experience with containerization and orchestration technologies (Docker and Kubernetes)
  • Deep understanding of monitoring and observability principles and tools (e.g., Prometheus, Grafana, ELK Stack, Datadog)
  • Experience with Infrastructure as Code (e.g., Terraform, CloudFormation)
  • Experience with CI/CD principles and tools (e.g., Jenkins, GitHub Actions, ArgoCD)
  • Strong Linux/Unix system administration fundamentals
Preferred Qualifications
  • Experience in a hybrid environment bridging cloud and on-premise/data center infrastructure.
  • Direct experience supporting ML/AI inferencing services in production.
  • Familiarity with GPU-accelerated computing and optimizing workloads for NVIDIA GPUs for purposes of mapping to RDUs.
  • Knowledge of model serving frameworks like vLLM, SGLang or Ray.
  • Understanding of MLOps principles and practices.
  • Experience with managing and tuning databases (SQL or NoSQL) and caching systems (Redis, Memcached)

Base Salary Range:

Base Pay Range
$210,000$280,000 USD

Submission Guidelines
Please note that in order to be considered an applicant for any position at SambaNova Systems, you must submit an application form for each position for which you believe you are qualified. 

EEO Policy
SambaNova Systems is an Equal Opportunity/Affirmative Action Employer. All qualified applicants will receive consideration for employment without regard basis of age (40 and over), color, disability, gender identity, genetic information, marital status, military or veteran status, national origin/ancestry, race, religion, creed, sex (including pregnancy, childbirth, breastfeeding), sexual orientation, and any other applicable status protected by federal, state, or local laws.

Benefits Summary for US-Based, Full-Time Employment Positions
SambaNova offers a competitive total rewards package, including the base salary, plus equity and benefits. We cover 95% premium coverage for employee medical insurance, and 77% premium coverage for dependents and offer a Health Savings Account (HSA) with employer contribution. We also offer Dental, Vision, Short/Long term Disability, Basic Life, Voluntary Life, and AD&D insurance plans in addition to Flexible Spending Account (FSA) options like Health Care, Limited Purpose, and Dependent Care. Our library of well-being benefits available to you and your dependents includes a full subscription to Headspace, Gympass+ membership with access to physical gyms, One Medical membership, counseling services with an Employee Assistance Program, and much more.

HQ

SambaNova Systems Palo Alto, California, USA Office

Our Palo Alto office is in a tech complex known for incubating research facilities and borders the Bay Trail along the Don Edwards Wildlife Refuge. Only a 5-minute walk, our employees often fly out of PAO for lunch with colleagues or enjoy happy hour at the nearby Palo Alto Country Club.

Similar Jobs

8 Days Ago
In-Office
121K-231K Annually
Senior level
121K-231K Annually
Senior level
Information Technology • Internet of Things • Other • Cybersecurity • Infrastructure as a Service (IaaS)
Lead development of Verizon’s AI-driven ATOM cloud automation platform for wireline and wireless networks. Design cloud-native, event-driven orchestration, self-healing service assurance, API-first workflows, and executable network procedures. Provide technical leadership to distributed engineering teams, translate telecom use cases into scalable capabilities, and oversee full-cycle development, testing, stability, and production support. Apply AI agents, infrastructure-as-code, CI/CD, and modern cloud technologies to automate network deployment, monitoring, configuration, and remediation.
Top Skills: 5G Network FunctionsAi AgentsAWSClaudeCloud-Native ArchitectureDjangoFastapiFastifyFlaskGCPGeminiGenerative AiGitGitlab Ci/CdJavaScriptJinja2JSONKubernetesLangchainLangflowLanggraphLinuxLlmsMariadbMcp ServersMongoDBOpenshiftOpenstackPostgresPythonRagTemporalVector DatabasesVerizon Cloud PlatformVueYaml
10 Days Ago
Remote or Hybrid
United States
130K-211K Annually
Senior level
130K-211K Annually
Senior level
Healthtech
Build and operate AWS and Azure cloud platform foundations, including landing zones, multi-account structures, account vending, Terraform modules, guardrails, CI/CD, centralized monitoring, and configuration-drift remediation. Lead and mentor platform engineers, establish reference architectures and self-service patterns, review application designs, author architecture decisions, and represent the platform in governance forums. The role supports regulated environments, AWS GovCloud, and large-scale data center migration.
Top Skills: Account Factory For TerraformAmazon CloudwatchAmazon EksAWSAws CloudtrailAws ConfigAws Control TowerAws GovcloudAws OrganizationsAzureCi/CdFedrampKubernetesOpen Policy AgentOpenshiftSentinelTerraformVMware
One Month Ago
Hybrid
Palo Alto, CA, USA
Expert/Leader
Expert/Leader
Financial Services
Lead technical direction for cloud platform engineering: design, build, and operate secure, scalable platforms supporting AI/ML and data initiatives. Drive DevEx, CI/CD, DevSecOps, observability, self-service developer workflows, and AI-assisted development including agent and MCP integrations.
Top Skills: Anthropic SdksAWSAzureBitbucketCrossplaneDatadogDockerEcsGCPGitGoGoogle AdkGrafanaJenkinsKroKubernetesLlm-Powered Development ToolsModel Context Protocol (Mcp)PrometheusPythonSpinnakerSplunkTerraform

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account