Tensormesh Logo

Tensormesh

Kubernetes Engineer

Reposted 14 Days Ago
In-Office or Remote
2 Locations
100K-200K Annually
Junior
In-Office or Remote
2 Locations
100K-200K Annually
Junior
Build and extend Kubernetes orchestration for AI inference: implement CRDs/controllers in Go, design high-performance ingress/load-balancing, implement GPU-aware autoscaling (HPA/VPA), tune scheduler/resource management, and ensure HA for platform components.
The summary above was generated by AI

About the company Tensormesh is building the next generation of AI inference infrastructure. Our mission is to make large language models faster, cheaper, and easier to deploy across any environment — cloud, on-prem, or hybrid. We help enterprises and AI teams optimize GPU utilization and scale inference workloads with up to 10× better performance.

About the role We are seeking a Kubernetes Engineer to build the core orchestration logic of the Tensormesh platform. Unlike a standard DevOps role, this position involves deep software development within the Kubernetes ecosystem. You will extend Kubernetes capabilities by writing custom operators and controllers that manage complex AI inference workloads, ensuring high availability and seamless auto-scaling.

What you’ll do

  • Develop Custom Operators: Design and implement Kubernetes Custom Resource Definitions (CRDs) and Controllers (using Golang/Kubebuilder) to manage the lifecycle of Tensormesh products.

  • Traffic Management: Architect and build high-performance ingress and load-balancing systems capable of handling high-throughput inference requests.

  • Resilience & Scaling: Develop logic for intelligent auto-scaling (HPA/VPA) based on GPU metrics and ensure High Availability (HA) for critical components.

  • K8s Optimization: Tune the scheduler and resource management configurations to maximize GPU utilization for inference tasks.

Ideal candidate credentials

  • 0-3 years of software engineering experience, with a focus on distributed systems or container orchestration.

  • Strong proficiency in Go (Golang); experience with the Kubernetes client-go library and Kubebuilder/Operator SDK is highly preferred.

  • Deep understanding of Kubernetes internals (API machinery, Controller runtime, Networking, CNI).

  • Experience with service meshes (Istio, Linkerd) or ingress controllers (Nginx, Traefik) is a plus.

  • Understanding of distributed consensus and state management.

Compensation & Benefits

  • Competitive base salary

  • Performance-based bonus

  • Equity options

  • Medical, dental, and vision insurance

  • 401(k) retirement plan

  • Paid time off

Why Join Tensormesh

  • Build infrastructure powering next-generation AI applications

  • Work alongside a highly technical and experienced engineering team

  • Make a direct impact in a fast-growing startup environment

  • Take ownership of challenging technical problems at scale

Similar Jobs

8 Days Ago
In-Office or Remote
113K-193K Annually
Senior level
113K-193K Annually
Senior level
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Lead design, automation, operation, and hardening of enterprise-scale OpenShift/Kubernetes platforms on HPC infrastructure. Administer clusters, implement IaC and monitoring, troubleshoot platform issues, enable HA/disaster recovery, support multi-cluster architectures, and drive platform modernization and operational excellence.
Top Skills: Advanced Cluster Manager (Acm)AlertmanagerAnsibleBashCi/CdF5Git ActionsGitopsGoGpfs (Ibm Spectrum Scale)GrafanaHaproxyIbm LinuxoneIbm ZInfrastructure-As-CodeKubernetesKubevirtLokiOc (Openshift Cli)OpenshiftOpenshift ConsolePrometheusPythonRed Hat Enterprise LinuxRed Hat OpenshiftS390XShell ScriptingThanos
2 Days Ago
Remote
USA
150K-175K Annually
Senior level
150K-175K Annually
Senior level
Artificial Intelligence • Big Data • Cloud • Information Technology
Design, build, and support cloud platforms and Kubernetes clusters on AWS/Azure; implement Terraform-based IaC and CI/CD pipelines; lead cloud migrations and modernization of legacy workloads; automate scaling and deployments; monitor and optimize system reliability, security, and cost; collaborate with stakeholders and contribute to runbooks, reviews, and knowledge sharing.
Top Skills: .NetAdfAirflowAWSAzureAzure DevopsCi/CdDatabasesDatastageETLGithub ActionsGlueIacJavaJenkinsKubernetesTerraform
2 Days Ago
Remote
USA
216K-264K Annually
Senior level
216K-264K Annually
Senior level
Security
Lead hands-on, forward-deployed customer engagements to take enterprise customers from evaluation to production Kubernetes deployments. Act as regional technical anchor, design deployment patterns and reference architectures, unblock complex platform and Kubernetes issues, mentor engineers, and feed field insights into product and roadmap decisions.
Top Skills: ArgocdCrdDatadogFluxGitopsGoGrafanaHelmKubernetesModel Context Protocol (Mcp)OpentelemetryOperatorsPrometheusToolhive

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account