Applied Industrial Technologies Logo

Applied Industrial Technologies

Senior Software Engineer - Cloud Infrastructure

Posted 10 Days Ago
Be an Early Applicant
In-Office
Sunnyvale, CA, USA
190K-270K Annually
Senior level
In-Office
Sunnyvale, CA, USA
190K-270K Annually
Senior level
Design, build, and operate multi-cluster Kubernetes infrastructure across cloud providers. Develop secure multi-tenant platform primitives, sandboxed execution environments, and scalable infrastructure for AI and distributed workloads. Act as a solutions architect for internal customers, improve developer tooling, debug distributed systems, lead production incidents and postmortems, and mentor engineers while driving platform reliability and security.
The summary above was generated by AI

Applied Intuition, Inc. is powering the future of physical AI. Founded in 2017 and now valued at $15 billion, the Silicon Valley company is creating the digital infrastructure needed to bring intelligence to every moving machine on the planet. Applied Intuition services the automotive, defense, trucking, construction, mining and agriculture industries in three core areas: tools and infrastructure, operating systems, and autonomy. Eighteen of the top 20 global automakers, as well as the United States military and its allies, trust the company’s solutions to deliver physical intelligence. Applied Intuition is headquartered in Sunnyvale, California, with offices in Washington, D.C.; San Diego; Ft. Walton Beach, Florida; Ann Arbor, Michigan; London; Stuttgart; Munich; Stockholm; Bangalore; Seoul; and Tokyo. Learn more at applied.co.

We are an in-office company, and our expectation is that full-time employees primarily work from their Applied Intuition office 5 days a week. However, we also recognize the importance of flexibility and trust our employees to manage their schedules responsibly. This may include occasional remote work, starting the day with morning meetings from home before heading to the office, or leaving earlier when needed to accommodate family commitments. This in-office expectation does not apply to contractor positions

About the role

Everything Applied ships runs on infrastructure we build. From the daily, large-scale simulations that test autonomous systems to the enterprise AI workloads behind Dana, our cloud infrastructure team builds and maintains the platform that product engineering uses to ship, operate, and scale their applications. We are cloud-native and cloud-agnostic across all cloud providers. We run on Kubernetes, manage everything as infrastructure as code, and own the full stack, including compute, networking, file system, blob storage, resource scaling, observability, security, and CI/CD. Most platform teams own a slice of that. We own the whole thing.

We also build the core infrastructure platform that powers Dana, Applied's AI platform for enterprises in physical industries, bringing apps, agents, and data together on one governed platform. You'll help build Dana's core infrastructure and everything customers touch to build, deploy, share, and govern apps. Dana is multi-tenant and serves critical enterprise workloads, so scale and reliability tradeoffs are part of the job. The compute and data generation scale of our workloads, from large-scale simulation to agentic execution, pushes the boundaries of standard cluster deployments, and you'll be at the forefront of building out this system and ensuring its reliability.

Owning the whole stack makes this as much a platform-building and solutions-architecture role as a hands-on infrastructure role. You'll work embedded with engineers across the entire org and support a variety of customer deployments, and you'll be trusted to make architectural calls, own incidents end-to-end, and drive infrastructure decisions rather than just execute tickets.

In this role, you will
  • Design, build, and operate our multi-cluster Kubernetes infrastructure (compute, networking, storage, autoscaling, observability, and security) with high reliability across all cloud providers

  • Build multi-tenant platform primitives: tenant-safe storage, shared secrets, RBAC, and workload identity

  • Design and orchestrate secure, sandboxed execution environments for agentic and untrusted workloads

  • Act as a solutions architect for internal platform customers, turning their scaling and reliability needs into concrete infrastructure designs and driving them to production

  • Improve developer effectiveness by spotting the friction nobody else has bothered to fix and building the tooling that removes it

  • Debug and profile the awkward edge cases in distributed systems, and own production end-to-end: incident command, postmortems, and the follow-through that prevents recurrence

  • Mentor engineers and raise the technical bar through design reviews and thoughtful collaboration

You may be a good fit if you
  • Have 5+ years of experience building and operating large-scale infrastructure, platform, SRE, or DevOps systems

  • Have experience with container orchestration frameworks such as Kubernetes

  • Have deep experience with at least one major cloud provider (AWS, GCP, Azure or OCI)

  • Write production-quality code and are an expert in at least one of Go, Python, Rust, or C++, and comfortable picking up whatever the problem needs

  • Are fluent with Infrastructure as Code and GitOps workflows (Terraform, OpenTofu, Pulumi, Crossplane, or similar)

  • Communicate and collaborate well across teams: aligning on interfaces, navigating tradeoffs, and driving cross-team execution

  • Hold a BS in Computer Science or a related field

Strong candidates may also have
  • Experience with serverless or scale-to-zero container platforms such as Knative, Cloud Run, or KEDA-based systems for running multi-tenant application workloads

  • Depth in sandboxing and workload isolation: Linux namespaces, cgroups, seccomp, gVisor, Firecracker/Kata, or comparable multi-tenant isolation designs

  • Depth in cluster and cloud networking: CNI (e.g., Cilium), eBPF, NetworkPolicy, service mesh, cross-cloud private connectivity

  • Deep multi-cluster or multi-region Kubernetes experience running diverse workloads at scale, from large batch and data-processing jobs to GPU and agentic workloads, including scheduling and autoscaling systems such as Karpenter, Kueue, or Volcano

  • Platform security experience: admission control, least-privilege IAM, workload identity (OIDC/SPIFFE), image provenance and supply-chain hardening

  • Incident command experience for customer-facing production systems

  • Experience building enterprise AI infrastructure or delivering platform solutions to internal or enterprise customers

  • Contributions to open source infrastructure tooling

Don’t meet every single requirement? If you’re excited about this role but your past experience doesn’t align perfectly with every qualification in the job description, we encourage you to apply anyway. You may be just the right candidate for this or other roles.

Applied Intuition is an equal opportunity employer and federal contractor or subcontractor. Consequently, the parties agree that, as applicable, they will abide by the requirements of 41 CFR 60-1.4(a), 41 CFR 60-300.5(a) and 41 CFR 60-741.5(a) and that these laws are incorporated herein by reference. These regulations prohibit discrimination against qualified individuals based on their status as protected veterans or individuals with disabilities, and prohibit discrimination against all individuals based on their race, color, religion, sex, sexual orientation, gender identity or national origin. These regulations require that covered prime contractors and subcontractors take affirmative action to employ and advance in employment individuals without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, protected veteran status or disability. The parties also agree that, as applicable, they will abide by the requirements of Executive Order 13496 (29 CFR Part 471, Appendix A to Subpart A), relating to the notice of employee rights under federal labor laws.

FOR US-BASED ROLES: Applied Intuition is committed to providing an accessible and inclusive application and interview experience to applicants who are disabled veterans and other applicants with disabilities or medical conditions. Reasonable accommodations are available, requesting an accommodation will not affect your candidacy in any way, and you are not required to disclose the nature of your disability or medical condition in order to make a request. If you require an accommodation please contact [email protected]. We will work with you!

Similar Jobs

2 Days Ago
In-Office
San Jose, CA, USA
124K-213K Annually
Senior level
124K-213K Annually
Senior level
Fintech • Payments
Designs, builds, and operates AWS cloud networking and routing infrastructure, including VPCs, transit gateways, load balancers, CDN, WAF, ingress, and service mesh systems. Develops Terraform infrastructure as code, troubleshoots network and TLS incidents, supports edge services through on-call rotation, and documents operational practices. Partners with SRE, security, and product engineering teams while mentoring engineers and driving technical best practices.
Top Skills: AlbAWSAws CloudfrontAws Transit GatewayAws VpcBashCdnContourDatadogDnsDockerEksElbEnvoyGithub ActionsGithub EnterpriseGoKubernetesNginxNlbPythonRoute 53Tcp/IpTerraformTls/SslWaf
22 Days Ago
In-Office
San Francisco, CA, USA
149K-246K Annually
Senior level
149K-246K Annually
Senior level
Cloud • Software
Lead DNS Operations for Salesforce Hyperforce: provide 24x7 incident response, troubleshoot production DNS outages, perform maintenance and builds, collaborate with engineering on stability improvements, automate DNS/infrastructure, and drive root-cause analysis and change management.
Top Skills: Ai-Assisted Development ToolsAnsibleAWSAws Route 53AzureChefDnsDnssecDockerGCPGitGrafanaHTTPJenkinsKubernetesLinuxLoad BalancingPuppetPythonSaltSpinnakerSplunkSQLSslTcp/IpTerraformTravis Ci
2 Days Ago
In-Office or Remote
Santa Clara, CA, USA
184K-357K Annually
Senior level
184K-357K Annually
Senior level
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
Design, build, and maintain large-scale AI/ML platform and infrastructure for training, inference, fine-tuning, and Agentic AI. Develop tools, APIs, reliability metrics, and root-cause analyses across application to hardware layers while improving efficiency, resiliency, monitoring, and observability.
Top Skills: C/C++Cloud-NativeDgx CloudDynamoElkGoInfinibandJaxKubernetesLokiNcclNvidia GpusObservability PlatformsPrometheusPythonPyTorchRayRdmaScripting LanguagesTensorFlow

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account