Palona AI Logo

Palona AI

AI Infrastructure Engineer

Posted 6 Hours Ago
Be an Early Applicant
In-Office
Los Altos, CA, USA
Mid level
In-Office
Los Altos, CA, USA
Mid level
Build and operate scalable, secure cloud infrastructure for real-time AI services. Improve reliability with SLOs, observability, capacity planning, and failure testing. Implement IaC, CI/CD, deployment safety, secrets and vulnerability management. Partner with product and AI teams, diagnose distributed failures, reduce cost, and automate platform patterns to enable faster, safer releases.
The summary above was generated by AI

Palona’s AI agents operate continuously in production, handle real-time guest interactions, integrate with restaurant systems, and face sharp traffic peaks. Infrastructure is therefore part of the product: latency, reliability, deployment safety, observability, security, and cost directly shape the guest and operator experience.

We are looking for an Infrastructure Engineer who combines cloud and reliability depth with strong software engineering judgment. You will build and operate the platform beneath Palona’s AI products, improve how engineers ship, and turn production signals into durable system improvements. This is not a ticket-driven IT or operations role. You will write production code, design systems, automate repetitive work, and own outcomes across the full service lifecycle.

Our current environment includes Python services, Docker, AWS and selected Azure services, ECS and Lambda workloads, API Gateway, load balancers, relational data systems, OpenTofu/Terraform, Datadog, and CI/CD automation. We value the ability to learn and make sound tradeoffs more than exact tool-for-tool matching.

What you will own:
  • Design, build, and evolve secure, scalable cloud infrastructure for real-time AI services and customer-facing applications.
  • Improve service reliability through clear SLOs, actionable observability, capacity planning, failure testing, and pragmatic incident prevention.
  • Build deployment and release systems that make production changes fast, repeatable, auditable, and safe.
  • Own infrastructure as code, environment consistency, and reusable platform patterns across development, staging, and production.
  • Partner with product and AI engineers on architecture, performance, data flows, and operational readiness for new capabilities.
  • Diagnose complex distributed-system failures across application, network, database, model-provider, and third-party integration boundaries.
  • Reduce infrastructure and model-serving cost without compromising customer experience or engineering velocity.
  • Strengthen secrets management, access controls, backup and recovery, vulnerability management, and other practical security foundations.
  • Build internal tooling and paved paths that let engineers ship and operate services with less manual work.
  • Participate in incident response and turn incidents into better systems, automation, documentation, and engineering judgment.

Requirements
  • 3+ years industrial experience in relevant technical domain.
  • Strong software engineering fundamentals and experience building or operating production distributed systems.
  • Hands-on experience with a major cloud platform; AWS experience is especially relevant.
  • Experience with containers, infrastructure as code, CI/CD, monitoring, alerting, and production debugging.
  • Ability to write reliable automation and services in Python or another modern programming language.
  • Sound judgment around availability, latency, scalability, security, and cost tradeoffs.
  • A track record of taking ambiguous operational problems from diagnosis through durable resolution.
  • Clear communication during architecture reviews, launches, and incidents.
  • AI-native working habits and curiosity about the operational behavior of LLM- and agent-powered systems.

Benefits
  • Competitive Salary and Stock Option Plan.
  • Medical, dental, vision, retirement, leave, and disability benefits as applicable.
  • Family Leave
  • Short Term & Long Term Disability
  • Paid time off and company holidays.
  • Learning and development support.
HQ

Palona AI Menlo Park, California, USA Office

Menlo Park, CA, United States, 94025

Similar Jobs

Mid level
Artificial Intelligence • Hardware • Software • Semiconductor
Design, implement, and maintain Python frameworks and services that orchestrate distributed engineering workflows across machines and clusters. Build scheduling, execution, resource management, failure recovery, and test infrastructure. Define APIs and abstractions, reason about concurrency and distributed-systems behavior, debug complex multi-system issues, write automated tests and documentation, and partner with platform, CI, release, QA, and product teams to deliver scalable infrastructure.
Top Skills: AsyncioBuild SystemsCi SystemsCluster SchedulersConcurrent.FuturesContainersKubernetesMultiprocessingPytestPythonRelease InfrastructureRemote Execution Systems
8 Days Ago
In-Office
130K-145K Annually
Senior level
130K-145K Annually
Senior level
Fintech • Payments • Financial Services
Design, build, and deploy Generative AI solutions (LLMs, RAG, vector DBs, agents) for financial use cases. Architect scalable AI infra, containerize models, implement CI/CD and MLOps, evaluate model performance, and collaborate with engineering and stakeholders to productionize safe, reliable GenAI systems.
Top Skills: AgentcoreAWSAws BedrockAzureCi/CdDockerFine-TuningGCPKubernetesLanggraphLlamaindexLlmsMlopsPgvectorPineconePrompt EngineeringPythonQdrantRag PipelinesSQLVector Databases
24 Days Ago
Hybrid
139K-155K Annually
Senior level
139K-155K Annually
Senior level
Healthtech • Software
Own reliability, observability, and security for AI/ML platforms (data processing, workspaces, labeling, model serving). Build IaC and automation, define SLOs/error budgets, run incident response and DR exercises, implement security controls, mentor engineers, and optimize cost, capacity, and operational standards across cloud environments.
Top Skills: Azure Ai (Azure Ml)Blue/Green DeploymentCanary DeploymentCi/CdContainer OrchestrationData Lineage ToolingDatabricksDistributed TracingEncryptionFinopsGitopsKey ManagementKubernetesLoggingObservability (MetricsSecrets ManagementSli/Slo FrameworksTerraformTraces)

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account