Lead the design, implementation, and support of scalable cloud-native platforms, Kubernetes environments, infrastructure-as-code modules, CI/CD pipelines, and DevSecOps automation. Manage cloud infrastructure, GitOps deployments, observability, security controls, reliability, and disaster recovery. Build self-service developer capabilities, internal developer portals, documentation, and automation while partnering with development, security, and SRE teams on incident response and continuous improvement.
Lead DevOps Platform Engineer
Job Summary
Location: Hybrid / Remote (Princeton, NJ)
We are seeking a highly skilled Senior DevOps Platform Engineer to design, build, and support scalable cloud-native platforms, CI/CD pipelines, infrastructure automation, and developer enablement solutions. The ideal candidate will have deep expertise in Platform Engineering, Kubernetes, Infrastructure as Code, DevSecOps practices, and cloud technologies while partnering closely with Development, Security, and SRE teams.
Key ResponsibilitiesPlatform Engineering & Infrastructure- Design, implement, and maintain scalable cloud-native platform solutions.
- Build and manage Kubernetes environments across development, staging, and production.
- Develop reusable Infrastructure as Code (IaC) modules using Terraform/OpenTofu.
- Implement platform standards, golden-path templates, and automation frameworks.
- Improve platform reliability, scalability, and operational efficiency.
- Design and maintain enterprise-grade CI/CD pipelines using GitHub Actions, GitLab CI, Jenkins, ArgoCD, or similar tools.
- Automate application deployment, testing, release management, and rollback processes.
- Implement GitOps practices and deployment automation.
- Manage artifact repositories and software package lifecycle processes.
- Integrate security controls into CI/CD pipelines.
- Implement SAST, SCA, container scanning, secrets detection, and policy-as-code solutions.
- Support software supply chain security initiatives including SBOM generation, artifact signing, and dependency management.
- Collaborate with Security teams to ensure compliance with enterprise security standards.
- Manage and optimize cloud environments across AWS, Azure, or GCP.
- Deploy and maintain Kubernetes clusters and containerized workloads.
- Implement monitoring, logging, observability, and alerting solutions.
- Support high availability, disaster recovery, and platform resilience initiatives.
- Build self-service developer capabilities and platform automation.
- Support Internal Developer Portal initiatives such as Backstage or similar platforms.
- Create technical documentation, runbooks, standards, and onboarding guides.
- Improve developer productivity through automation and platform enhancements.
- Implement observability solutions using Prometheus, Grafana, ELK, Datadog, Splunk, or similar tools.
- Monitor platform health, availability, and performance metrics.
- Participate in incident response, root cause analysis, and continuous improvement activities.
- 7+ years of experience in DevOps, Platform Engineering, Cloud Engineering, or Site Reliability Engineering.
- Strong expertise in Kubernetes administration and container orchestration.
- Hands-on experience with Terraform/OpenTofu and Infrastructure as Code.
- Experience designing and supporting CI/CD pipelines using GitHub Actions, GitLab CI, Jenkins, ArgoCD, or Tekton.
- Strong Linux system administration and scripting skills (Python, Bash, PowerShell).
- Experience with AWS, Azure, or Google Cloud Platform.
- Knowledge of GitOps practices and deployment automation.
- Experience integrating security tools such as Semgrep, Snyk, Checkmarx, Trivy, Prisma Cloud, Gitleaks, or HashiCorp Vault.
- Working knowledge of Policy-as-Code using OPA/Rego or Kyverno.
- Familiarity with SBOM, SLSA, software supply chain security, and artifact signing.
- Experience implementing monitoring and observability solutions.
- Knowledge of DORA metrics and platform performance measurement.
- Strong troubleshooting and incident management skills.
- Experience with Backstage or Internal Developer Portals.
- Exposure to AI-assisted development tools such as GitHub Copilot, Cursor, or Agentic workflows.
- Experience in regulated industries such as Financial Services, Healthcare, or Government.
- Knowledge of Service Mesh technologies (Istio, Linkerd).
- Experience with eBPF security tools such as Falco or Tetragon.
- Cloud certifications (AWS, Azure, GCP, Kubernetes).
- Bachelor's degree in Computer Science, Engineering, Information Technology, or equivalent practical experience.
- Relevant certifications in Cloud, Kubernetes, DevOps, or Security are highly preferred.
Similar Jobs
Artificial Intelligence • Software
The DevOps Engineer will enhance internal compute infrastructure, manage integration of external technologies, and support system reliability for evolving use cases.
Top Skills:
AWSAzureCi/CdDockerGceGoKubernetesPythonRust
Cloud • Information Technology • Machine Learning
Lead technical programs delivering GPU and CPU compute fleets from infrastructure handover through provisioning, validation, operational acceptance, and steady-state operations. Translate customer and business requirements into delivery plans, manage dependencies and blockers, assess capacity gaps, drive prioritization, establish readiness criteria, and coordinate cross-functional engineering and operations teams. Define delivery metrics, improve forecasting, and identify tooling, automation, and process improvements to increase provisioning throughput and operational readiness.
Top Skills:
Computer NetworkingCpuData Center InfrastructureGpuHardware/Software IntegrationInfrastructure AutomationServer ProvisioningSQL
Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Build and deploy production-grade AI agents, assistants, models, and reusable skills. Develop secure agentic workflows with guardrails, isolation, and controls; integrate frameworks, tools, APIs, cloud services, and models; establish rigorous evaluation metrics for quality, safety, and reliability; rapidly prototype solutions and partner with engineering to harden them for production. Document solutions and contribute to AI Center of Excellence consulting and best practices.
Top Skills:
AWSAzureClaude SdkDockerGCPGoogle Ai SdkKubernetesOpenai SdkPythonSkill.Md
What you need to know about the San Francisco Tech Scene
San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.
Key Facts About San Francisco Tech
- Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
- Major Tech Employers: Google, Apple, Salesforce, Meta
- Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
- Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
- Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
- Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine


