Perficient Logo

Perficient

Lead Site Reliability Engineer (AI & Cloud Operations)

Posted Yesterday
In-Office or Remote
Hiring Remotely in United States
111K-145K Annually
Senior level
In-Office or Remote
Hiring Remotely in United States
111K-145K Annually
Senior level
Leads SRE and cloud operations for highly available, scalable platforms and AI-powered solutions. Designs AWS infrastructure, Kubernetes environments, CI/CD pipelines, Infrastructure as Code, observability, automation, and incident response processes. Develops AI-driven operational capabilities, supports MLOps and model lifecycle management, establishes reliability metrics and SLOs, and mentors engineering teams. Partners with software, data, machine learning, security, and product teams to improve platform performance, resilience, and operational excellence.
The summary above was generated by AI

We are seeking a highly skilled Lead Site Reliability Engineer (AI & Cloud Operations) to drive the reliability, scalability, automation, and operational excellence of our cloud-native platforms and AI-powered solutions. This role will serve as a technical leader responsible for building resilient infrastructure, implementing modern DevOps and SRE practices, and enabling enterprise AI capabilities through automation, observability, and operational intelligence.

The ideal candidate combines deep expertise in AWS cloud technologies, Kubernetes, infrastructure automation, CI/CD, and incident management with hands-on experience supporting AI/ML and Generative AI platforms. This individual will partner closely with software engineering, data engineering, machine learning, security, and product teams to establish highly available systems, streamline deployments, optimize platform performance, and accelerate innovation through AI-driven operations.

Responsibilities
  • Lead the design, implementation, and continuous improvement of Site Reliability Engineering (SRE) practices to ensure highly available, scalable, and resilient cloud platforms.
  • Architect, deploy, and support AWS-based infrastructure and services, including containerized and serverless environments.
  • Build, maintain, and optimize CI/CD pipelines and Infrastructure as Code (IaC) solutions to accelerate and standardize deployments.
  • Develop automation solutions, operational tooling, and self-healing capabilities using Python, Shell, and modern DevOps technologies.
  • Manage Kubernetes and container platforms, ensuring performance, scalability, and operational stability.
  • Establish and enhance observability through monitoring, logging, alerting, and performance management tools to proactively identify and resolve issues.
  • Lead incident response, root cause analysis, problem management, and service reliability improvement initiatives.
  • Partner with engineering, data, AI/ML, and security teams to support enterprise applications, analytics platforms, and cloud-native solutions.
  • Design and implement AI-driven operational capabilities, including intelligent monitoring, automated remediation, predictive analytics, and chatbot-enabled support workflows.
  • Support MLOps and AI platform operations, including model deployment, monitoring, governance, and lifecycle management.
  • Define and track reliability metrics, service-level objectives (SLOs), and operational KPIs to drive continuous improvement.
  • Mentor and provide technical leadership to engineering teams while promoting best practices in reliability, automation, cloud operations, and AI-enabled innovation.


 

Qualifications
  • 5+ years of experience in Site Reliability Engineering (SRE), DevOps, or Cloud Operations.
  • Strong hands-on experience with AWS services including EC2, EKS, ECS, Lambda, S3, RDS, IAM, CloudWatch, and VPC.
  • Experience building and managing CI/CD pipelines using Jenkins, GitHub Actions, GitLab CI/CD, Azure DevOps, or similar platforms.
  • Strong scripting and automation skills using Python, Shell, or similar languages.
  • Experience with Infrastructure as Code (IaC) tools such as Terraform or CloudFormation.
  • Expertise in containerization and orchestration technologies (Docker, Kubernetes).
  • Experience with observability tools such as Prometheus, Grafana, Datadog, Splunk, or ELK Stack.
  • Understanding of analytics platforms, data pipelines, and operational data analysis.
  • Strong troubleshooting, problem-solving, and incident management skills.
  • AI & Automation Experience
  • Experience implementing AI/ML or Generative AI solutions within enterprise environments.
  • Familiarity with AI platforms such as Azure OpenAI, AWS Bedrock, Amazon SageMaker, OpenAI APIs, LangChain, or NVIDIA AI ecosystem.
  • Experience building AI-assisted operational workflows, chatbots, intelligent monitoring, predictive analytics, or automated remediation solutions.
  • Understanding of MLOps concepts, model deployment, monitoring, and governance.

WHAT WE BELIEVE

At Perficient, we promise to challenge, champion, and celebrate our people. You will experience a unique and collaborative culture that values every voice. Join our team, and you’ll become part of something truly special.

 

We believe in developing a workforce that is as diverse and inclusive as the clients we work with. We’re committed to actively listening, learning, and acting to further advance our organization, our communities, and our future leaders… and we’re not done yet.

 

Perficient, Inc. proudly provides equal employment opportunities (EEO) to all employees and applicants for employment without regard to race, color, religion, gender, sexual orientation, national origin, age, disability, genetic information, marital status, amnesty, or status as a protected veteran in accordance with applicable federal, state and local laws. Perficient, Inc. complies with applicable state and local laws governing non-discrimination in employment in every location in which the company has facilities. This policy applies to all terms and conditions of employment, including, but not limited to, hiring, placement, promotion, termination, layoff, recall, transfer, leaves of absence, compensation, and training. Perficient, Inc. expressly prohibits any form of unlawful employee harassment based on race, color, religion, gender, sexual orientation, national origin, age, genetic information, disability, or covered veterans. Improper interference with the ability of Perficient, Inc. employees to perform their expected job duties is absolutely not tolerated.

 

Disability Accommodations:

 

Perficient is committed to providing a barrier-free employment process with reasonable accommodations for qualified individuals with disabilities and disabled veterans in our job application procedures. If you need assistance or accommodation due to a disability, please contact us.

 

Applications will be accepted until the position is filled or the posting is removed.

 

The salary range for this position takes into consideration a variety of factors, including but not limited to skill sets, level of experience, applicable office location, training, licensure and certifications, and other business and organizational needs. The new hire salary range displays the minimum and maximum salary targets for this position across all US locations, and the range has not been adjusted for any specific state differentials. It is not typical for a candidate to be hired at or near the top of the range for their role, and compensation decisions are dependent on the unique facts and circumstances regarding each candidate. A reasonable estimate of the current salary range for this position is $111,300 to $144,600. Please note that the salary range posted reflects the base salary only and does not include benefits or any potential variable compensation programs. Information regarding the benefits available for this position are in our benefits overview.

 

Disclaimer:  The above statements are not intended to be a complete statement of job content, rather to act as a guide to the essential functions performed by the employee assigned to this classification.  Management retains the discretion to add or change the duties of the position at any time. 

#LI-RS1

 

About UsPerficient is the global AI and technology consulting firm disrupting the traditional consulting model. Powered by our 7,000+ advisors, engineers, and designers, Perficient implements AI-first solutions that break conventions and deliver outcomes that matter. Proudly serving clients that represent the world’s most innovative brands, and in collaboration with our powerful technology partner ecosystem, we bring deep industry expertise and data-driven design to redefine how businesses run and succeed. Perficient is different. For real. Learn more at perficient.com.

Perficient Oakland, California, USA Office

Oakland, United States

Perficient San Francisco, California, USA Office

345 California St, San Francisco, CA, United States, 94104

Similar Jobs

41 Minutes Ago
Remote or Hybrid
USA
145K-220K Annually
Expert/Leader
145K-220K Annually
Expert/Leader
Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Lead technical marketing for Falcon Secure Access by creating demos, labs, workshops, and technical content that translate secure access, browser security, Zero Trust, and AI security capabilities into market-facing messaging, sales enablement, and go-to-market strategy. Collaborate cross-functionally to drive adoption and competitive differentiation.
Top Skills: AIBrowser ExtensionsBrowser SecurityCasbCrowdstrike FalconFalcon Secure AccessGenaiIdentity-Aware AccessPolicy-Based Access ControlSafe BrowsingSaseSseSwgVdiVpnZero TrustZtna
An Hour Ago
Easy Apply
Remote or Hybrid
Easy Apply
130K-144K Annually
Mid level
130K-144K Annually
Mid level
Food • Software
Sell ChowNow’s restaurant technology to independent restaurants across an assigned territory. Generate revenue through inbound, partner, outbound, and referral leads; manage the full sales cycle from qualification through close; conduct product demonstrations; build customer relationships and referral networks; manage pipeline and KPIs in Salesforce; and meet or exceed quotas. The role requires approximately 60% territory travel, restaurant drop-ins, a valid driver’s license, an acceptable driving record, and access to a reliable insured vehicle.
Top Skills: Salesforce (Sfdc)
An Hour Ago
Easy Apply
Remote or Hybrid
Easy Apply
130K-144K Annually
Mid level
130K-144K Annually
Mid level
Food • Software
Sell ChowNow’s restaurant technology to independent restaurants through in-person territory sales. Responsibilities include qualifying leads, managing a sales pipeline, developing outbound and referral strategies, conducting product demos, closing full-cycle deals, building customer relationships, and meeting quotas. The role requires approximately 60% travel throughout the assigned territory, regular restaurant drop-ins, Salesforce pipeline management, and collaboration with sales leaders and peers.
Top Skills: Salesforce (Sfdc)

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account