Intel Logo

Intel

Reliability Engineer

Posted Yesterday
Be an Early Applicant
In-Office
Santa Clara, CA, USA
122K-232K Annually
Mid level
In-Office
Santa Clara, CA, USA
122K-232K Annually
Mid level
Define and own pod-level reliability specs for compute, memory, storage, network, power, and cooling. Translate SLAs to subsystem requirements, lead FMEA/root-cause analysis and fleet failure analytics, architect RAS features, partner on power/cooling redundancy, and establish HALT/HASS, burn-in, qualification processes while tracking field returns and KPIs.
The summary above was generated by AI
Job Details:

Job Description: 

Join us to help build the next generation of AI hardware solutions. You will be part of a highly skilled, agile team developing cutting-edge hardware for the AI domain, where we push the boundaries of what silicon can do for emerging AI workloads. With a startup-like culture, we move quickly and give engineers the opportunity to drive significant technical and business impact. 

We are continuously developing modern and effective working methods, including hands-on adoption of AI tools throughout the chip development flow.  

Mission: Define and own the pod-level reliability specifications that ensure the availability, resilience, and serviceability of a large-scale data center across hardware, thermal, and operational dimensions.

Responsibilities:

  • Define and maintain pod-level reliability/availability specs and targets (MTBF, AFR, RAS) for compute, memory, storage, network, power, and cooling subsystems.

  • Translate system/SLA requirements into pod and subsystem level reliability specs; flow requirements down to silicon, platform, and facilities teams.

  • Lead FMEA, root-cause analysis, and pod fleet failure-data analytics to drive corrective actions and spec updates.

  • Architect RAS features (ECC, memory mirroring, predictive failure, telemetry) and graceful degradation/redundancy against pod-level specs.

  • Partner with facilities on pod power/cooling redundancy (N+1, 2N), thermal margins, and disaster-recovery readiness.

  • Establish HALT/HASS, burn-in, qualification processes; track field returns and KPIs against pod spec.

Qualifications:

Minimum Qualifications:

  • BS/MS/PhD in EE/ME Reliability or related; and/or at least 4-6 yrs experience.

  • Experience authoring and owning reliability specs and requirement flow-down.

  • Strong RAS, FMEA, statistical reliability (Weibull, FIT) skills.

  • Experience with large-scale fleet telemetry and thermal/power redundancy.

Preferred Qualifications:

  • AI cluster operations, data analytics (Python/SQL).

          

Job Type:Experienced Hire

Shift:Shift 1 (United States of America)

Primary Location: US, Massachusetts, Beaver Brook

Additional Locations:US, California, Santa Clara

Posting Statement:All qualified applicants will receive consideration for employment without regard to race, color, religion, religious creed, sex, national origin, ancestry, age, physical or mental disability, medical condition, genetic information, military and veteran status, marital status, pregnancy, gender, gender expression, gender identity, sexual orientation, or any other characteristic protected by local law, regulation, or ordinance.Position of TrustN/ABenefits

We offer a total compensation package that ranks among the best in the industry. It consists of competitive pay, stock bonuses, and benefit programs which include health, retirement, and vacation. Find out more about the benefits of working at Intel.



Annual Salary Range for jobs which could be performed in the US: $122,440.00-232,190.00 USD

The range displayed on this job posting reflects the minimum and maximum target compensation for the position across all US locations. Within the range, individual pay is determined by work location and additional factors, including job-related skills, experience, and relevant education or training. Your recruiter can share more about the specific compensation range for your preferred location during the hiring process.

Work Model for this Role

This role will require an on-site presence. * Job posting details (such as work model, location or time type) are subject to change.

*

ADDITIONAL INFORMATION: Intel is committed to Responsible Business Alliance (RBA) compliance and ethical hiring practices. We do not charge any fees during our hiring process. Candidates should never be required to pay recruitment fees, medical examination fees, or any other charges as a condition of employment. If you are asked to pay any fees during our hiring process, please report this immediately to your recruiter.
HQ

Intel Santa Clara, California, USA Office

Robert Noyce Building, Santa Clara, CA, United States, 95052

Intel Mountain View, California, USA Office

Mountain View, United States

Intel San Jose, California, USA Office

San Jose, United States

Intel Santa Clara, California, USA Office

2200 Mission College Blvd. , Santa Clara, CA, United States, 95054

Intel Sunnyvale, California, USA Office

Sunnyvale, United States

Similar Jobs

Yesterday
Remote or Hybrid
United States
112K-140K Annually
Junior
112K-140K Annually
Junior
Digital Media • Gaming • Information Technology • Software • Sports • Esports • Big Data Analytics
Support and improve reliability, scalability, and performance of large-scale databases across cloud and on-prem. Build automation, Kubernetes operators, GitOps workflows, and tooling in Go or Python; implement observability, capacity planning, failover, backups, schema migrations, and self-healing systems; partner with application teams and leverage AI to improve operations and engineering productivity.
Top Skills: AerospikeArgocdAurora MysqlClaudeCursorEksFluxcdGithub CopilotGitopsGkeGoKubernetesKubernetes OperatorsMcpMongoDBMySQLPersistent VolumesPostgresPulumiPythonRedisScylladbStatefulsetsTerraform
3 Days Ago
In-Office
107K-193K Annually
Mid level
107K-193K Annually
Mid level
Aerospace • Information Technology • Software • Cybersecurity • Design • Defense • Manufacturing
Drive supplier quality and material reliability across supply chain, engineering, and manufacturing. Lead containment and root-cause analysis for nonconformances, track yields and escapes, implement corrective actions, improve supplier capability, and incorporate lessons learned into NPI to meet program reliability goals.
Top Skills: Advanced CompositesApqpAs9100Avionics IntegrationCnc MachiningControl PlansDesign Of ExperimentsFluid System Pressure TestingLeanMeasurement Systems AnalysisMetrologyPfmeaStatistical Process ControlWelding
5 Days Ago
Remote or Hybrid
Santa Clara, CA, USA
167K-291K Annually
Senior level
167K-291K Annually
Senior level
Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Design, build, and operate cloud-native reliability, release, and test platforms. Integrate test pipelines, observability, and automation into CI/CD and GitOps workflows. Build Kubernetes-based scalable test infrastructure, progressive delivery, chaos/resilience testing, and validation for deployment and operational health. Mentor engineers and drive platform reliability, automation-first solutions, and developer self-service environments.
Top Skills: AnsibleArgo CdArgo WorkflowsAws EksAzure AksCi/CdCypressFluxGateway ApiGitlab Ci/CdGitopsGoGoogle GkeHelmIngressIstioJavaJunitKubernetesKustomizeLinkerdOpentelemetryPlaywrightPrometheusPytestPythonRest AssuredRubySeleniumTerraformTestng

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account