Design and build AI-driven automation, observability platforms, internal tools, APIs, and telemetry to improve infrastructure operations. Develop self-service solutions, monitoring and dashboards, automate operational workflows, troubleshoot production issues, perform root cause analysis, and partner with infrastructure and operations teams to improve reliability and scalability.
Position Summary...What you'll do...Role summary:
As a Software Engineer III, you will design and develop AI-driven automation, observability platforms, and internal engineering tools that improve the efficiency, reliability, and scalability of enterprise infrastructure operations. You will build self-service solutions, telemetry platforms, APIs, and intelligent workflows that reduce manual effort, accelerate incident resolution, and enhance operational visibility. Working closely with infrastructure, operations, and engineering teams, you will leverage software engineering, data-driven insights, and AI technologies to modernize operational processes, improve service reliability, and enable sustainable growth across Microsoft's large-scale compute infrastructure.
About the team:
The Enterprise Compute Frontline (ECF) team manages over 70,000 physical servers supporting critical compute, storage, and enterprise platforms. The team oversees infrastructure operations, incident response, hardware recovery, provisioning, observability, and service reliability at scale. ECF is advancing from manual operations to software-driven, AI-enabled infrastructure management. Engineers develop internal platforms, self-service tools, telemetry systems, dashboards, and intelligent automation solutions to enhance visibility, reduce operational effort, accelerate incident resolution, and improve service reliability. By integrating software engineering, observability, and AI-driven automation, the team ensures scalable and reliable service delivery across enterprise infrastructure.
What you'll do:
What you'll bring:
Eligibility requirements apply to some benefits and may depend on your job classification and length of employment. Benefits are subject to change and may be subject to a specific plan or program terms.
For information about benefits and eligibility, see One.Walmart.
The annual salary range for this position is $90,000.00 - $180,000.00 Additional compensation includes annual or quarterly performance bonuses. Additional compensation for certain positions may also include :
- Stock
Option 2: 4 years’ experience in software engineering or related area.Preferred Qualifications...
As a Software Engineer III, you will design and develop AI-driven automation, observability platforms, and internal engineering tools that improve the efficiency, reliability, and scalability of enterprise infrastructure operations. You will build self-service solutions, telemetry platforms, APIs, and intelligent workflows that reduce manual effort, accelerate incident resolution, and enhance operational visibility. Working closely with infrastructure, operations, and engineering teams, you will leverage software engineering, data-driven insights, and AI technologies to modernize operational processes, improve service reliability, and enable sustainable growth across Microsoft's large-scale compute infrastructure.
About the team:
The Enterprise Compute Frontline (ECF) team manages over 70,000 physical servers supporting critical compute, storage, and enterprise platforms. The team oversees infrastructure operations, incident response, hardware recovery, provisioning, observability, and service reliability at scale. ECF is advancing from manual operations to software-driven, AI-enabled infrastructure management. Engineers develop internal platforms, self-service tools, telemetry systems, dashboards, and intelligent automation solutions to enhance visibility, reduce operational effort, accelerate incident resolution, and improve service reliability. By integrating software engineering, observability, and AI-driven automation, the team ensures scalable and reliable service delivery across enterprise infrastructure.
What you'll do:
- Build internal tools, APIs, and self-service platforms that improve operational productivity and reduce manual effort.
- Develop AI-driven automation solutions for incident triage, workflow orchestration, and operational decision support.
- Design and maintain observability, telemetry, monitoring, and dashboard solutions that improve infrastructure visibility.
- Automate repetitive operational processes to improve efficiency and reduce alert fatigue.
- Troubleshoot production issues, perform root cause analysis, and drive continuous service improvements.
- Partner with infrastructure, operations, and engineering teams to improve reliability and scalability.
- Drive innovation through AI, automation, and data-driven engineering practices.
What you'll bring:
- Strong software engineering experience with Python, APIs, React, and distributed systems.
- Experience building internal platforms, automation solutions, and self-service tools.
- Knowledge of observability, telemetry, monitoring, dashboards, and operational analytics.
- Experience with AI/ML technologies, agentic AI, or intelligent automation frameworks.
- Ability to design scalable APIs and backend services supporting infrastructure operations.
- Strong troubleshooting, debugging, and root cause analysis skills.
- Understanding of cloud-native architectures and CI/CD practices.
- Passion for improving operational efficiency, reliability, and scalability through automation and data-driven engineering.
Eligibility requirements apply to some benefits and may depend on your job classification and length of employment. Benefits are subject to change and may be subject to a specific plan or program terms.
For information about benefits and eligibility, see One.Walmart.
The annual salary range for this position is $90,000.00 - $180,000.00 Additional compensation includes annual or quarterly performance bonuses. Additional compensation for certain positions may also include :
- Stock
ㅤ
ㅤ
ㅤ
ㅤ
Minimum Qualifications...Outlined below are the required minimum qualifications for this position. If none are listed, there are no minimum qualifications.
Option 1: Bachelor's degree in computer science, computer engineering, computer information systems, software engineering, or related area and 2 years’ experience in software engineering or related area.Option 2: 4 years’ experience in software engineering or related area.Preferred Qualifications...
Outlined below are the optional preferred qualifications for this position. If none are listed, there are no preferred qualifications.
Master’s degree in Computer Science, Computer Engineering, Computer Information Systems, Software Engineering, or related area, We value candidates with a background in creating inclusive digital experiences, demonstrating knowledge in implementing Web Content Accessibility Guidelines (WCAG) 2.2 AA standards, assistive technologies, and integrating digital accessibility seamlessly. The ideal candidate would have knowledge of accessibility best practices and join us as we continue to create accessible products and services following Walmart’s accessibility standards and guidelines for supporting an inclusive culture.Masters: Computer SciencePrimary Location...2403 Se J St, Bentonville, AR 72716, United States of AmericaWalmart and its subsidiaries are committed to maintaining a drug-free workplace and has a no tolerance policy regarding the use of illegal drugs and alcohol on the job. This policy applies to all employees and aims to create a safe and productive work environment.Walmart Global Tech Sunnyvale, California, USA Office
840 W California Ave, Sunnyvale, CA, United States, 94086
Similar Jobs
Big Data • Cloud • Logistics • Machine Learning • Retail
Design, implement, and support SAP security and access solutions (S/4HANA, Fiori, GRC). Create/maintain roles and authorizations, enforce SoD, troubleshoot security issues, automate security processes, support audits, optimize access management, and collaborate with cross-functional teams on implementations, migrations, and upgrades to ensure compliant, scalable enterprise access controls.
Top Skills:
Ai And AutomationAraArmBrmCloud ApplicationsEamIdentity And Access ManagementItsmSap Businessobjects (Bobj)Sap Bw/BiSap EccSap Enterprise PortalSap FioriSap Grc Access ControlSap HrSap S/4Hana
Big Data • Cloud • Logistics • Machine Learning • Retail
Design, develop, test, and maintain scalable catalog microservices and platform integrations. Troubleshoot production issues, perform root cause analysis, implement CI/CD and telemetry, write automation scripts, and mentor junior/offshore engineers while collaborating with cross-functional teams to deliver secure, maintainable solutions.
Top Skills:
Api DesignAutomation ScriptingAWSCi/CdMicroservicesAzureObject-Oriented ProgrammingTelemetry
Big Data • Cloud • Logistics • Machine Learning • Retail
Design and implement infrastructure automation, self-service platforms, and IaC/CI-CD pipelines to manage provisioning, hardware lifecycle, and fleet-scale operations. Troubleshoot production incidents, drive root-cause analysis, improve observability and reliability, collaborate with datacenter operations and vendors, and participate in on-call rotations to reduce MTTR and operational toil.
Top Skills:
AnsibleCi/CdCloud-NativeInfrastructure As Code (Iac)KubernetesLinuxNutanixPythonVMware
What you need to know about the San Francisco Tech Scene
San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.
Key Facts About San Francisco Tech
- Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
- Major Tech Employers: Google, Apple, Salesforce, Meta
- Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
- Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
- Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
- Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine
