Lead ownership of service lifecycle including architecture, deployment, operations, and optimization. Define production readiness, monitoring, CI/CD governance, automation-first practices, incident response, capacity planning, and cross-team alignment. Mentor engineers, drive resiliency and reliability improvements, and influence roadmaps to reduce toil and improve system performance.
Our Purpose
Mastercard powers economies and empowers people in 200+ countries and territories worldwide. Together with our customers, we're helping build a sustainable economy where everyone can prosper. We support a wide range of digital payments choices, making transactions secure, simple, smart and accessible. Our technology and innovation, partnerships and networks combine to deliver a unique set of products and services that help people, businesses and governments realize their greatest potential.
Title and Summary
Lead Site Reliability Engineer
Overview:
The role of Business Operations Organization is to be the production readiness steward for Mastercard products. As a Business Operations we are responsible for ensuring that our platform is stable and healthy. We break down barriers to run our products by fostering developer run ownership and empowering developers to build resilient products. We support our developers during the application build phase in software run principals that includes operational design, automation, capacity planning, monitoring that leads to fault-tolerant, scalable products. We see the big picture and help create and enforce operations standards while facilitating an agile and learning culture.
We accomplish this transformation through supporting daily operations with a hyper focus on triage and then root cause by understanding the business impact of our products. The goal of every biz ops team is to shift left to be more proactive and upfront in the development process, and to proactively manage production and change activities to maximize customer experience and increase the overall value of supported applications. Biz Ops teams also focus on risk management by tying all our activities together with an overarching responsibility for compliance and risk mitigation across all our environments. A biz ops focus is also on streamlining and standardizing traditional application specific support activities and centralizing points of interaction for both internal and external partners by communicating effectively with all key stakeholders.
Ultimately, the role of biz ops is to align Product and Customer Focused priorities with Operational needs. We regularly review our run state not only from an internal perspective but also understanding and providing the feedback loop to our development partners on how we can improve the customer experience of our applications.
Key Responsibilities• Lead and own the full lifecycle of services-from architecture and design through deployment, operations, and continuous optimization-ensuring scalability, reliability, and alignment with business objectives.• Analyze platform-level ITSM performance and proactively establish feedback loops with engineering teams, influencing roadmap prioritization to address systemic gaps and improve resiliency.• Define and drive production readiness standards, including operational design reviews, capacity planning, and launch governance, ensuring services meet reliability and scalability benchmarks before go-live.• Define and evolve monitoring frameworks for availability, latency, and system health, leveraging metrics and telemetry to proactively prevent incidents and improve service performance.• Champion automation-first principles to scale systems efficiently, reducing manual toil while improving deployment velocity and overall system reliability.• Lead the design and governance of CI/CD pipelines, implementing robust validation, operational gates, and best practices to drive consistency, quality, and speed across environments.• Drive best-in-class incident response practices, including rapid mitigation, stakeholder communication, and blameless postmortems, ensuring continuous improvement and resilience.• Take a holistic, system-wide approach during critical incidents, connect• Collaborate effectively across distributed, global teams, ensuring alignment, continuity, and high performance across time zones and technology hubs.• Act as a technical leader and mentor, developing junior engineers, promoting best practices, and raising the overall bar for engineering excellence within the organization.
All about you• Bachelor's degree in computer science, Engineering, or a related technical field (e.g., Physics, Mathematics), or equivalent practical experience.• 8-15 years of relevant experience in Site Reliability Engineering, Infrastructure, or DevOps roles, with a combination of hands-on technical expertise and early leadership responsibilities.• Strong technical foundation across enterprise platforms, Linux/UNIX systems, operating systems, and database environments (Oracle/SQL, DBA), with the ability to provide technical guidance and support to the team.• Experience with observability and monitoring tools (e.g., Splunk, Dynatrace), driving improved system visibility, performance, and reliability.• Solid experience in DevOps and CI/CD practices, with the ability to support and guide automation, deployment pipelines, and operational improvements.• Proficiency in one or more programming or scripting languages such as Python, Java, Go, C/C++, Perl, or Ruby, with practical application in automation or system• Strong foundation in Security and/or Enterprise Monitoring environments, with exposure to coding and system-level design.• Experience designing, analyzing, and troubleshooting large-scale distributed systems, with a strong focus on reliability, scalability, and performance optimization.• Strong program management capabilities, with a track record of successfully leading large-scale, cross-functional initiatives from concept through execution.• Extensive experience working across development, operations, and product teams to prioritize initiatives, build strong partnerships, and deliver end-to-end solutions.• Practical knowledge of cloud platforms, preferably AWS, with familiarity in cloud-native architectures and operational best practices.• Ability to critically assess existing processes and challenge the status quo, identifying opportunities to improve efficiency, scalability, and overall business impact.
We are seeking site reliability engineers with an appetite for change and who can push the boundaries of what can be completed through automation, while managing service levels for some of Mastercard's most critical security services.
Corporate Security Responsibility
All activities involving access to Mastercard assets, information, and networks comes with an inherent risk to the organization and, therefore, it is expected that every person working for, or on behalf of, Mastercard is responsible for information security and must:
Mastercard powers economies and empowers people in 200+ countries and territories worldwide. Together with our customers, we're helping build a sustainable economy where everyone can prosper. We support a wide range of digital payments choices, making transactions secure, simple, smart and accessible. Our technology and innovation, partnerships and networks combine to deliver a unique set of products and services that help people, businesses and governments realize their greatest potential.
Title and Summary
Lead Site Reliability Engineer
Overview:
The role of Business Operations Organization is to be the production readiness steward for Mastercard products. As a Business Operations we are responsible for ensuring that our platform is stable and healthy. We break down barriers to run our products by fostering developer run ownership and empowering developers to build resilient products. We support our developers during the application build phase in software run principals that includes operational design, automation, capacity planning, monitoring that leads to fault-tolerant, scalable products. We see the big picture and help create and enforce operations standards while facilitating an agile and learning culture.
We accomplish this transformation through supporting daily operations with a hyper focus on triage and then root cause by understanding the business impact of our products. The goal of every biz ops team is to shift left to be more proactive and upfront in the development process, and to proactively manage production and change activities to maximize customer experience and increase the overall value of supported applications. Biz Ops teams also focus on risk management by tying all our activities together with an overarching responsibility for compliance and risk mitigation across all our environments. A biz ops focus is also on streamlining and standardizing traditional application specific support activities and centralizing points of interaction for both internal and external partners by communicating effectively with all key stakeholders.
Ultimately, the role of biz ops is to align Product and Customer Focused priorities with Operational needs. We regularly review our run state not only from an internal perspective but also understanding and providing the feedback loop to our development partners on how we can improve the customer experience of our applications.
Key Responsibilities• Lead and own the full lifecycle of services-from architecture and design through deployment, operations, and continuous optimization-ensuring scalability, reliability, and alignment with business objectives.• Analyze platform-level ITSM performance and proactively establish feedback loops with engineering teams, influencing roadmap prioritization to address systemic gaps and improve resiliency.• Define and drive production readiness standards, including operational design reviews, capacity planning, and launch governance, ensuring services meet reliability and scalability benchmarks before go-live.• Define and evolve monitoring frameworks for availability, latency, and system health, leveraging metrics and telemetry to proactively prevent incidents and improve service performance.• Champion automation-first principles to scale systems efficiently, reducing manual toil while improving deployment velocity and overall system reliability.• Lead the design and governance of CI/CD pipelines, implementing robust validation, operational gates, and best practices to drive consistency, quality, and speed across environments.• Drive best-in-class incident response practices, including rapid mitigation, stakeholder communication, and blameless postmortems, ensuring continuous improvement and resilience.• Take a holistic, system-wide approach during critical incidents, connect• Collaborate effectively across distributed, global teams, ensuring alignment, continuity, and high performance across time zones and technology hubs.• Act as a technical leader and mentor, developing junior engineers, promoting best practices, and raising the overall bar for engineering excellence within the organization.
All about you• Bachelor's degree in computer science, Engineering, or a related technical field (e.g., Physics, Mathematics), or equivalent practical experience.• 8-15 years of relevant experience in Site Reliability Engineering, Infrastructure, or DevOps roles, with a combination of hands-on technical expertise and early leadership responsibilities.• Strong technical foundation across enterprise platforms, Linux/UNIX systems, operating systems, and database environments (Oracle/SQL, DBA), with the ability to provide technical guidance and support to the team.• Experience with observability and monitoring tools (e.g., Splunk, Dynatrace), driving improved system visibility, performance, and reliability.• Solid experience in DevOps and CI/CD practices, with the ability to support and guide automation, deployment pipelines, and operational improvements.• Proficiency in one or more programming or scripting languages such as Python, Java, Go, C/C++, Perl, or Ruby, with practical application in automation or system• Strong foundation in Security and/or Enterprise Monitoring environments, with exposure to coding and system-level design.• Experience designing, analyzing, and troubleshooting large-scale distributed systems, with a strong focus on reliability, scalability, and performance optimization.• Strong program management capabilities, with a track record of successfully leading large-scale, cross-functional initiatives from concept through execution.• Extensive experience working across development, operations, and product teams to prioritize initiatives, build strong partnerships, and deliver end-to-end solutions.• Practical knowledge of cloud platforms, preferably AWS, with familiarity in cloud-native architectures and operational best practices.• Ability to critically assess existing processes and challenge the status quo, identifying opportunities to improve efficiency, scalability, and overall business impact.
We are seeking site reliability engineers with an appetite for change and who can push the boundaries of what can be completed through automation, while managing service levels for some of Mastercard's most critical security services.
Corporate Security Responsibility
All activities involving access to Mastercard assets, information, and networks comes with an inherent risk to the organization and, therefore, it is expected that every person working for, or on behalf of, Mastercard is responsible for information security and must:
- Abide by Mastercard's security policies and practices;
- Ensure the confidentiality and integrity of the information being accessed;
- Report any suspected information security violation or breach, and
- Complete all periodic mandatory security trainings in accordance with Mastercard's guidelines.
Mastercard San Francisco, California, USA Office
123 Mission Street, San Francisco, CA, United States, 94105
Similar Jobs at Mastercard
Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Design next-generation B2B and B2B2C product experiences for Mastercard’s Global SME team. Responsibilities include contributing to design systems, creating reusable components, prototyping AI-driven experiences, developing high-fidelity visual designs, mapping user journeys, optimizing user engagement, and collaborating with product, UX research, engineering, and behavioral science teams. The role applies human-centered design, design thinking, atomic design, and rapid design-to-code workflows.
Top Skills:
Ai-Assisted Design ToolingAtomic DesignDesign SystemsDesign TokensFigmaFigma MakeUi KitsVibe Coding
Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Contributes to client consulting engagements by analyzing business issues, developing strategies and recommendations, preparing presentations, identifying risks, supporting proposals, and building client relationships. The role also contributes to intellectual capital and identifies opportunities for senior consulting staff. Strong communication, analytical, teamwork, and independent-working abilities are required, along with an undergraduate degree.
Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Own production readiness and reliability for platform services: design for scalability, automate deployments and recovery, monitor availability and performance, run incident response and postmortems, consult on capacity and launch reviews, and mentor junior engineers while collaborating across global development and product teams.
Top Skills:
AutomationCC++Ci/CdDevOpsDynatraceGoItsmJavaOraclePerlPythonRubyScriptingSplunkSQLUnix/Linux
What you need to know about the San Francisco Tech Scene
San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.
Key Facts About San Francisco Tech
- Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
- Major Tech Employers: Google, Apple, Salesforce, Meta
- Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
- Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
- Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
- Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

