Lead the Site Reliability Engineering team focusing on network infrastructure, ensuring performance, capacity, and reliability of Mastercard applications, and leading incident reviews and proactive issue detection.
Our Purpose
Mastercard powers economies and empowers people in 200+ countries and territories worldwide. Together with our customers, we're helping build a sustainable economy where everyone can prosper. We support a wide range of digital payments choices, making transactions secure, simple, smart and accessible. Our technology and innovation, partnerships and networks combine to deliver a unique set of products and services that help people, businesses and governments realize their greatest potential.
Title and Summary
Lead Engineer, Site Reliability Engineering
Our Purpose:
Mastercard powers economies and empowers people across more than 200 countries and territories worldwide.
We are committed to building an inclusive, digital economy that benefits everyone, everywhere-by making transactions safe, simple, smart, and accessible. Through secure data, trusted networks, strong partnerships, and relentless innovation, we help individuals, financial institutions, governments, and businesses unlock their greatest potential.
About the Role:
Mastercard's Program aligned Site Reliability Engineering (SRE) teams are dedicated to delivering a seamless experience for our customers. We achieve this by maintaining every aspect of our Programs infrastructure and technology ecosystem to the highest standards, ensuring compliance with rigorous security requirements.
Within Mastercard, SRE focuses on the reliability and performance of core infrastructure, networks, and foundational services that power our applications. Our mission is to ensure these components operate with excellence, enabling applications to deliver an outstanding customer experience.
In this role, you will join our Payments Network SRE team and take ownership of continuously assessing and elevating the end to end service quality of our platform. You will leverage data to drive root cause analysis and deliver strategic insights to key stakeholders on resource utilization, capacity forecasting, and performance trends-ensuring the availability, scalability, and resilience of our network.
Key Responsibilities:
Lead continuous assessments of the application infrastructure supporting critical Mastercard applications, focusing on health, performance, monitoring and alerting, and capacity analysis. Collaborate with Product and Development teams to forecast growth requirements and ensure scalability and resiliency.
Champion observability as a core principle for infrastructure services by assessing environments and technologies to uncover gaps in monitoring and alerting. Design and implement strategies to close these gaps, ensuring all infrastructure telemetry is integrated into a unified, single-pane-of-glass view. Build custom dashboards to investigate and perform root cause analysis on complex issues.
Lead regular incident reviews with internal support teams to ensure root causes are identified. When patterns of failure or compatibility issues between software and infrastructure emerge, develop and implement strategies to remediate or mitigate risks.
Leverage automation and AI technologies to enhance proactive issue detection, enable self-healing capabilities, reducing Mean Time to Detect (MTTD) and Mean Time to Mitigate (MTTM).
Develop testing and validation plans for new environment builds, disaster recovery exercises and post-maintenance activities to certify environment readiness before customer traffic is routed to it.
Champion continuous learning, development, and knowledge sharing across networking and other infrastructure disciplines to strengthen multi-disciplinary SRE team capabilities. Lead training initiatives for team members and Product and Development on networking aspects of the platforms.
Evaluate vendor hardware, firmware, and software upgrade roadmaps, and conduct proof-of-concept (POC) testing to identify potential risks and opportunities for improvement in upcoming releases.
All about you:
5-10 years of experience in an SRE or SRE related operations role, including 3+ years supporting e commerce, financial services, or large scale SaaS platforms.
Excellent infrastructure troubleshooting and analytical problem solving skills.
Strong hands on experience with observability and monitoring tools such as Splunk, Dynatrace, or equivalent, with a proven ability to triage and investigate complex issues.
Familiarity with network telemetry tools such as SolarWinds and NetScout.
Proficiency in packet level debugging, including capturing traffic with tools like tcpdump and analyzing packets using Wireshark.
Broad understanding of end to end infrastructure supporting payment platforms-spanning platform services, networking, databases, and storage.
Experience with automation and Infrastructure as Code tools such as Chef, Ansible, and Terraform, as well as structured data formats (JSON/YAML).
Excellent communication skills with the ability to coordinate cross functional troubleshooting efforts and lead RCA processes to closure.
Demonstrated ability to troubleshoot complex production issues, perform root cause analysis, and drive long term corrective actions.
Experience partnering with development teams to shape architecture, define SLIs/SLOs, and embed reliability into services from design through operation.
Strong understanding of monitoring and observability ecosystems, including Prometheus, Grafana, ELK/EFK, Splunk, and OpenTelemetry.
Effective incident management skills with a structured, analytical approach to problem solving.
The Payments Network SRE team is responsible for the runtime availability of some of Mastercard's most critical core payment systems, which support national infrastructure and operate 24/7 year-round. As a result, this role will include periodic on-call responsibilities when required.
Corporate Security Responsibility
All activities involving access to Mastercard assets, information, and networks comes with an inherent risk to the organization and, therefore, it is expected that every person working for, or on behalf of, Mastercard is responsible for information security and must:
• Abide by Mastercard's security policies and practices;• Ensure the confidentiality and integrity of the information being accessed;• Report any suspected information security violation or breach, and• Complete all periodic mandatory security trainings in accordance with Mastercard's guidelines.
Corporate Security Responsibility
All activities involving access to Mastercard assets, information, and networks comes with an inherent risk to the organization and, therefore, it is expected that every person working for, or on behalf of, Mastercard is responsible for information security and must:
Mastercard powers economies and empowers people in 200+ countries and territories worldwide. Together with our customers, we're helping build a sustainable economy where everyone can prosper. We support a wide range of digital payments choices, making transactions secure, simple, smart and accessible. Our technology and innovation, partnerships and networks combine to deliver a unique set of products and services that help people, businesses and governments realize their greatest potential.
Title and Summary
Lead Engineer, Site Reliability Engineering
Our Purpose:
Mastercard powers economies and empowers people across more than 200 countries and territories worldwide.
We are committed to building an inclusive, digital economy that benefits everyone, everywhere-by making transactions safe, simple, smart, and accessible. Through secure data, trusted networks, strong partnerships, and relentless innovation, we help individuals, financial institutions, governments, and businesses unlock their greatest potential.
About the Role:
Mastercard's Program aligned Site Reliability Engineering (SRE) teams are dedicated to delivering a seamless experience for our customers. We achieve this by maintaining every aspect of our Programs infrastructure and technology ecosystem to the highest standards, ensuring compliance with rigorous security requirements.
Within Mastercard, SRE focuses on the reliability and performance of core infrastructure, networks, and foundational services that power our applications. Our mission is to ensure these components operate with excellence, enabling applications to deliver an outstanding customer experience.
In this role, you will join our Payments Network SRE team and take ownership of continuously assessing and elevating the end to end service quality of our platform. You will leverage data to drive root cause analysis and deliver strategic insights to key stakeholders on resource utilization, capacity forecasting, and performance trends-ensuring the availability, scalability, and resilience of our network.
Key Responsibilities:
Lead continuous assessments of the application infrastructure supporting critical Mastercard applications, focusing on health, performance, monitoring and alerting, and capacity analysis. Collaborate with Product and Development teams to forecast growth requirements and ensure scalability and resiliency.
Champion observability as a core principle for infrastructure services by assessing environments and technologies to uncover gaps in monitoring and alerting. Design and implement strategies to close these gaps, ensuring all infrastructure telemetry is integrated into a unified, single-pane-of-glass view. Build custom dashboards to investigate and perform root cause analysis on complex issues.
Lead regular incident reviews with internal support teams to ensure root causes are identified. When patterns of failure or compatibility issues between software and infrastructure emerge, develop and implement strategies to remediate or mitigate risks.
Leverage automation and AI technologies to enhance proactive issue detection, enable self-healing capabilities, reducing Mean Time to Detect (MTTD) and Mean Time to Mitigate (MTTM).
Develop testing and validation plans for new environment builds, disaster recovery exercises and post-maintenance activities to certify environment readiness before customer traffic is routed to it.
Champion continuous learning, development, and knowledge sharing across networking and other infrastructure disciplines to strengthen multi-disciplinary SRE team capabilities. Lead training initiatives for team members and Product and Development on networking aspects of the platforms.
Evaluate vendor hardware, firmware, and software upgrade roadmaps, and conduct proof-of-concept (POC) testing to identify potential risks and opportunities for improvement in upcoming releases.
All about you:
5-10 years of experience in an SRE or SRE related operations role, including 3+ years supporting e commerce, financial services, or large scale SaaS platforms.
Excellent infrastructure troubleshooting and analytical problem solving skills.
Strong hands on experience with observability and monitoring tools such as Splunk, Dynatrace, or equivalent, with a proven ability to triage and investigate complex issues.
Familiarity with network telemetry tools such as SolarWinds and NetScout.
Proficiency in packet level debugging, including capturing traffic with tools like tcpdump and analyzing packets using Wireshark.
Broad understanding of end to end infrastructure supporting payment platforms-spanning platform services, networking, databases, and storage.
Experience with automation and Infrastructure as Code tools such as Chef, Ansible, and Terraform, as well as structured data formats (JSON/YAML).
Excellent communication skills with the ability to coordinate cross functional troubleshooting efforts and lead RCA processes to closure.
Demonstrated ability to troubleshoot complex production issues, perform root cause analysis, and drive long term corrective actions.
Experience partnering with development teams to shape architecture, define SLIs/SLOs, and embed reliability into services from design through operation.
Strong understanding of monitoring and observability ecosystems, including Prometheus, Grafana, ELK/EFK, Splunk, and OpenTelemetry.
Effective incident management skills with a structured, analytical approach to problem solving.
The Payments Network SRE team is responsible for the runtime availability of some of Mastercard's most critical core payment systems, which support national infrastructure and operate 24/7 year-round. As a result, this role will include periodic on-call responsibilities when required.
Corporate Security Responsibility
All activities involving access to Mastercard assets, information, and networks comes with an inherent risk to the organization and, therefore, it is expected that every person working for, or on behalf of, Mastercard is responsible for information security and must:
• Abide by Mastercard's security policies and practices;• Ensure the confidentiality and integrity of the information being accessed;• Report any suspected information security violation or breach, and• Complete all periodic mandatory security trainings in accordance with Mastercard's guidelines.
Corporate Security Responsibility
All activities involving access to Mastercard assets, information, and networks comes with an inherent risk to the organization and, therefore, it is expected that every person working for, or on behalf of, Mastercard is responsible for information security and must:
- Abide by Mastercard's security policies and practices;
- Ensure the confidentiality and integrity of the information being accessed;
- Report any suspected information security violation or breach, and
- Complete all periodic mandatory security trainings in accordance with Mastercard's guidelines.
Mastercard San Francisco, California, USA Office
123 Mission Street, San Francisco, CA, United States, 94105
Similar Jobs at Mastercard
Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Supports technology regulatory execution, resilience, risk management, governance, and reporting initiatives. Gathers and documents business, process, data, and reporting requirements; maps controls and processes to regulatory standards; analyzes data and metrics; validates dashboards and reports; tracks remediation activities; and identifies process improvements. Collaborates with business, technology, risk, regulatory, and operational stakeholders while preparing clear documentation, presentations, and recommendations.
Top Skills:
APIsCopilot StudioData ModelsEnterprise Reporting PlatformsExcelPower AppsPower AutomatePower BISQL
Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Leads Mastercard’s Advanced Analytics sales strategy across Asia Pacific’s retail and commerce sectors. Develops executive relationships, builds strategic pipelines, originates and closes complex enterprise opportunities, and sells Test & Learn and related analytics solutions. Advises clients on customer acquisition, loyalty, marketing effectiveness, digital transformation, and data-driven growth. Collaborates with account, product, consulting, and customer success teams while representing Mastercard at industry and executive events.
Top Skills:
Advanced AnalyticsAICustomer DataMarketing TechnologySaaSTest & Learn
Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Manage multi-country financial close, forecasting, budgeting, variance analysis, reporting, and KPI development. Partner with finance and business stakeholders on planning cycles, revenue targets, strategic reviews, and leadership presentations. Improve financial processes through automation, reporting solutions, financial tools, and analytical models. Provide business insights, communicate risks and opportunities, and support regional and global initiatives in a matrixed environment.
Top Skills:
AlteryxExcelHyperionOraclePower BITableauVBA
What you need to know about the San Francisco Tech Scene
San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.
Key Facts About San Francisco Tech
- Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
- Major Tech Employers: Google, Apple, Salesforce, Meta
- Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
- Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
- Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
- Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

