Ai & Technology Governance, Technology Risk Contract Type:
Permanent Job Reference:
2988837242-2 Apply for this job now Job Description Mastercard is seeking a BizOps Engineer II to join our Information Technology & Data Management team in Financial Services. In this role, you will optimize and support mission-critical platforms that power secure, high-volume transactions worldwide. You'll design and implement robust monitoring, automation, and incident management solutions to ensure high availability, performance, and scalability. Collaborating closely with software engineers, data engineers, and product teams, you will troubleshoot complex issues, analyze system metrics, and drive continuous improvement across infrastructure and applications. This position offers the opportunity to work with cutting-edge cloud and data technologies while influencing operational best practices and reliability standards. You'll contribute to incident response, root-cause analysis, and post-incident reviews, helping to build more resilient systems. Mastercard's culture emphasizes innovation, collaboration, and continuous learning, giving you room to experiment with new tools and approaches. If you are passionate about system reliability, automation, and bridging the gap between development and operations in a dynamic, global environment, this role provides a chance to make a tangible impact on secure digital payments worldwide. Responsibilities Design, implement, and maintain monitoring, alerting, and observability for mission-critical applications and infrastructure. Automate operational tasks, deployments, and remediation workflows to improve reliability and reduce manual intervention. Collaborate with software and data engineering teams to optimize system performance, scalability, and resilience. Lead and participate in incident response, troubleshooting complex production issues, and driving timely resolution. Conduct root-cause analysis and implement long-term fixes to prevent recurrence of incidents. Support CI/CD pipelines and release processes to enable safe, rapid, and reliable deployments. Analyze system metrics and logs to identify bottlenecks, trends, and optimization opportunities. Contribute to reliability best practices, runbooks, and operational documentation. Partner with security and compliance teams to ensure systems meet regulatory and security requirements. Mentor junior team members and help foster a culture of continuous improvement and learning. Required Skills Site Reliability Engineering (SRE) practices Cloud platforms (AWS, GCP, or Azure; AWS preferred) Infrastructure as Code (Terraform, Cloud Formation, or similar) Linux systems administration Containerization and orchestration (Docker, Kubernetes) CI/CD pipelines (Jenkins, Git Lab CI, or similar) Monitoring and observability (Splunk, Prometheus, Grafana, Cloud Watch) Scripting/programming (Python, Bash, or similar) SQL and basic data querying/analysis Incident management and root-cause analysis