About Us
Visa is a world leader in payments technology, facilitating transactions between consumers, merchants, financial institutions and government entities across more than 200 countries and territories, dedicated to uplifting everyone, everywhere by being the best way to pay and be paid.
At Visa, you'll have the opportunity to create impact at scale — tackling meaningful challenges, growing your skills and seeing your contributions impact lives around the world.
Join Visa and do work that matters – to you, to your community, and to the world. Progress starts with you.
Job Description
Job Summary
The Senior Manager, Site Reliability Engineering (SRE) – Data Platforms leads teams responsible for ensuring the reliability, scalability, security, performance, and operational excellence of data platforms, data infrastructure, and data-intensive production systems. This role supports the Data area by partnering with data engineering, analytics, machine learning, product, risk, compliance, security, and infrastructure teams to operate and improve the systems that capture, process, store, move, and serve critical business data.
This role oversees the development and execution of reliability strategies, infrastructure automation, observability practices, incident management processes, and operational standards for data pipelines, data platforms, distributed processing systems, cloud data services, streaming platforms, orchestration tools, and data storage environments. The Senior Manager ensures that data services meet expectations for availability, latency, throughput, recoverability, security, compliance, and data integrity.
This role is responsible for establishing processes, automation, monitoring, alerting, service-level objectives, incident response practices, disaster recovery capabilities, capacity planning, and operational excellence frameworks based on business and technical requirements. The Senior Manager guides teams in improving platform reliability, reducing toil, troubleshooting complex production issues, and supporting resilient data operations across high-volume, distributed environments.
The Senior Manager acts as a leader for SRE, infrastructure, platform, or operations engineering teams supporting the Data organization by removing blockers, enabling productivity, improving operational maturity, and fostering an inclusive, high-performance engineering culture. The role also involves coaching and developing engineers to adapt to evolving cloud platforms, data technologies, automation tools, observability frameworks, reliability practices, and AI-driven workflows. This leader partners closely with software engineering, data engineering, product management, QA, risk, compliance, security, infrastructure, and business stakeholders to deliver safe, resilient, transparent, and customer-centric data solutions.
All roles require digital fluency, including the ability to work with emerging technologies such as Generative AI tools — for example, ChatGPT and Claude Cod — to support everyday work, improve operational efficiency, accelerate documentation, assist with troubleshooting, and enhance engineering productivity.
Key Responsibilities
Lead the design, development, deployment, and operation of scalable, reliable, secure, and highly available SRE solutions supporting data platforms, data pipelines, and data infrastructure.
Oversee the creation and maintenance of automation, monitoring, alerting, observability, incident response, disaster recovery, and reliability engineering practices for data systems across multiple platforms and services.
Guide teams in troubleshooting and resolving complex production issues involving data pipelines, distributed processing systems, storage platforms, orchestration tools, streaming systems, infrastructure services, performance bottlenecks, and service degradations.
Drive the adoption of automation, CI/CD, infrastructure as code, quality engineering, observability, reliability testing, and operational excellence practices across the Data area.
Coach and mentor SRE, infrastructure, platform, or operations engineering teams supporting Data, fostering a culture of continuous improvement, accountability, learning, reliability, data integrity, and technical excellence.
Collaborate with data engineering, analytics, machine learning, product, security, risk, compliance, QA, infrastructure, and business partners to ensure solutions meet business priorities, customer expectations, regulatory requirements, security standards, data governance expectations, and operational risk controls.
Ensure adherence to reliability, resilience, security, compliance, access management, disaster recovery, change management, and data protection standards throughout the technology and data lifecycle.
Lead the implementation of service-level indicators, service-level objectives, error budgets, capacity planning, performance optimization, incident metrics, reliability dashboards, and operational health checks for data platforms and data services.
Partner with data engineering and platform teams to improve the reliability of batch pipelines, real-time streaming workloads, machine learning pipelines, data ingestion frameworks, transformation jobs, metadata processes, and reporting or analytics services.
Oversee the development and maintenance of technical documentation, runbooks, architecture diagrams, operational procedures, post-incident reviews, production-readiness checklists, and best practices for code quality, change review, and data platform supportability.
This is a hybrid position. Expectation of days in the office will be confirmed by your Hiring Manager.
Qualifications
Basic Qualifications
8+ years of relevant work experience and a Bachelor's degree, OR 11+ years of relevant work experience.
Experience in leading teams in the design, development, and deployment of large-scale engineering solutions.
Experience with cloud infrastructure, distributed systems, observability platforms, and reliability engineering practices.
Experience in building and maintaining scalable infrastructure, automation frameworks, incident response processes, and service reliability programs.
Experience with programming, scripting, and infrastructure tools such as Python, Go, Bash, Terraform, Kubernetes, CI/CD platforms, and version control systems such as Git.
Experience troubleshooting and resolving issues in complex, high-volume, highly available production environments.
Experience implementing security, compliance, access control, disaster recovery, and operational risk management standards.
Experience coaching, mentoring, and developing SRE, infrastructure, platform, or operations engineering teams.
Experience collaborating with engineering, product, security, infrastructure, and business stakeholders to improve reliability, scalability, and operational excellence.
Experience developing and maintaining technical documentation, runbooks, post-incident reviews, service-level objectives, operational procedures, and reliability dashboards.
Preferred Qualifications
9 or more years of relevant work experience with a Bachelor’s Degree, or 7 or more years of relevant experience with an Advanced Degree, or 3 or more years of experience with a PhD.
Experience supporting large-scale data platforms and data infrastructure in production environments.
Experience with cloud platforms and services, such as AWS, Azure, or GCP, used to operate reliable and scalable data systems.
Experience with Site Reliability Engineering practices, including monitoring, alerting, incident management, service-level objectives, error budgets, and post-incident reviews.
Experience improving the reliability of data pipelines, distributed processing systems, streaming platforms, orchestration tools, storage platforms, and analytics environments.
Experience troubleshooting complex production issues across data platforms, infrastructure, applications, and distributed systems.
Experience with automation, CI/CD, infrastructure as code, observability tools, containerization, and orchestration platforms.
Experience supporting systems requiring high availability, low latency, scalability, resiliency, and disaster recovery.
Experience partnering with data engineering, analytics, machine learning, product, security, risk, compliance, QA, and infrastructure teams.
Experience leading cross-functional SRE, platform, infrastructure, operations, or data engineering teams in a matrixed organization.
Visa is an EEO Employer
Qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability or protected veteran status. Visa will also consider for employment qualified applicants with criminal histories in a manner consistent with EEOC guidelines and applicable local law.