AWS Infrastructure Services owns the design, planning, delivery, and operation of all AWS global infrastructure. In other words, we’re the people who keep the cloud running. We support all AWS data centers and all of the servers, storage, networking, power, and cooling equipment that ensure our customers have continual access to the innovation they rely on. We work on the most challenging problems, with thousands of variables impacting the supply chain — and we’re looking for talented people who want to help.
You’ll join a diverse team of technical program managers, design engineers, construction managers, network engineers, supply chain specialists, security experts, operations managers, and other vital roles. You’ll collaborate with people across AWS to help us deliver the highest standards for safety and security while providing seemingly infinite capacity at the lowest possible cost for our customers. And you’ll experience an inclusive culture that welcomes bold ideas and empowers you to own them to completion.
The ML Capacity Delivery Team (MLZ) is seeking a Principal Technical Infrastructure Program Manager to define and drive the technical vision, strategy, and roadmap for the systems and tools that enable ML capacity delivery at global scale. In this role, you will serve as the single-threaded technical leader responsible for building, evolving, and scaling the platforms and automation that underpin how we plan, track, and deliver ML infrastructure; from demand signal through rack-level installation.
You will work at the intersection of infrastructure delivery, data engineering, and tooling while partnering with systems engineers, software development teams, Business Intelligence Engineers (BIE), operations leaders, and cross-functional stakeholders to translate complex operational requirements into scalable, reliable systems. The ideal candidate is a seasoned technical leader who thrives in ambiguity, can articulate a compelling 1–3 year technical roadmap, and has a track record of driving cross-organizational alignment to deliver platforms that fundamentally improve how teams operate at scale.
Key job responsibilities
• Own the technical vision and multi-year roadmap (1–3 years) for ML Capacity Delivery systems and tools, aligning investments to business priorities and capacity delivery goals across core and Gen AI platforms.
• Drive cross-functional alignment across systems engineering, software development, BIE, operations, and planning teams to define requirements, prioritize capabilities, and deliver integrated tooling solutions.
• Architect and deliver scalable platforms that automate and optimize capacity delivery workflows; including demand planning, build tracking, supply chain visibility, and operational reporting.
• Serve as the technical authority for systems and tools within MLZ, influencing Director/VP-level leadership on investment decisions, build-vs-buy trade-offs, and platform strategy through data-driven business cases.
• Identify and eliminate operational friction by mapping end-to-end workflows, quantifying inefficiencies, and delivering automation that reduces manual effort and accelerates delivery velocity.
• Establish engineering excellence; define standards for system reliability, data quality, observability, and documentation that enable teams to build and operate with confidence at scale.
• Force-multiply across the organization by creating reusable frameworks, mechanisms, and best practices that elevate the effectiveness of multiple teams and programs simultaneously.
• Drive clarity in highly ambiguous environments where the business strategy, architectural approach, and problem definition may not yet exist. While decomposing complex challenges into actionable execution plans.
• Build and maintain strategic partnerships with internal platform teams, data engineering organizations, and tooling providers to ensure MLZ systems integrate seamlessly with the broader AWS infrastructure ecosystem.
• Establish metrics, dashboards, and reporting mechanisms to provide stakeholders and senior leadership clear insight into platform health, delivery timelines, adoption, and business impact.
About the team
*Why AWS*
Amazon Web Services (AWS) is the world’s most comprehensive and broadly adopted cloud platform. We pioneered cloud computing and never stopped innovating — that’s why customers from the most successful startups to Global 500 companies trust our robust suite of products and services to power their businesses.
*Diverse Experiences*
Amazon values diverse experiences. Even if you do not meet all of the preferred qualifications and skills listed in the job description, we encourage candidates to apply. If your career is just starting, hasn’t followed a traditional path, or includes alternative experiences, don’t let it stop you from applying.
*Work/Life Balance*
We value work-life harmony. Achieving success at work should never come at the expense of sacrifices at home, which is why we strive for flexibility as part of our working culture. When we feel supported in the workplace and at home, there’s nothing we can’t achieve in the cloud.
*Inclusive Team Culture*
Here at AWS, it’s in our nature to learn and be curious. Our employee-led affinity groups foster a culture of inclusion that empower us to be proud of our differences. Ongoing events and learning experiences, including our Conversations on Race and Ethnicity (CORE) and AmazeCon conferences, inspire us to never stop embracing our uniqueness.
*Mentorship and Career Growth*
We’re continuously raising our performance bar as we strive to become Earth’s Best Employer. That’s why you’ll find endless knowledge-sharing, mentorship and other career-advancing resources here to help you develop into a better-rounded professional.
Basic Qualifications
- 10+ years of program management, tech program management, or related experience
- 10+ years of working directly with engineering teams experience
- Experience defining roadmap strategy and prioritizing deliverables for your team products
- 4+ years of designing or architecting (design patterns, reliability and scaling) of new and existing systems experience
- Bachelor's degree or above in Computer Science, Computer Engineering, Information Management, Information Systems, or other related discipline
- • Experience managing programs across cross-functional teams, building processes and coordinating release schedules
Preferred Qualifications
- Experience working in data centers or critical infrastructure
- Experience in automation or monitoring frameworks, deployment or development
- Experience performing complex business case analysis to justify technical decisions and presenting the justification to management in a high-level review
- Experience communicating and presenting to senior leadership
- • Experience with AI/ML infrastructure, GPU/accelerator deployments, or high-performance computing environments
- • Track record of driving technical standardization and process improvement across globally distributed teams
- • MBA or Master's degree in Engineering, Computer Science, or related field
Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.
Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit
https://amazon.jobs/content/en/how-we-hire/accommodations for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.
The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits.
USA, VA, Herndon - 176,900.00 - 239,400.00 USD annually
USA, WA, Seattle - 176,900.00 - 239,400.00 USD annually