A NVIDIA UltraServer isn't a server: it's 18 compute nodes, 72 Blackwell GPUs, and an NVLink fabric that behaves as one giant, memory-coherent computer. When any part of it fails, you have to diagnose across racks, orchestrate hardware swaps with data center technicians, rebuild the NVLink fabric, re-run burn-in, and hand healthy capacity back to customers training frontier AI models. Every hour an UltraServer sits idle is some of the most sought-after compute on Earth going to waste.
That's the problem our team owns, and we're looking for a Software Development Manager to lead it.
The EC2 UltraServer Availability team builds the automated repair and recovery systems for Amazon's GB200 and GB300 fleet. One of the fastest-growing and most visible infrastructure domains at AWS. This is early-days territory: many repair flows that are fully automated for traditional EC2 hosts are still being invented for UltraServers. You'll set the technical direction that turns manual, expert-driven recovery into deterministic, self-healing automation, and you'll see your team's impact directly in fleet availability numbers that leadership watches weekly.
Key job responsibilities
You'll lead a two-pizza team of 8–12 engineers automating end-to-end repair of entire UltraServers: detection, control-plane teardown, chassis swap orchestration with hardware engineering and data center operations, fabric rebuild, burn-in testing, and return to customer.
Your north-star metrics are fleet availability and repair dwell time: how fast a broken UltraServer gets back to serving customers.
Examples of projects your team is working on
1. Automating deterministic repair: building the system that maps a failure signature directly to the right repair action, removing humans from the loop for well-understood failures
2. Orchestrating multi-team repair procedures that today require coordination across hardware engineering, data center technicians, and EC2 control-plane services
3. Cutting repair dwell time by parallelizing teardown, physical repair, and validation steps that currently run serially
4. Extending recovery workflows into new regions, including air-gapped (ADC) environments
A day in the life
You'll spend most of your time where an SDM should: growing engineers (1:1s, career development, hiring), making roadmap and architecture calls with your senior engineers, and unblocking the team. Because we sit at the intersection of software, hardware, and data center operations, you'll regularly partner with EC2 ML Supercomputing, hardware engineering, and DCO leadership. You'll also drive operational reviews - this is a fleet-facing team with an on-call rotation, and reducing that operational load through automation is itself a core part of the charter. You'll communicate progress and strategy to senior leadership through narratives and business reviews.
About the team
The EC2 UltraServer Availability team (part of EC2 Nitro) maintains the availability of NVIDIA-based ML infrastructure at scale. We own end-to-end recovery and repair for GB200 and GB300 UltraServers: from detecting an availability event through repair, testing, and return to service. We work closely with hardware engineering, data center operations, EC2 ML Supercomputing, and EC2 capacity teams. We value engineers who are curious about the full stack, from NVLink cabling to control-plane APIs, and we invest in reducing our own operational toil through automation.
Basic Qualifications
- 3+ years of engineering team management experience
- 7+ years of working directly within engineering teams experience
- 3+ years of designing or architecting (design patterns, reliability and scaling) of new and existing systems experience
- 8+ years of leading the definition and development of multi tier web services experience
- Knowledge of engineering practices and patterns for the full software/hardware/networks development life cycle, including coding standards, code reviews, source control management, build processes, testing, certification, and livesite operations
- Experience partnering with product or program management teams
Preferred Qualifications
- Experience in communicating with users, other technical teams, and senior leadership to collect requirements, describe software product features, technical designs, and product strategy
- Experience in recruiting, hiring, mentoring/coaching and managing teams of Software Engineers to improve their skills, and make them more effective, product software engineers
Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.
Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit
https://amazon.jobs/content/en/how-we-hire/accommodations for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.
The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits.
USA, WA, Seattle - 184,900.00 - 250,200.00 USD annually