Join Deloitte’s AI & Engineering team to help transform technology platforms, modernize mission-critical operations, and drive innovation for leading organizations. In this HMS SRE Lead role, you will build and operate automation, observability, and cloud engineering capabilities that improve reliability, reduce operational effort, and enable AI-assisted incident investigation, triage, and resolution. You’ll work with cross-functional teams to deliver engineering-led solutions that help clients improve performance, accelerate digital outcomes, and support business growth.
AI & Engineering leverages innovative engineering capabilities to build, deploy, and operate integrated/verticalized sector solutions in software, data, AI, network, and hybrid cloud infrastructure. These solutions are powered by engineering for business advantage, transforming mission-critical operations. We enable clients to stay ahead with the latest advancements by transforming engineering teams and modernizing technology and data platforms. Our delivery models are tailored to meet each client's unique requirements.
Recruiting for this role ends on 09/04/2026.
Work you'll do
As an HMS SRE Lead on the Hybrid Managed Services team, you will be responsible for:
- Design, build, and maintain automated remediation pipelines that reduce manual effort and improve uptime across distributed client environments
• Embed AI-enabled automation into incident, event, and problem management workflows in collaboration with AI Enterprise Architects and delivery teams
• Define service level indicators and service level objectives and drive measurable improvements in mean time to resolution and major incident volume
• Lead root cause analysis, post-incident reviews, and continuous improvement actions across automation, detection, and operational processes
• Design and implement secure cloud infrastructure and continuous integration and continuous deployment pipelines using infrastructure-as-code
• Advise client teams on observability and IT service management tooling, including integrations, dashboards, alerting, and short-term prototyping support
A successful candidate would possess these skills:
- Ability to work independently and collaborate as part of a team
• Effective written and verbal communication skills
• Meticulous attention to detail and quality of work product
• Ability to build and sustain professional relationships
• Ability to lead projects or workstreams
• Ability to manage and prioritize multiple tasks in a fast-paced and dynamic environment
• Strong interpersonal skills and professional demeanor
• Ability to meet deadlines
• Ability to mentor and provide clear guidance to others
The team
AI & Engineering helps organizations transform technology platforms, modernize critical operations, and drive innovation at scale. Professionals across the practice work with clients to improve performance, launch new digital capabilities, and increase productivity through engineering-led solutions.
Qualifications
Required:
- Bachelor's degree in Computer Science, Engineering, Information Systems, or a technical field; alternatively, equivalent practical experience
• 10+ years of IT infrastructure, operations, site reliability engineering, or platform engineering experience, including 3+ years in a dedicated site reliability engineering or platform engineering role
• Experience designing and operating Amazon Web Services infrastructure, including Elastic Compute Cloud, Simple Storage Service, Virtual Private Cloud, Identity and Access Management, Lambda, CloudWatch, and CloudTrail
• Experience building automated remediation pipelines, infrastructure-as-code deployments, and continuous integration and continuous deployment workflows using tools such as Terraform, CloudFormation, Ansible, Puppet, or Jenkins
• Experience with at least 3 of the following platforms: ServiceNow IT Operations Management or Event Management, Splunk, SolarWinds, Grafana, Dynatrace, AppDynamics, Datadog, New Relic, or PagerDuty
• Ability to travel 50%, on average, based on the work you do and the clients and industries/sectors you serve.
• Limited immigration sponsorship may be available.
Preferred:
- AWS Certified Solutions Architect - Associate; AWS Certified Solutions Architect - Professional preferred
• Experience integrating artificial intelligence agents, copilots, or large language model-enabled automation into incident management or IT service management workflows
• Experience with Python, shell scripting, YAML, and JSON for automation and tooling integrations
• Experience evaluating observability, event management, synthetic monitoring, or real user monitoring platforms across enterprise environments
• Experience supporting managed services or consulting delivery models across multiple client accounts
• Experience with VMware, Nutanix, Red Hat Enterprise Linux, IBM Power Systems, or AIX environments
For individuals assigned and/or hired to work in Colorado, Deloitte is required by law to include a reasonable estimate of the compensation range for this role. This compensation range is specific to Colorado and takes into account the wide range of factors that are considered in making compensation decisions including but not limited to skill sets; experience and training; licensure and certifications; and other business and organizational needs. At Deloitte, it is not typical for an individual to be hired at or near the top of the range for their role and compensation decisions are dependent on the facts and circumstances of each case. A reasonable estimate of the current range is $144,000 to $241,000.
You may also be eligible to participate in a discretionary annual incentive program, subject to the rules governing the program, whereby an award, if any, depends on various factors, including, without limitation, individual and organizational performance.