Job Description
You belong to the top echelon of talent in your field. At JPMorganChase, infrastructure isn't just a foundation — it's a competitive advantage. This is your opportunity to bring deep storage expertise to a team that operates at global scale, where your contributions directly impact the stability and performance of critical financial services.
As an Infrastructure Engineer III at JPMorganChase within Enterprise Technology (Infrastructure Platforms), you apply strong knowledge of software, applications, and technical processes within the infrastructure engineering discipline. In this role, you will hold end-to-end accountability for the administration, stability, and resilience of the storage technology estate — spanning reactive incident management through to proactive automation and toil reduction. You will operate on a structured shift rotation, including weekend day coverage, to ensure continuity of service and operational excellence across the infrastructure landscape.
Job responsibilities
- Apply technical knowledge and problem-solving methodologies to storage infrastructure projects of moderate scope, ensuring end-to-end monitoring, performance, and resilience of storage services running at scale
- Use enterprise-authorized AI capabilities to accelerate infrastructure analysis, monitoring, and capacity documentation, validating outputs and handling operational data according to sensitivity and security requirements
- Apply reuse-first, AI-assisted practices within delivery and automation routines to identify recurring issues, improve remediation workflows, and ensure changes are traceable, auditable, and aligned to resiliency and security expectations
- Operate and enhance block, file, and object storage platforms across on-premises and cloud environments, including performance tuning, capacity planning, lifecycle management, and resiliency testing such as failover and disaster recovery validation
- Lead incident response for storage outages and performance degradations, drive root cause analyses, and implement preventative actions to reduce recurrence
- Build and maintain automation for provisioning, patching, upgrades, replication, backup and restore, and compliance checks to reduce toil and improve operational consistency
- Create and maintain runbooks, escalation paths, and standardized operational procedures to support on-call readiness and team knowledge continuity
- Partner with infrastructure, network, operating system, database, and application teams to meet workload requirements and reliability targets
- Implement AI-driven observability and AIOps capabilities — including telemetry correlation, anomaly and regression detection, and large language model-assisted incident and runbook workflows — with a focus on accuracy, auditability, and safe rollout
- Own and continuously improve service level objectives, service level indicators, error budgets, and on-call readiness for storage services
Required qualifications, capabilities, and skills
- Formal training or certification on infrastructure engineering concepts and 3+ years applied experience
- Demonstrated experience using enterprise-authorized AI capabilities within the work environment to support infrastructure engineering workflows, with strong validation habits and awareness of data sensitivity
- Ability to review and validate AI-assisted recommendations before implementation, escalating when uncertain and ensuring outcomes align to resiliency, security, and auditability expectations
- Strong knowledge of storage fundamentals including RAID and erasure coding, replication, snapshots, tiering and caching, IOPS and latency, multipathing, SAN and NAS, and object storage semantics
- Hands-on experience with at least one major storage ecosystem such as NetApp, Dell EMC PowerStore or Isilon, Pure Storage, Hitachi, Ceph, IBM, or cloud-native storage services
- Strong scripting or programming proficiency in one or more of Python, Go, or Bash
- Solid Linux fundamentals including system performance, networking basics, and kernel and storage-stack concepts
- Experience with observability stacks such as Prometheus and Grafana, Elastic or OpenSearch, Splunk, Datadog, or OpenTelemetry
- Proven incident management skills and ability to operate effectively within an on-call rotation
- Practical skills in AI and data operations including anomaly detection, forecasting, correlation, classification, feature extraction, and integrating AI into production tooling and continuous integration and delivery pipelines with safe large language model use, guardrails, and human-in-the-loop review
Preferred qualifications, capabilities, and skills
- Experience with Kubernetes storage using the Container Storage Interface, stateful workloads, and container platform operations
- Proficiency with Infrastructure as Code tools such as Terraform or CloudFormation, and configuration management tools such as Ansible, Chef, or Puppet
- Familiarity with streaming and queue tooling for telemetry and event pipelines such as Kafka
- Experience with IT service management and event management platforms such as ServiceNow
- Knowledge of backup and disaster recovery products and strategy design, including recovery point objective and recovery time objective tradeoffs
- Experience with security controls for data platforms including key management services, hardware security modules, secrets management, and key rotation