Be an integral part of an agile team that's constantly pushing the envelope to enhance, build, and deliver top-notch technology products.
As a Lead Software Engineer at JPMorgan Chase within Corporate - AIML Data Platforms team , you will design, build, and operate the foundational cloud infrastructure that enables data scientists and machine learning engineers to develop, train, and deploy intelligent solutions across the firm. In this role you will serve as a technical leader, driving platform reliability, scalability, and automation while collaborating with cross-functional teams to solve complex infrastructure challenges. Your work will directly accelerate the firm’s AI/ML capabilities, enabling faster experimentation and production-grade deployments that create measurable business impact.
Job Responsibilities
- Builds and maintains reusable AI/ML platform infrastructure and shared services to support development, deployment, and operations at scale.
- Architects, deploys, and operates secure cloud and container-based environments for training and inference, including GPU-intensive workloads.
- Design and implement platform tooling, automation, and infrastructure-as-code solutions to streamline model deployment, environment provisioning, release management, and operational support.
- Develops and maintains production-grade services, APIs, SDK integrations, and workflows that support model training, serving, evaluation pipelines, and AI application lifecycle management.
- Partners with data science, ML engineering, and application teams to translate model and compute requirements into platform standards and deployment patterns.
- Optimizes platform reliability, scalability, latency, and cost through orchestration, scheduling, and hardware acceleration.
- Establishes operational best practices including monitoring, logging, observability, access controls, incident response, and production troubleshooting.
- Supports enterprise LLM operationalization, including fine-tuning workflows, inference Drives team adoption of enterprise-authorized AI-assisted engineering practices within the work environment to improve code quality, delivery speed, and operational outcomes (e.g., AI-assisted code review/refactoring, test strategy acceleration, incident/root-cause analysis support), while establishing consistent validation standards (secure coding, peer review, automated testing) and promoting reuse of effective patterns across the team.
- Applies knowledge of tools within the Software Development Life Cycle toolchain, including enterprise-authorized AI-assisted development and automation capabilities, to improve the value realized by automation.
Required Qualifications, Capabilities, and Skills
- Formal training or certification on software engineering concepts and 5+ years applied experience
- Experience delivering secure, production-quality code in Python or Java.
- Strong foundations in distributed systems, microservices, and platform architecture/design principles.
- Proven ability to architect and operate cloud-native infrastructure on AWS (compute, networking, storage, security) and other major clouds.
- Demonstrated expertise with infrastructure-as-code tooling, specifically Terraform, in large-scale cloud environments.
- Hands-on experience with Docker and Kubernetes, including AWS EKS operations.
- Experience building or supporting production AI/ML platforms (training, deployment, and model serving/inference), including GPU infrastructure/tooling.
- Strong DevOps/platform engineering practices: CI/CD, release automation, automated testing, and observability (monitoring/logging/tracing).
- Experience with SQL/NoSQL databases and data integration; strong Linux, scripting, and networking fundamentals.
- Demonstrated experience leading effective use of approved AI-assisted software development tools (e.g., for coding, code review, test acceleration, troubleshooting) with the ability to set team expectations for validating AI outputs for correctness, performance, and security.
- Strong understanding of responsible AI use in engineering workflows, including data sensitivity considerations, secure handling of inputs/outputs, and adherence to resiliency and security expectations; experience coaching engineers on safe, compliant adoption within delivery practices
Preferred Qualifications, Capabilities, and Skills
- Proficiency in Go or Python for automation, tooling development, or platform service implementation.
- Experience with MLOps frameworks and tools such as Kubeflow, MLflow, or similar AI/ML lifecycle management platforms.
- Working knowledge of ML frameworks (PyTorch, TensorFlow, Hugging Face, scikit-learn) for model integration and operationalization.
- Exposure to multi-cloud or hybrid cloud architectures and platform portability strategies.