EA's Fan Growth organization elevates the experiences customers get from our entertainment and maximizing their fun. We're obsessed with the experiences our fans have with EA. We value people who bring new ideas to make Fan Growth a great place to be and be from. We celebrate inclusion. It's just as important for us to create great experiences for our people as for our players. We're looking for people who will show up excited to work and ready for fun.
The Site Reliability Engineer will report to the Director of Engineering
Responsibilities
Infrastructure Automation & IaC: You will architect, automate, and maintain cloud infrastructure on AWS using Infrastructure as Code
Kubernetes Platform Development: You will containerize applications using Docker and extend Kubernetes cluster capability by writing custom tooling and controllers to support production workloads.
CI/CD Pipeline Automation: You will develop automated CI/CD pipelines with GitLab CI to ensure reliable, high-frequency deployment of backend microservices.
Observability & Live Service Incident Management: Implement monitoring, logging, and alerting systems using Grafana, Prometheus, ELK, and PagerDuty; support live production services, and drive incident response and root-cause analysis.
Qualifications
Minimum 5 years of professional experience in web or backend software development, alongside experience running and supporting live production services.
Proficiency in at least one backend programming language Golang, Python, or Ruby.
Hands-on proficiency with AWS, Docker, Kubernetes (including writing custom tooling/controllers), and Infrastructure as Code (Terraform, CloudFormation, or Helm).
3+ years of experience with pipeline automation (GitLab CI), version control (Git), and observability suites (Grafana, Prometheus, ELK, and PagerDuty).
Experience with a data warehouse like Snowflake, Redshift, or Spark
Experience operating and optimizing CDN/edge delivery infrastructure (e.g., Akamai, CloudFront), including caching strategies, configuration management, performance optimization, and troubleshooting production traffic at the edge.