Key Responsibility
• Operate and support production ETL pipelines, including extractors, loaders, batch jobs, streaming ingestion, and data consolidation workflows.
• Support loader jobs running in an EKS Airflow cluster that load data into Snowflake.
• Support Spark jobs running on EKS that load data into the Datalake or Lakehouse.
• Own platform and pipeline reliability, including uptime, throughput, job success rate, SLA adherence, and data freshness.
• Triage and resolve production incidents such as failed jobs, delayed loads, OOMKilled pods, CrashLoopBackOff pods, stalled pipelines, driver/executor failures, credential issues, and network/connectivity problems.
• Perform root cause analysis for incidents and drive corrective actions through automation, configuration changes, capacity tuning, or code/platform improvements.
• Tune Kubernetes workloads, including CPU/memory requests and limits, autoscaling, pod scheduling, service accounts, secrets, config maps, and workload health.
• Tune Spark workloads, including executor count, cores, memory, memory overhead, shuffle partitions, dynamic allocation, adaptive query execution, spill reduction, and failed stage analysis.
• Monitor and troubleshoot Airflow DAGs, task retries, scheduling issues, dependency failures, executor/resource constraints, and SLA misses.
• Diagnose data latency across the full pipeline path, such as source event, Kafka or messaging layer, staging/object storage, ETL processing, Snowflake, and Datalake availability.
• Support Kafka or equivalent streaming systems, including consumer lag, offsets, partitions, topic health, producer delay, and SASL/TLS client configuration.
• Manage platform and job configuration through approved source-of-truth or GitOps-style processes instead of ad hoc production changes.
• Manage secrets, certificates, truststores, JDBC credentials, Snowflake credentials, object-storage credentials, and key rotations following least-privilege practices.
• Build and maintain dashboards, alerts, log queries, and runbooks for job health, freshness, backlog, failures, infrastructure usage, and incident response.
• Write scripts and automation using Python, Bash, or similar tools to reduce toil and improve recovery speed.
• Coordinate with application engineering, data engineering, platform, security, network, Snowflake, and Datalake teams.
• Participate in on-call or production support rotation for China business hours and critical incidents.