We are looking for a Software Systems Engineer III to help build, operate and evolve the infrastructure platforms that power AI services, automation systems and AI-enabled operational capabilities.
This is primarily an infrastructure and platform engineering role.
You will work across Kubernetes, cloud infrastructure, GitOps, CI/CD, observability, infrastructure automation and production reliability. You will help ensure that the platforms used by our engineering and operations teams are scalable, reliable, secure and easy to operate.
The ideal candidate is comfortable working hands-on in production environments, troubleshooting complex infrastructure problems and taking ownership of technical problems from investigation through implementation and operationalization.
This role is a strong fit for an engineer who enjoys the intersection of cloud infrastructure, platform engineering, DevOps/SRE practices and automation.
What You Will Do
Build, operate and improve Kubernetes-based infrastructure supporting production services and internal platforms.
Manage and evolve cloud infrastructure across development, staging and production environments.
Develop and maintain infrastructure-as-code using tools such as Terraform.
Build and improve CI/CD and GitOps workflows using technologies such as Argo CD, GitHub Actions or Azure DevOps.
Improve platform reliability through monitoring, alerting, observability and proactive operational improvements.
Troubleshoot infrastructure, networking, deployment, configuration and application-platform integration issues.
Support containerized workloads and platform services running across Kubernetes environments.
Build automation that reduces manual operational work and improves engineering productivity.
Partner with software engineers, SREs and operations teams to design reliable deployment and runtime architectures.
Participate in incident response, root-cause analysis and implementation of preventative improvements.
Improve infrastructure security, access controls, secrets management and deployment safety.
Document platform architecture, operational procedures and troubleshooting guidance.
Evaluate new infrastructure and platform technologies and introduce them where they provide meaningful operational value.
Support infrastructure used by AI, ML and automation workloads where required.
What We Are Looking For
Strong hands-on experience with Kubernetes and containerized workloads.
Practical experience operating infrastructure in Azure, AWS or another major cloud platform.
Experience with Terraform or equivalent infrastructure-as-code technologies.
Experience designing or maintaining CI/CD and GitOps pipelines.
Working knowledge of tools such as Argo CD, Helm and GitHub Actions/Azure DevOps.
Experience with monitoring and observability platforms such as Grafana, Prometheus, Datadog or equivalent technologies.
Strong troubleshooting skills across distributed systems, infrastructure and application environments.
Ability to read and write automation or tooling code using Python, Go, PowerShell, Bash or similar languages.
Understanding of networking, DNS, certificates, identity, secrets management and cloud security fundamentals.
Experience supporting production systems and participating in incident resolution.
Ability to independently investigate technical problems, evaluate alternatives and recommend an implementation approach.
Strong written and verbal communication skills when working across engineering and operational teams.
Preferred Experience
Experience in one or more of the following areas would be valuable:
Kubernetes platform administration and cluster lifecycle management.
Argo CD and GitOps-based deployment models.
Azure Kubernetes Service or similar managed Kubernetes environments.
Kubeflow, ML platforms or infrastructure supporting AI/ML workloads.
Platform engineering or internal developer platforms.
SRE practices and production reliability engineering.
Distributed tracing, telemetry and advanced observability.
Infrastructure security and DevSecOps.
Building automation for operational workflows.
Supporting globally distributed engineering teams.