Why join Upscale AI
Upscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.
We focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.
If you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.
About the role
AI training and inference clusters live or die on the network fabric. ScaleUp builds the software that programs our high-performance Ethernet switch silicon — SDK, SAI, simulation, and the CI that keeps that stack shippable.
We’re looking for a DevOps engineer at the intersection of AI infrastructure, data-center networking, and silicon-aware software: reliable pipelines for an ASIC SDK, fast feedback for developers, and release paths worthy of cloud- and AI-scale deployment.
You’ll partner with SDK, SAI, and QA engineers, and work in GitHub + Jira day to day. Success means CI is trusted, simulation and software gates are clear, and infrastructure never blocks the next generation of AI networking features.
Responsibility
Own CI/CD for the ScaleUp stack that powers AI/HPC Ethernet switching (build, unit/integration, gating, artifacts, release promotion).
Scale pipelines across multi-repo dependencies: switch SDK, SAI adapters, shared test frameworks, and switch simulation / model targets.
Improve build caching, shared runners, and nightlies so large C/C++ and Python SDK builds stay fast and predictable.
Own release / promote workflows (branch policies, artifact publish, pre-gate vs post-gate steps) so SDK drops are repeatable.
Shorten commit → green for teams building programmable data-plane, QoS, ACL, RoCE-era fabrics, and AI-workload traffic patterns.
Add and maintain quality gates: lint, unit tests, model smoke, coverage where useful, overnight soak — high signal, low flake.
Manage artifacts and caches (CI artifacts, object storage as needed) with clear retention and failure handling.
Operate Linux build/test farms for C/C++ SDKs and Python harnesses (toolchains, containers, lab/sim hosts).
Improve pipeline observability (dashboards, alerts, runbooks, postmortems).
Secure secrets and access for GitHub Actions, artifact stores, and internal services.
Partner with QA on PR / nightly / soak jobs that validate end-to-end switch behavior before customer and cloud deployments.
Publish reusable pipeline templates so new ScaleUp / AI-networking repos onboard quickly.
Jira: keep engineering work traceable — epics/stories/bugs for CI and infra, link PRs and releases to tickets, support sprint and release planning with SDK/SAI/QA, tighten ticket hygiene (status, components, labels) so blockers and CI debt are visible.
Qualification
Strong CI/CD experience (GitHub Actions and/or Jenkins; GitLab CI also fine) in multi-repo environments.
Deep Linux fluency: shells, packaging, toolchains, debugging native builds and shared-library / SDK load paths.
Automation in Python and bash; enough CMake/Make to unblock switch-SDK CI.
Containers (Docker) and self-hosted or cloud runners at meaningful scale.
Proven work reducing flaky tests and improving CI signal-to-noise.
Hands-on Jira (or equivalent): workflows, boards, linking commits/PRs to issues, release/version fields — comfortable driving process with developers, not only “keeping the lights on.”
Clear communication with hardware-adjacent software teams.
Excitement about AI data centers, Ethernet fabrics, and silicon-software co-design — not only generic cloud DevOps.
Nice to have
Background in networking ASICs, switch SDKs, SAI, DPDK, or NIC/SmartNIC software.
Exposure to AI/ML cluster networking (GPU fabrics, RoCE/RDMA, congestion control, telemetry).
Simulation / hardware-in-the-loop CI, pytest at scale, junit/Allure-style reporting.
Build-cache and large-monorepo / multi-repo release patterns; backport or branch-gating workflows.
Static analysis / lint gates in CI (e.g. cppcheck, language linters).
IaC (Terraform/Ansible) and cloud runners (AWS/GCP).
Atlassian suite beyond Jira (Confluence runbooks, Jira + GitHub automation).
Release hygiene: versioning, artifacts, SBOM/signing, branch protection / merge gates.
Why ScaleUp
Your pipelines sit under real AI networking product software — not a side CRUD service.
You influence how fast we ship switch SDK + SAI features that AI clusters depend on.
Small team, high ownership: changes land and developers feel them the next day.
Where you fall within that range depends on your experience, skills, and impact—we benchmark against internal levels to keep things fair and consistent.
Equal Opportunity
Upscale AI is building a team that reflects a wide range of perspectives, backgrounds, and experiences. We’re proud to be an Equal Opportunity Employer and consider all qualified applicants regardless of race, color, religion, national origin, sex, sexual orientation, gender identity, disability, or veteran status.
Accessibility & Accommodations
We’re committed to making our hiring process accessible to everyone. If you need accommodations at any stage, just reach out to us at hiring@upscaleai.com—we’re happy to help. Note: This inbox is only for accommodation requests.