Upscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.
We focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.
If you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.
Role Overview
We are looking for a Technical Support Lead Engineer with 12+ years of experience who is passionate about building reliable networking switches and solving challenging infrastructure problems.
Own end-to-end network bring-up and performance for GPU AI clusters during customer PoCs. You’ll be the leader for switching ASIC pipeline tuning, platform readiness, and fabric performance (Ethernet/RoCEv2 and/or InfiniBand), turning requirements into reproducible configurations that hit utilization and tail-latency targets at scale.
As part of the AI Networking team, you'll debug software, develop automation, validate networking platforms, troubleshoot complex system issues, and work closely with engineering teams and customers to ensure the best customer experience. You'll have the opportunity to work on technologies powering next-generation GPU clusters and AI data centers while learning from experienced engineers in a fast-paced startup environment.
Design, develop, and maintain features and enhancements for the SONiC NOS platform.
· Debug, troubleshoot, and resolve issues on SONiC platforms.
· Develop and execute debugging and troubleshoot infrastructure.
· Collaborate closely with cross-functional teams including hardware engineers and Test teams.
· Participate in code reviews, architecture discussions, and documentation efforts.
· Develop support strategies to root-cause Networking ASICs and Networking Systems issues.
· Develop debugging tools, for Traffic monitoring and Performance measurements.
· Debug issues across software, Linux systems, networking stacks, and distributed infrastructure.
· Be a point of contact for customer deployments, integration testing, proof-of-concepts, and field issue resolution.
· Build tools that improve deployment efficiency, observability, telemetry, and automated testing.
· Collaborate with software, infrastructure, QA, and product teams to identify root causes and deliver robust solutions.
· Contribute to backend services, APIs, orchestration components, and infrastructure automation.
· Document technical findings and communicate effectively with engineering teams and customers.
Required Qualifications
· Bachelor’s or master’s degree in computer science, Electrical Engineering, or a related field.
· Minimum of 12 years of work experience is required, with at least 3 years of hands-on SONiC or equivalent Network Operating System (NOS) development experience preferred.
· Strong programming skills in Python, Go, or a similar language.
· Solid understanding of Linux, TCP/IP networking, routing, switching, VLANs, and network troubleshooting.
· Experience with PTF (Packet Test Framework) and SPyTest for network validation.
· Familiarity with Linux internals, docker containers.
· Excellent problem-solving skills and ability to work in a fast-paced, collaborative environment.
· Knowledge of network ASICs and switch hardware architecture is mandatory.
· Excellent written and verbal communication skills.
· Ability to thrive in a collaborative, fast-paced startup environment with a strong sense of ownership.
Preferred Qualifications
· Data-center networking with hands-on switch ASIC tuning and platform bring-up.
· Proven RoCEv2 deployments at 200/400/800 G: ECN/PFC design, DCQCN tuning, DSCP/PCP mapping, queue/WRR shaping.
· Deep buffer/queueing knowledge (headroom, xon/xoff, dynamic thresholds, VOQ vs shared pools).
· SONiC (buffers.json, qos.json, pfcwd, ecn), Enterprise-OS (class-map/type qos & network-qos, policy-map, queuing).
· Optics & PHY: PAM4 signal integrity, FEC modes, DOM/RS-FEC counters, AN/LT quirks, DAC/AOC selection.
· Tooling: ethtool, devlink, mlnx_qos, perfquery, switch telemetry (INT/sFlow/ERSPAN), gNMI/REST, Prometheus.
· Benchmarking: perftest, nccl-tests, iPerf3, traffic generators; reading queue stats, ECN marks, and CNP behavior.
· EVPN/VXLAN leaf-spine for AI pods; flowlet or latency-aware hashing.
· BlueField DPU offloads, GPUDirect RDMA; NIC QoS (DSCP-to-TC, PFCx, GEARBOX/FW).
· Storage fabrics for AI (NFS-RDMA, NVMe-oF) and their QoS interactions.
· Python/Ansible for templated ASIC profiles; Git workflows for config promotion.
Where you fall within that range depends on your experience, skills, and impact—we benchmark against internal levels to keep things fair and consistent.
Equal Opportunity
Upscale AI is building a team that reflects a wide range of perspectives, backgrounds, and experiences. We’re proud to be an Equal Opportunity Employer and consider all qualified applicants regardless of race, color, religion, national origin, sex, sexual orientation, gender identity, disability, or veteran status.
Accessibility & Accommodations
We’re committed to making our hiring process accessible to everyone. If you need accommodations at any stage, just reach out to us at hiring@upscaleai.com—we’re happy to help. Note: This inbox is only for accommodation requests.