Why join Upscale AI
Upscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.
We focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.
If you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.
The role
You will set the multi-year technical vision for the platform — and prove it works by building it. This is not a strategy-only role. You perform competency analysis across the architecture, build working prototypes to validate direction, write production code, and lead cross-functional delivery. You are accountable to stakeholders for the quality and viability of what ships. You work hand-in-hand with Product to ensure the technical vision maps to customer value.
Responsibilities
What you'll work on
- Distributed control plane spanning cloud services and on-prem edge appliances connected via mTLS gRPC streams
- Observability and telemetry at scale: OpenTelemetry collection pipelines, stream processing (Kafka), time-series storage (ClickHouse/VictoriaMetrics), real-time fabric state views, packet-event analysis, and fleet-wide health aggregation
- Agentic AI operations: design and build autonomous infrastructure agents that evaluate prerequisites, orchestrate multi-step workflows (onboarding, upgrades, drift remediation), handle failure recovery, and interact with the control plane through tool-use patterns (LangGraph, MCP)
- Fleet orchestration: aggregate health/compliance/drift APIs, cross-site campaign execution, template promotion workflows, parallel site onboarding
- Data architecture across Postgres, ArangoDB, ClickHouse, Redis, Kafka, and Git-backed content stores
- Multi-tenant SaaS with site-scoped RBAC, session lifecycle, and enterprise IdP integration
- Device lifecycle management: zero-touch provisioning, enrollment protocols, config push via edge relay, drift detection and remediation
- Intent compilation engine: workspace management, merge request lifecycle, change request execution with gate-based verification and rollback
Requirements
- 18–20+ years building and operating distributed systems serving enterprise customers across cloud and on-prem environments
- Deep proficiency in Go (or comparable systems language) with strong distributed systems fundamentals
- Significant experience designing observability and telemetry platforms: collection agents, stream processing, time-series databases, alerting pipelines, and real-time dashboards at scale
- Production experience with microservices architecture: gRPC, Protocol Buffers, spec-first REST APIs (OpenAPI)
- Hands-on with multiple storage paradigms: relational, graph, time-series, key-value, and streaming
- Track record building multi-tenant platforms with tenant isolation, RBAC, and identity federation
- Experience with Kubernetes, Helm, and hybrid cloud/on-prem deployment models
- Strong API design sense: versioning, backward compatibility, contract-first development
- Proven ability to evaluate technical landscape, identify gaps through competency analysis, and validate solutions with working prototypes before committing the organization
- Track record of making high-stakes build-vs-integrate decisions that proved right under production pressure
- Ability to communicate architectural decisions clearly through writing and diagrams to engineering, product, and executive audiences
- Cross-functional by default — you work hand-in-hand with Product to shape strategy, with Customer Engineering to understand deployment reality, and across every engineering team to drive alignment
- Solution-oriented — when the path is unclear, you build a POC and prove it. Vision without working code is just an opinion.
- Accountable for delivery — you own the quality and viability of what ships to stakeholders. The buck stops with you on the hardest technical problems.
- Hands-on always — you write code, build prototypes, review critical PRs, and stay deep enough in the system to make sound architectural calls
Nice to have
- AI/ML agent architectures for infrastructure operations: LangGraph, AutoGen, MCP tool-use, human-in-the-loop gating, and autonomous workflow orchestration
- Network automation or infrastructure management platforms
- Datacenter networking experience: OpenConfig, gNMI, or fabric management at scale
- Hub-and-spoke / edge computing / control-plane-data-plane separation architectures
- OpenTelemetry contributor experience or deep familiarity with the collector ecosystem
- Device enrollment or zero-touch provisioning systems
- Open-source contributions or published work in distributed systems
Where you fall within that range depends on your experience, skills, and impact—we benchmark against internal levels to keep things fair and consistent.
Equal Opportunity
Upscale AI is building a team that reflects a wide range of perspectives, backgrounds, and experiences. We’re proud to be an Equal Opportunity Employer and consider all qualified applicants regardless of race, color, religion, national origin, sex, sexual orientation, gender identity, disability, or veteran status.
Accessibility & Accommodations
We’re committed to making our hiring process accessible to everyone. If you need accommodations at any stage, just reach out to us at hiring@upscaleai.com—we’re happy to help. Note: This inbox is only for accommodation requests.