Why join Upscale AI
Upscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.
We focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.
If you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.
Location: India · Experience: 12–15 years · Type: Full-time, IC
About us
We build a distributed control plane for datacenter network fabrics. Our platform orchestrates intent-driven configuration, change management, drift remediation, and full-stack observability across cloud-hosted services and on-prem edge appliances deployed in customer datacenters worldwide.
The role
You will own critical platform subsystems end-to-end — from design through production. You'll build and ship core distributed systems components, drive the technical quality of the codebase, and unblock the team through direct hands-on contribution. You work closely with Product, QA, and Customer Engineering to deliver quality product to stakeholders.
What you'll work on
- Distributed control plane spanning cloud services and on-prem edge appliances connected via mTLS gRPC streams
- Observability and telemetry at scale: OpenTelemetry collection pipelines, stream processing (Kafka), time-series storage (ClickHouse/VictoriaMetrics), real-time fabric state views, packet-event analysis, and fleet-wide health aggregation
- Agentic AI operations: design and build autonomous infrastructure agents that evaluate prerequisites, orchestrate multi-step workflows (onboarding, upgrades, drift remediation), handle failure recovery, and interact with the control plane through tool-use patterns (LangGraph, MCP)
- Fleet orchestration: aggregate health/compliance/drift APIs, cross-site campaign execution, template promotion workflows, parallel site onboarding
- Data architecture across Postgres, ArangoDB, ClickHouse, Redis, Kafka, and Git-backed content stores
- Multi-tenant SaaS with site-scoped RBAC, session lifecycle, and enterprise IdP integration
- Device lifecycle management: zero-touch provisioning, enrollment protocols, config push via edge relay, drift detection and remediation
- Intent compilation engine: workspace management, merge request lifecycle, change request execution with gate-based verification and rollback
What you bring
- 12–15 years building and operating distributed systems serving enterprise customers across cloud and on-prem environments
- Deep proficiency in Go (or comparable systems language) with strong distributed systems fundamentals
- Significant experience designing observability and telemetry platforms: collection agents, stream processing, time-series databases, alerting pipelines, and real-time dashboards at scale
- Production experience with microservices architecture: gRPC, Protocol Buffers, spec-first REST APIs (OpenAPI)
- Hands-on with multiple storage paradigms: relational, graph, time-series, key-value, and streaming
- Track record building multi-tenant platforms with tenant isolation, RBAC, and identity federation
- Experience with Kubernetes, Helm, and hybrid cloud/on-prem deployment models
- Strong API design sense: versioning, backward compatibility, contract-first development
- Ability to communicate architectural decisions clearly through writing and diagrams
How we expect you to work
- Cross-functional by default — you work closely with Product, QA, Design, and Customer Engineering, not in isolation
- Solution-oriented — when the team is stuck, you unblock them by driving toward answers and building what's needed
- Accountable for delivery — you take personal responsibility for shipping quality product to stakeholders on time
- Hands-on always — you write code daily and prove ideas by building them
Nice to have
- AI/ML agent architectures for infrastructure operations: LangGraph, AutoGen, MCP tool-use, human-in-the-loop gating, and autonomous workflow orchestration
- Network automation or infrastructure management platforms (Apstra, NSO, Terraform, Crossplane)
- Datacenter networking experience: OpenConfig, gNMI, or fabric management at scale
- Hub-and-spoke / edge computing / control-plane-data-plane separation architectures
- OpenTelemetry contributor experience or deep familiarity with the collector ecosystem
- Device enrollment or zero-touch provisioning systems
- Open-source contributions or published work in distributed systems
Where you fall within that range depends on your experience, skills, and impact—we benchmark against internal levels to keep things fair and consistent.
Equal Opportunity
Upscale AI is building a team that reflects a wide range of perspectives, backgrounds, and experiences. We’re proud to be an Equal Opportunity Employer and consider all qualified applicants regardless of race, color, religion, national origin, sex, sexual orientation, gender identity, disability, or veteran status.
Accessibility & Accommodations
We’re committed to making our hiring process accessible to everyone. If you need accommodations at any stage, just reach out to us at hiring@upscaleai.com—we’re happy to help. Note: This inbox is only for accommodation requests.