Job opening
Principal AI Architect, Data-Center Platform
Filed under Semiconductor Manufacturing
Full job description
Principal AI Architect, Data-Center Platform
OXMIQ Labs · Senior technical leadership, hands-on
About the role
We are building the foundational software platform that turns racks of accelerated compute into a production AI serving and agent platform — from bare metal and GPU runtimes up through orchestration, inference serving, and agent frameworks, twelve teams in all.
As Principal Architect, you are the single technical authority across planned twelve teams — a leadership role held as a builder, not a manager. Every team will have its own architect/tech lead who owns their layer; you own the thing none of them can — end-to-end design coherence, anchored in a deep understanding of the GPU itself. You know what the silicon can actually do, and you make sure every layer above it is designed around that reality: how a request travels from a user's keystroke down to a GPU kernel and back with no wasted cycles along the way.
This is a rare, force-multiplier role. It is currently held by the VP of AI & Infrastructure and will be handed to you — you'll join early enough to shape the platform's foundations and team, with direct partnership from the VP from day one.
What you'll own
- End-to-end architecture of the platform: cross-layer design decisions, interface contracts between the 12 teams, and the technical narrative that keeps 8 platform layers and 4 cross-cutting functions coherent as one product.
- The GPU efficiency budget for the whole system: the full token path — kernel efficiency, memory hierarchy and KV-cache behavior, quantization strategy, batching and scheduling, serving runtime, and the network/storage I/O that feeds it — treated as one latency/throughput/cost envelope.
- GPU software strategy: driver and runtime choices, kernel-level optimization priorities, serving-stack decisions (vLLM, SGLang, diffusion runtimes), and how new silicon capabilities get exposed up the stack.
- Build-vs-adopt strategy: where we differentiate with proprietary layers versus ride open source, and how we stay upgrade-safe on the OSS we adopt.
- Architecture review and technical escalation across all 12 architect/tech leads; you are the tiebreaker on cross-team design disputes.
- Hands-on contribution: this is not a review-only position. You prototype the risky seams yourself, write the design docs that matter, and land code where the hardest problems live — usually at the GPU boundary.
What we're looking for
- 12+ years building systems software, with deep GPU-level expertise as the anchor: GPU architecture (SMs/CUs, memory hierarchy, interconnects), kernel development and optimization (CUDA, ROCm, Triton, or equivalent), GPU drivers and compute runtimes, and performance analysis at the hardware boundary.
- Proven depth in LLM inference performance: serving runtimes (vLLM, SGLang, TensorRT-LLM or similar), quantization (AWQ/GPTQ/FP8), KV-cache management, speculative decoding, continuous batching — you've made these fast on real hardware, not just read the papers.
- Working fluency across the rest of the modern data-center stack — orchestration/Kubernetes, virtualization, networking, storage, bare-metal provisioning — enough to design sound interfaces to teams that own those layers, without needing to be the expert in each.
- A track record as the architect of record for a large multi-team platform — you've owned coherence across an org, not just a service.
- Strong open-source judgment: you've made consequential adopt/fork/build calls and lived with the results.
- The credibility and communication skills to be technical authority to twelve senior tech leads without line authority over any of them — influence through design quality, not org position.
- Hands-on to the core. You still ship.
Nice to have
- Experience at a GPU/accelerator company, or bringing up software platforms for novel silicon.
- Contributions to relevant OSS (vLLM, Triton, PyTorch, Kubernetes, KVM/QEMU, etc.).
- Experience with multi-GPU and multi-node inference (tensor/pipeline parallelism, NCCL or equivalent collectives).
- Prior work on agent frameworks, model endpoints, or AI developer experience.
Why this role, why now
You'd be a founding member of the platform org, with a flat structure, hands-on leads, and a mandate to over-invest in the differentiated layers. The architecture is not yet ossified — the person in this seat will set the shape of the entire platform.
OXMIQ Labs is an equal opportunity employer. Location open; compensation commensurate with the scope of the role.