Find Jobs
Find Jobs Near You – Available Work in Your Location
Member of Technical Staff, Inference Systems
Job Description
About the Company Our client is a seed-stage, stealth-mode startup building a high-performance AI inference platform from the ground up. The founding team brings deep AI-infrastructure experience and is tackling the hardest problems in LLM serving — scheduling, KV cache management, request routing, and the runtime systems that power model inference at scale — with Rust at the core of the stack. This is a ground-floor opportunity to shape the architecture of a fast, reliable inference system without legacy constraints. Recently founded
• Small founding team
•
Industry:
AI infrastructure / LLM inference The Role You'll join a small, fast-moving team building a new inference system from scratch. This is a role for a systems engineer who lives and breathes inference internals — attention, KV cache, batching, scheduling — and wants to own the whole stack rather than a narrow slice. What you'll be doing Build a new inference runtime from scratch in Rust, owning batching, scheduling, request routing, and the full serving stack. Design and implement KV cache management, prefix caching, and optimizations that cut latency and cost per token. Scale serving across GPUs and nodes, tackling multi-GPU and multi-node challenges directly. Profile, benchmark, and ship performance improvements across the entire inference pipeline. Work closely with the founding team on the core architectural decisions that define the platform.
Tech stack:
Rust, Python, PyTorch, C++, Go, vLLM, SGLang, TensorRT-LLM, CUDA, Triton, NCCL Requirements 2-10 years of experience as a backend or distributed systems engineer. Hands-on experience building, operating, or optimizing LLM inference or serving systems at the engine, router, or runtime layer — beyond just calling hosted APIs. A deep working knowledge of transformer inference internals: attention, KV cache, batching, scheduling, and where the real bottlenecks are. Performance-critical backend or distributed systems experience where latency, throughput, and cost were first-order concerns. Hands-on time with a production inference engine (vLLM, SGLang, or TensorRT-LLM) plus strong systems-language skills (Rust, C++, Go, or systems-level Python/PyTorch). If you haven't used Rust yet, you should be able to become productive in it within a few weeks of joining. Nice to Haves Time on an inference team at a model provider, accelerator vendor, or research lab, or open-source contributions to vLLM, SGLang, or Dynamo. A CS or systems degree from a strong program. Production Rust experience, CUDA/Triton kernel work, multi-GPU or multi-node serving (NCCL, NVLink, RDMA), prefix caching, speculative decoding, or prefill/decode disaggregation. Why Join Own the architecture of a brand-new inference system with no legacy constraints. Solve the hardest problems in LLM serving, from KV cache management to request routing. A performance engineer's dream: latency and cost per token are the whole game. High visibility on a small, elite team where your code powers the core engine.
Details Location:
Palo Alto, CA Work policy: Full-time, on-site five days a week
Compensation:
$230K-$350K + equity Visa sponsorship: Open to visa transfers (OPT, H-1B) and new sponsorships (new
H-1B, TN
) Employment type: Full-time