Find Jobs
Find Jobs Near You – Available Work in Your Location
Skip to job details
CA
Crux AI
Senior Site Reliability & Software Engineering Manager
Career Insights for Software Development / Engineering Manager
See where this job fits in the broader career landscape. Knowing your career path helps you see what's possible from here.
Scorecard
Based on California data
Review key factors to help you decide if this role fits your goals. How is this calculated?
What they do
A Software Development or Engineering Manager leads teams of software developers who design or improve computer software. Manages and oversees the software development process and directs the work of software engineers. May be primary contact with customers or users; may manage software development project budgets and hire or train staff.
$190,032 / year median in California
Job Description
Built to set the gold standard for integrated AI infrastructure Crux AI is a newly formed, U.S.-based integrated AI infrastructure company created to remove the physical and operational constraints on consequential AI ambitions. Crux brings together power, high-density data centers, TPU silicon, networking, orchestration software, and ongoing operations as one integrated system. Crux is being capitalized to plan every layer together, develop each one to demanding standards, and operate the whole system with efficiency and reliability. That gives hyperscalers, frontier AI labs, sovereign customers, enterprises, and AI-native companies greater freedom to pursue the AI they are here to create. Crux AI is led by CEO, Ben Treynor Sloss , who spent over two decades in executive technical leadership at Google and founded the Site Reliability Engineering (SRE) disciplin e. At Crux AI, we treat operations fundamentally as a software engineering problem.
WHAT YOU'LL DO
We are recruiting founding Senior Site Reliability and Software Engineering Managers to build and lead our initial fleet reliability engineering teams in Palo Alto, CA. In this organization, there is no separate software development team. Your team owns the software, control plane, telemetry, and automated remediation controllers that keep multi-gigawatt TPU clusters provisioned, resilient, and continuously executing customer AI workloads. This is a true hands-on, builder seat, not a supervisory position. In the early days, you will write the first remediation controllers, set reliability baselines, and take initial pages yourself to stay close to the work before expanding your team. You will lead an elite group of unusually senior software and reliability engineers: engineers with significantly greater technical and software depth than traditional operational SRE orgs. You must thrive on independence and revel in ambiguity, turning unknowns into concrete engineering priorities in a fast-paced, high-growth environment. While you will devise and participate in initial on-call rotations, your core mandate is to combine SRE disciplines with extensive AI/ML automation to drive operational pages down to zero. In this role, you will:Build & Lead Senior Engineering Teams:
Recruit, lead, and mentor an initial team of senior software and reliability engineers across Palo Alto and Europe as fleet capacity ramps rapidly. Automate Pages toZero:
Own fleet availability end-to-end; establish on-call rotations while relentlessly developing self-healing systems and predictive remediation to eliminate manual pages. Embed AI/ML intoSRE Disciplines:
Apply agentic techniques, machine learning models, and automated diagnostic workflows extensively to telemetry collection, root-cause analysis, and predictive cluster recovery.Own Bare-Metal & Fleet Lifecycle Software:
Drive software engineering for bare-metal node provisioning, firmware deployment, thermal/stress burn-in validation, host/TPU health monitoring, and decommissioning.Control Plane & Fabric Reliability:
Own software reliability for cluster orchestration, scheduling, capacity allocation APIs, and high-performance TPU host/interconnect networks.Define Observability & SLOs:
Establish customer-facing SLIs/SLOs (job goodput, time-to-detect, node availability) and build the telemetry pipelines serving operators, executives, and customers.SIGNALS OF SUCCESS
After 60 days in this role: Baseline reliability framework defined (v1 SLOs, severity structure, change management); first AI/ML-driven automated remediation controller shipped to production; recruiting active for senior engineering hires in Palo Alto. After 6 months: First TPU cluster brought online under your team's automated acceptance criteria; automated remediation pipeline running with pass rates tracked; observability v1 in daily production use; core senior team onboarded After 1 year: TPU fleet operating against published customer SLOs; >90% of node/fabric faults automatically quarantined and remediated without human paging; team scaled ahead of rapid capacity ramps.EXPERIENCES, ATTRIBUTES AND MINDSET THAT INDICATE A GOOD MATCH
Experiences 10+ years of software or infrastructure engineering experience, with 3+ years managing engineering teams owning direct production SLAs and on-call.Deep SRE Discipline:
Grounded in foundational SRE principles (SLOs, error budgets, blameless postmortems) paired with a strict "code over heroics" mindset. Hands-On Technical Depth (SRE + SWE): Track record shipping production code in Go, Python, or C++, with hands-on systems expertise across Linux OS kernels, bare-metal provisioning, firmware, and/or high-performance networking fabrics. Extensive experience with distributed systems and open source software.Extensive AI/ML Adoption:
Active utilization of AI agents and automated LLM/ML workflows in modern software engineering and diagnostic operations.Senior Talent Magnet:
Track record of attracting, evaluating, developing and leading unusually senior software engineers who thrive in fast-paced, high-stakes environments. Attributes Possess a high tolerance for ambiguity. The first clusters will carry customer workloads while the SLOs are still being defined and the team is still being hired. You absorb that, translate unknowns into concrete near-term priorities, and never manufacture false certainty about reliability the data does not support. Understands that the customer's job is the unit of reliability. A node that is "up" while a training run stalls on a flapping link is down. You measure what customers experience — goodput, time-to-recover, lost progress — and hold the whole stack, and Google, to it.Mindset Builder Mindset & Ambiguity:
A true "builder, not supervisory" orientation; comfortable operating with high autonomy, navigating ambiguity, and establishing structure amidst rapid growth.AI-Agentic First:
Fluent with AI agents — or committed to becoming so quickly — and you embed them as first principles in how you and your team work, defaulting to agentic workflows before adding headcount or process. Nice to have (Preferred, not required): Hyperscaler /Neocloud Scale:
SRE or fleet leadership at a hyperscaler (Google, AWS, Meta, MSFT) or neocloud (CoreWeave, Lambda, Nebius, Nscale) during rapid fleet ramps.Accelerated Compute:
Direct TPU experience or large-scale GPU cluster ops (NCCL collective debugging, RDMA/GPU-Direct, Slurm/Kubernetes AI schedulers).Custom Fleet Tooling:
Hands-on experience building custom remediation controllers, event-driven fleet management software (Go, NetBox/DCIM), or OpenTelemetry/Prometheus pipelines.Facility Telemetry & Thermal Signals:
Familiarity with high-density, liquid-cooled environments and integrating facility telemetry (power, thermal, flow) into compute health signals.Customer SLAs & Reporting:
Proven experience constructing customer SLAs/SLOs, credit mechanics, and executive/customer-facing reliability reviews. Salary Range Information The annual salary range for this position has been estimated based on market data and other factors. However, a salary higher or lower than this range may be appropriate for a candidate whose qualifications differ meaningfully from those listed in the job description. About Crux We offer generous base, bonus and additional incentive based compensation Health, dental, and vision coverage for you and your dependents Company-paid life insurance and disability Full suite of other optional benefits 401(k) Plan with 4% company match (USA employees)Benefits
- 401(k) Plans
- Health Insurance
- Dental Insurance
- Vision Insurance