An Engineering Manager oversees engineering and technical design of large projects, and manages communication and coordination among different types of engineers working on a project. Supervises project schedule, budget, and communications with stakeholders.
Anywhere in Country At EY, we're all in to shape your future with confidence. We'll help you succeed in a globally connected powerhouse of diverse teams and take your career wherever you want it to go. Join EY and help to build a better working world. The opportunity We are seeking AI Systems Engineers to build and operate the foundational substrate that powers EY's AI-native platform. This role owns the infrastructure and cloud-native platform layers of the Hybrid AI Multi-Environment Runtime ( HAI ), from bare-metal and GPU infrastructure through Kubernetes, cluster fabric, and multi-tenant scaling. You will be responsible for building and managing EY Fabric environments across cloud, on-prem, edge, and air-gapped targets. This is the substrate on which EY Agentic AI capabilities run. This role is ideal for a full-stack infrastructure leader who is equally comfortable with bare-metal and GPU systems and production Kubernetes at scale, who treats reliability and portability as non-negotiable in regulated client contexts, and who understands that the substrate is a product in its own right, measured by the velocity, safety, and portability it unlocks for every team building above it. Your key responsibilities + Own the cluster & cloud-native platform: compute, Kubernetes and scheduling, cluster fabric/networking, multi-tenancy, and distributed compute, as the substrate for Agentic AI workflows and tooling. + Own the infrastructure foundation: Ubuntu/OS, BMC /bare-metal, DPU architecture, and NVAIE ( GPU /Network/ DCGM ), ensuring the physical and virtual bedrock is provisioned, patched, and production-ready. + Stand up and manage EY Agentic AI environments across cloud ( EKS / AKS / GKE ), on-prem AI Factory (RKE2/ NVAIE ), edge (K3s), and air-gapped deployment modes, maintaining one consistent stack contract across all targets. + Deliver foundational platform capabilities such as Infrastructure Management, Kubernetes & Scheduling, and Cluster Fabric Management, so downstream runtime, data, and execution services can run safely and consistently. + Own cluster lifecycle, autoscaling, GPU pooling/virtualization, and multi-tenancy boundaries (vCluster/Crossplane/Karpenter), providing isolated, elastic capacity per tenant and engagement. + Own secure execution and inference: Ray Serve, v
LLM/ NIM
/Triton, and NVIDIA Dynamo, with sandboxed execution (gVisor/Firecracker for hosted, NVIDIA OpenShell/vNode for on-prem) for isolated, To view full details and how to apply, please login or create a Job Seeker account