Find Jobs
Find Jobs Near You – Available Work in Your Location
Staff ML Engineer
Career Insights for Machine Learning Engineer
See where this job fits in the broader career landscape. Knowing your career path helps you see what's possible from here.
Scorecard
Based on California data
Review key factors to help you decide if this role fits your goals. How is this calculated?
What they do
A Machine Learning Engineer specializes in designing, building, and deploying machine learning models. They utilize statistical and mathematical techniques, parallelizing processing, hyperparameter tuning, and other optimization methodologies to improve model performance. Responsibilities also include collecting and preprocessing large datasets, conducting exploratory data analysis, working closely with data engineers to understand data requirements, and engineer input variables for machine learning models.
$168,439 / year median in California
Job Description
Some examples include:
session cookies needed to transmit the website, authentication cookies, and security cookies. The website will not function properly without these cookies." Please see our Privacy Policy Read Full Privacy Message Decline Accept Cookies Sign In Home Search for Jobs Join Our Talent Community! Staff ML Engineer page is loaded Staff ML Engineer remote type Hybrid locations Pune- Panchshil
- India (Office) Bangalore
- India (Office) time type Full time posted on Posted Yesterday time left to apply
End Date:
November 30, 2026 (30+ days left to apply) job requisition id R04445 Cohesity is the leader in AI-powered data security. Over 13,600 enterprise customers, including over 85 of the Fortune 100 and nearly 70% of the Global 500, rely on Cohesity to strengthen their resilience while providing Gen AI insights into their vast amounts of data. Formed from the combination of Cohesity with Veritas' enterprise data protection business, the company's solutions secure and protect data on-premises, in the cloud, and at the edge. Backed by NVIDIA, IBM, HPE, Cisco, AWS, Google Cloud, and others, Cohesity is headquartered in Santa Clara, CA, with offices around the globe. We've been named a Leader by multiple analyst firms and have been globally recognized for Innovation, Product Strength, and Simplicity in Design , and our culture. Want to join the leader in AI-powered data security?Staff ML Engineer Level:
Staff Software Engineer (Level 5)Team:
Gaia Emblem Engine Location:
India |Type:
Full-Time About the Role Cohesity is looking for a Staff ML Engineer (Level 5) to help design and scale the Gaia Emblem engine, the intelligence layer powering our next-generation data platform. In this role, you will architect and build large-scale, production-grade ML/LLM systems — spanning model serving, retrieval-augmented generation (RAG), GPU infrastructure, and distributed backend services. As a Staff-level engineer, you will set technical direction, mentor senior engineers, and partner closely with Product and Data Engineering to bring Gaia Emblem's roadmap to life. Key Responsibilities- Architect and drive the technical roadmap for the Gaia Emblem engine, including LLM integration, RAG pipelines, and prompt-engineering frameworks.
- Design and operate scalable, GPU-aware infrastructure on Kubernetes (vanilla K8s, OpenShift, EKS, GKE, or equivalent), including GPU scheduling, autoscaling, and costefficient resource utilization.
- Build resilient, high-throughput distributed systems and microservices exposed via gRPC and REST APIs.
- Own data engineering pipelines feeding ML workflows, including ingestion, transformation, and storage across SQL, NoSQL, and search systems (PostgreSQL, Redis, Elasticsearch).
- Define and enforce engineering best practices: CI/CD, test automation, observability, and infrastructure-as-code (Terraform, Helm).
- Lead design reviews, set coding and architectural standards, and mentor engineers across the organization.
- Partner with Product, Applied ML, and SRE teams to translate business requirements into scalable technical solutions.
- Drive root-cause debugging and performance optimization across distributed, GPUbacked services.
- Evaluate and introduce new tools, frameworks, and infrastructure patterns to keep Gaia Emblem at the forefront of applied ML engineering. Minimum Qualifications Experience
- 10+ years of professional software engineering experience, including 8+ years building distributed, production-grade backend systems.
- 3+ years of hands-on experience with LLM-based systems, RAG architectures, or applied ML infrastructure.
- Demonstrated experience operating services on Kubernetes (vanilla K8s, OpenShift, EKS, GKE, or similar) at production scale, including GPU-backed workloads.
- Track record of technical leadership: driving architecture decisions, leading crossteam initiatives, and mentoring engineers at Staff/Senior level. Education
- Bachelor's degree in Computer Science, Engineering, or a related technical field required.
- Master's degree or PhD in Computer Science, Machine Learning, or a related field preferred. Required Skills
Languages:
Python, Java, Go, JavaScriptML/AI:
LLMs, Prompt Engineering, Retrieval-Augmented Generation (RAG)Infrastructure & Cloud:
Kubernetes (platform-agnostic — vanilla K8s, OpenShift, EKS, GKE, or equivalent), GPU Scheduling, Docker, Helm Charts, Terraform, AWS, Cloud ComputingDistributed Systems & APIs:
Distributed Systems, Microservices, gRPC, REST
APIsData & Storage:
Data Engineering, SQL, NoSQL, PostgreSQL, Redis, Elasticsearch, DatabasesEngineering Practices:
CI/CD, DevOps, Test Automation, Git, Debugging, ObjectOriented Programming (OOP) Preferred / Nice-to-Have Skills- Node.js, Back-End Web Development
- Jenkins and other CI/CD tooling
- Experience with formal verification or correctness-focused engineering practices
- Prior experience in enterprise data protection, storage, or infrastructure software Expectations from the Role
- Operate with Staff-level ownership: define the "what" and "how" for a major component of Gaia Emblem, not just execute tickets.
- Balance long-term architectural vision with pragmatic, incremental delivery.
- Act as a force multiplier — raise the technical bar of the team through design reviews, documentation, and mentorship.
- Communicate complex technical trade-offs clearly to both engineering and nonengineering stakeholders.
- Be hands-on: this role is expected to write and review production code regularly, not purely architect from a distance.