Must Have Technical/Functional Skills Strong cloud engineering and AI/ML platform engineering experience. Kubernetes, containers, Infrastructure as Code, CI/CD, and API technologies. Python and software engineering experience. Experience with ML/LLM deployment and GenAI frameworks. Knowledge of vector databases, embedding models, model APIs, and AI orchestration. Experience with monitoring, telemetry, security, and production support.
Programming:
Advanced proficiency in Python; additional experience in Go, C++, or Rust is often preferred.
Infrastructure & Cloud:
Hands-on expertise with Kubernetes, Docker, and cloud platforms like AWS (SageMaker, Bedrock) or Azure (AI Foundry, OpenAI).
AI Frameworks:
Familiarity with orchestration and LLM tooling such as LangChain, LangGraph, Ray, or Kubeflow. Infrastructure-as-Code (IaC): Knowledge of automation tools like Terraform and Ansible. Roles & Responsibilities
- Build and maintain enterprise AI/ML and GenAI platform capabilities.
- Implement model endpoints, gateways, orchestration layers, and AI services.
- Build infrastructure supporting LLMs, SLMs, RAG, embeddings, vector stores, and AI agents.
- Develop automated deployment and CI/CD pipelines for AI applications and models.
- Implement LLMOps/MLOps capabilities covering deployment, monitoring, evaluation, versioning, and observability.
- Implement model-routing and inference-management capabilities.
- Support token, compute, latency, and inference-cost optimization.
- Integrate AI platforms with enterprise identity, security, logging, monitoring, networking, and API-management services.
- Automate infrastructure provisioning and platform configuration.
- Establish production reliability, scalability, resilience, and operational standards.
- Partner with AI architects, application engineers, data engineers, and security teams.
Platform Architecture:
Design and deploy scalable, cloud-native or bare-metal infrastructure (using Kubernetes, OpenShift, or AWS/Azure) to support large language models (LLMs) and machine learning workloads.
- Model Operations (MLOps/LLMOps): Optimize model serving, inference performance, GPU utilization, and automated pipelines for fine-tuning and deploying AI models.
Agentic & GenAI Integration:
Build and maintain shared services like AI gateways, Retrieval-Augmented Generation (RAG) frameworks, and autonomous agent orchestration platforms.
Governance & Security:
Enforce enterprise data protection, compliance, auditability, and cybersecurity standards across all AI tooling.
Developer Enablement:
Create self-service platforms, APIs, and monitoring/observability tools so internal software and data science teams can safely adopt AI capabilities.
TCS Employee Benefits Summary:
Discretionary Annual Incentive.
Comprehensive Medical Coverage:
Medical & Health, Dental & Vision, Disability Planning & Insurance, Pet Insurance Plans.
Family Support:
Maternal & Parental Leaves.
Insurance Options:
Auto & Home Insurance, Identity Theft Protection.
Convenience & Professional Growth:
Commuter Benefits & Certification & Training Reimburseme nt.
Time Off:
Vacation, Time Off, Sick Leave & Holidays.
Legal & Financial Assistance:
Legal Assistance, 401K Plan, Performance Bonus, College Fund, Student Loan Refinancing. Salary Range $110,000 - $130,000 a year