Job Description One of our largest Telecom customers is seeking a skilled DevOps Engineer to join their Data Platform team within the Chief Data Office. This Engineer will play a key role in building, operating, and scaling a production-grade AI inferencing platform that supports high-throughput, low-latency large language model (LLM) workloads. The position combines software engineering, platform engineering, and reliability engineering responsibilities, with a focus on developing Python-based services, automating model deployments, and optimizing model serving performance using vLLM, Kubernetes, KServe, and Knative. The engineer will be responsible for containerization with Docker and Podman, maintaining and tuning PostgreSQL data models, troubleshooting production issues, and driving platform reliability through observability and automation. Working closely with ML engineers and platform architects, this individual will help scale mission-critical AI services, support new model rollouts, and continuously improve the infrastructure that powers enterprise AI applications. We are a company committed to creating diverse and inclusive environments where people can bring their full, authentic selves to work every day. We are an equal opportunity/affirmative action employer that believes everyone matters. Qualified candidates will receive consideration for employment regardless of their race, color, ethnicity, religion, sex (including pregnancy), sexual orientation, gender identity and expression, marital status, national origin, ancestry, genetic factors, age, disability, protected veteran status, military or uniformed service member status, or any other status or characteristic protected by applicable laws, regulations, and ordinances. If you need assistance and/or a reasonable accommodation due to a disability during the application or recruiting process, please send a request to HR@insightglobal.com.
To learn more about how we collect, keep, and process your private information, please review
Insight Global's Workforce Privacy Policy:
https://insightglobal.com/workforce-privacy-policy/. Skills and Requirements
- Strong hands-on python coding & development experience as well as experience working with FastAPI
- Expertise with Docker & Kubernetes, including deployments, services, autoscaling, and troubleshooting workloads in production.
- Practical experience with Postgres, including query optimization and schema design.
- Cloud platform experience, such as Azure/AKS, AWS/EKS, or GCP/GKE.
- Strong debugging skills across the stack — from application code to infrastructure
- Direct experience with vLLM or LLM inferencing
- Experience with GPU-aware scheduling and resource management in Kubernetes.
- Familiarity with LLM-specific optimizations: speculative decoding, quantization (FP8/INT8), KV cache management, and continuous batching.
- Experience with API gateways or LLM routing layers, such as LiteLLM or similar.
- Background in SRE practices: SLOs/SLAs, incident response, and on-call tooling.
- Experience with Helm, ArgoCD, or other GitOps deployment tooling.
- Familiarity with observability stacks, including Prometheus, Grafana, and OpenTelemetry.