Skip to main content
Tallo logoTallo logo

Find Jobs

Find Jobs Near You – Available Work in Your Location

Skip to job details

Back to Results

Apply for this opportunity

To apply for this job, you'll continue to an external website or email application.

VST Consulting, Inc

Senior Engineer AI Workload & Infrastructure Validation

Review key factors to help you decide if the role fits your goals.
Pay Growth
?
out of 5
Not enough data
Not enough info to score pay or growth
Job Security
?
out of 5
Not enough data
Calculating job security score...
Total Score
76
out of 100
Average of individual scores

Were these scores useful?

Job Description

Job Title:
Senior Engineer AI Workload & Infrastructure Validation Client:
LTTS Location:
Plano, TX Hybrid Employment Type:
FTE Experience:
5+ Years ML Engineering, MLOps, or Performance Engineering with
Production GPU Workloads Interview Mode:
Virtual Practice:
AI Infrastructure / GPU-as-a-Service About the Role We are building a GPU-as-a-Service and AI Factory practice supporting enterprise and industrial customers. This role combines AI workload engineering, distributed training, production inference, GPU infrastructure validation, benchmarking, and customer acceptance testing . You will enable and tune distributed training and inference workloads while owning an independent validation offering that determines whether GPU infrastructure performs according to the reference architecture. Key Responsibilities Enable and tune distributed training using PyTorch DDP, FSDP, DeepSpeed/ZeRO, tensor parallelism, pipeline parallelism, and NCCL tuning . Measure and improve Model FLOPs Utilization (MFU) across distributed GPU workloads. Architect production inference serving using Triton, NVIDIA NIM, vLLM, and TensorRT-LLM . Implement and optimize quantization, continuous batching, KV-cache, paged attention, and autoscaling to meet latency SLOs. Size inference platforms based on customer time-to-first-token, inter-token latency, concurrency, and context-length requirements . Support GenAI workloads including RAG, fine-tuning, agentic pipelines, and vector database integration . Build and maintain a GPU infrastructure validation suite covering fabric validation, NCCL scaling, GPU burn/thermal/power testing, storage throughput, GPUDirect verification, and workload MFU baselines. Deliver GPU infrastructure acceptance validation engagements and prepare customer-facing acceptance reports. Design continuous validation processes covering driver/firmware certification and performance drift detection . Develop benchmark methodologies, reports, and evidence for customer engagements. Own PoC workload design and benchmark execution . Required Qualifications 5+ years of ML engineering, MLOps, or performance engineering experience with production GPU workloads. Hands-on experience with distributed training on multi-node GPU clusters . Production LLM inference serving experience with Triton, vLLM, or TensorRT-LLM. Strong benchmarking and performance-analysis skills with the ability to profile systems and explain GPU utilization and performance results. Strong Python development experience. Experience with containers and Kubernetes-native workflows , including Kubeflow, Ray, or Argo Workflows. Strong technical writing skills for producing customer-facing validation reports. Nice to Have NVIDIA NIM and NeMo experience. MLPerf or formal benchmark program experience. Test-and-validation engineering background. Large-scale fine-tuning using LoRA/QLoRA . Customer-facing PoC delivery experience.