Skip to main content
Tallo logoTallo logo

Find Jobs

Find Jobs Near You – Available Work in Your Location

Skip to job details
Apply for this opportunity

To apply for this job, you'll continue to an external website or email application.

dicedemo

AI Infrastructure Engineer

Career Insights for Artificial Intelligence Engineer (General)

See where this job fits in the broader career landscape. Knowing your career path helps you see what's possible from here.

Scorecard

Based on Connecticut data

Review key factors to help you decide if this role fits your goals. How is this calculated?

Were these scores useful?

What they do

An Artificial Intelligence Engineer develops, tests, and deploys artificial intelligence models. May work closely with data software engineers and data professionals to train and implement AI models into existing systems or develop new applications.

$145,672 / year median in Connecticut

Explore Career

Job Description

AI Infrastructure Engineer dicedemo Boston, CT Job Details Full-time $130,000 a year 4 hours ago Qualifications Containerization systems Performance monitoring Software deployment Infrastructure as Code (IaC) Infrastructure architecture design IT monitoring tools Cloud networking Monitoring system implementation Cloud Architecture Design (Architecture design skills) Linux Cloud automation Distributed computing DevOps automation Cloud monitoring Full Job Description Job Description AI Infrastructure Engineer Position Overview We are seeking an AI Infrastructure Engineer to design, build, and scale the infrastructure that powers our artificial intelligence and machine learning workloads. This role sits at the intersection of AI/ML, cloud infrastructure, distributed systems, and DevOps/MLOps. The ideal candidate has experience building highly available, scalable infrastructure for training, deploying, and operating machine learning and generative AI applications. You will partner closely with Machine Learning Engineers, Data Scientists, Software Engineers, and Platform Engineering teams to ensure AI workloads can run reliably, securely, and efficiently at scale. Key Responsibilities Design, build, and maintain scalable infrastructure for AI, machine learning, and Generative AI workloads Build and manage cloud infrastructure across AWS, Azure, and/or Google Cloud Platform Deploy and operate GPU-based compute environments for model training and inference Design infrastructure supporting LLMs, model training, fine-tuning, inference, and AI applications Build and manage containerized workloads using Docker and Kubernetes Develop infrastructure-as-code using tools such as Terraform, CloudFormation, or Pulumi Build CI/CD and MLOps pipelines supporting model development and deployment Optimize GPU/CPU utilization, infrastructure performance, scalability, and cloud costs Implement monitoring, logging, observability, and alerting for AI infrastructure and services Support distributed training and high-performance computing environments Build secure, highly available systems capable of supporting production AI workloads Partner with ML Engineers and Data Scientists to move models from experimentation into production Troubleshoot infrastructure, networking, performance, and deployment issues Evaluate emerging AI infrastructure technologies and recommend improvements to the platform Required Qualifications 3+ years of experience in Cloud Infrastructure, DevOps, Platform Engineering, SRE, MLOps, or AI/ML Infrastructure Strong experience with at least one major cloud platform: AWS, Azure, or GCP Experience with Kubernetes and Docker Experience with Infrastructure-as-Code tools such as Terraform Strong scripting/programming skills in Python, Bash, Go, or similar languages Experience building CI/CD pipelines using tools such as GitHub Actions, GitLab CI, Jenkins, or similar Knowledge of networking, Linux systems, distributed computing, and cloud architecture Experience implementing monitoring and observability solutions Understanding of machine learning development and deployment workflows Preferred Qualifications Experience managing GPU infrastructure, including NVIDIA GPUs and CUDA environments Experience with AI/ML frameworks such as PyTorch, TensorFlow, JAX, or Hugging Face Experience supporting LLM training, fine-tuning, RAG, or inference workloads Experience with MLOps platforms such as MLflow, Kubeflow, SageMaker, Vertex AI, or Azure Machine Learning Experience with distributed training technologies such as Ray, DeepSpeed, PyTorch Distributed, or Horovod Familiarity with AI inference technologies such as v
LLM, NVIDIA
Triton, or TensorRT Experience managing Kubernetes-based GPU clusters Understanding of model serving, vector databases, and modern Generative AI architecture Experience optimizing infrastructure for performance and cloud/GPU cost efficiency What Success Looks Like In this role, you will help create the infrastructure foundation that allows AI teams to experiment faster, train models efficiently, deploy AI applications reliably, and scale them into production. You will reduce friction between AI development and production while improving reliability, performance, security, and infrastructure cost. , Required Skills 3+ years of experience in Cloud Infrastructure, DevOps, Platform Engineering, SRE, MLOps, or AI/ML Infrastructure Strong experience with at least one major cloud platform: AWS, Azure, or GCP Experience with Kubernetes and Docker Experience with Infrastructure-as-Code tools such as Terraform Strong scripting/programming skills in Python, Bash, Go, or similar languages Experience building CI/CD pipelines using tools such as GitHub Actions, GitLab CI, Jenkins, or similar Knowledge of networking, Linux systems, distributed computing, and cloud architecture Experience implementing monitoring and observability solutions Understanding of machine learning development and deployment workflows , Desired Skills Experience managing GPU infrastructure, including NVIDIA GPUs and CUDA environments Experience with AI/ML frameworks such as PyTorch, TensorFlow, JAX, or Hugging Face Experience supporting LLM training, fine-tuning, RAG, or inference workloads Experience with MLOps platforms such as MLflow, Kubeflow, SageMaker, Vertex AI, or Azure Machine Learning Experience with distributed training technologies such as Ray, DeepSpeed, PyTorch Distributed, or Horovod Familiarity with AI inference technologies such as v
LLM, NVIDIA
Triton, or TensorRT Experience managing Kubernetes-based GPU clusters Understanding of model serving, vector databases, and modern Generative AI architecture Experience optimizing infrastructure for performance and cloud/GPU cost efficiency , About dicedemo New Company