An Artificial Intelligence Engineer develops, tests, and deploys artificial intelligence models. May work closely with data software engineers and data professionals to train and implement AI models into existing systems or develop new applications.
Monday-Friday 8:00AM-4:30PM| Hybrid Position| Weekly Earned Wage Access is an option for this position.
Job Purpose or Goals:
The AI Operations and Monitoring Engineer is responsible for ensuring the reliability, performance, documentation, and compliance of production AI/ML systems, keeping them stable and functioning as intended. They play a key role in maintaining system uptime and driving effective incident response to minimize outages and protect overall service quality.
Tasks:
1. Monitor AI/ML system health, drift, and anomalies by continuously tracking model performance, identifying unexpected behavior, and ensuring systems operate within expected parameters. 2. Build dashboards, alerts, and observability tooling to give teams real-time insight into system performance and enable rapid detection of issues before they impact operations. 3. Implement CI/CD pipelines for models to streamline deployments, maintain version control, and ensure that model updates move reliably from development to production. 4. Maintain training and deployment environments by keeping infrastructure stable, updated, and optimized so models can be trained and deployed efficiently. 5. Develop automation for scaling, retraining and reducing manual workloads and ensuring models automatically adapt to new data or demand. 6. Respond to incidents and outages to restore services quickly, conduct rootcause analysis, and prevent future disruptions to AI/ML workloads. 7. Collaborate on security, compliance, and governance to ensure AI systems follow organizational standards and regulatory requirements. 8. Write runbooks and maintain audit logs to document procedures, support operational consistency, and provide traceability for all system actions. 9. Performs other duties as may be assigned.
Job Requirements:
Bachelor's degree in computer science or related field, or 4 years relevant professional experience 3 years' experience in DevOps, SRE, or MLOps Proficiency with tools like Kubernetes, Docker, and cloud services Experience with observability tools Familiarity with ML deployment workflows Bachelor's degree in computer science or related field, or 4 years relevant professional experience 3 years' experience in DevOps, SRE, or MLOps Proficiency with tools like Kubernetes, Docker, and cloud services Experience with observability tools Familiarity with ML deployment workflows 1. Monitor AI/ML system health, drift, and anomalies by continuously tracking model performance, identifying unexpected behavior, and ensuring systems operate within expected parameters. 2. Build dashboards, alerts, and observability tooling to give teams real-time insight into system performance and enable rapid detection of issues before they impact operations. 3. Implement CI/CD pipelines for models to streamline deployments, maintain version control, and ensure that model updates move reliably from development to production. 4. Maintain training and deployment environments by keeping infrastructure stable, updated, and optimized so models can be trained and deployed efficiently. 5. Develop automation for scaling, retraining and reducing manual workloads and ensuring models automatically adapt to new data or demand. 6. Respond to incidents and outages to restore services quickly, conduct rootcause analysis, and prevent future disruptions to AI/ML workloads. 7. Collaborate on security, compliance, and governance to ensure AI systems follow organizational standards and regulatory requirements. 8. Write runbooks and maintain audit logs to document procedures, support operational consistency, and provide traceability for all system actions. 9. Performs other duties as may be assigned.