Skip to main content
Tallo logoTallo logo

Find Jobs

Find Jobs Near You – Available Work in Your Location

Skip to job details
Apply for this opportunity

To apply for this job, you'll continue to an external website or email application.

Choctaw Nation of Oklahoma

AI Operations and Monitoring Engineer

Career Insights for Artificial Intelligence Engineer (General)

See where this job fits in the broader career landscape. Knowing your career path helps you see what's possible from here.

Scorecard

Based on Oklahoma data

Review key factors to help you decide if this role fits your goals. How is this calculated?

Were these scores useful?

What they do

An Artificial Intelligence Engineer develops, tests, and deploys artificial intelligence models. May work closely with data software engineers and data professionals to train and implement AI models into existing systems or develop new applications.

$132,541 / year median in Oklahoma

Explore Career

Job Description

Monday-Friday 8:00AM-4:30PM| Hybrid Position| Weekly Earned Wage Access is an option for this position.
Job Purpose or Goals:
The AI Operations and Monitoring Engineer is responsible for ensuring the reliability, performance, documentation, and compliance of production AI/ML systems, keeping them stable and functioning as intended. They play a key role in maintaining system uptime and driving effective incident response to minimize outages and protect overall service quality.
Tasks:
1. Monitor AI/ML system health, drift, and anomalies by continuously tracking model performance, identifying unexpected behavior, and ensuring systems operate within expected parameters. 2. Build dashboards, alerts, and observability tooling to give teams real-time insight into system performance and enable rapid detection of issues before they impact operations. 3. Implement CI/CD pipelines for models to streamline deployments, maintain version control, and ensure that model updates move reliably from development to production. 4. Maintain training and deployment environments by keeping infrastructure stable, updated, and optimized so models can be trained and deployed efficiently. 5. Develop automation for scaling, retraining and reducing manual workloads and ensuring models automatically adapt to new data or demand. 6. Respond to incidents and outages to restore services quickly, conduct rootcause analysis, and prevent future disruptions to AI/ML workloads. 7. Collaborate on security, compliance, and governance to ensure AI systems follow organizational standards and regulatory requirements. 8. Write runbooks and maintain audit logs to document procedures, support operational consistency, and provide traceability for all system actions. 9. Performs other duties as may be assigned.
Job Requirements:
Bachelor's degree in computer science or related field, or 4 years relevant professional experience 3 years' experience in DevOps, SRE, or MLOps Proficiency with tools like Kubernetes, Docker, and cloud services Experience with observability tools Familiarity with ML deployment workflows Bachelor's degree in computer science or related field, or 4 years relevant professional experience 3 years' experience in DevOps, SRE, or MLOps Proficiency with tools like Kubernetes, Docker, and cloud services Experience with observability tools Familiarity with ML deployment workflows 1. Monitor AI/ML system health, drift, and anomalies by continuously tracking model performance, identifying unexpected behavior, and ensuring systems operate within expected parameters. 2. Build dashboards, alerts, and observability tooling to give teams real-time insight into system performance and enable rapid detection of issues before they impact operations. 3. Implement CI/CD pipelines for models to streamline deployments, maintain version control, and ensure that model updates move reliably from development to production. 4. Maintain training and deployment environments by keeping infrastructure stable, updated, and optimized so models can be trained and deployed efficiently. 5. Develop automation for scaling, retraining and reducing manual workloads and ensuring models automatically adapt to new data or demand. 6. Respond to incidents and outages to restore services quickly, conduct rootcause analysis, and prevent future disruptions to AI/ML workloads. 7. Collaborate on security, compliance, and governance to ensure AI systems follow organizational standards and regulatory requirements. 8. Write runbooks and maintain audit logs to document procedures, support operational consistency, and provide traceability for all system actions. 9. Performs other duties as may be assigned.