Skip to main content
Tallo logoTallo logo

Find Jobs

Find Jobs Near You – Available Work in Your Location

Skip to job details

Back to Results

Apply for this opportunity

To apply for this job, you'll continue to an external website or email application.

Hoonify Technologies Inc.

AI Infrastructure Engineer

Job Description

Between $120k and $175k Per Year DOE (Depends on Experience) Position range in Albuquerque MSA $77k - $158k Per Year AI Infrastructure Engineer Hoonify Technologies Inc.
Occupation:
Computer Systems Engineers/Architects
Location:
Albuquerque, NM - 87106
Job Type:
Regular, Full Time (30 Hours or More), Permanent Employment
Posted:
08/21/2026 Positions available: 1
Source:
New Mexico Jobs
Web Site:
New Mexico Jobs Onsite /
Remote:
Not Specified
Updated:
08/24/2026
Expires:
09/25/2026 Job #: 961417 Job Requirements and Properties Help for Job Requirements and Properties. Opens a new window. Work Onsite Full Time Education Bachelor's Degree Experience 36 Month(s) Language English, Very Well Schedule Full Time Job Type Regular Duration Permanent Employment Public Transit Available Benefits Job Description Help for Job Description. Opens a new window. We are seeking an AI Infrastructure Engineer to help build, deploy, and operate the LLM serving infrastructure underpinning our AI/ML platform. This role focuses on implementation, automation, and optimization of production inference systems, tuning model serving runtimes for performance and scale across NVIDIA and AMD GPU fleets, working under the technical direction of senior engineering leadership and the platform's established architectural patterns. You'll ship well-engineered, well-tested infrastructure changes and grow your depth in GPU-backed workloads, distributed model serving, observability, and continuous delivery. You'll work directly with senior engineers on real production systems, receive code and design review on everything you ship, and have a clear path to expanded scope and ownership as your experience deepens. What You'll Do Deploy, tune, and optimize high-performance LLM inference pipelines on GPU infrastructure, improving throughput, latency, and cost efficiency within established design patterns. Analyze, profile, and optimize model serving workloads across inference frameworks such as vLLM, SGLang, and TensorRT-LLM, and across different model families and hardware architectures. Build and operate scalable, production-grade API services for model inference, including request routing, multi-tenant isolation, usage metering, and observability. Develop benchmarking harnesses, monitoring infrastructure, and automation tooling that make serving performance measurable and reproducible. Scale inference workloads across multi-GPU, multi-node environments spanning NVIDIA and AMD accelerators. Evaluate, prototype, and integrate model fine-tuning workflows and frameworks. Collaborate closely with engineering and product teams to align infrastructure capabilities with customer-facing services. Investigate and resolve issues across the stack, including container, node, network, and accelerator-level problems, escalating appropriately when scope exceeds the role. Write clear documentation, including runbooks, internal references, and design notes for the changes you ship. Participate in code and design reviews, both as author and reviewer, and incorporate feedback from senior engineers into your work. Required Qualifications Bachelor's degree in Computer Science, Computer Engineering, Applied Math, or Data Science, plus three (3) years relevant work experience or equivalent combination of education and relevant experience. Professional software engineering experience, with at least some of it touching ML systems, GPU workloads, or high-performance backend services. Working knowledge of Kubernetes in a production context, including writing and debugging manifests, understanding core resource types, and operating production workloads. Hands-on experience serving or deploying LLMs you've run vLLM, SGLang, TGI, TensorRT-LLM, or similar. Comfort working in a Linux environment and with standard developer tooling, including Git-based workflows. Familiarity with CI/CD systems and the basic mechanics of automated build, test, and deployment pipelines. Strong proficiency in Python with familiarity at least one programming or scripting language used for infrastructure work (Go, Rust, C++ or Bash). Preferred Qualifications Experience building or operating retrieval-augmented generation (RAG) pipelines, including vector databases, embedding models, and retrieval serving at scale. AMD/ROCm experience. Experience cleaning and curating datasets for LLM training and fine tuning. Fine-tuning experience of any depth: LoRA/QLoRA, full fine-tunes, dataset curation, or evaluation design. Experience with usage metering, billing systems, or multi-tenant API platforms. Experience instrumenting services and consuming observability data, including writing Prometheus queries, building Grafana dashboards, or working with distributed traces. Experience with HPC batch schedulers and MPI based workloads. Experience with alternative compute architectures for inference (RISC-V, FPGA, ASIC). Why Join Hoonify You'll have a direct line to leadership and genuine influence over the company's growth trajectory. This is a rare opportunity to build a cutting-edge multi-cloud computational platform at a company doing meaningful work in AI with the autonomy and resources to make it your own. About Our Team Hoonify delivers secure, sovereign AI infrastructure designed for the next generation of inference workloads. Powered by TurbOS , our platform enables organizations and NeoCloud/data center operators to transform CPU/GPU infrastructure into production-ready AI environmentssupporting local LLMs, agentic copilots, RAG, and embeddings. We empower teams with robust model lifecycle management, multi-tenant controls, usage metering, and fully auditable operations. Hoonify is an equal opportunity employer. We welcome applicants from all backgrounds and are committed to building a diverse and inclusive team. Must be eligible to obtain and maintain a US government security clearance.

Career Insights for Artificial Intelligence Engineer (General)

See where this job fits in the broader career landscape. Knowing your career path helps you see what's possible from here.

Scorecard

Based on New Mexico data

Review key factors to help you decide if this role fits your goals. How is this calculated?

Were these scores useful?

What they do

An Artificial Intelligence Engineer develops, tests, and deploys artificial intelligence models. May work closely with data software engineers and data professionals to train and implement AI models into existing systems or develop new applications.

$131,394 / year median in New Mexico

Explore Career