Skip to main content
Tallo logoTallo logo

Find Jobs

Find Jobs Near You – Available Work in Your Location

Skip to job details

Back to Results

Apply for this opportunity

To apply for this job, you'll continue to an external website or email application.

Propio LS LLC

Senior Machine Learning Engineer, Speech & LLM Training Data

Career Insights for Machine Learning Engineer

See where this job fits in the broader career landscape. Knowing your career path helps you see what's possible from here.

Scorecard

Based on Kansas data

Review key factors to help you decide if this role fits your goals. How is this calculated?

Were these scores useful?

What they do

A Machine Learning Engineer specializes in designing, building, and deploying machine learning models. They utilize statistical and mathematical techniques, parallelizing processing, hyperparameter tuning, and other optimization methodologies to improve model performance. Responsibilities also include collecting and preprocessing large datasets, conducting exploratory data analysis, working closely with data engineers to understand data requirements, and engineer input variables for machine learning models.

$133,257 / year median in Kansas

Explore Career

Job Description

Senior Machine Learning Engineer, Speech & LLM Training Data Propio

LS LLC - 5.0

Overland Park, KS Job Details Full-time 17 hours ago Qualifications AI models AWS IAM Databricks Big data projects Data augmentation PyTorch Tooling Spark Amazon SageMaker AI platforms (beyond public GPTs) Computational framework Maintaining data pipelines SQL Machine learning cloud services Machine intelligence Docker AWS Glue AWS Step Functions Model training S3 Linux Machine learning libraries Model evaluation Machine learning frameworks Full Job Description Overland Park, KS

• Engineering Job Type Full-time Description Propio Language Services is one of the top 5 providers high-quality, real-time multilingual interpretation, translation, and localization services, operating at 9-figure scale across healthcare, legal, and other industries. We are driven by a passion for cutting-edge technology and exceptional service, building seamless experiences that bridge communication gaps across languages, cultures, and modalities. Propio is hiring a Senior Machine Learning Engineer, Speech & LLM Training Data to transform large volumes of multilingual conversational audio into high-quality training and evaluation datasets. This hands-on role owns audio processing, dataset curation, annotation and QA workflows, model training, and evaluation for our multilingual speech, translation, and conversational AI systems.

Key Responsibilities:

Define the data roadmap for multilingual speech, translation, multimodal LLMs, and conversational AI. Build audio-processing pipelines covering resampling, channel handling, VAD, diarization, language identification, transcription, alignment, and quality filtering. Build dataset pipelines for cleaning, deduplication, PII/PHI redaction, quality scoring, sampling, balancing, versioning, and lineage. Design annotation guidelines, QA rubrics, golden datasets, and reviewer workflows. Build evaluation datasets, analyze model failures, and translate performance gaps into targeted data improvements. Run training, fine-tuning, post-training, and evaluation experiments, including SFT, preference data, DPO/RLHF-style workflows, and synthetic data generation. Productionize secure, traceable, and reproducible data and ML workflows on AWS.

Requirements Qualifications:

Bachelor's or Master's degree in Computer Science, Machine Learning, Data Science, Electrical Engineering, Computational Linguistics, or a related field, or equivalent practical experience. 5+ years of experience in ML engineering, speech/audio ML, ML data engineering, NLP, or LLM training-data workflows. Strong hands-on experience with Python, SQL, Linux, Git, and Docker. Experience training or evaluating models using PyTorch, Hugging Face, or comparable ML frameworks. Experience with FFmpeg and audio-processing libraries such as TorchCodec, torchaudio, librosa, or equivalent tools. Experience with speech-processing tasks such as VAD, diarization, ASR, forced alignment, language identification, and audio-quality analysis. Experience with Databricks/Spark, Parquet/Arrow, and large-scale dataset pipelines. Working knowledge of AWS S3, SageMaker, Glue, Step Functions, IAM, and KMS. Experience with an annotation platform such as Labelbox, Label Studio, Scale AI, Prodigy, Argilla, or custom internal tooling. Experience with experiment tracking and data versioning tools such as MLflow, Weights & Biases, DVC, Delta Lake, or LakeFS. Experience with multilingual speech, translation, annotation workflows, and evaluation datasets.

Preferred Qualifications:

Experience with multilingual telephony, healthcare, interpretation, or call-center audio. Experience with tools such as Silero VAD, pyannote, WhisperX, NeMo, Kaldi, or equivalent speech technologies. Experience with distributed processing or training using Ray, PySpark, or similar frameworks. Experience with

HIPAA, PHI/PII

redaction, and secure data governance. Experience with low-resource languages, accents, dialects, and code-switching. Experience with synthetic data, active learning, weak supervision, or LLM-as-judge evaluation. #LI-JS1 Notice of AI Use in Job Application Review As part of our commitment in creating a fair, efficient, and consistent hiring process we may use artificial intelligence (AI) to help our recruiting teams organize, summarize, and analyze information provided by candidates, including resumes, application responses, and other materials submitted during the application process.



AI may be used to identify patterns, highlight relevant skills, and experience, and assist in comparing a candidate's qualifications with the requirement of a specific role. These tools are to improve efficiency and consistency while supporting more informed hiring decisions, which will ultimately be made by the hiring team.

Benefits

  • Dental Insurance