Find Jobs
Find Jobs Near You – Available Work in Your Location
Skip to job details
L
Lilly
Senior Scientific Data Curator
Career Insights for Clinical Data Specialist
See where this job fits in the broader career landscape. Knowing your career path helps you see what's possible from here.
Scorecard
Based on Indiana data
Review key factors to help you decide if this role fits your goals. How is this calculated?
What they do
A Clinical Data Specialist manages data and information gathered during a clinical trial, including studies of medical treatments and drugs.
$76,356 / year median in Indiana
+4% projected growth
Job Description
At Lilly, the work is demanding because patients are waiting. We unite caring with discovery to help make life better for people around the world, knowing that every decision, every detail, and every day matters. Headquartered in Indianapolis, Indiana, our over 50,000 employees around the globe take on complex challenges to discover and deliver life-changing medicines, strengthen how health is understood and managed, and support the communities we serve. This is hard, urgent, selfless work—but it's work worth doing. If you're driven by purpose and ready to bring your best to work that truly matters for patients, we invite you to join us. Position Summary The Senior Scientific Data Curator will lead the systematic discovery, assessment, harmonization, and quality assurance of Lilly's scientific datasets across the full breadth of TuneLab's modeling domains—small-molecule ADME/ADMET, safety and secondary pharmacology, in vivo pharmacokinetics and toxicology, antibody and biologics developability, and clinical PK/PD—in support of a strategic, cross-modality data unification initiative. This role sits at the intersection of biological and pharmacological domain expertise and data science, translating decades of fragmented, heterogeneous datasets spanning discovery through the clinic into a unified, AI-ready data infrastructure. The curator will partner closely with computational scientists, DMPK scientists, pharmacometricians, antibody engineers, and external consortium collaborators to ensure that the data substrate underpinning TuneLab's federated AI/ML models is comprehensive, well-documented, and scientifically sound. Core Responsibilities Data Discovery & Assessment Conduct comprehensive inventory of historical and ongoing datasets across TuneLab's modeling domains—small-molecule ADME/ADMET, safety and secondary pharmacology, in vivo PK and toxicology, antibody and biologics developability, and clinical PK/PD—spanning therapeutic areas (oncology, immunology, metabolic diseases, neuroscience, etc.) and 20+ years of discovery, preclinical, and clinical data Assess and score data quality, completeness, and integration feasibility for each dataset, accounting for the distinct data structures of each domain, including assay and dose-response measurements, concentration-time profiles and dosing regimens, in vivo study readouts, sequence- and structure-derived features for biologics, and biomarker and clinical covariate data Map metadata gaps across legacy systems and source platforms, documenting study contexts, assay and protocol methods, protocol deviations, data quality flags, and provenance information Develop automated pipelines (including LLM-assisted extraction where appropriate) to identify and extract domain-relevant data from internal documents, assay databases, and study reports into standardized, model-ready formats Produce a prioritized data assessment report recommending which domains, therapeutic areas, and indications to integrate first, based on data volume, complexity, portfolio relevance, and model feasibility Data Harmonization & Integration Design and implement standardized, extensible schemas for the integrated multi-domain database, working with computational partners to ensure AI/ML readiness across small-molecule and biologics modalities Build and maintain data harmonization pipelines: label normalization, unit and assay-condition standardization across studies and sources, time-point and dose alignment, sequence and structure normalization for biologics, and covariate encoding Apply domain-driven quality control practices—sequence validation, hidden duplicate detection, cross-source discrepancy resolution, and cross-species dataset integration using allometric scaling where applicable Develop and execute outlier detection protocols, flagging and adjudicating anomalous values in collaboration with clinical pharmacologists, DMPK scientists, toxicologists, and antibody engineers as appropriate to the domain Create reproducible data quality assurance workflows with documented acceptance criteria and audit trails Curate and enrich metadata to enable cross-study and cross-domain querying—linking compound and molecule identifiers, sequence and construct identifiers, assay methods, formulation details, and study design parameters Cross-functional Partnership Serve as the primary data domain expert for external consortium partners working within Lilly's controlled cloud environment Collaborate with pharmacometricians, DMPK scientists, toxicologists, and antibody engineers to validate harmonized datasets against legacy models and established analyses (e.g., NONMEM/Monolix outputs for the clinical PK/PD domain) Work with the TuneLab ML team to ensure curated datasets meet the input specifications for the platform's multi-task ML models, representation and foundation-model embeddings, and mechanistic/hybrid PK/PD frameworks (e.g., Neural ODE, SINDy) Contribute to platform deployment by supporting the development of data dictionaries, user documentation, and training materials for internal and consortium end users Required Qualifications M.S. or Ph.D. in Computational Data Science, Pharmacometrics, Pharmaceutical Sciences, Computational Biology, Cheminformatics, Biomedical Informatics, or a related quantitative discipline 1+ year of hands-on experience curating, harmonizing, or building analysis-ready datasets from biological, chemical, or clinical data sources Demonstrated skill in scientific dataset construction with domain-driven