
Richard J. Young, Ph.D.
AI Safety & Evaluation Researcher
I evaluate how foundation models fail at deployment scale.
I evaluate how foundation models fail under adversarial, clinical, and decision-support pressure — and ship the benchmarks, datasets, models, and production systems needed to deploy them safely at scale.
What I Believe
Foundation-model safety is empirical, not philosophical. We learn what models actually do under pressure by running large, replicable evaluations — and shipping the benchmarks, datasets, and models openly so others can do the same. The most consequential alignment progress comes from being honest about failure modes at deployment scale.
I work where AI methodology meets stakes that matter: clinical care, cybersecurity, behavioral health, and decision-support systems that touch real people. The methods are general; the domains aren't. That's the point.
Selected projects · 10
TEMPEST
adversarial attack success on 6 of 10 frontier LLMs. 97K queries across 8 vendors.
CARDIOEMBED
retrieval accuracy on clinical cardiology — +1452% over base embedding models.
CODE SAFETY
consensus-labeled prompts (Fleiss κ=0.876). 4-paper program with Dr. G.D. Moody on a new evaluation axis in AI safety.
EQUITRIAGE
fairness audit of gender bias in LLM-based emergency department triage.
CXS INSIGHTS
member records + 26.4M call transcripts vector-indexed. 53 LLM-routed tools in production at UHG.
ASK THAUR
replaced 2–3 FTE analysts (~$150K each) plus $70K/yr compute over 150M+ records with a self-serve conversational UI. 60× latency, sub-3-second warm queries, 80% query-cache hit rate.
ASK RICHARD
data scientists adopted my internal RAG + LLM training platform across UHG / Optum.
OPEN RELEASES
80+ public model releases, 6 datasets, 9 Hugging Face Spaces. Companion artifact for nearly every paper.
KIMI-K2 · MLX
6 bit-depth quantizations (2/3/4/5/6/8-bit) of trillion-parameter Kimi-K2 for Apple Silicon inference.
IRB OVERSIGHT
lives, 70,000+ physicians. Co-Chair, UHG Enterprise IRB (UnitedHealth Group, Optum, Reliant Medical, international affiliates).
Five threads
AI Safety & Evaluation
The science of how foundation models fail.
I run adversarial multi-turn evaluations at frontier scale, measure instruction adherence and guardrail robustness across hundreds of models, and study chain-of-thought faithfulness in open-weight reasoning systems. The benchmarks ship as open datasets so the field can stand on them.
Code Safety & Cybersecurity
A new evaluation axis in AI safety.
Four-paper research program with Dr. Gregory D. Moody (UNLV Lee, Director of Cybersecurity Programs) operationalizing a new evaluation axis: malicious code generation versus defensive security knowledge. A 1,554-prompt consensus-validated benchmark (Fleiss' κ = 0.876), a multi-vendor behavioral study across 10 coding LLMs, a mechanistic test of whether code-safety and content-safety are separable directions in activation space, and a 13-corpus systematic review.
Clinical AI at Deployment Scale
What happens when foundation models meet real patient data.
PHI leakage in medical OCR. Fairness audits of LLM-based emergency-department triage (EQUITRIAGE). Domain-specialized clinical embeddings — CardioEmbed reaches 99.6% retrieval, +15.94pp over the prior state of the art. Paired with Co-Chair oversight of the UHG Enterprise IRB across 110M+ lives.
Production AI Systems
Research without deployment is theater.
I build production conversational-analytics systems on Databricks at UnitedHealth Group: CxS Insights (47.5M members + 26.4M call transcripts, 53 LLM-routed tools), Ask Lucky (member insights, 27 tools, Postgres+pgvector), Ask Thaur (replaced ~$450K/yr of analyst labor + compute with a self-serve conversational UI), and Ask Richard (used daily by 80+ data scientists).
Open-Source Ecosystem
Every paper ships with an artifact.
80+ public model releases, 6 datasets, 9 Spaces across Hugging Face and Ollama. ~33,000 cumulative downloads and pulls. Companion artifacts for nearly every paper I publish — code, data, model weights — so the work is verifiable and reusable.
In flight
- Four-paper code-safety program with Dr. Moody (Paper 1 published; Paper 4 submission-ready; Papers 2 and 3 in preparation)
- Nineteen-paper AI-safety series on reasoning models: chain-of-thought faithfulness, sandbagging detection, MCP protocol safety, steganographic encoding, machine unlearning
- NSF SBIR pitch on EQUITRIAGE in preparation
- The Neuroscience of Artificial Intelligence — 38-section NeuroAI Handbook
- Healthcare Analytics and AI: Building Systems That Actually Work — primary text for UNLV Lee Business School graduate course, Fall 2026
The short version
I'm a Ph.D. computational neuroscientist working at the intersection of foundation-model evaluation and consequential applied domains. I lead AI/ML research at UnitedHealth Group, where I built four production AI systems serving Optum and UHC. I co-chair the UHG Enterprise IRB, with scientific oversight covering UnitedHealth Group, Optum, Reliant Medical, and international affiliates. I'm an Assistant Professor-in-Residence in Information Systems at UNLV's Lee Business School, where I teach business analytics and a Fall 2026 graduate course on healthcare AI. Portland, Oregon and Las Vegas, Nevada.
Warm regards,

Richard Young, Ph.D.
Senior AI Research Scientist · UHG Enterprise IRB Co-Chair · UNLV Assistant Professor-in-Residence
Selected Talks
- Keynote, John Snow Labs Applied AI Healthcare (50,000+ data scientists, 2026)
- Commencement Speaker, Concorde College (May 2026)
- Applied Healthcare AI Summit (“When the Safety Net Becomes the Attack Vector”)
- 35+ invited talks and keynotes total, audiences from 250 to 50,000
Available for
Available for research collaborations, advisory roles, speaking engagements, and faculty conversations.
The best way to reach me is [email protected].