The work

Research.

AI + Safety

Five threads, one pattern. The paper archive lives in the Library.

Active threads · 5

Five threads.

AI Safety & Evaluation

13 papers

The science of how foundation models fail.

I run adversarial multi-turn evaluations at frontier scale, measure instruction adherence and guardrail robustness across hundreds of models, and study chain-of-thought faithfulness in open-weight reasoning systems. The benchmarks ship as open datasets so the field can stand on them.

Code Safety & Cybersecurity

4 papers

A new evaluation axis in AI safety.

Four-paper research program with Dr. Gregory D. Moody (UNLV Lee, Director of Cybersecurity Programs) operationalizing a new evaluation axis: malicious code generation versus defensive security knowledge. A 1,554-prompt consensus-validated benchmark (Fleiss' κ = 0.876), a multi-vendor behavioral study across 10 coding LLMs, a mechanistic test of whether code-safety and content-safety are separable directions in activation space, and a 13-corpus systematic review.

Clinical AI at Deployment Scale

19 papers

What happens when foundation models meet real patient data.

PHI leakage in medical OCR. Fairness audits of LLM-based emergency-department triage (EQUITRIAGE). Domain-specialized clinical embeddings — CardioEmbed reaches 99.6% retrieval, +15.94pp over the prior state of the art. Paired with Co-Chair oversight of the UHG Enterprise IRB across 110M+ lives.

Production AI Systems

5 metrics

Research without deployment is theater.

I build production conversational-analytics systems on Databricks at UnitedHealth Group: CxS Insights (47.5M members + 26.4M call transcripts, 53 LLM-routed tools), Ask Lucky (member insights, 27 tools, Postgres+pgvector), Ask Thaur (replaced a $5–6K/mo + 2–3 FTE workflow with a $100/mo conversational UI), and Ask Richard (used daily by 80+ data scientists).

Open-Source Ecosystem

4 metrics

Every paper ships with an artifact.

80+ public model releases, 6 datasets, 9 Spaces across Hugging Face and Ollama. ~33,000 cumulative downloads and pulls. Companion artifacts for nearly every paper I publish — code, data, model weights — so the work is verifiable and reusable.

See the full archive of 34 papers in the Library, grouped by year with DOI and PDF links.