Detecting LLM hallucinations, plus 150+ customer discovery interviews.
A project on detecting when large language models make things up (February to May 2025). I built an evaluation harness that scores whether an LLM’s answer is actually grounded in its source.
Measured on annotated benchmarks.
I took part in 150+ NSF I-Corps interviews with AI leaders, engineers, researchers and product managers. A key finding: enterprises need fast, low-latency guardrails, not slow "LLM-as-a-judge" re-checks, which redirected the design from LLM-as-a-judge re-checks to sub-100 ms scoring based on the model’s own logits and entropy.