Research and engineering notes on hallucination detection and agent-native verification.
Research · 40 min read · August 2026
How often language models make things up, why scale didn't fix it, what fabricated answers already cost in court, every detection method compared — and why the most reliable check is the one your model can't run on itself.
Read the guide →