Researchers find sparse neuron subsets can predict factual hallucination in several open models
A Tsinghua research team found that very small, model-specific groups of feed-forward neurons predicted factual hallucinations across several tests in six open-weight models. Later studies suggest that the signal is domain-specific and is not, by itself, a reliable control handle.
The Reality Check
Context
The H-Neurons study used sparse logistic probes over neuron-contribution features to distinguish faithful and hallucinated answers. Across six open-weight models and several factual-question settings, selected subsets usually outperformed matched random subsets. The reported subsets were approximately 0.001–0.035% of feed-forward neurons, depending on the model and test. The same indices retained predictive signal when applied to corresponding base models, which supports persistence of the signal across model stages but does not prove a causal origin.
The study’s activation-scaling experiments changed behaviours such as accepting false premises, following misleading context, sycophancy, and harmful compliance. They did not establish that scaling those neurons changes factual hallucination rates. An independent cross-domain study found mean AUROC of 0.783 within a trained domain but 0.563 after transfer to another domain, indicating that detectors require domain-specific calibration. A separate medical study found that hallucination signals can be distributed and redundant, with detectability not reliably translating into neuron-level control.
A detector can read a signal without controlling the behaviour that produced it.
Earlier work on internal truthfulness, entity awareness, and shared uncertainty circuits is consistent with the broader idea that models encode information about factual reliability internally, but it does not establish a universal hallucination circuit.
THE TAKEAWAY
The evidence supports a useful, model-specific measurement result: sparse internal features can help predict factual hallucination in tested settings. It does not establish a universal circuit, show that fewer than 0.1% of neurons are solely responsible for hallucinations, or provide a reliable correction mechanism. Detection should be evaluated by domain, and factual reliability still requires external verification and task-specific testing.
Continue the Thread
AI Model HallucinationsTracks evidence about why AI models produce factually incorrect or unsupported outputs and how reliably those failures can be detected, predicted, reduced, or prevented.
Sources
H-Neurons: On the Existence, Impact, and Origin of Hallucination-Associated Neurons in LLMs
arXiv / Tsinghua University researchers
Used for: Model set; neuron-selection method; training-data construction; selected-neuron ratios; detection results; intervention benchmarks; base-model transfer; parameter-drift analysis; and acknowledged mitigation trade-offs.
Do Hallucination Neurons Generalize? Evidence from Cross-Domain Transfer in LLMs
arXiv / independent researchers
Used for: Independent reconstruction of H-Neuron detection; within-domain and cross-domain AUROC; model and domain coverage; and the negative factual-hallucination activation-scaling result.
H-Neurons official implementation
THUNLP / GitHub
Used for: Inspectable data collection, answer-token extraction, CETT activation extraction, sparse-classifier and intervention code, example data, and reproducibility boundaries.
Do I Know This Entity? Knowledge Awareness and Hallucinations in Language Models
ICLR 2025
Used for: Previous sparse-autoencoder frontier for entity recognition, causal steering of refusal and hallucination, and reuse of base-model features after chat fine-tuning.
LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations
ICLR 2025
Used for: Previous hidden-state detection frontier, concentration of truthfulness information at particular tokens, and earlier evidence that detectors fail to generalize universally across datasets.
Readable but Not Controllable: Neuron-Level Evidence for Medical LLM Hallucination
arXiv / independent researchers
Used for: Separate medical-domain evidence for strong neuron-level detectability, distributed and redundant predictive signals, and the gap between decoding and reliable intervention.
SCIURus: Shared Circuits for Interpretable Uncertainty Representations in Language Models
NAACL 2025 / ACL Anthology
Used for: Evidence that factual recall and expressed uncertainty rely on overlapping network regions across eight models and five datasets, without establishing an impossibility result for self-verification.
Why Language Models Hallucinate
arXiv / OpenAI researchers and collaborators
Used for: Learning-theoretic account of factual errors during next-token pre-training and the separate role of post-training evaluation incentives that reward guessing.
Last checked Methodology 2.0.0