Researchers find sparse neuron subsets can predict factual hallucination in several open models
A Tsinghua research team found that very small, model-specific groups of feed-forward neurons predicted factual hallucinations across several tests in six open-weight models. Later studies suggest that the signal is domain-specific and is not, by itself, a reliable control handle.