AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

NeuroActiSep: Detecting Factual Hallucinations from Feed-Forward Neurons in a Single Pass

arXiv · AI, language, vision and robotics · article · Sep 13, 2026 · UTC

Hallucination in large language models reduces their reliability and slows adoption. Various white-box studies have used internal representations to detect patterns of truthfulness and factuality. A less-studied approach is to identify feed-forward neurons correlated with hallucination. We propose a method to rank feed-forward neurons at the final prompt token using a custom neuron selection dataset. We transfer the selected neuron identities to train hallucination classifiers on other factual question answering datasets. Our work provides empirical evidence that probes trained using the featu

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T12:21:05.240Z. This is not the publication date.