Grammatical ``grandmother neurons'' are rare in LLMs
Linyang He ⋅ Nima Mesgarani
Abstract
Understanding how Large Language Models (LLMs) encode linguistic structures remains a fundamental challenge in interpretability research. While diagnostic classifiers (or "probes") are the standard tool for this task, they face significant methodological criticism: training auxiliary classifiers introduces capacity confounds and calibration issues, often making it difficult to distinguish the model's intrinsic representations from the probe's ability to learn the task. To address these limitations, we introduce a probe-free framework for localizing linguistic selectivity at the individual neuron level. Leveraging the controlled contrasts of linguistic minimal pairs, we propose a Neuron Separability Index (NSI), a metric that directly quantifies how reliably single neurons differentiate grammatical from ungrammatical constructions without parameter updates. By applying this framework to the Qwen3 model, we found: (1) the layer-wise pattern of raw effect sizes closely matches standard probing results, revealing an earlier peak for syntactic distinctions than for semantic ones; (2) after permutation normalization, the metric indicates that grammaticality-sensitive neurons are extremely sparse, accounting for only 3.6\% of neurons; (3) we find little evidence for grammatical ``grandmother neurons'': neurons with very strong selectivity (e.g., $z$-score $> 2$) are essentially absent; and (4) selectivity is highly domain-specific, with genuinely highly-selective neurons rare and few neurons generalizing across paradigms.
Successful Page Load