MLPs are Hebbians: Constructing Efficient Fact-Storing MLPs for Transformers
Abstract
Large language models (LLMs) store factual knowledge in their parameters, yet the underlying mechanism remains unclear. Recent work has found that factual knowledge resides in LLM MLPs as key-value mappings and has provided closed-form fact-storing MLP constructions. However, existing constructive and mechanistic interpretability approaches fail to capture three empirically observed properties of MLPs in LLMs: optimal fact-storage capacity scaling, ability to handle arbitrary embedding geometries, and usability within Transformer blocks. Towards understanding fact-storage in LLMs, we analyze the decoding margin of MLPs, whereas prior work only studies MLP fact storage. Our analysis allows us to develop the first closed-form MLP construction that realizes all three properties. Under isotropic embeddings, our construction achieves optimal storage capacity scaling, reducing the capacity gap to trained MLPs by 8-16x over prior constructions. Further, our theoretical margin scaling bounds precisely characterize how embedding geometry governs empirical margin (R^2 > 0.95) and illustrate our construction’s ability to store fact-sets under arbitrary embedding geometries. Moreover, we show that MLPs can be used within Transformer blocks for factual recall tasks at optimal capacity scaling, reducing the capacity gap to trained MLPs by up to 100x over prior constructions. Finally, as a proof-of-concept, we show that fact-storing MLPs enable modular fact editing by swapping a Transformer's MLP with a new one.