Discovering Hierarchical Latent Capabilities of Language Models via Causal Representation Learning
Abstract
Deriving post-training insights from language model evaluations is difficult because benchmark outcomes reflect both base-model differences and downstream adaptation, while controlled studies remain expensive. Public LLM leaderboards contain rich observational evidence about post-training, but this evidence is strongly confounded by base-model family. We study capability transfer: whether improving one capability tends to improve another, and whether that carryover differs across related families. Our goal is to study this question mainly from cheap observational evaluation data rather than from large collections of expensive training runs. We propose Hierarchical Component Analysis (HCA), a family-aware latent-factor method for studying capability transfer from observational data alone. Applied to Llama and Qwen families containing 1500+ models from the Open LLM Leaderboard, HCA identifies a three-factor structure whose most stable pattern is a positive link from an instruction-aligned factor to a math-aligned factor. We then verify this finding via real fine-tuning with external dataset, showing that it is not merely a benchmark artifact. Overall, this work shows that observational data can already suggest concrete post-training predictions.