Grounding latent algorithm routing in transformer reasoning
Abstract
A central question in the in-context learning literature is whether transformers rely on a single broad meta-learner or instead first infer which inductive-bias family is appropriate for the current episode. We study the latter hypothesis, which we call latent algorithm routing. We argue that routing should be established not by answer accuracy alone, but by three stronger criteria: preferred solver-family identity should change when the latent data-generating regime changes while prompt form is held fixed, remain stable under nuisance formatting perturbations, and be selectively altered by targeted activation interventions without collapsing answer quality. To study these conditions, we introduce ROUTEBENCH, a controlled benchmark in which matched prompt syntax is paired with latent regimes that differentially favor global shrinkage, sparsity, robustness, and locality. In experiments, we operationalize these inductive-bias families using ridge-like, lasso-like, Huber-like, and kNN-like family representatives. Across decoder-only transformers from 44M to 612M parameters, we find converging evidence for routing: models shift family preference under true regime changes while remaining relatively invariant to formatting changes; a 306M model closes 80.6% of the oracle routing gap and achieves route F1 84.1; and probes and activation patching recover and causally flip route variables while preserving over 96% answer quality. Together, these results support latent routing over canonical solver families rather than a monolithic in-context learner.