From Geometry to Behavior: How Instruction Exposure Unlocks Latent Rhetorical Directions in LLMs
Abstract
Language models do not merely memorize facts---they absorb rhetorical strategies from pretraining data, including the human tendency to reason by analogy. We ask: is this absorbed capability geometrically accessible before any instruction tuning, and what determines whether it can be activated at inference time? We extract an "Analogy Vector'' from the residual stream via difference-in-means and show it enables steerable explanation style across seven instruction-tuned models (Gemma-2, Llama-3.1, Qwen-2.5; 2B--32B), improving harmonic mean from 3.22 to 3.94 (+22\%) for Gemma-2-9B-IT on 300 held-out test prompts. Our central finding, from seven matched base-instruct pairs, is a gradient of causal wiring: the analogy direction exists in every base model, but its behavioral accessibility scales continuously with instruction exposure during training. Gemma-2 base models (strict pretraining/IT separation) show flat steering landscapes---present but causally inert. Qwen-2.5 base models (instruction data woven into pretraining) are already partially wired before any explicit fine-tuning. This challenges the binary ``base vs. instruct'' framing: the relevant variable is how much instruction exposure a model has received across its full training pipeline. We position activation steering as a complementary interpretability tool and a training-free inference-time lever for amplifying latent rhetorical strategies---a reflection of human cognitive patterns absorbed from text.