Manifold Steering Reveals the Shared Geometry of Neural Network Representation and Behavior
Abstract
What is the geometric relationship between neural network beliefs---predictive distributions over latent concepts---and representations? We fit manifolds to geometric structures in activation space and belief space and explore this question via intervention. First, we show that interventions that steer along the representation manifold yield output trajectories that are natural (follow the belief manifold). In contrast, linear interpolation produces unnatural trajectories. Second, we use the geometry of belief space to optimize interventions in activation space; finding paths whose induced belief trajectories follow paths along the belief manifold. We show that these pullback-optimal activation paths follow the representation manifold more closely than a linear path. We demonstrate these findings for reasoning over cyclic concepts with circular geometries, e.g., "What is three days after Tuesday?", and in-context learning tasks over graph geometries. In sum, our results show that representation geometry provides a blueprint for control protocols via interventions that respect belief geometry.