Beyond Semantics: Rediscovering Spatial Awareness in Vision-Language Models
Abstract
Vision and Language Models (VLMs) recognize objects well but often remain weak on spatial reasoning. In VLMs with a projector, spatial information can reach the decoder either through the content of vision tokens or through decoder positional encoding (RoPE), yet random permutation of vision tokens often causes only minor performance change. To measure positional use more directly, we introduce the Position Sensitivity Index (PSI), a RoPE sensitivity probe, and the 2D Synthetic Spatial (2DS) benchmark, a controlled benchmark without semantic shortcuts. We trace the weak use of RoPE to a large norm imbalance between vision and text at the projector interface, and find that this imbalance suppresses positional effects in the decoder. A broader survey of twelve models shows that this imbalance is common across VLMs with a projector. Our targeted interventions then reveal a dissociation: balancing norms increases positional reliance, while enriching the content of vision tokens yields even higher spatial accuracy, with lower PSI. These results show that RoPE in VLMs with a projector is suppressed rather than totally absent, and that positional reliance and spatial accuracy can be improved through independent mechanisms.