TRAPSBench: Vision-Language Models Encode but Fail to Express Epistemic Restraint
Fnu pramono ⋅ John Cai ⋅ Sourabh Kulkarni
Abstract
When visual evidence is occluded or chaotic, Vision-Language Models (VLMs) should abstain rather than fabricate outcomes. We introduce \textbf{TRAPSBench}, a procedurally generated video benchmark of matched answerable/unanswerable physics pairs, and the \textbf{Penalized Epistemic Calibration Score (PECS)}, a text-output metric that jointly penalizes overconfidence and false abstention. Evaluating 15 VLMs across three prompt regimes, we find: (1)spontaneous epistemic restraint is poor (best PECS\,$=$\,0.165); (2)models possess latent epistemic capability that guided prompting unlocks (median $3.7\times$ AbsRec gain) but fail to exercise spontaneously; (3)VLMs detect textual impossibility much more readily than visual information gaps (median $4{-}7\times$ under standard prompting); and (4)chain-of-thought reasoning can \emph{degrade} calibration by amplifying confabulations rather than catching them. Probing Qwen3-VL-8B's hidden states reveals a domain-general epistemic signal that the model encodes but fails to surface behaviorally. Activation steering confirms this gap is \emph{causal}: injecting the learned void direction at a single layer controls abstention across physics domains, but a unidirectional output constraint under standard prompting blocks abstention induction while permitting suppression. This representation--output gap---not a lack of perceptual understanding---is the core bottleneck for safe VLM deployment.
Successful Page Load