Wiener Filtering for VLM Hallucination Suppression
Abstract
Vision-language models (VLMs) excel at open-ended captioning and visual QA but often describe objects, attributes, or relations absent from the image, a phenomenon known as object hallucination. We propose a training-free, post-hoc correction that operates in the representation space of the language backbone. By modeling hidden states as a superposition of truthful and hallucination-associated components, we derive a Wiener-type estimator whose optimal gains are given in closed form from the covariances of paired truthful and hallucinated representations. An eigendecomposition yields mode-wise attenuation that respects a stability criterion, i.e., the filter responds continuously to estimation noise. The correction is applied once to the feed-forward output projections of selected deeper layers; at inference time, the model runs unchanged and at the same speed. Experiments on LLaVA-1.5, MiniGPT-4, and mPLUG-Owl2 demonstrate consistent reductions in object hallucination on CHAIR, POPE, and MME while maintaining caption fluency and overall response quality. To further show the generality of our approach beyond autoregressive models, we also evaluate it on discrete diffusion language models for grounded dialogue, demonstrating that representation filtering reduces hallucinations even in multi-step, sequence-wide denoising settings.