Restoring Generalization in Fine-tuned Multimodal LLMs via Geometric Alignment
Abstract
Fine-tuning Multimodal Large Language Models for downstream tasks requires balancing task adaptation with retention of broad pre-trained capabilities. In practice, full fine-tuning often improves in-distribution performance but degrades zero-shot behavior, while parameter-efficient tuning is more stable yet may leave a substantial adaptation gap. We present \textbf{Restoring Generalization Alignment (ReGA)}, a simple two-stage post-tuning procedure. ReGA first applies standard LoRA to obtain a task-adapted checkpoint, and then performs a short alignment stage that combines a low-loss path constraint with a soft anchor to the zero-shot model. This second stage encourages a nearby solution that preserves task performance while remaining closer to the pre-trained operating point. Experiments on LLaVA-1.5, Qwen2.5-VL, and InternVL-3.5 show that ReGA consistently improves the trade-off between in-distribution adaptation and broad zero-shot generalization relative to standard PEFT, regularized baselines, and merging-based baselines. The same recipe also remains effective in replay-free continual learning.