arXiv:2602.12533cs.LG2026-02被引 3

让多模态模型按需调整视觉与文本偏好,避免过度依赖任一模态。

AMPS: Adaptive Modality Preference Steering via Functional Entropy

  • 基于功能熵设计动态诊断指标,识别每样本对调优的敏感度。
  • 针对敏感样本自动降低调优强度,保持生成错误率稳定。
  • 可学习模块实现个性化调优,适合需要精准控制的场景。

多模态大语言模型常表现出显著的模态偏好,即在输入变化时过度依赖语言先验或视觉显著性,而忽视文本事实。现有方法采用统一强度调优,但过强会损害标准推理并提升错误率,过弱则无效。由于不同样本对调优的敏感性差异大,全局固定强度难以适配。为此,我们提出一种实例感知的诊断指标,量化各模态的信息贡献,并揭示样本对调优的敏感性。基于此,设计自适应缩放策略,对敏感样本降低调优强度,并引入可学习模块推断缩放模式,实现实例级模态偏好控制。实验表明,该方法在调节模态偏好方面优于传统方法,在保持低生成错误率的前提下实现有效调优。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) often exhibit significant modality preference, which is a tendency to favor one modality over another. Depending on the input, they may over-rely on linguistic priors relative to visual evidence, or conversely over-attend to visually salient but facts in textual contexts. Prior work has applied a uniform steering intensity to adjust the modality preference of MLLMs. However, strong steering can impair standard inference and increase error rates, whereas weak steering is often ineffective. In addition, because steering sensitivity varies substantially across multimodal instances, a single global strength is difficult to calibrate. To address this limitation with minimal disruption to inference, we introduce an instance-aware diagnostic metric that quantifies each modality's information contribution and reveals sample-specific susceptibility to steering. Building on these insights, we propose a scaling strategy that reduces steering for sensitive samples and a learnable module that infers scaling patterns, enabling instance-aware control of modality preference. Experimental results show that our instance-aware steering outperforms conventional steering in modulating modality preference, achieving effective adjustment while keeping generation error rates low.

多模态模型调优偏好控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。