用语言指令控制多模态模型依赖的证据,提升推理可靠性。
GUIDE: Guiding Internal Evidence with Language Instructions

- 通过指令调控证据路径,让模型按需选择信息源。
- 在多个数据集上增强对特定证据扰动的鲁棒性。
- 适合研究多模态可解释性与可控推理的学者。
大型多模态模型能遵循生成指令,但未必遵循证据使用指令,仍可能依赖捷径线索。我们提出GUIDE框架,通过语言指令控制内部证据使用。GUIDE结合分组参数高效适配与指令条件门控,在推理和生成中调节多模态证据路径。我们还引入路径级评估框架,通过依赖敏感性、受控扰动分析、路径调制和自回归解码动态,刻画指令引导的证据调节。在GQA、TextVQA、MM-IMDb、CREMA-D、RAVDESS和Flickr30K上的实验表明,GUIDE实现了结构化且与指令对齐的证据依赖重分配,同时基本保持任务性能。它在多种多模态场景下提升了对抗性证据扰动的鲁棒性,并实现可控调节,表明多模态指令遵循可从输出控制扩展至调节不同证据源的贡献方式。
原文摘要 · Abstract (English)
Large multimodal models follow instructions about what to generate, but not necessarily about what evidence to rely on. Hence, models may continue to depend on shortcut-associated cues even when instructions suggest otherwise. We introduce GUIDE, a framework for controlling internal evidence usage through language instructions. GUIDE combines grouped parameter-efficient adaptation with instruction-conditioned gating to modulate multimodal evidence pathways during reasoning and generation. We further introduce a pathway-level evaluation framework that characterizes instruction-conditioned evidence modulation through reliance sensitivity, controlled perturbation analysis, pathway modulation, and autoregressive decoding dynamics. Across multimodal reasoning, classification, and generation, GUIDE induces structured and instruction-aligned redistribution of evidence reliance while largely preserving task behavior. Experiments on GQA, TextVQA, MM-IMDb, CREMA-D, RAVDESS, and Flickr30K show that GUIDE improves robustness under targeted evidence perturbations and enables controllable modulation across diverse multimodal settings. This suggests that multimodal instruction following can extend beyond output control toward regulating how different evidence sources contribute to model predictions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。