arXiv:2608.04509cs.AI2026-08

让视觉语言模型在矛盾证据中做出可靠判断并适时放弃回答

CARGO-VL: Counterfactual Arbitration with Risk-Constrained Group Optimization for Vision-Language Models

论文配图:CARGO-VL: Counterfactual Arbitration with Risk-Constrained Group Optimization for Vision-Language Models
图 1 · 摘自论文原文
  • 通过联合优化四种证据状态,实现对多模态信息的统一判断
  • 在冲突场景下减少错误回答,同时避免过度回避,提升决策平衡性
  • 适合追求模型可靠性与鲁棒性的研究者和工业应用

视觉语言系统融合图像与检索文本,但二者可能矛盾或均无法支持答案。可靠模型需识别可信来源并在无足够依据时放弃回答。现有后训练目标独立评分,无法保证反事实变化下的行为一致性。本文提出CARGO-VL,一种面向四类证据状态(对齐、图像正确、文本正确、均错误)的组相对框架,其目标耦合条件正确性与转移奖励,实现答案不变性、源等变性及答案到弃答切换;同时引入原语义-双元控制器,在安全回答与过度退避间动态平衡。我们还构建了XMC(eXtended Modal Conflict)作为四条件冲突训练数据集,并在CMC-Bench和Modality-Bias上验证迁移效果。多种子实验表明,相比点式基线,CARGO-VL显著提升冲突处理能力、非支持回答规避率与模态均衡性。消融实验验证了关系转移信号与自适应风险控制的互补优势,支持反事实一致性作为可靠多模态证据仲裁的实用目标。

原文摘要 · Abstract (English)

Vision-language systems combine images with retrieved text, but these sources can disagree or jointly fail to support an answer. Reliable models must identify the trustworthy source and abstain when neither is adequate. Existing post-training objectives score instances independently and therefore do not enforce coherent behavior under counterfactual evidence changes. We introduce CARGO-VL, a group-relative framework that optimizes matched variants covering aligned, image-correct, text-correct, and both-wrong (A/V/T/N) evidence states as one bundle. Its objective couples condition-wise correctness with transition rewards for answer invariance, source equivariance, and answer-to-abstention switching, while a primal-dual controller balances unsafe answers against excessive deferral. We also contribute XMC (eXtended Modal Conflict), a four-condition conflict training resource, and evaluate transfer on CMC-Bench and Modality-Bias. Across multiple seeds, CARGO-VL improves conflict handling, unsupported-answer avoidance, and modality balance over pointwise baselines. Ablations identify complementary benefits from relational transition signals and adaptive risk control, supporting counterfactual consistency as a practical objective for reliable multimodal evidence arbitration.

多模态模型可靠性视觉语言反事实学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。