arXiv:2505.14627cs.AIcs.CL2025-05

让弱模型通过辩论,提升强视觉语言模型的推理能力。

Debating for Better Reasoning: An Unsupervised Multimodal Approach

  • 双专家模型在图像问答中辩论答案,文本裁判仅评估论据质量。
  • 辩论机制使模型表现超越单个专家,且弱模型也能提升强模型。
  • 无需角色扮演,聚焦分歧点,适合低资源场景下的模型优化。

随着大型语言模型在多领域和多模态任务中能力增强,可扩展的监督机制面临挑战,尤其当模型能力超过人类评估者时。辩论成为一种有前景的监督方式。本文将辩论范式拓展至多模态场景,探索弱模型如何监督并提升强模型性能。以视觉问答(VQA)为例,两个具备视觉能力的专家模型就答案展开辩论,而一个仅能处理文本的“盲判官”依据论证质量进行裁决。专家仅辩护其信念一致的答案,避免了显式的角色扮演,使辩论集中于专家间分歧。实验表明,该框架在多个多模态任务上持续优于单个专家模型。此外,弱模型的判断可通过微调帮助视觉语言模型习得更优推理能力。

原文摘要 · Abstract (English)

As Large Language Models (LLMs) gain expertise across diverse domains and modalities, scalable oversight becomes increasingly challenging, particularly when their capabilities may surpass human evaluators. Debate has emerged as a promising mechanism for enabling such oversight. In this work, we extend the debate paradigm to a multimodal setting, exploring its potential for weaker models to supervise and enhance the performance of stronger models. We focus on visual question answering (VQA), where two "sighted" expert vision-language models debate an answer, while a "blind" (text-only) judge adjudicates based solely on the quality of the arguments. In our framework, the experts defend only answers aligned with their beliefs, thereby obviating the need for explicit role-playing and concentrating the debate on instances of expert disagreement. Experiments on several multimodal tasks demonstrate that the debate framework consistently outperforms individual expert models. Moreover, judgments from weaker LLMs can help instill reasoning capabilities in vision-language models through finetuning.

多模态辩论推理增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。