arXiv:2509.13919cs.CV2025-09ICML被引 1

让视觉语言模型的推理过程与答案对齐,提升回答准确性。

Towards Rationale-Answer Alignment of LVLMs via Self-Rationale Calibration

  • 通过自推理校准框架,自动调整模型输出的推理与答案一致性。
  • 在多个基准上显著提升感知、推理与泛化能力,表现优于基线。
  • 适合关注模型可解释性与推理质量的研究者和开发者。

大型视觉语言模型(LVLMs)展现出强大的视觉问答能力,但仍存在推理过程与答案不一致的问题。本文提出自推理校准(SRC)框架,通过轻量级推理微调使模型在无显式提示下先生成推理再得出答案。随后,从微调后的模型中采样多种候选回答,并利用专用评分模型R-Scorer评估其推理质量与事实一致性。基于置信度加权偏好筛选,将对齐校准转化为偏好微调方式,在多个基准上显著提升模型在感知、推理与泛化能力的表现。结果强调了以推理为导向对齐在挖掘LVLM潜力中的重要性。

原文摘要 · Abstract (English)

Large Vision-Language Models (LVLMs) have manifested strong visual question answering capability. However, they still struggle with aligning the rationale and the generated answer, leading to inconsistent reasoning and incorrect responses. To this end, this paper introduces the Self-Rationale Calibration (SRC) framework to iteratively calibrate the alignment between rationales and answers. SRC begins by employing a lightweight "rationale fine-tuning" approach, which modifies the model's response format to require a rationale before deriving an answer without explicit prompts. Next, SRC searches for a diverse set of candidate responses from the fine-tuned LVLMs for each sample, followed by a proposed pairwise scoring strategy using a tailored scoring model, R-Scorer, to evaluate both rationale quality and factual consistency of candidates. Based on a confidence-weighted preference curation process, SRC decouples the alignment calibration into a preference fine-tuning manner, leading to significant improvements of LVLMs in perception, reasoning, and generalization across multiple benchmarks. Our results emphasize the rationale-oriented alignment in exploring the potential of LVLMs.

视觉语言模型推理对齐模型校准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。