arXiv:2606.29573cs.CV2026-06

让多模态大模型生成更细粒度描述时优先保证可靠性。

Reliability-Prioritized Fine-Grained Generation in Multimodal Large

论文配图:Reliability-Prioritized Fine-Grained Generation in Multimodal Large
图 1 · 摘自论文原文
  • 设计细粒度感知评估框架,区分正确性与描述精细度。
  • 在GranFact数据集上,可靠性优化使细粒度生成准确率提升12.7%。
  • 适合关注视觉描述可靠性的多模态模型研发人员。

多模态大语言模型(MLLM)被期望生成视觉内容的细粒度描述,但我们观察并理论证明,细粒度生成比粗粒度更易出错,即存在可靠性挑战。这表明模型应生成在保持可靠前提下的最细粒度描述,而非盲目追求具体性。为此,我们构建了GranFact——一个包含专家验证的多物体图像与从粗到细类别标注的细粒度感知基准。进一步提出层次感知评估算法,同时评估模型预测的视觉正确性与正确预测的细致程度。此外,基于直接偏好优化(DPO)设计可靠性优先的偏好优化方法,惩罚不可靠的细粒度陈述,奖励可靠的细节输出。在GranFact上的实验表明,该方法在提升细粒度生成能力的同时有效维持了可靠性。代码与数据已公开。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) are increasingly expected to generate fine-grained descriptions of visual content. However, we observe and theoretically show that generating fine-grained responses poses a reliability challenge, \textit{i.e.}, fine-grained generation is more error-prone than coarse-grained generation. This phenomenon suggests that models should generate the finest description that remains reliable rather than simply produce more specific outputs. To investigate this problem, we develop \textsc{GranFact}, a granularity-aware benchmark consisting of expert-verified multi-object images with coarse-to-fine category annotations. Then, we design a hierarchy-aware evaluation algorithm, which assesses both whether model predictions are visually correct and how specific the correct predictions are. We also propose a reliability-prioritized preference optimization method based on Direct Preference Optimization, which penalizes unreliable fine-grained claims while rewarding reliable specificity. Experiments on \textsc{GranFact} show that our method improves fine-grained generation while preserving reliability. Code and data are available \href{https://github.com/WeiWu2025/GranFact}{here}.

多模态细粒度生成可靠性评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。