提出新评估框架与反馈方法,显著提升图像描述细节准确性。
Painting with Words: Elevating Detailed Image Captioning with Benchmark and Alignment Learning
- 将描述拆解为最小信息单元,精准评估细节与幻觉。
- 新指标DCScore更贴近人工判断,优于现有评估方式。
- 适用于优化视觉语言模型,尤其适合追求高精度描述的场景。
图像描述是视觉理解的关键任务,近年来视觉-语言模型(VLMs)在生成详细描述方面取得显著进展。然而,由于评估指标陈旧和标注粗略,详细图像描述的评估仍不充分。本文提出DeCapBench及新指标DCScore,专为详细描述任务设计。DCScore通过将回答分解为最小自洽单元——原始信息单元,并逐个评估其真实性与细粒度完整性,有效识别幻觉并衡量描述全面性。实验表明,DCScore比其他基于规则或模型的指标更贴近人类判断。同时,DeCapBench与VLM竞技场在描述任务上高度相关,超越现有基准。此外,我们提出自动细粒度反馈收集方法FeedQuill,基于新指标进行偏好优化,在自动生成偏好数据上表现出强泛化能力。多组VLM实验证明,该方法显著降低幻觉,提升多个基准表现,整体性能超越GPT-4o。
原文摘要 · Abstract (English)
Image captioning has long been a pivotal task in visual understanding, with recent advancements in vision-language models (VLMs) significantly enhancing the ability to generate detailed image captions. However, the evaluation of detailed image captioning remains underexplored due to outdated evaluation metrics and coarse annotations. In this paper, we introduce DeCapBench along with a novel metric, DCScore, specifically designed for detailed captioning tasks. DCScore evaluates hallucinations and fine-grained comprehensiveness by deconstructing responses into the smallest self-sufficient units, termed primitive information units, and assessing them individually. Our evaluation shows that DCScore aligns more closely with human judgment than other rule-based or model-based metrics. Concurrently, DeCapBench exhibits a high correlation with VLM arena results on descriptive tasks, surpassing existing benchmarks for vision-language models. Additionally, we present an automatic fine-grained feedback collection method, FeedQuill, for preference optimization based on our advanced metric, showing robust generalization capabilities across auto-generated preference data. Extensive experiments on multiple VLMs demonstrate that our method not only significantly reduces hallucinations but also enhances performance across various benchmarks, achieving superior detail captioning performance while surpassing GPT-4o.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。