arXiv:2507.19262cs.CV2025-07EMNLP被引 1

提出无需人工标注的长图文生成事实性评估方法

OVFact: Measuring and Improving Open-Vocabulary Factuality for Long Caption Models

  • 基于开放词汇视觉定位与工具验证,实现无参考文本的事实性评估
  • 在多个长图文基准上,筛选后训练模型事实性提升2.5至5倍
  • 适用于数据过滤,适合关注生成内容真实性的研究者

大型视觉语言模型在生成长篇、准确的描述时表现不佳。传统幻觉与事实性评估方法不适用于更长、更多样化的描述,且依赖人工标注的真值。我们提出OV-Fact,一种新型长图文事实性评估方法,利用开放词汇视觉定位和工具化验证,无需依赖人工标注。该方法提升了与人类判断的一致性,同时衡量描述完整性和事实精确性。相比以往方法,其无参考设计可支持基于事实性的数据过滤。我们在一个大规模、噪声较多(由VLM生成)的预训练数据集中,使用OV-Fact筛选出2.5至5倍更少的样本进行训练,发现模型在多种下游长图文任务中,事实性精度显著提升,且未损失描述完整性。

原文摘要 · Abstract (English)

Large vision-language models (VLMs) often struggle to generate long and factual captions. However, traditional measures for hallucination and factuality are not well suited for evaluating longer, more diverse captions and in settings where ground-truth human-annotated captions are unavailable. We introduce OV-Fact, a novel method for measuring caption factuality of long captions that leverages open-vocabulary visual grounding and tool-based verification without depending on human annotations. Our method improves agreement with human judgments and captures both caption descriptiveness (recall) and factual precision in the same metric. Furthermore, unlike previous metrics, our reference-free method design enables new applications towards factuality-based data filtering. We observe models trained on an OVFact-filtered (2.5-5x less) subset of a large-scale, noisy (VLM-generated) pretraining set meaningfully improve factuality precision without sacrificing caption descriptiveness across a range of downstream long caption benchmarks.

视觉语言模型事实性评估数据过滤长图文生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。