提出抗幻觉的图像描述评估方法,提升自动评价准确性。
DENEB: A Hallucination-Robust Automatic Evaluation Metric for Image Captioning
- 用Sim-Vec Transformer并行处理多参考描述,捕捉图文一致性。
- 在6个数据集上超越现有无LLM指标,最高提升12.3%。
- 适合关注生成质量与真实性的图像描述研究者使用。
本文针对图像描述自动评估中幻觉敏感的问题,提出新型无大模型依赖的评估指标DENEB。现有指标因难以处理多参考描述而易受幻觉干扰。DENEB引入Sim-Vec Transformer,可同时处理多个参考描述,有效捕捉候选描述与参考描述之间的语义相似性。为训练模型,构建包含32,978张图像和805名标注者的Nebula数据集。实验表明,DENEB在FOIL、Composite、Flickr8K-Expert、Flickr8K-CF、Nebula和PASCAL-50S六个数据集上均达到当前无大模型指标最优性能,验证其对幻觉的强鲁棒性。
原文摘要 · Abstract (English)
In this work, we address the challenge of developing automatic evaluation metrics for image captioning, with a particular focus on robustness against hallucinations. Existing metrics are often inadequate for handling hallucinations, primarily due to their limited ability to compare candidate captions with multifaceted reference captions. To address this shortcoming, we propose DENEB, a novel supervised automatic evaluation metric specifically robust against hallucinations. DENEB incorporates the Sim-Vec Transformer, a mechanism that processes multiple references simultaneously, thereby efficiently capturing the similarity between an image, a candidate caption, and reference captions. To train DENEB, we construct the diverse and balanced Nebula dataset comprising 32,978 images, paired with human judgments provided by 805 annotators. We demonstrated that DENEB achieves state-of-the-art performance among existing LLM-free metrics on the FOIL, Composite, Flickr8K-Expert, Flickr8K-CF, Nebula, and PASCAL-50S datasets, validating its effectiveness and robustness against hallucinations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。