构建首个综合评估图像细粒度描述质量与偏见的基准榜单
LOTUS: A Leaderboard for Detailed Image Captioning from Quality to Societal Bias and User Preferences
- 设计多维度评估体系,覆盖描述准确性、细节丰富度与社会偏见
- 发现模型越详细描述,越易出现幻觉与性别偏见
- 支持按用户偏好定制评估,适配不同应用场景需求
大型视觉语言模型(LVLMs)已将图像描述从简洁文本转向详尽描述。我们提出LOTUS,一个用于评估详细图像描述的基准榜单,旨在解决现有评估体系三大不足:缺乏标准化标准、缺乏偏见感知评估、忽略用户偏好。LOTUS全面评估多个方面,包括描述质量(如对齐度、描述性)、风险(如幻觉)和社会偏见(如性别偏见),同时支持根据用户偏好定制评估标准。对近期LVLMs的分析显示,没有单一模型在所有指标上表现最优,且描述详细程度与偏见风险之间存在相关性。偏好导向评估表明,最优模型选择取决于用户的具体需求。
原文摘要 · Abstract (English)
Large Vision-Language Models (LVLMs) have transformed image captioning, shifting from concise captions to detailed descriptions. We introduce LOTUS, a leaderboard for evaluating detailed captions, addressing three main gaps in existing evaluations: lack of standardized criteria, bias-aware assessments, and user preference considerations. LOTUS comprehensively evaluates various aspects, including caption quality (e.g., alignment, descriptiveness), risks (\eg, hallucination), and societal biases (e.g., gender bias) while enabling preference-oriented evaluations by tailoring criteria to diverse user preferences. Our analysis of recent LVLMs reveals no single model excels across all criteria, while correlations emerge between caption detail and bias risks. Preference-oriented evaluations demonstrate that optimal model selection depends on user priorities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。