arXiv:2504.01225cs.CLcs.AI2025-04ACL被引 2

用置信控制框架提升CLIPScore的细粒度错误检测与不确定性校准

A Conformal Risk Control Framework for Granular Word Assessment and Uncertainty Calibration of CLIPScore Quality Estimates

  • 基于模型无关的置信风险控制,生成可校准的CLIPScore分布
  • 简单掩码方法生成的分数分布性能媲美复杂方法
  • 能精准定位错误词,且不确定性估计更贴近真实误差

本研究探讨现有图像描述评价指标的局限性:无法对描述中的错误进行细粒度评估,且仅依赖单点质量估计而忽略不确定性。为此,我们提出一种简单但有效的策略,用于生成并校准CLIPScore值的分布。通过模型无关的置信风险控制框架,对特定任务变量校准CLIPScore,解决上述问题。实验表明,使用置信风险控制,仅通过输入掩码等简单方法生成的分数分布,性能可媲美更复杂的方案。该方法能有效识别错误词汇,同时提供符合预定风险水平的形式化保障。此外,其不确定性估计与预测误差的相关性显著提升,增强了描述评价指标的整体可靠性。

原文摘要 · Abstract (English)

This study explores current limitations of learned image captioning evaluation metrics, specifically the lack of granular assessments for errors within captions, and the reliance on single-point quality estimates without considering uncertainty. To address the limitations, we propose a simple yet effective strategy for generating and calibrating distributions of CLIPScore values. Leveraging a model-agnostic conformal risk control framework, we calibrate CLIPScore values for task-specific control variables, tackling the aforementioned limitations. Experimental results demonstrate that using conformal risk control, over score distributions produced with simple methods such as input masking, can achieve competitive performance compared to more complex approaches. Our method effectively detects erroneous words, while providing formal guarantees aligned with desired risk levels. It also improves the correlation between uncertainty estimations and prediction errors, thus enhancing the overall reliability of caption evaluation metrics.

CLIPScore不确定性细粒度评估置信控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。