arXiv:2606.29997cs.CV2026-06

用自蒸馏技术提升图文视频描述评估的准确性。

Rigel: Self-Distilled Score Adaptation for Image and Video Captioning Evaluation

论文配图:Rigel: Self-Distilled Score Adaptation for Image and Video Captioning Evaluation
图 1 · 摘自论文原文
  • 从冻结大模型中提取任务对齐的评分头,避免词汇量不匹配问题。
  • 在3338个视频上训练,参考与候选描述对比测试,比现有方法高10分以上。
  • 适合需要精准评估生成内容质量的研究者和开发者使用。

自动评估图像和视频描述对多模态系统基准至关重要,但标准评估指标与人类判断的对齐度有限。近期基于大语言模型(LLM)的方法(常称为LLM-as-a-Judge)虽提升了对齐度,但仍存在大词汇量语言建模与小标签集评估之间的不匹配问题。为此,我们提出Rigel,一种基于自蒸馏得分适配的图像和视频描述自动评估指标。该方法采用从冻结的LLM中蒸馏出的特定评估评分头,在任务对齐空间中捕捉判断信号,无需依赖大词汇量词元集合。随后,利用人类判断数据微调LLM主干。为训练Rigel,我们构建了Vid-Lepus数据集,包含3,338个视频片段、33,380条参考描述和5,637条候选描述。在多个基准上的实验表明,Rigel优于当前最优指标,在无参考设置下于ActivityNet-Fact上取得超过10分的提升。

原文摘要 · Abstract (English)

Automatic evaluation of image and video captioning is essential for benchmarking multimodal systems, although standard evaluation metrics show limited alignment with human judgments. Recent approaches using large language models (LLMs), commonly referred to as LLM-as-a-Judge, have improved alignment with human judgments but still suffer from a mismatch between large-vocabulary language modeling and evaluation over a small label set. To address this, we propose Rigel, an automatic evaluation metric for image and video captioning, based on self-distilled score adaptation. The metric employs an evaluation-specific scoring head distilled from a frozen LLM, which captures judgment signals in a task-aligned space without relying on large-vocabulary token sets. We then refine the LLM backbone with human judgment data. To train Rigel, we constructed the Vid-Lepus dataset, which contains 3,338 video clips, 33,380 reference captions, and 5,637 candidate captions. Experiments on multiple benchmarks show that Rigel outperforms state-of-the-art metrics, achieving over 10-point improvements on ActivityNet-Fact in the reference-free setting.

自动评估大模型视频描述自蒸馏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。