轻量级自信度指标实现音频字幕无参考评估
Resource-Efficient Reference-Free Evaluation of Audio Captions
- 用模型自身置信度代替参考文本进行评估
- 在资源受限场景下保持评估准确性
- 适合边缘设备部署,无需大模型
为验证自动生成音频、图像和视频字幕系统的可靠性,现有无参考评估指标依赖大型预训练模型,难以在资源受限环境中使用。为此,我们提出基于模型自身置信度的评估指标,通过测试其与依赖参考字幕的正确性度量之间的校准性,评估这些指标的有效性。分析表明,某些置信度指标与特定正确性度量更匹配,且温度缩放能有效提升其性能。主要贡献是提供了一套适用于资源受限场景的、校准良好的轻量级无参考字幕评估指标。
原文摘要 · Abstract (English)
To establish the trustworthiness of systems that automatically generate text captions for audio, images and video, existing reference-free metrics rely on large pretrained models which are impractical to accommodate in resource-constrained settings. To address this, we propose some metrics to elicit the model's confidence in its own generation. To assess how well these metrics replace correctness measures that leverage reference captions, we test their calibration with correctness measures. We discuss why some of these confidence metrics align better with certain correctness measures. Further, we provide insight into why temperature scaling of confidence metrics is effective. Our main contribution is a suite of well-calibrated lightweight confidence metrics for reference-free evaluation of captions in resource-constrained settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。