arXiv:2506.10019cs.CLcs.AI2025-06综述被引 6

系统梳理文本、图像、语音生成的自动评估方法,构建统一分类框架。

A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations

  • 提出跨模态自动评估的五类核心范式,整合现有方法。
  • 覆盖文本、图像、音频生成评估,验证框架普适性。
  • 适合研究生成模型评估或跨模态研究者参考。

深度学习的进展显著提升了文本、图像和音频生成的能力,但自动评估生成内容的质量仍面临持续挑战。尽管已有众多自动评估方法,现有研究缺乏对跨文本、视觉和语音模态的系统性框架。为此,本文全面回顾并构建了跨三模态生成内容的自动评估方法统一分类体系,识别出五种刻画现有评估方法的基本范式。分析从文本生成评估入手,该领域技术最成熟;随后将框架扩展至图像与音频生成,证明其广泛适用性。最后,探讨未来跨模态评估方法的潜在研究方向。

原文摘要 · Abstract (English)

Recent advances in deep learning have significantly enhanced generative AI capabilities across text, images, and audio. However, automatically evaluating the quality of these generated outputs presents ongoing challenges. Although numerous automatic evaluation methods exist, current research lacks a systematic framework that comprehensively organizes these methods across text, visual, and audio modalities. To address this issue, we present a comprehensive review and a unified taxonomy of automatic evaluation methods for generated content across all three modalities; We identify five fundamental paradigms that characterize existing evaluation approaches across these domains. Our analysis begins by examining evaluation methods for text generation, where techniques are most mature. We then extend this framework to image and audio generation, demonstrating its broad applicability. Finally, we discuss promising directions for future research in cross-modal evaluation methodologies.

自动评估生成模型跨模态综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。