为大模型生成的语音摘要设计更可靠的真人评估方案
Optimizing the role of human evaluation in LLM-based spoken document summarization systems
- 引入社会科学研究方法构建适配生成式AI的评估框架
- 提出可复现且可信的人类评估标准与实践指南
- 在科技公司落地案例验证方法有效性
大型语言模型(LLMs)的兴起推动了语音文档抽象摘要的范式变革。尽管其创造力、流畅表达及从海量语料中提炼信息的能力极具价值,但这些特性也带来了内容评估的新挑战。目前快速低成本的自动评估方法如ROUGE和BERTScore尚未达到与人类评估相媲美的表现。本文借鉴社会科学方法,提出一种专为生成式AI内容设计的语音文档摘要人类评估范式,提供详细的评估标准与最佳实践指南,确保实验设计稳健性、可复现性与可信度。此外,还包含两个案例研究,展示该方法在一家美国科技巨头中的实际应用。
原文摘要 · Abstract (English)
The emergence of powerful LLMs has led to a paradigm shift in abstractive summarization of spoken documents. The properties that make LLMs so valuable for this task -- creativity, ability to produce fluent speech, and ability to abstract information from large corpora -- also present new challenges to evaluating their content. Quick, cost-effective automatic evaluations such as ROUGE and BERTScore offer promise, but do not yet show competitive performance when compared to human evaluations. We draw on methodologies from the social sciences to propose an evaluation paradigm for spoken document summarization explicitly tailored for generative AI content. We provide detailed evaluation criteria and best practices guidelines to ensure robustness in the experimental design, replicability, and trustworthiness of human evaluation studies. We additionally include two case studies that show how these human-in-the-loop evaluation methods have been implemented at a major U.S. technology company.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。