用音频+迭代互评,让人工总结对话更准更全。
Beyond Transcripts: Iterative Peer-Editing with Audio Unlocks High-Quality Human Summaries of Conversational Speech

- 用音频配合多人轮流修改,提升总结质量。
- 音频总结经迭代后与转录稿效果相当。
- 适合无转录文本时的高质量数据收集。
语音摘要任务缺乏成熟基准。构建新基准需人工标注,因大模型可能引入系统性错误和偏见。我们测试了十种标注流程,包括输入模态(音频、转录稿或两者结合)和编辑方式(自检或互评),研究人类标注在总结语音时的质量权衡。对比基于音频和转录稿的人工摘要,评估不同信息模态对摘要质量的影响;同时与四个LLM基准(三个文本型,一个音频型)比较,检验人工摘要是否不如自动化输出。结果发现,仅用音频生成的摘要信息量更少且更压缩;但通过迭代式同伴互评可显著缓解这一差距,使音频摘要的信息量与转录稿及大模型输出相当。该结果验证了人类标注中迭代互评的有效性,能融合词汇与语调信息,即使在无转录文本的情况下也能构建高质量数据集。
原文摘要 · Abstract (English)
There are not enough established benchmarks for the task fo speech summarization. Creating new benchmarks demands human annotation, as LLMs could embed systemic errors and bias into datasets. We test ten annotation workflows varying input modality (audio, transcript, or both) and the inclusion of editing (self or peer-editing) to investigate potential quality tradeoffs from using human annotators to summarize audio. We compare human audio-based summaries to human transcript-based summaries to track the impact of the different information modalities on summary quality. We also compare the human outputs against four LLM benchmarks (three text, one audio) to examine whether human-written summaries are less informative than highly fluent automated outputs. We find that audio-based summaries are less informative and more compressed than transcript summaries. However, iterative peer-editing with audio mitigates this difference, enabling audio-based summaries to be as informative as their transcript counterparts and LLM summaries. These findings validate iterative peer-editing among human annotators for the creation of benchmarks informed by both lexical and prosodic information. This enables crucial dataset collection even in setting where transcripts are unavailable.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。