系统梳理演讲自动化辅导技术,揭示研究空白与挑战。
A Survey of Automated Presentation Coaching: Systems, Methods, and Open Challenges
- 构建五维评估体系,覆盖发音、重音、语调、节奏和内容忠实度。
- 发现现有系统在跨语言口音公平反馈上严重不足。
- 适合语音处理、教育科技领域研究者参考。
自动化演讲辅导融合语音合成、韵律建模与发音训练,但此前缺乏系统性综述。本文分类整理了发音导师、流利度与韵律教练、多模态训练工具及会议问答练习系统,并提出涵盖音段发音、词汇重音、超音段韵律、语速与内容忠实度的五维任务分类框架,明确标注各系统覆盖情况,揭示数据缺口。进一步分析其核心技术:基于语音合成的示范生成与发音、韵律、流利度诊断方法。关键挑战包括高质量演讲语料库稀缺、跨母语背景口音公平反馈难,以及实时演练中的低延迟诊断实现。
原文摘要 · Abstract (English)
Automated coaching for oral presentations sits at the intersection of computer-assisted pronunciation training (CAPT), prosody modeling, and speech synthesis, yet no prior work has systematically surveyed and compared existing systems along these dimensions. This survey reviews and categorizes automated presentation coaching systems, spanning pronunciation tutors, fluency and prosody coaches, multimodal trainers, and conference Q&A practice tools. We introduce a five-dimensional task taxonomy - covering segmental pronunciation, lexical stress, suprasegmental prosody, pacing, and content faithfulness - and explicitly map surveyed systems onto it to reveal coverage gaps. We further review the core technical methods these systems employ: TTS-based exemplar generation and diagnostic methods for pronunciation, prosody, and fluency assessment. Key open challenges include the scarcity of annotated presentation corpora, achieving accent-fair feedback across diverse L1 backgrounds, and delivering low-latency diagnostics for real-time rehearsal.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。