用大模型生成多样化口语表达,提升小样本情况下的评分准确性。
A Novel Data Augmentation Approach for Automatic Speaking Assessment on Opinion Expressions
- 用大模型生成指定水平的口语回答,再合成语音用于训练。
- 通过动态权重调整,使合成语音与真实语音特征更匹配。
- 适合资源少但需跨模态评估口语表达的场景。
针对意见表达类自动口语评估中标注语音数据稀缺的问题,本文提出一种新训练范式:利用大语言模型生成特定水平的多样化口语回复,通过说话人感知的文本转语音技术将其合成为语音,并采用动态重要性损失根据合成语音与真实语音的特征分布差异自适应重标训练样本权重。随后,多模态大语言模型融合对齐后的文本特征与语音信号,直接预测口语水平得分。在LTTC数据集上的实验表明,该方法优于依赖真实数据或传统增强的方法,有效缓解低资源限制,实现基于跨模态信息的意见表达自动评估。
原文摘要 · Abstract (English)
Automated speaking assessment (ASA) on opinion expressions is often hampered by the scarcity of labeled recordings, which restricts prompt diversity and undermines scoring reliability. To address this challenge, we propose a novel training paradigm that leverages a large language models (LLM) to generate diverse responses of a given proficiency level, converts responses into synthesized speech via speaker-aware text-to-speech synthesis, and employs a dynamic importance loss to adaptively reweight training instances based on feature distribution differences between synthesized and real speech. Subsequently, a multimodal large language model integrates aligned textual features with speech signals to predict proficiency scores directly. Experiments conducted on the LTTC dataset show that our approach outperforms methods relying on real data or conventional augmentation, effectively mitigating low-resource constraints and enabling ASA on opinion expressions with cross-modal information.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。