arXiv:2508.06277cs.CLcs.LG2025-08中稿 · KONVENS 2025被引 1

用大模型生成德国老年人语音意图识别数据,提升系统泛化能力。

Large Language Model Data Generation for Enhanced Intent Recognition in German Speech

  • 用三款大模型生成合成文本,训练适配老年德语的意图识别模型。
  • 合成数据使模型在跨数据集测试中准确率显著提升,对新词汇和口音更鲁棒。
  • 小而专的LeoLM生成数据质量优于超大规模的ChatGPT,适合低资源场景。

语音指令中的意图识别(IR)对人工智能助手至关重要,但现有方法多局限于短指令,且主要面向英语。本文聚焦老年德语使用者的语音意图识别问题,提出一种新方法:结合在老年德语语音数据集(SVC-de)上微调的Whisper ASR模型,与基于三款主流大语言模型(LeoLM、Llama3、ChatGPT)生成的合成文本训练的Transformer模型。通过文本转语音生成合成语音,并开展跨数据集测试以评估鲁棒性。结果表明,大模型生成的合成数据显著提升分类性能与对不同发音风格及未见词汇的适应能力。值得注意的是,参数量仅130亿的领域专用模型LeoLM生成的数据质量优于参数达1750亿的ChatGPT。本研究证明生成式AI可有效弥补低资源领域的数据缺口。我们公开了完整的数据生成与训练流程文档,确保透明性与可复现性。

原文摘要 · Abstract (English)

Intent recognition (IR) for speech commands is essential for artificial intelligence (AI) assistant systems; however, most existing approaches are limited to short commands and are predominantly developed for English. This paper addresses these limitations by focusing on IR from speech by elderly German speakers. We propose a novel approach that combines an adapted Whisper ASR model, fine-tuned on elderly German speech (SVC-de), with Transformer-based language models trained on synthetic text datasets generated by three well-known large language models (LLMs): LeoLM, Llama3, and ChatGPT. To evaluate the robustness of our approach, we generate synthetic speech with a text-to-speech model and conduct extensive cross-dataset testing. Our results show that synthetic LLM-generated data significantly boosts classification performance and robustness to different speaking styles and unseen vocabulary. Notably, we find that LeoLM, a smaller, domain-specific 13B LLM, surpasses the much larger ChatGPT (175B) in dataset quality for German intent recognition. Our approach demonstrates that generative AI can effectively bridge data gaps in low-resource domains. We provide detailed documentation of our data generation and training process to ensure transparency and reproducibility.

语音识别大模型生成德语NLP老年人交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。