让机器听懂人类的'嗯'、'呃'等语气词,提升对话自然度
Beyond Words: Interjection Classification for Improved Human-Computer Interaction
- 专为语气词分类设计新任务,解决语音识别忽略非词汇表达的问题
- 构建首个专用语气词数据集,经节奏/音高增强后分类准确率显著提升
- 开源数据集与工具包,适合人机交互、语音理解方向研究者使用
在人机交互领域,实现自然对话至关重要。然而,诸如'mmm'、'hmm'等语气词虽常用于表达同意、犹豫或请求信息,却常被自动语音识别(ASR)系统视为'非词语'而忽略。为此,我们提出一项全新的语气词分类任务,据我们所知为该领域的开创性工作。该任务因语气词持续时间短、跨说话人和同说话人间差异大而具有挑战性。本文构建并发布了一个专用于语气词分类的数据集,并在此基础上训练和评估了一个基线深度学习模型。通过采用节奏和音高变换等数据增强技术,显著提升了模型性能,使其更具鲁棒性。相关语气词数据集、Python 增强工具库、基线模型及评估脚本已向研究社区公开。
原文摘要 · Abstract (English)
In the realm of human-computer interaction, fostering a natural dialogue between humans and machines is paramount. A key, often overlooked, component of this dialogue is the use of interjections such as "mmm" and "hmm". Despite their frequent use to express agreement, hesitation, or requests for information, these interjections are typically dismissed as "non-words" by Automatic Speech Recognition (ASR) engines. Addressing this gap, we introduce a novel task dedicated to interjection classification, a pioneer in the field to our knowledge. This task is challenging due to the short duration of interjection signals and significant inter- and intra-speaker variability. In this work, we present and publish a dataset of interjection signals collected specifically for interjection classification. We employ this dataset to train and evaluate a baseline deep learning model. To enhance performance, we augment the training dataset using techniques such as tempo and pitch transformation, which significantly improve classification accuracy, making models more robust. The interjection dataset, a Python library for the augmentation pipeline, baseline model, and evaluation scripts, are available to the research community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。