用多模态模型检测对话中的听觉困难时刻,提升助听设备响应速度。
Identifying Hearing Difficulty Moments in Conversational Audio
- 采用音频语言模型进行多模态推理,捕捉听觉困难线索。
- 性能显著优于语音识别关键词和Wav2Vec微调方法。
- 适合助听技术、实时语音辅助系统研发者参考。
人们在日常对话中常遭遇听觉困难时刻。在助听技术领域,及时识别这些时刻对实现实时辅助至关重要。本文提出并比较了多种机器学习方案,用于连续检测对话音频中体现听觉困难的语句。结果表明,凭借多模态推理能力,音频语言模型在此任务上表现优异,显著优于基于ASR关键词的简单启发式方法以及使用Wav2Vec(一种先进的仅音频输入架构)进行微调的常规方法。
原文摘要 · Abstract (English)
Individuals regularly experience Hearing Difficulty Moments in everyday conversation. Identifying these moments of hearing difficulty has particular significance in the field of hearing assistive technology where timely interventions are key for realtime hearing assistance. In this paper, we propose and compare machine learning solutions for continuously detecting utterances that identify these specific moments in conversational audio. We show that audio language models, through their multimodal reasoning capabilities, excel at this task, significantly outperforming a simple ASR hotword heuristic and a more conventional fine-tuning approach with Wav2Vec, an audio-only input architecture that is state-of-the-art for automatic speech recognition (ASR).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。