为听障人士打造能听懂模糊发音的多模态对话助手
HI-TransPA: Hearing Impairments Translation Personal Assistant
- 融合语音与唇动信息,统一建模听障者模糊发音
- 在自建数据集上实现最优字面准确率与语义保真度
- 适合听障辅助技术、多模态交互研究者参考
听障人士因发音不清常面临日常沟通障碍。为此,我们引入全模型范式,提出指令驱动的音视频个人助手HI-TransPA,将模糊语音与唇部动态融合,在单一多模态框架中实现翻译与对话。针对听障语音特有发音模式及现有模型适应性差的问题,设计了多模态预处理与筛选流程:检测面部关键点,稳定唇区,并量化样本质量。质量评分指导课程学习策略,先训练高置信度清洁样本,逐步加入难例以增强模型鲁棒性。架构上采用新型统一3D-Resampler高效编码唇动信息,对精准理解至关重要。在自建的HI-Dialogue数据集上的实验表明,HI-TransPA在字面准确率和语义保真度上均达到当前最优水平。本工作为全模型在辅助通信技术中的应用奠定基础,提供了端到端建模框架与关键技术工具,推动后续研究。
原文摘要 · Abstract (English)
Hearing-impaired individuals often face significant barriers in daily communication due to the inherent challenges of producing clear speech. To address this, we introduce the Omni-Model paradigm into assistive technology and present HI-TransPA, an instruction-driven audio-visual personal assistant. The model fuses indistinct speech with lip dynamics, enabling both translation and dialogue within a single multimodal framework. To address the distinctive pronunciation patterns of hearing-impaired speech and the limited adaptability of existing models, we develop a multimodal preprocessing and curation pipeline that detects facial landmarks, stabilizes the lip region, and quantitatively evaluates sample quality. These quality scores guide a curriculum learning strategy that first trains on clean, high-confidence samples and progressively incorporates harder cases to strengthen model robustness. Architecturally, we employs a novel unified 3D-Resampler to efficiently encode the lip dynamics, which is critical for accurate interpretation. Experiments on purpose-built HI-Dialogue dataset show that HI-TransPA achieves state-of-the-art performance in both literal accuracy and semantic fidelity. Our work establishes a foundation for applying Omni-Models to assistive communication technology, providing an end-to-end modeling framework and essential processing tools for future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。