arXiv:2512.14083eess.AScs.CL2025-12

构建鲁棒可扩展的音视频语音识别系统,应对真实环境干扰。

Scalable Frameworks for Real-World Audio-Visual Speech Recognition

  • 分三层优化:表示、架构、系统,提升抗干扰能力
  • 统一模型学习跨环境鲁棒特征,无需额外模块
  • 融合大模型增强认知与生成能力,适合实际部署

音视频语音识别(AVSR)在真实环境中因不可预测的噪声和视觉干扰导致性能显著下降。本文提出系统化、分层的方法,在表示、架构和系统三个层面实现鲁棒可扩展性。在表示层面,构建统一模型,学习对多种现实污染具有内在鲁棒性的音视频特征,实现对新环境的泛化而无需专用模块。在架构层面,探索高效扩展模型容量的方法,通过智能分配计算资源,根据输入特性自适应使用多模态信息。在系统层面,通过模块化集成大规模基础模型,利用其强大的认知与生成能力,最大化最终识别准确率。本研究系统性地解决三层次挑战,旨在构建下一代高可靠性的真实场景音视频语音识别系统。

原文摘要 · Abstract (English)

The practical deployment of Audio-Visual Speech Recognition (AVSR) systems is fundamentally challenged by significant performance degradation in real-world environments, characterized by unpredictable acoustic noise and visual interference. This dissertation posits that a systematic, hierarchical approach is essential to overcome these challenges, achieving the robust scalability at the representation, architecture, and system levels. At the representation level, we investigate methods for building a unified model that learns audio-visual features inherently robust to diverse real-world corruptions, thereby enabling generalization to new environments without specialized modules. To address architectural scalability, we explore how to efficiently expand model capacity while ensuring the adaptive and reliable use of multimodal inputs, developing a framework that intelligently allocates computational resources based on the input characteristics. Finally, at the system level, we present methods to expand the system's functionality through modular integration with large-scale foundation models, leveraging their powerful cognitive and generative capabilities to maximize final recognition accuracy. By systematically providing solutions at each of these three levels, this dissertation aims to build a next-generation, robust, and scalable AVSR system with high reliability in real-world applications.

音视频识别鲁棒性可扩展性多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。