arXiv:2412.19005eess.AScs.AI2024-12中稿 · AAAI被引 3

通过双焦点偏好优化,提升真实场景下音视频语音识别准确率。

Enhancing Audiovisual Speech Recognition through Bifocal Preference Optimization

  • 从音频和视觉输入两端模拟常见错误,构建偏好数据。
  • 在多领域真实视频上,显著超越现有最先进模型。
  • 适合需要鲁棒音视频识别的工业落地场景。

音视频自动语音识别(AV-ASR)通过结合视觉信号提升语音识别准确率,但在噪声环境、即兴口语及视觉信息使用不确定的真实场景中仍具挑战性。以往方法多在音视频数据集上微调纯音频模型,仅优化传统语音识别目标,忽视视觉特征与真实视频中的常见错误。本文提出一种偏好优化策略,通过双重聚焦——分别操纵音频或视觉输入并重写转录文本——构建偏好数据。进一步提出BPO-AVASR方法,利用输入侧与输出侧的双重偏好信息优化模型。大量实验表明,该方法在多个真实视频领域显著提升识别准确率,优于现有最先进模型。

原文摘要 · Abstract (English)

Audiovisual Automatic Speech Recognition (AV-ASR) aims to improve speech recognition accuracy by leveraging visual signals. It is particularly challenging in unconstrained real-world scenarios across various domains due to noisy acoustic environments, spontaneous speech, and the uncertain use of visual information. Most previous works fine-tune audio-only ASR models on audiovisual datasets, optimizing them for conventional ASR objectives. However, they often neglect visual features and common errors in unconstrained video scenarios. In this paper, we propose using a preference optimization strategy to improve speech recognition accuracy for real-world videos. First, we create preference data via simulating common errors that occurred in AV-ASR from two focals: manipulating the audio or vision input and rewriting the output transcript. Second, we propose BPO-AVASR, a Bifocal Preference Optimization method to improve AV-ASR models by leveraging both input-side and output-side preference. Extensive experiments demonstrate that our approach significantly improves speech recognition accuracy across various domains, outperforming previous state-of-the-art models on real-world video speech recognition.

音视频识别偏好优化真实场景

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。