arXiv:2503.17551cs.MMcs.AI2025-03KDD被引 16

通过扩展潜在空间提升音频增强的多模态模型性能

Audio-Enhanced Vision-Language Modeling with Latent Space Broadening for High Quality Data Expansion

  • 用kNN扩展潜在空间,改进主动学习效率
  • 融合音频信息,提升视觉语言模型在短视频场景表现
  • 已部署于生产系统,带来显著业务收益

基于Transformer的多模态模型广泛应用于工业级推荐、搜索和广告系统中,用于内容理解与相关性排序。提升标注数据质量和跨模态融合能显著改善模型性能,影响质量观看率和广告收入等关键指标。高质量标注对内容建模至关重要,但传统基于统计的主动学习(AL)方法存在局限:难以检测过度自信的误分类,且在深度神经网络中区分语义相似项效果不佳。此外,音频信息在短视频平台中作用日益重要,但多数预训练多模态架构仍主要依赖文本和图像。尽管可从头训练三模态模型,却牺牲了利用现有预训练视觉-语言(VL)和音频模型的优势。为此,我们提出基于kNN的潜在空间扩展(LSB),以提升AL效率,并设计了一种中层融合的音频增强视觉语言建模方法(VLMAE)。该系统已在生产环境中部署,带来显著业务增长。

原文摘要 · Abstract (English)

Transformer-based multimodal models are widely used in industrial-scale recommendation, search, and advertising systems for content understanding and relevance ranking. Enhancing labeled training data quality and cross-modal fusion significantly improves model performance, influencing key metrics such as quality view rates and ad revenue. High-quality annotations are crucial for advancing content modeling, yet traditional statistical-based active learning (AL) methods face limitations: they struggle to detect overconfident misclassifications and are less effective in distinguishing semantically similar items in deep neural networks. Additionally, audio information plays an increasing role, especially in short-video platforms, yet most pre-trained multimodal architectures primarily focus on text and images. While training from scratch across all three modalities is possible, it sacrifices the benefits of leveraging existing pre-trained visual-language (VL) and audio models. To address these challenges, we propose kNN-based Latent Space Broadening (LSB) to enhance AL efficiency and Vision-Language Modeling with Audio Enhancement (VLMAE), a mid-fusion approach integrating audio into VL models. This system deployed in production systems, leading to significant business gains.

多模态音频增强主动学习生产部署

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。