arXiv:2509.15837cs.CL2025-09被引 2

视觉信息对语音与文本模型的影响不同,语音模型仍以发音为主。

The Curious Case of Visual Grounding: Different Effects for Speech- and Text-based Language Encoders

  • 对比语音与文本模型的内部表示,发现视觉增强对两者影响不同。
  • 视觉训练提升语音模型的词义识别,但不改善语义区分能力。
  • 研究结果有助于提升语音模型的语义理解效率,适合语音处理研究者。

视觉信息如何影响基于音频和文本的深度学习模型中的语言处理?我们探讨了视觉引导对模型内部词表示的影响,发现语音与文本语言编码器存在显著差异。首先,全局表示比较显示,视觉引导增强了口语与书面语表示之间的对齐,但这一效应主要源于词身份的强化而非语义。随后,通过定向聚类分析,我们检验了模型表示中语音与语义的可区分性。语音模型在视觉引导下仍以语音特征为主导,而文本模型虽有改进,但视觉引导并未提升其语义判别能力。这些发现可为更高效地引入视觉信息以增强语音模型的语义理解提供指导。

原文摘要 · Abstract (English)

How does visual information included in training affect language processing in audio- and text-based deep learning models? We explore how such visual grounding affects model-internal representations of words, and find substantially different effects in speech- vs. text-based language encoders. Firstly, global representational comparisons reveal that visual grounding increases alignment between representations of spoken and written language, but this effect seems mainly driven by enhanced encoding of word identity rather than meaning. We then apply targeted clustering analyses to probe for phonetic vs. semantic discriminability in model representations. Speech-based representations remain phonetically dominated with visual grounding, but in contrast to text-based representations, visual grounding does not improve semantic discriminability. Our findings could usefully inform the development of more efficient methods to enrich speech-based models with visually-informed semantics.

视觉接地语音模型语义区分多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。