用手语注释桥接文本、音频与图像,提升多模态理解能力。
Bridging the Gap between Text, Audio, Image, and Any Sequence: A Novel Approach using Gloss-based Annotation
- 通过手语符号化表示文本和音频,简化语义复杂性以对齐图像。
- 在多个数据集上超越现有模型,显著提升跨模态匹配精度。
- 适合研究多模态对齐、具身认知与通用视觉语言模型的学者。
本文提出一种名为BGTAI的新方法,通过手语注释作为中间表征,实现文本与音频到图像的更好对齐。由于文本和音频具有动态时序特征,包含影响语义的各类谓词形容词,而图像为静态场景,因此将文本和音频转换为省略复杂语义细节的手语符号,有望改善与图像的对齐效果。本研究首次提出Langue2Gloss模型,并将其集成至多模态模型UniBriVL中进行联合训练。为增强手语表征与文本/音频的适应性,并解决多模态训练中的效率与不稳定性问题,我们设计了DS-Net(数据对选择网络)、结果过滤模块及新型SP-Loss函数。实验表明,该方法在主流数据集上优于现有模型,有效提升了多模态表示能力,增强了文本、音频、视觉及任意序列模态间的兼容性。
原文摘要 · Abstract (English)
This paper presents an innovative approach called BGTAI to simplify multimodal understanding by utilizing gloss-based annotation as an intermediate step in aligning Text and Audio with Images. While the dynamic temporal factors in textual and audio inputs contain various predicate adjectives that influence the meaning of the entire sentence, images, on the other hand, present static scenes. By representing text and audio as gloss notations that omit complex semantic nuances, a better alignment with images can potentially be achieved. This study explores the feasibility of this idea, specifically, we first propose the first Langue2Gloss model and then integrate it into the multimodal model UniBriVL for joint training. To strengthen the adaptability of gloss with text/audio and overcome the efficiency and instability issues in multimodal training, we propose a DS-Net (Data-Pair Selection Network), an Result Filter module, and a novel SP-Loss function. Our approach outperforms previous multimodal models in the main experiments, demonstrating its efficacy in enhancing multimodal representations and improving compatibility among text, audio, visual, and any sequence modalities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。