跨模态语法推断,用图像语音文本联合建模语法结构
Grammar Induction from Visual, Speech and Text
- 设计多模态递归自编码框架,融合视觉语音文本特征
- 在双数据集上实现新最优性能,验证多模态互补性
- 提出无文本设置,仅凭视听信号推导语法结构
语法推断可受益于丰富的异构信号,如文本、视觉和声学信息。不同模态的特征在该过程中发挥互补作用。本文提出一个全新的无监督视觉-音频-文本语法推断任务(称为VAT-GI),旨在从并行的图像、文本和语音输入中推导出成分语法树。基于语言语法天然存在于文本之外这一观察,我们主张文本不必是语法推断中的主导模态。因此,进一步引入了无文本设置的VAT-GI,即仅依赖视觉和听觉输入完成任务。为解决该问题,我们提出一种视觉-音频-文本内外递归自编码器(VaTiora)框架,有效利用各模态特有及互补的特征进行语法解析。此外,构建了一个更具挑战性的基准数据集,用于评估VAT-GI系统的泛化能力。在两个基准数据集上的实验表明,所提出的VaTiora系统在融合多种多模态信号方面更为有效,并在VAT-GI任务上达到了新的最先进性能。
原文摘要 · Abstract (English)
Grammar Induction could benefit from rich heterogeneous signals, such as text, vision, and acoustics. In the process, features from distinct modalities essentially serve complementary roles to each other. With such intuition, this work introduces a novel \emph{unsupervised visual-audio-text grammar induction} task (named \textbf{VAT-GI}), to induce the constituent grammar trees from parallel images, text, and speech inputs. Inspired by the fact that language grammar natively exists beyond the texts, we argue that the text has not to be the predominant modality in grammar induction. Thus we further introduce a \emph{textless} setting of VAT-GI, wherein the task solely relies on visual and auditory inputs. To approach the task, we propose a visual-audio-text inside-outside recursive autoencoder (\textbf{VaTiora}) framework, which leverages rich modal-specific and complementary features for effective grammar parsing. Besides, a more challenging benchmark data is constructed to assess the generalization ability of VAT-GI system. Experiments on two benchmark datasets demonstrate that our proposed VaTiora system is more effective in incorporating the various multimodal signals, and also presents new state-of-the-art performance of VAT-GI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。