用原子命题统一表达多模态数据,实现可解释的跨模态理解。
From Modalities to Propositions: A Language-Centric Framework for Multimodal Intelligence

- 将图像、视频等数据转化为原子命题组成的集合,构建统一语义空间。
- 在自动驾驶和开放世界数据上实现细粒度事实到高层概念的联合推理。
- 适合需要可解释性与复杂结构检索的多模态应用开发者。
我们提出一种多模态数据的语言表示框架,将任意观测(如图像、视频或文本)表达为关于场景中实体、动作和关系的原子命题集合。一个全局语义代码本将这些命题统一为一组标准原子命题,使所有模态和观测都进入同一可解释的空间,涵盖从细粒度事实到高层概念,并支持组合生成更复杂的语义。该框架实现了可解释性与推理能力、跨模态理解与检索,以及组合性,从而支持复杂多模态理解、丰富数据整理与结构化检索。我们在自动驾驶和开放世界数据上验证了该方法的有效性。
原文摘要 · Abstract (English)
We propose a language representation for multimodal data in which any observation, whether image, video, or text, is expressed as a bag of atomic propositions, simple statements about the entities, actions, and relations in a scene. A global semantic codebook unifies these into a shared vocabulary of canonical atomic propositions, placing every modality and observation into one interpretable space that spans fine grained facts to high level concepts and composes into richer ones. This brings interpretability with reasoning, cross-modal understanding and retrieval, and compositionality that enables complex multimodal understanding, rich data curation and complex structured retrieval. We demonstrate the framework on autonomous driving and open-world data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。