研究视觉语言模型如何识别信息来源,提升多模态系统可靠性。
Source-Modality Monitoring in Vision-Language Models
- 通过语法和语义线索追踪输入信息的来源
- 当模态差异大时,语义信号比语法更关键
- 适合关注多模态系统鲁棒性的研究人员
我们定义并研究了源模态监控——多模态模型跟踪并传达信息输入来源的能力。将其视为更广泛的绑定问题的一个实例,评估模型在将用户提示中的词语(如'图像')与实际输入组件(如真实图像)关联时,对句法与语义信号的利用程度。在11个执行目标模态信息检索任务的视觉语言模型上展开实验,发现句法与语义信号均起重要作用,但在模态分布差异显著时,语义信号往往占主导。这些发现对模型鲁棒性及日益复杂的多模态智能体系统具有重要意义。
原文摘要 · Abstract (English)
We define and investigate source-modality monitoring -- the ability of multimodal models to track and communicate the input source from which pieces of information originate. We consider source-modality monitoring as an instance of the more general binding problem, and evaluate the extent to which models exploit syntactic vs. semantic signals in order to bind words like image in a user-provided prompt to specific components of their input and context (i.e., actual images). Across experiments spanning 11 vision-language models (VLMs) performing target-modality information retrieval tasks, we find that both syntactic and semantic signals play an important role, but that the latter tend to outweigh the former in cases when modalities are highly distinct distributionally. We discuss the implications of these findings for model robustness, and in the context of increasingly multimodal agentic systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。