通过电话游戏揭示多模态系统的隐藏语言与概念关联。
Saying the Unsaid: Revealing the Hidden Language of Multimodal Systems Through Telephone Games
- 用多轮电话游戏放大系统偏好偏差,捕捉概念间隐含连接。
- 在1万+概念对数据集上构建全局概念连接图谱,识别稳定路径。
- 结合推理大模型发现超越视觉/文本相似性的意外关联,适合可解释性研究者。
近期闭源多模态系统取得了显著进展,但其理解世界所依赖的隐藏语言因黑箱架构而难以解析。本文利用系统在压缩图像为文本、再重建图像过程中的偏好偏差,研究其隐藏语言:输入图像中多个概念的共现关系在输出中被偏好偏差扭曲。我们采用多轮电话游戏策略,通过观察概念共现频率,定量分析多模态系统理解中的概念连接强度,即“隐藏语言”。我们还构建了包含10,000+概念对的Telescope数据集,作为电话游戏框架的基础数据库。该方法具备测试时可扩展性:通过迭代运行电话游戏,可构建系统理解中的全局概念连接图谱。在此基础上,我们能识别训练继承的偏好偏差,评估泛化能力提升,并发现更稳定的脆弱概念连接路径。此外,借助推理型大模型(Reasoning-LLMs),我们揭示了超越文本与视觉相似性的意外概念关系,推断多模态系统如何理解并模拟世界。本研究为多模态系统的可解释性与可控性研究提供了新视角。
原文摘要 · Abstract (English)
Recent closed-source multimodal systems have made great advances, but their hidden language for understanding the world remains opaque because of their black-box architectures. In this paper, we use the systems' preference bias to study their hidden language: During the process of compressing the input images (typically containing multiple concepts) into texts and then reconstructing them into images, the systems' inherent preference bias introduces specific shifts in the outputs, disrupting the original input concept co-occurrence. We employ the multi-round "telephone game" to strategically leverage this bias. By observing the co-occurrence frequencies of concepts in telephone games, we quantitatively investigate the concept connection strength in the understanding of multimodal systems, i.e., "hidden language." We also contribute Telescope, a dataset of 10,000+ concept pairs, as the database of our telephone game framework. Our telephone game is test-time scalable: By iteratively running telephone games, we can construct a global map of concept connections in multimodal systems' understanding. Here we can identify preference bias inherited from training, assess generalization capability advancement, and discover more stable pathways for fragile concept connections. Furthermore, we use Reasoning-LLMs to uncover unexpected concept relationships that transcend textual and visual similarities, inferring how multimodal systems understand and simulate the world. This study offers a new perspective on the hidden language of multimodal systems and lays the foundation for future research on the interpretability and controllability of multimodal systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。