arXiv:2602.19562cs.AIcs.CV2026-02

用简单模型实现人类级视觉语言对齐,效率更高、准确率翻倍。

A Multimodal Framework for Aligning Human Linguistic Descriptions with Visual Perceptual Data

  • 结合视觉特征与语言处理,模拟人类在模糊场景中指代理解
  • 仅需人类65%的语句就达成稳定对齐,单句识别正确率达41.66%
  • 适合研究具身认知、跨模态理解与人机交互的学者参考

建立自然语言与视觉感知之间的稳定映射是认知科学与人工智能的基础问题。人类常在噪声和模糊的感知环境中进行语言指代,但其背后机制尚不明确。本文提出一种计算框架,通过整合语言表达与大规模众包图像的感知表征,模拟人类指代理解的核心过程。系统采用尺度不变特征变换(SIFT)与通用质量指数(UQI)在认知合理特征空间中量化相似性,并结合语言预处理与查询转换操作捕捉指代表达的语用变异性。在斯坦福重复指代游戏语料库(15,000条语句配对拓扑图刺激)上评估,该框架表现稳健:所需语句数仅为人类的65%即可达成稳定映射,且能以41.66%的准确率从单个表达中正确识别目标对象(人类为20%)。结果表明,相对简单的感知-语言对齐机制即可在经典认知基准上达到人类水平,为具身沟通、感知推理与跨模态概念形成提供新见解。代码已公开于https://anonymous.4open.science/r/metasequoia-9D13/README.md。

原文摘要 · Abstract (English)

Establishing stable mappings between natural language expressions and visual percepts is a foundational problem for both cognitive science and artificial intelligence. Humans routinely ground linguistic reference in noisy, ambiguous perceptual contexts, yet the mechanisms supporting such cross-modal alignment remain poorly understood. In this work, we introduce a computational framework designed to model core aspects of human referential interpretation by integrating linguistic utterances with perceptual representations derived from large-scale, crowd-sourced imagery. The system approximates human perceptual categorization by combining scale-invariant feature transform (SIFT) alignment with the Universal Quality Index (UQI) to quantify similarity in a cognitively plausible feature space, while a set of linguistic preprocessing and query-transformation operations captures pragmatic variability in referring expressions. We evaluate the model on the Stanford Repeated Reference Game corpus (15,000 utterances paired with tangram stimuli), a paradigm explicitly developed to probe human-level perceptual ambiguity and coordination. Our framework achieves robust referential grounding. It requires 65\% fewer utterances than human interlocutors to reach stable mappings and can correctly identify target objects from single referring expressions 41.66\% of the time (versus 20\% for humans).These results suggest that relatively simple perceptual-linguistic alignment mechanisms can yield human-competitive behavior on a classic cognitive benchmark, and offers insights into models of grounded communication, perceptual inference, and cross-modal concept formation. Code is available at https://anonymous.4open.science/r/metasequoia-9D13/README.md .

视觉语言对齐认知建模指代理解多模态学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。