arXiv:2410.09921cs.CV2024-10被引 2

用多模态语义相关性指标揭示视觉认知中的上下文作用

The Roles of Contextual Semantic Relevance Metrics in Human Visual Processing

  • 融合视觉与语言模型计算上下文语义相关性
  • 联合指标显著提升注视点预测精度
  • 适合认知科学与人机交互研究者阅读

语义相关性度量可捕捉单个物体的内在语义及其在视觉场景中与其他元素的关系。以往研究已表明这些度量会影响人类视觉处理,但常未充分考虑上下文信息或采用最新深度学习模型进行精确计算。本研究通过引入上下文语义相关性度量,从视觉和语言双视角评估目标对象与其周围环境的语义关系。基于大规模视觉理解眼动数据集,采用先进深度学习技术计算度量,并通过高级统计模型分析其对注视行为的影响。该度量能模拟自上而下与自下而上的感知过程。研究进一步将视觉与语言度量融合为新型联合度量,填补了此前研究中视觉与语义相似性分离处理的空白。结果表明,所有度量均可精确预测视觉感知与处理中的注视行为,且各有不同作用;联合度量表现最优,支持语义与视觉信息交互影响感知的理论。这一发现契合多模态信息处理在人类认知中重要性的共识,深化了对视觉处理认知机制的理解,对认知科学与人机交互领域计算模型的发展具有重要意义。

原文摘要 · Abstract (English)

Semantic relevance metrics can capture both the inherent semantics of individual objects and their relationships to other elements within a visual scene. Numerous previous research has demonstrated that these metrics can influence human visual processing. However, these studies often did not fully account for contextual information or employ the recent deep learning models for more accurate computation. This study investigates human visual perception and processing by introducing the metrics of contextual semantic relevance. We evaluate semantic relationships between target objects and their surroundings from both vision-based and language-based perspectives. Testing a large eye-movement dataset from visual comprehension, we employ state-of-the-art deep learning techniques to compute these metrics and analyze their impacts on fixation measures on human visual processing through advanced statistical models. These metrics could also simulate top-down and bottom-up processing in visual perception. This study further integrates vision-based and language-based metrics into a novel combined metric, addressing a critical gap in previous research that often treated visual and semantic similarities separately. Results indicate that all metrics could precisely predict fixation measures in visual perception and processing, but with distinct roles in prediction. The combined metric outperforms other metrics, supporting theories that emphasize the interaction between semantic and visual information in shaping visual perception/processing. This finding aligns with growing recognition of the importance of multi-modal information processing in human cognition. These insights enhance our understanding of cognitive mechanisms underlying visual processing and have implications for developing more accurate computational models in fields such as cognitive science and human-computer interaction.

视觉认知语义相关性多模态眼动研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。