arXiv:2604.08494cs.CVcs.CL2026-04中稿 · ETRA 2026 GenAI wo…被引 3

用视觉语言模型分析注视路径语义相似性,突破传统空间对齐局限。

What They Saw, Not Just Where They Looked: Semantic Scanpath Similarity via VLMs and NLP metric

  • 通过视觉上下文编码注视区域,生成简洁文本描述
  • 语义相似度能捕捉空间差异下的内容一致性,独立于几何对齐
  • 适合关注认知理解与视觉注意力关联的研究者

注视路径相似性度量是眼动研究的核心,但现有方法多关注空间与时间对齐,忽视被注视图像区域的语义等价性。本文提出一种融合视觉语言模型(VLMs)的语义注视路径相似性框架:将每个注视点在受控视觉上下文中(基于区块和标记策略)编码,并转化为简明文本描述,再聚合为路径级表征。采用基于嵌入和词汇的NLP度量计算语义相似性,并与MultiMatch、DTW等经典空间度量对比。在自由观看眼动数据上的实验表明,语义相似性可捕捉部分独立于几何对齐的变异,揭示了空间发散但内容高度一致的情形。进一步分析了上下文编码对描述保真度与度量稳定性的影。结果表明,多模态基础模型可实现可解释、内容感知的注视路径分析扩展,为ETRA社区提供互补维度。

原文摘要 · Abstract (English)

Scanpath similarity metrics are central to eye-movement research, yet existing methods predominantly evaluate spatial and temporal alignment while neglecting semantic equivalence between attended image regions. We present a semantic scanpath similarity framework that integrates vision-language models (VLMs) into eye-tracking analysis. Each fixation is encoded under controlled visual context (patch-based and marker-based strategies) and transformed into concise textual descriptions, which are aggregated into scanpath-level representations. Semantic similarity is then computed using embedding-based and lexical NLP metrics and compared against established spatial measures, including MultiMatch and DTW. Experiments on free-viewing eye-tracking data demonstrate that semantic similarity captures partially independent variance from geometric alignment, revealing cases of high content agreement despite spatial divergence. We further analyze the impact of contextual encoding on description fidelity and metric stability. Our findings suggest that multimodal foundation models enable interpretable, content-aware extensions of classical scanpath analysis, providing a complementary dimension for gaze research within the ETRA community.

眼动分析视觉语言模型语义相似性多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。