arXiv:2606.30035cs.CVcs.HC2026-06

通过共识聚类分析自由观看眼动数据,揭示人类信息交互的注意力模式。

Consensus Clustering of Free-Viewing Gaze Data: New Insights into Human-Information Interaction

论文配图:Consensus Clustering of Free-Viewing Gaze Data: New Insights into Human-Information Interaction
图 1 · 摘自论文原文
  • 提出端到端无监督集成学习系统EnsembleGaze,融合多种聚类方法
  • 图像刺激分组在不同方法中高度一致,体现环境与聚焦注视模式差异
  • 用户分组依赖图像上下文,仅双步条件与谱双聚类可捕捉该结构

自由观看眼动数据为理解人类视觉注意提供了丰富的、无需任务的视角。传统探索性数据分析通过注视点和兴趣区域揭示用户注意模式,但对人-信息交互(HII)规律的研究仍不足。本文提出一种基于用户与刺激特征的共识聚类方法,构建了全新的端到端无监督集成学习系统EnsembleGaze。通过基于注视分布统计描述符的特征工程,结合多个聚类方法的共识投票生成共关联矩阵。以用户与刺激分别聚类为基线,进一步提出两种高维聚类策略:共识子空间聚类与谱双聚类,实现联合用户与图像特征的注视簇识别。使用标准指标评估聚类性能,并结合图像级属性进行解释。结果表明,不同方法对图像刺激的分组高度一致,反映出稳定的环境性与聚焦性注视模式;而用户分组受图像上下文影响,仅双聚类与两步条件方法能有效恢复此结构。在公开数据集上的测试揭示了数据集特异性模式,各聚类策略提供互补见解。

原文摘要 · Abstract (English)

Free-viewing gaze data provides a rich, task-free window into human visual attention. Conventional exploratory data analysis of the data provides user attention patterns through fixations and areas of interest. However, despite the richness of this gaze data, its human-information interaction (HII) patterns are understudied. We address this gap using consensus clustering of gaze data with respect to users and stimulus characteristics. We present a novel end-to-end unsupervised ensemble learning system for consensus clustering of free-viewing gaze datasets, EnsembleGaze. With a goal of characterizing the user behavior and stimulus type, we propose a feature engineering step based on statistical descriptors of fixation-based distributions. EnsembleGaze involves consensus voting of selected clustering methods implemented on the feature vector to compute the co-association matrix. Using the separate consensus clustering of users and stimuli as a baseline, we further propose two high-dimensional clustering strategies for determining gaze clusters based on joint user and image characterization. They are consensus subspace clustering and spectral biclustering. Clustering performance is evaluated using selected standard metrics and is further interpreted through image-level properties. Our system provides a replicable method for the unsupervised analysis of fixation behavior in scene perception research. Our results show that image stimuli groupings are highly consistent across methods, reflecting a robust ambient-versus-focal viewing mode distinction, whereas user groupings are image-context-dependent, a structure that only biclustering and the two-step conditional approaches are architecturally capable of recovering. Testing on the publicly available datasets revealed dataset-specific patterns, with each offering complementary insights through distinct clustering strategies.

眼动分析聚类人机交互无监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。