arXiv:2507.18252cs.HCcs.AI2025-07被引 2

用大模型分析眼动数据,提升认知模式识别准确率

Multimodal Behavioral Patterns Analysis with Eye-Tracking and LLM-Based Reasoning

  • 分阶段处理眼动数据,结合大模型推理挖掘潜在线索
  • 融合专家判断与模型输出,信任度评分提升解释性
  • 时序+语义双模检测,对学习难度预测准确率达50%

眼动数据能揭示用户认知状态,但因其结构化、非语言特性难以分析。尽管大语言模型在文本推理上表现优异,却难以处理时间序列和数值数据。本文提出一种多模态人机协同框架,以增强从眼动信号中提取认知模式的能力:(1) 采用水平与垂直分割的多阶段流程,结合大模型推理发现潜在注视模式;(2) 设计专家-模型联合评分模块,融合专家判断与模型输出生成行为解释的信任分数;(3) 构建混合异常检测模块,结合LSTM时序建模与大模型语义分析。在多个大模型与提示策略下测试,结果表明该方法在一致性、可解释性和性能上均有提升,学习难度预测任务最高达50%准确率。该方案为认知建模提供了可扩展、可解释的解决方案,适用于自适应学习、人机交互与教育分析等领域。

原文摘要 · Abstract (English)

Eye-tracking data reveals valuable insights into users' cognitive states but is difficult to analyze due to its structured, non-linguistic nature. While large language models (LLMs) excel at reasoning over text, they struggle with temporal and numerical data. This paper presents a multimodal human-AI collaborative framework designed to enhance cognitive pattern extraction from eye-tracking signals. The framework includes: (1) a multi-stage pipeline using horizontal and vertical segmentation alongside LLM reasoning to uncover latent gaze patterns; (2) an Expert-Model Co-Scoring Module that integrates expert judgment with LLM output to generate trust scores for behavioral interpretations; and (3) a hybrid anomaly detection module combining LSTM-based temporal modeling with LLM-driven semantic analysis. Our results across several LLMs and prompt strategies show improvements in consistency, interpretability, and performance, with up to 50% accuracy in difficulty prediction tasks. This approach offers a scalable, interpretable solution for cognitive modeling and has broad potential in adaptive learning, human-computer interaction, and educational analytics.

眼动分析大模型认知建模人机协同

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。