arXiv:2604.09585cs.HCcs.AI2026-04中稿 · IEEE PacificVis 20…

用可视化图像提升大模型对眼动数据的识别效率

Evaluating Visual Prompts with Eye-Tracking Data for MLLM-Based Human Activity Recognition

论文配图:Evaluating Visual Prompts with Eye-Tracking Data for MLLM-Based Human Activity Recognition
图 1 · 摘自论文原文
  • 将眼动数据转为时间线、热力图、注视路径三类图像输入多模态大模型
  • 在三个公开数据集上验证,视觉提示在不同时间窗口下均降低令牌开销
  • 适合做物联网中高频率传感器数据的高效建模,尤其关注眼动分析场景

大型语言模型(LLMs)已成为物联网中人体活动识别(HAR)的基础模型。然而,直接处理高频多维传感器数据(如眼动追踪数据)会导致信息损失和高昂的令牌成本。为此,本文研究了一种视觉提示策略:将传感器信号转换为数据可视化图像,作为多模态大模型(MLLMs)的输入。我们在三个公开的眼动追踪数据集上,对三种可视化类型(时间线、热力图、注视路径)在不同时间窗口下的MLLM-HAR表现进行了系统评估。结果表明,视觉提示能以更少的令牌实现高效且可扩展的表示,展现了其在物联网场景中让大模型有效处理高频传感器信号的巨大潜力。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have emerged as foundation models for IoT applications such as human activity recognition (HAR). However, directly applying high-frequency and multi-dimensional sensor data, such as eye-tracking data, leads to information loss and high token costs. To mitigate this, we investigate a visual prompting strategy that transforms sensor signals into data visualization images as an input to multimodal LLMs (MLLMs) using eye-tracking data. We conducted a systematic evaluation of MLLM-based HAR across three public eye-tracking datasets using three visualization types of timeline, heatmap, and scanpath, under varying temporal window sizes. Our findings suggest that visual prompting provides a token-efficient and scalable representation for eye-tracking data, highlighting its potential to enable MLLMs to effectively reason over high-frequency sensor signals in IoT contexts.

视觉提示眼动追踪多模态模型物联网

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。