用眼动生成自然语言描述人类目标,突破固定标签限制。
Gaze-to-text Generation: Beyond Categorical Decoding of Human Attention

- 基于多模态大模型,将眼动轨迹转为自由文本描述。
- 通过合成思维讲解提升模型对目标动态的捕捉能力。
- 适用于多种视觉任务,可非侵入式推断人类意图。
我们提出一种新型学习问题:将眼动解码为跨多样化视觉任务中人类目标的自然语言描述。与以往将眼动解码视为预定义类别上的判别任务不同,本文将其建模为生成式学习问题——训练模型生成能捕捉人类意图丰富细节和开放性特征的自由形式描述。为此,我们提出了首个眼动到文本的解码框架Gazette。Gazette基于多模态大语言模型(MLLM),学习将眼动轨迹转化为超出固定标签范畴、需自然语言表达的目标描述。为帮助模型过滤个体差异并学习目标相关的时空动态,我们提出一种新策略:利用大语言模型的百科知识与推理能力,合成名为“思考自述”(think-aloud transcripts)的自然语言解释,以描述目标导向的眼动行为。通过对这些合成叙述进行指令微调,Gazette在多个任务上达到最先进的眼动解码性能,展现出良好的泛化性和适应性,使眼动成为多样场景中非侵入式推断人类目标与意图的强大信号。
原文摘要 · Abstract (English)
We introduce a novel learning problem: decoding gaze into natural language descriptions of human goals across diverse visual tasks. Unlike prior work, which frames gaze decoding as a discriminative task over predefined categories, we formulate it as a generative learning problem: training a model to produce free-form descriptions that capture the rich nuances and open-ended nature of human intentions beyond fixed labels. To this end, we introduce Gazette, the first gaze-to-text decoding framework. Based on multimodal large language models (MLLMs), Gazette learns to decode gaze scanpaths into natural language for goals that may extend beyond categorical labels and require articulation in natural language. To help Gazette filter out individual differences in gaze behavior and learn the goal-specific spatiotemporal dynamics crucial for generating accurate natural language goal descriptions, we propose a novel strategy that leverages the encyclopedic knowledge and reasoning abilities of a large language model to synthesize natural language explanations of goal-directed attentional behavior called think-aloud transcripts. Instruction tuning on these synthetic narratives allows Gazette to achieve state-of-the-art performance in gaze decoding across multiple tasks, demonstrating its generalizability and versatility, thereby enabling gaze to serve as a powerful, non-intrusive cue for inferring human goals and intentions in diverse scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。