用眼动信号提升语言模型,让AI更懂人类认知。
Integrating Cognitive Processing Signals into Language Models: A Review of Advances, Applications and Future Directions
- 引入眼动等认知信号增强语言模型与多模态模型
- 显著改善视觉问答表现并减少幻觉现象
- 适合关注人机对齐与高效训练的研究者
近年来,认知神经科学与自然语言处理的融合受到广泛关注。本文系统回顾了利用认知信号(尤其是眼动信号)增强语言模型与多模态大语言模型的最新进展。通过引入以用户为中心的认知信号,该方法有效缓解了数据稀缺问题,降低了大规模模型训练的环境成本。认知信号可实现高效数据增强、加速模型收敛,并提升模型与人类意图的一致性。研究特别强调眼动数据在视觉问答任务中的价值,以及在缓解多模态大模型幻觉方面的潜力。最后,文章讨论了当前挑战与未来研究方向。
原文摘要 · Abstract (English)
Recently, the integration of cognitive neuroscience in Natural Language Processing (NLP) has gained significant attention. This article provides a critical and timely overview of recent advancements in leveraging cognitive signals, particularly Eye-tracking (ET) signals, to enhance Language Models (LMs) and Multimodal Large Language Models (MLLMs). By incorporating user-centric cognitive signals, these approaches address key challenges, including data scarcity and the environmental costs of training large-scale models. Cognitive signals enable efficient data augmentation, faster convergence, and improved human alignment. The review emphasises the potential of ET data in tasks like Visual Question Answering (VQA) and mitigating hallucinations in MLLMs, and concludes by discussing emerging challenges and research trends.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。