arXiv:2410.01532cs.CLcs.AI2024-10ICLR被引 13

用眼动数据提升大模型对齐人类偏好的能力

Seeing Eye to AI: Human Alignment via Gaze-Based Response Rewards for Large Language Models

  • 引入眼动追踪数据作为隐式反馈优化奖励模型
  • 在标准偏好数据集上显著提升奖励模型准确率
  • 适合关注人机交互与认知反馈的AI研究者

自然语言处理的进步催生了GPT、Llama、Claude和Gemini等大型语言模型,它们在多种任务中表现优异,但需大量微调才能使其输出符合人类预期。目前广泛采用的强化学习人类反馈(RLHF)方法虽有效,但在建模人类偏好方面仍存在挑战。本文提出GazeReward框架,将眼动追踪(ET)数据作为隐式反馈融入奖励模型(RM)。通过消融实验验证不同整合方式、语言模型及眼动生成模型的效果,结果表明该方法显著提升了奖励模型在标准人类偏好数据集上的准确性。本研究推动了人工智能与人类价值观对齐的讨论,探索了认知数据在未来NLP研究中的潜力。

原文摘要 · Abstract (English)

Advancements in Natural Language Processing (NLP), have led to the emergence of Large Language Models (LLMs) such as GPT, Llama, Claude, and Gemini, which excel across a range of tasks but require extensive fine-tuning to align their outputs with human expectations. A widely used method for achieving this alignment is Reinforcement Learning from Human Feedback (RLHF), which, despite its success, faces challenges in accurately modelling human preferences. In this paper, we introduce GazeReward, a novel framework that integrates implicit feedback -- and specifically eye-tracking (ET) data -- into the Reward Model (RM). In addition, we explore how ET-based features can provide insights into user preferences. Through ablation studies we test our framework with different integration methods, LLMs, and ET generator models, demonstrating that our approach significantly improves the accuracy of the RM on established human preference datasets. This work advances the ongoing discussion on optimizing AI alignment with human values, exploring the potential of cognitive data for shaping future NLP research.

大模型对齐眼动追踪强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。