arXiv:2510.16368cs.AIcs.HC2025-10NeurIPS被引 2

用户如何用小动作让算法听自己的?

The Burden of Interactive Alignment with Inconsistent Preferences

  • 用双系统模型模拟用户决策:理性判断是否参与,冲动决定停留时长
  • 存在关键时间阈值,越有远见的用户越能成功引导算法
  • 只需一次额外点击,就能大幅降低对长期策略的依赖

从媒体平台到聊天机器人,算法影响着人们的互动、学习与信息发现方式。用户与算法的交互通常分多步进行,战略型用户可通过选择性参与内容,引导算法更贴合自身真实兴趣。然而,用户常表现出不一致偏好:花费大量时间在低长期价值内容上,无意中传递了错误信号。本文从用户角度出发,提出核心问题:具有不一致偏好的用户需付出何种代价才能实现算法对齐?我们构建双系统模型——系统1(冲动)决定停留时长,系统2(理性)决定是否参与。在此基础上,建立多领导者-单追随者扩展型斯塔克尔伯格博弈,用户(系统2)先行承诺参与策略,算法根据观测行为最优响应。定义对齐负担为用户必须优化的最小时间跨度。研究发现存在临界时间阈值:足够前瞻的用户可实现对齐,否则将被算法目标反向对齐。该临界期可能很长,造成显著负担。但仅需一个微小、高成本信号(如一次额外点击),即可显著缩短所需期限。整体框架揭示了具备不一致偏好的用户如何在斯塔克尔伯格均衡下,借助微量信号完成对基于互动的算法对齐,既说明挑战也提出可行缓解路径。

原文摘要 · Abstract (English)

From media platforms to chatbots, algorithms shape how people interact, learn, and discover information. Such interactions between users and an algorithm often unfold over multiple steps, during which strategic users can guide the algorithm to better align with their true interests by selectively engaging with content. However, users frequently exhibit inconsistent preferences: they may spend considerable time on content that offers little long-term value, inadvertently signaling that such content is desirable. Focusing on the user side, this raises a key question: what does it take for such users to align the algorithm with their true interests? To investigate these dynamics, we model the user's decision process as split between a rational system 2 that decides whether to engage and an impulsive system 1 that determines how long engagement lasts. We then study a multi-leader, single-follower extensive Stackelberg game, where users, specifically system 2, lead by committing to engagement strategies and the algorithm best-responds based on observed interactions. We define the burden of alignment as the minimum horizon over which users must optimize to effectively steer the algorithm. We show that a critical horizon exists: users who are sufficiently foresighted can achieve alignment, while those who are not are instead aligned to the algorithm's objective. This critical horizon can be long, imposing a substantial burden. However, even a small, costly signal (e.g., an extra click) can significantly reduce it. Overall, our framework explains how users with inconsistent preferences can align an engagement-driven algorithm with their interests in a Stackelberg equilibrium, highlighting both the challenges and potential remedies for achieving alignment.

用户对齐博弈论行为建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。