用户自述的写作助手需求偏好与实际行为严重不符,导致系统设计失效。
Users Mispredict Their Own Preferences for AI Writing Assistance
- 通过实验发现,用户决策主要受任务复杂度影响,而非紧急程度。
- 用户自述优先考虑紧急性,但行为数据表明紧急性几乎无影响。
- 基于真实行为数据的系统比基于用户自述的系统准确率高3.6个百分点。
主动式AI写作助手需要预测用户何时需要帮助,但我们对驱动偏好的因素缺乏实证理解。通过对50名参与者进行因子化情景研究,共完成750次两两比较,发现任务复杂度是主导因素(ρ=0.597),而紧急性几乎无预测力(ρ≈0)。更关键的是,用户存在显著的感知-行为偏差:他们在自述中将紧急性排在首位,但其真实行为中紧急性是最弱的驱动因素,形成完全偏好反转。这种偏差带来实际后果:基于用户自述偏好设计的系统准确率仅为57.7%,甚至低于朴素基线;而基于行为模式设计的系统准确率达61.3%(p < 0.05),显著更高。结果表明,依赖用户自我报告会误导系统优化,对主动自然语言生成(NLG)系统设计具有直接启示。
原文摘要 · Abstract (English)
Proactive AI writing assistants need to predict when users want drafting help, yet we lack empirical understanding of what drives preferences. Through a factorial vignette study with 50 participants making 750 pairwise comparisons, we find compositional effort dominates decisions ($ρ= 0.597$) while urgency shows no predictive power ($ρ\approx 0$). More critically, users exhibit a striking perception-behavior gap: they rank urgency first in self-reports despite it being the weakest behavioral driver, representing a complete preference inversion. This misalignment has measurable consequences. Systems designed from users' stated preferences achieve only 57.7\% accuracy, underperforming even naive baselines, while systems using behavioral patterns reach significantly higher 61.3\% ($p < 0.05$). These findings demonstrate that relying on user introspection for system design actively misleads optimization, with direct implications for proactive natural language generation (NLG) systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。