构建个体化人类推理模拟基准,揭示大模型预测个人信念变化的短板。
HugAgent: A Human Simulation Benchmark for Individual-Level Reasoning
- 基于问卷与思维自述采集真实个体信念与推理轨迹
- 模型能较好还原个体初始信念,但预测信念更新能力差
- 适合研究人类认知差异与个性化智能系统设计
在开放任务中模拟人类推理是人工智能与认知科学的核心目标。尽管大语言模型可规模化近似人类响应,但其仍偏向群体共识,忽视个体推理风格与信念演化。为此,我们提出HugAgent(HUman-Grounded AGENT Benchmark),从平均化转向个体化、从行为模仿转向认知对齐、从片段测试转向开放数据。该基准评估模型在部分先验信息下,预测特定个体在分布外场景中的行为反应与信念更新的能力。通过结构化问卷与半结构化思维自述访谈,收集人类参与者的生态有效信念状态、信念更新及推理轨迹。实验显示:模型能较准确恢复个体初始信念,但在干预下的信念更新预测表现不佳。跨人、跨域对照表明此差距源于主题内关联匹配,而非身份一致的推理,提示进展需更强的变更检测机制,而非更多上下文。基准聚焦医疗、监控、用地三类政策领域,完整数据采集流程与配套聊天机器人已开源。
原文摘要 · Abstract (English)
Simulating human reasoning in open-ended tasks has long been a central aspiration in AI and cognitive science. While large language models now approximate human responses at scale, they remain tuned to population-level consensus, often erasing the individuality of reasoning styles and belief trajectories. To advance the vision of more human-like reasoning in machines, we introduce HugAgent (HUman-Grounded AGENT Benchmark), which rethinks human reasoning simulation along three dimensions: (i) from averaged to individualized reasoning, (ii) from behavioral mimicry to cognitive alignment, and (iii) from vignette-based to open-ended data. The benchmark evaluates whether a model can predict a specific person's behavioral responses and the underlying reasoning dynamics in out-of-distribution scenarios, given partial evidence of their prior views. HugAgent combines structured questionnaires with semi-structured think-aloud interviews to collect ecologically valid belief states, belief updates, and reasoning traces from human participants. Our experiments reveal a clear asymmetry: models recover a person's belief state from their own context reasonably well, but struggle to predict belief updates under intervention. Cross-person and cross-domain controls trace this gap to associative matching within a topic rather than identity-consistent reasoning, suggesting that progress requires better-calibrated change detection, not simply more context. We scope the benchmark to self-reported belief reasoning in three policy domains: healthcare, surveillance, and zoning. The benchmark, along with its complete data collection pipeline and companion chatbot, is open-sourced as HugAgent (https://github.com/jajamoa/HugAgent) and TraceYourThinking (https://github.com/jajamoa/trace-your-thinking).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。