用可演化知识图谱模拟用户,评估大模型的信息匹配能力。
KnowSim: Evaluating Information Calibration in LLM Assistants with User Simulators that Learn

- 构建基于知识图谱的用户模拟器,动态追踪用户认知变化。
- 在705次人机交互中验证,与人工判断一致性达73%-74%。
- 发现最优模型随用户知识水平变化,揭示个性化适配潜力。
为在知识密集型任务中有效协作,大语言模型需具备信息校准能力:根据用户不断演变的理解水平和认知容量调整内容输出。然而,现有用户模拟器未显式建模用户知识,无法生成跨知识层级的真实互动,也无法反映知识演进过程中的交互动态。为此,我们提出KNOWSIM,一个基于显式知识状态建模的评估框架,其用户模拟器以信息单元图表示知识结构,并通过学习理论驱动的更新规则实现演化。KNOWSIM从知识状态轨迹计算三项指标(知识增量、传递校准度、认知过载),反映信息校准的核心机制。在两个领域共705次人机会话中验证,其排名与人工判断显著一致(73-74%符号一致率),优于三种基线模拟器。应用于9个LLM,结果表明最佳模型随用户知识水平而异,揭示了标准评估无法捕捉的适性-治疗交互效应。
原文摘要 · Abstract (English)
To effectively collaborate with users on knowledge-intensive tasks, Large Language Models (LLMs) must perform information calibration: matching content to a user's evolving understanding and cognitive capacity. Yet user simulators used to evaluate and train LLMs do not explicitly model user knowledge so they neither produce realistic interactions across knowledge levels nor reflect how interactions unfold as that knowledge evolves. To close this gap, we introduce KNOWSIM, an evaluation framework built around a user simulator that maintains explicit knowledge states, represented as a graph of Information Units with prerequisite relationships, that evolve under update rules grounded in learning theory. KNOWSIM computes three metrics (Knowledge Gain, Delivery Calibration, Cognitive Overload) directly from the knowledge state trajectory, reflecting key mechanistic aspects of information calibration. We validate KNOWSIM against 705 human-AI sessions across two domains, stratified by knowledge level: its rankings align significantly with human judgments (73-74% sign agreement), outperforming three baseline simulators. Applied to 9 LLMs, KNOWSIM reveals that the best model shifts by user knowledge level, revealing aptitude-treatment interactions invisible to standard evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。