arXiv:2506.05579cs.AIcs.CL2025-06NeurIPS被引 5

测试AI能否把推理能力有效教给人类,发现模型强不等于好沟通。

When Models Know More Than They Can Explain: Quantifying Knowledge Transfer in Human-AI Collaboration

  • 设计新框架KITE,通过双阶段实验分离模型解释对人类的影响
  • 118人实验显示模型性能与人类学习效果关联弱,存在明显异常值
  • 揭示行为策略是知识迁移成功的关键,适合关注人机协作的研究者

近期人工智能推理能力的进展显著提升了各类任务表现。一个关键开放问题在于:这些提升是否也带来了更好的知识转移——即模型能否以人类可理解、可应用、可学习的方式传递推理过程。为探究此问题,我们提出知识整合与转移评估框架(KITE),并开展首个大规模人类实验(N=118),专门测量人- AI 知识转移能力。实验采用两阶段设计:人类先与AI共同构思解题策略,再独立实现解决方案,从而隔离模型解释对人类理解的影响。结果表明,尽管模型基准性能与协作成果相关,但这种关系显著不一致,存在大量异常值,说明知识转移需要专门优化。分析识别出影响成功知识转移的行为与战略因素。我们开源代码、数据集与评估框架,以支持未来具备可解释性的模型研究。

原文摘要 · Abstract (English)

Recent advancements in AI reasoning have driven substantial improvements across diverse tasks. A critical open question is whether these improvements also yields better knowledge transfer: the ability of models to communicate reasoning in ways humans can understand, apply, and learn from. To investigate this, we introduce Knowledge Integration and Transfer Evaluation (KITE), a conceptual and experimental framework for Human-AI knowledge transfer capabilities and conduct the first large-scale human study (N=118) explicitly designed to measure it. In our two-phase setup, humans first ideate with an AI on problem-solving strategies, then independently implement solutions, isolating model explanations' influence on human understanding. Our findings reveal that although model benchmark performance correlates with collaborative outcomes, this relationship is notably inconsistent, featuring significant outliers, indicating that knowledge transfer requires dedicated optimization. Our analysis identifies behavioral and strategic factors mediating successful knowledge transfer. We release our code, dataset, and evaluation framework to support future work on communicatively aligned models.

人机协作知识转移可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。