收集人类试错过程数据,助力AI学习真实问题解决策略。
TEC: A Collection of Human Trial-and-error Trajectories for Problem Solving
- 构建平台记录用户多轮试错轨迹与反馈反思
- 获5370条试错轨迹,人类准确率显著高于LLMs
- 适合研究人机交互、强化学习与认知建模的学者
试错是人类解决复杂问题的根本策略,也是现实环境中人工智能系统所需能力。尽管近年已有多种试错型AI方法提出,但大多依赖研究人员设计的简单启发式规则,性能提升有限。核心原因在于缺乏真实数据:现有模型无法从人类实际试错过程的详细记录中学习。为此,我们开发了一个数据标注平台及对应数据集,名为试错收集(TEC)。该平台记录用户在多轮尝试中的完整轨迹,并收集其在收到错误反馈后的反思。通过此平台,我们在58个任务上记录了46名参与者的问题解决过程,共获得5,370条试错轨迹及41,229个网页上的错误反思。实验表明,人类在试错中的准确率显著高于现有大语言模型。我们认为,TEC平台与数据集为理解人类试错行为、开发更强大的AI系统提供了重要基础。平台与数据集已公开可用。
原文摘要 · Abstract (English)
Trial-and-error is a fundamental strategy for humans to solve complex problems and a necessary capability for Artificial Intelligence (AI) systems operating in real-world environments. Although several trial-and-error AI techniques have recently been proposed, most of them rely on simple heuristics designed by researchers and achieve limited performance gains. The core issue is the absence of appropriate data: current models cannot learn from detailed records of how humans actually conduct trial-and-error in practice. To address this gap, we introduce a data annotation platform and a corresponding dataset, termed Trial-and-Error Collection (TEC). The platform records users' complete trajectories across multiple trials and collects their reflections after receiving error feedback. Using this platform, we record the problem-solving processes of 46 participants on 58 tasks, resulting in 5,370 trial trajectories along with error reflections across 41,229 webpages. With this dataset, we observe that humans achieve substantially higher accuracy compared to LLMs, which demonstrates that humans are more effective in trial-and-error than LLMs. We believe that the TEC platform and dataset provide a valuable foundation for understanding human trial-and-error behavior and for developing more capable AI systems. Platform and dataset are publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。