用交互数据训练模型,能提升泛化能力并模拟专家行为。
Representing expertise accelerates learning from pedagogical interaction data

- 通过合成师生互动数据,训练Transformer模型学习导航任务。
- 交互数据训练的模型在多种场景下表现更稳健,优于仅学专家示范的模型。
- 模型具备区分不同认知主体的能力,即使少见专家行为也能模仿其表现。
认知科学与人工智能研究指出,让学习智能体接触多方交互痕迹可提升性能,但交互中哪些因素促成改进仍不明确。本研究采用受控范式,精确区分了交互与专家单独行为的差异。我们生成了专家与新手在空间导航任务中的合成交互数据集,并用Transformer模型进行训练,评估不同数据集下的性能。实验表明,基于教学交互数据训练的模型在多种场景下更具鲁棒性,优于仅依赖专家示范的模型;且具备表征认知上不同主体能力的模型,即使在专家行为罕见时,也能表现出类似专家的行为。
原文摘要 · Abstract (English)
Work in cognitive science and artificial intelligence has suggested that exposing learning agents to traces of interaction between multiple individuals can improve performance in a variety of settings, yet it remains unknown which features of interactions contribute to this improvement. We examined the factors that support the effectiveness of interaction data, using a controlled paradigm that allowed us to precisely operationalize key distinctions between interaction and an expert acting alone. We generated synthetic datasets of simple interactions between an expert and a novice in a spatial navigation task, and then trained transformer models on those datasets, evaluating performance after exposure to different datasets. Our experiments showed that models trained on pedagogical interactions were more robust across a variety of scenarios compared to models trained only on expert demonstrations, and that having the ability to represent epistemically distinct agents led to expert-like behavior even when expert behavior was rarely observed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。