让AI学会像人一样指导其他AI解题,效果更优且更省资源。
CoLa: Learning to Interactively Collaborate with Large Language Models
- 通过模仿人类引导行为训练自动化引导模型
- 小模型引导表现超越GPT-4,跨任务稳定领先
- 自适应策略优于人类,适合构建高效协作系统
大型语言模型(LLMs)在处理多种语言任务上的卓越能力,为人类与AI协同解决问题开辟了新路径。它们能以规模化方式应用人类的直觉与推理策略,放大人类能力。本文探索是否可通过泛化人类示范,模拟人类引导者,协助AI解决复杂语言问题。我们提出CoLa——一种新型自引导学习范式,用于训练自动化引导模型,并在两个问答数据集、一个谜题求解任务和一个受限文本生成任务上进行评估。实证结果表明,CoLa在所有领域均持续优于现有方法。此外,小型训练过的引导模型在作为引导者时,表现超过GPT-4等强模型。通过在问答数据集上开展人类研究,比较人类与自动化引导者的策略,结果显示自动化引导者能根据推理者能力动态调整策略,表现出更优性能。定性分析揭示了引导策略的显著差异。
原文摘要 · Abstract (English)
LLMs' remarkable ability to tackle a wide range of language tasks opened new opportunities for collaborative human-AI problem solving. LLMs can amplify human capabilities by applying their intuitions and reasoning strategies at scale. We explore whether human guides can be simulated, by generalizing from human demonstrations of guiding an AI system to solve complex language problems. We introduce CoLa, a novel self-guided learning paradigm for training automated $\textit{guides}$ and evaluate it on two QA datasets, a puzzle-solving task, and a constrained text generation task. Our empirical results show that CoLa consistently outperforms competitive approaches across all domains. Moreover, a small-sized trained guide outperforms a strong model like GPT-4 when acting as a guide. We compare the strategies employed by humans and automated guides by conducting a human study on a QA dataset. We show that automated guides outperform humans by adapting their strategies to reasoners' capabilities and conduct qualitative analyses highlighting distinct differences in guiding strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。