研究指令微调大模型在空间推理中的泛化能力,发现复杂任务下表现显著下降。
From Templates to Natural Language: Generalization Challenges in Instruction-Tuned LLMs for Spatial Reasoning
- 用合成指令微调模型,测试其对人类真实指令的适应能力
- 简单任务泛化良好,复杂任务性能大幅下滑
- 揭示了模型在真实场景指令理解上的关键缺陷
指令微调的大语言模型在多种任务中表现出色,但在从合成指令到人类编写的实际指令的泛化上仍面临挑战。本文研究了在具身环境中进行空间定位的任务,模型需解读并翻译指令,在2.5D网格上构建物体排列。我们仅使用合成指令对大模型进行微调,并在包含合成与人类撰写指令的基准数据集上评估其性能。结果表明,模型在简单任务上泛化良好,但在更复杂的任务中性能显著下降。本文还进行了详细的错误分析,揭示了指令泛化中的差距。
原文摘要 · Abstract (English)
Instruction-tuned large language models (LLMs) have shown strong performance on a variety of tasks; however, generalizing from synthetic to human-authored instructions in grounded environments remains a challenge for them. In this work, we study generalization challenges in spatial grounding tasks where models interpret and translate instructions for building object arrangements on a $2.5$D grid. We fine-tune LLMs using only synthetic instructions and evaluate their performance on a benchmark dataset containing both synthetic and human-written instructions. Our results reveal that while models generalize well on simple tasks, their performance degrades significantly on more complex tasks. We present a detailed error analysis of the gaps in instruction generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。