用合成数据训练小模型,让机器人在救灾中具备物理常识推理能力。
FRIDA to the Rescue! Analyzing Synthetic Data Effectiveness in Object-Based Common Sense Reasoning for Disaster Response
- 结合专家知识生成高质量提示,合成少量数据用于微调。
- 仅用物体状态与功能数据的精简模型超越全量数据训练版本。
- 适合资源受限场景下的机器人智能推理研究者使用。
在灾难救援的人机交互中,大语言模型(LLMs)具有潜在的物理推理能力以辅助任务目标实现。然而,这类能力通常仅存在于大型模型中,而当前难以部署于受尺寸限制的机器人系统。为满足实际需求,我们提出一个数据集与流程,构建了面向现场推理与指令解码的代理模型(FRIDA)。该流程由领域专家与语言学家协作,设计高质量少样本提示,生成用于微调的合成数据。我们手工构建了少样本提示数据集及评估数据集,以提升LLM在通用及灾情特定物体上的推理能力。同时开展消融实验,分析不同合成数据对性能的影响。我们微调多个小型指令微调模型,发现仅基于物体物理状态与功能数据训练的精简版FRIDA模型,在评估中优于全量合成数据训练的FRIDA模型及基础模型。结果表明,该FRIDA流程可仅用少量数据注入物理常识推理能力。
原文摘要 · Abstract (English)
During Human Robot Interactions in disaster relief scenarios, Large Language Models (LLMs) have the potential for substantial physical reasoning to assist in mission objectives. However, these reasoning capabilities are often found only in larger models, which are not currently reasonable to deploy on robotic systems due to size constraints. To meet our problem space requirements, we introduce a dataset and pipeline to create Field Reasoning and Instruction Decoding Agent (FRIDA) models. In our pipeline, domain experts and linguists combine their knowledge to make high-quality, few-shot prompts used to generate synthetic data for fine-tuning. We hand-curate datasets for this few-shot prompting and for evaluation to improve LLM reasoning on both general and disaster-specific objects. We concurrently run an ablation study to understand which kinds of synthetic data most affect performance. We fine-tune several small instruction-tuned models and find that ablated FRIDA models only trained on objects' physical state and function data outperformed both the FRIDA models trained on all synthetic data and the base models in our evaluation. We demonstrate that the FRIDA pipeline is capable of instilling physical common sense with minimal data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。