arXiv:2605.21242cs.RO2026-05

用合成数据训练小模型,精准匹配任务与机器人能力

To Select or not to Select, that is the Question: Distilling Robot Skill Prediction into a Small Ensemble

论文配图:To Select or not to Select, that is the Question: Distilling Robot Skill Prediction into a Small Ensemble
图 1 · 摘自论文原文
  • 用大模型生成+人工审核构建合成数据集
  • 133万参数小模型在200任务上达83.5%准确率
  • 适合需要高效任务分配的机器人集群系统

随着机器人集群日益多样化,包括人形机器人、探测车、四足机和无人机,为任务选择合适机器人成为核心系统问题。本文研究机器人技能预测:将自然语言任务描述映射到执行所需物理能力,如飞行、轮式移动、腿部行走、水面/水下移动及操作手。由于缺乏标注的任务-技能数据,我们利用大模型辅助生成并针对性审核标签,构建合成数据集。在此数据上训练的约133万参数双编码器集成模型(mpnet + MiniLM),在分层200任务数据集上达到83.5%的任务-技能匹配率,优于Kimi K2(1T MoE)的72.0%、GPT-OSS-120B的71.5%和Llama-4-Scout-17B的69.0%,均在相同零样本提示条件下。结果表明,在固定技能分类体系下,基于合成数据训练的小型专用模型可超越更大通用大模型,适用于舰队级任务调度。

原文摘要 · Abstract (English)

As robot fleets become more heterogeneous, including humanoids, rovers, quadrupeds, and drones, selecting the right robot for a task becomes a core systems problem. We study robot skill prediction: mapping a natural-language task description to the physical capabilities required to execute it, such as fly, wheels, legs, surface water, under water and hands. Since labelled data that maps natural-language task descriptions to robot's physical capabilities does not exist, we construct a synthetic task-to-skill dataset using LLM-assisted generation and targeted label auditing. Trained on this data, a ~133M-parameter ensemble of two fine-tuned sentence encoders (mpnet + MiniLM) reaches 83.5% task-to-skill matching on a stratified 200 task dataset, outperforming Kimi K2 (1T MoE) at 72.0%, GPT-OSS-120B at 71.5%, and Llama-4-Scout-17B at 69.0% under the same zero-shot prompt. These results suggest that, for fixed robot skill taxonomies, small specialized models trained on synthetic data can outperform much larger general-purpose LLMs for fleet-level task routing.

机器人调度技能预测小模型合成数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。