小模型微调后可高效完成机器人角色识别,适合边缘部署。
Evaluating Zero-Shot and One-Shot Adaptation of Small Language Models in Leader-Follower Interaction
- 用微调替代提示工程,提升角色分类准确率
- 零样本微调达86.66%准确率,单次推理仅22.2毫秒
- 对话越复杂,小模型表现越差,需权衡可靠性
领导-跟随交互是人机交互中的重要范式,但资源受限的移动与辅助机器人实时分配角色仍具挑战。尽管大语言模型在自然交流中表现优异,其规模与延迟限制了本地部署。小语言模型(SLMs)提供了潜在替代方案,但其在人机交互角色分类中的有效性尚未系统评估。本文构建了一个针对领导-跟随通信的SLM基准,引入基于公开数据库并经合成样本增强的新数据集,以捕捉交互特异性动态。研究对比了提示工程与微调两种适配策略,在零样本与单样本交互模式下的表现,并与未训练基线进行比较。实验基于Qwen2.5-0.5B模型显示,零样本微调达到86.66%准确率,单样本推理延迟低至22.2毫秒,显著优于基线与提示工程方法。然而,在单样本模式下,上下文长度增加导致模型架构容量不足,引发性能下降。结果表明,微调后的小模型能有效实现直接角色分配,但也揭示了对话复杂性与分类可靠性之间的关键权衡。
原文摘要 · Abstract (English)
Leader-follower interaction is an important paradigm in human-robot interaction (HRI). Yet, assigning roles in real time remains challenging for resource-constrained mobile and assistive robots. While large language models (LLMs) have shown promise for natural communication, their size and latency limit on-device deployment. Small language models (SLMs) offer a potential alternative, but their effectiveness for role classification in HRI has not been systematically evaluated. In this paper, we present a benchmark of SLMs for leader-follower communication, introducing a novel dataset derived from a published database and augmented with synthetic samples to capture interaction-specific dynamics. We investigate two adaptation strategies: prompt engineering and fine-tuning, studied under zero-shot and one-shot interaction modes, compared with an untrained baseline. Experiments with Qwen2.5-0.5B reveal that zero-shot fine-tuning achieves robust classification performance (86.66% accuracy) while maintaining low latency (22.2 ms per sample), significantly outperforming baseline and prompt-engineered approaches. However, results also indicate a performance degradation in one-shot modes, where increased context length challenges the model's architectural capacity. These findings demonstrate that fine-tuned SLMs provide an effective solution for direct role assignment, while highlighting critical trade-offs between dialogue complexity and classification reliability on the edge.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。