arXiv:2603.08817cs.RO2026-03

构建首个多模态按摩机器人数据集并提出分层控制框架

HMR-1: Hierarchical Massage Robot with Vision-Language-Model for Embodied Healthcare

  • 分层架构:高层用多模态大模型定位穴位,底层生成运动轨迹
  • 发布12,190张图像、17.4万组问答的MedMassage-12K数据集
  • 基于Qwen-VL微调验证框架有效性,适合康复机器人研究者

具身智能的快速发展为物理治疗与康复带来新机遇,但缺乏标准化评估基准和开放的多模态穴位按摩数据集仍是关键挑战。为此,我们构建了MedMassage-12K——一个包含12,190张图像和174,177组问答对的多模态数据集,涵盖多样光照条件与背景。同时提出一种分层具身按摩框架,包含高层穴位定位模块与低层控制模块:前者利用多模态大语言模型理解语言指令并识别穴位位置,后者生成规划轨迹。基于此框架,我们评估现有多模态大模型,并建立具身按摩任务基准。通过微调Qwen-VL模型,验证了框架有效性。实物实验进一步证明其实际应用可行性。数据集与代码已开源:https://github.com/Xiaofeng-Han-Res/HMR-1。

原文摘要 · Abstract (English)

The rapid advancement of Embodied Intelligence has opened transformative opportunities in healthcare, particularly in physical therapy and rehabilitation. However, critical challenges remain in developing robust embodied healthcare solutions, such as the lack of standardized evaluation benchmarks and the scarcity of open-source multimodal acupoint massage datasets. To address these gaps, we construct MedMassage-12K - a multimodal dataset containing 12,190 images with 174,177 QA pairs, covering diverse lighting conditions and backgrounds. Furthermore, we propose a hierarchical embodied massage framework, which includes a high-level acupoint grounding module and a low-level control module. The high-level acupoint grounding module uses multimodal large language models to understand human language and identify acupoint locations, while the low-level control module provides the planned trajectory. Based on this, we evaluate existing MLLMs and establish a benchmark for embodied massage tasks. Additionally, we fine-tune the Qwen-VL model, demonstrating the framework's effectiveness. Physical experiments further confirm the practical applicability of the framework.Our dataset and code are publicly available at https://github.com/Xiaofeng-Han-Res/HMR-1.

具身智能按摩机器人多模态数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。