用开源大模型实现医疗翻译机器人精准识别语意并生成自然手势。
Vision-Language System using Open-Source LLMs for Gestures in Medical Interpreter Robots
- 基于本地部署的开源大模型,通过少样本提示识别医疗对话中的同意与指令。
- 语意识别准确率达90%,加权F1分数达0.91,优于基线。
- 专为隐私保护设计,适合医疗场景中需自然互动的机器人应用。
医疗沟通中跨语言障碍下,非语言线索和手势至关重要。本文提出一种隐私保护的视觉-语言框架,用于医疗翻译机器人检测特定言语行为(同意与指令),并生成相应机械手势。系统基于本地部署的开源大模型,采用少样本提示的大型语言模型进行意图识别。我们还构建了一个标注了言语行为并配对手势视频的临床对话新数据集。识别模块在测试中达到0.90的准确率、0.93的加权精确率和0.91的加权F1分数。该方法显著提升计算效率,在用户研究中,生成的手势在自然度上优于语音-手势基线,且适切性相当。
原文摘要 · Abstract (English)
Effective communication is vital in healthcare, especially across language barriers, where non-verbal cues and gestures are critical. This paper presents a privacy-preserving vision-language framework for medical interpreter robots that detects specific speech acts (consent and instruction) and generates corresponding robotic gestures. Built on locally deployed open-source models, the system utilizes a Large Language Model (LLM) with few-shot prompting for intent detection. We also introduce a novel dataset of clinical conversations annotated for speech acts and paired with gesture clips. Our identification module achieved 0.90 accuracy, 0.93 weighted precision, and a 0.91 weighted F1-Score. Our approach significantly improves computational efficiency and, in user studies, outperforms the speech-gesture generation baseline in human-likeness while maintaining comparable appropriateness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。