仅用音频和加速度计数据,实现隐私保护的实时操作指导。
Proactive Conversational Assistant for a Procedural Manual Task based on Audio and IMU
- 通过音频与可穿戴设备的加速度数据理解用户上下文
- 微调语言模型使回答准确率提升150%,对话精准度提高50%
- 可在无云支持的边缘设备上运行,适合隐私敏感场景
针对流程性手动任务的实时对话助手通常依赖视频输入,计算开销大且影响用户隐私。本文首次提出一种仅使用可穿戴设备采集的轻量级、隐私保护模态(音频与IMU)的实时对话助手,实现对家具组装和烹饪等任务的全程引导,并能主动提供分步指令并回答用户问题。文中展示了数据生成方法与系统设计。实验发现,现成语言模型虽对话活跃但答错率高;通过微调,其回答召回率提升150%,对话精准度提高50%。此外,该助手已实现在无云端依赖的边缘设备上运行。
原文摘要 · Abstract (English)
Real-time conversational assistants for procedural manual tasks often depend on video input, which can be computationally expensive and compromise user privacy. For the first time, we propose a real-time conversational assistant that provides comprehensive guidance for procedural manual tasks using only lightweight privacy-preserving modalities such as audio and IMU inputs from a user's wearable device to understand the context. Using a furniture assembly task and a cooking task, we show how this assistant proactively communicates step-by-step instructions to a user performing a procedural task, and answers user questions. We illustrate the data generation method and the system design to achieve such an assistant. On observing that an off-the-shelf language model is a talkative assistant but is not always able to answer questions correctly, we demonstrate how finetuning the model improves its ability to limit unnecessary dialogues with a 50% increase in the precision, while also improving its ability to answer questions correctly, measured by a 150% increase in the recall of answers. We further describe how such an assistant is implemented on an edge device with no dependence on the cloud.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。