用可穿戴传感器实现日常活动对话式追踪,兼顾隐私与准确性
Sensor2Text: Enabling Natural Language Interactions for Daily Activity Tracking Using Wearable Sensors
- 融合传感器与语言模型,通过迁移学习提升数据利用效率
- 在活动识别和问答任务中表现优于或媲美视觉语言模型
- 适合关注隐私的医疗健康监测场景,尤其适用于老年人
视觉问答技术虽已进步显著,但视频监控存在隐私泄露和视野受限问题。本文提出Sensor2Text,首个基于可穿戴传感器实现日常活动追踪与自然语言对话的模型。针对传感器数据信息密度低、单设备感知不足及对话能力弱等挑战,采用迁移学习与师生网络机制,从视觉-语言模型中迁移知识;设计编码器-解码器架构联合处理多模态传感器与语言数据,并引入大语言模型增强交互能力。实验表明,该模型能准确识别活动并进行问答对话,在生成与对话任务中表现媲美甚至超越现有视觉语言模型。这是首个可对传感器数据进行自然语言交互的系统,为解决视觉方案的隐私与视野限制提供了创新路径。
原文摘要 · Abstract (English)
Visual Question-Answering, a technology that generates textual responses from an image and natural language question, has progressed significantly. Notably, it can aid in tracking and inquiring about daily activities, crucial in healthcare monitoring, especially for elderly patients or those with memory disabilities. However, video poses privacy concerns and has a limited field of view. This paper presents Sensor2Text, a model proficient in tracking daily activities and engaging in conversations using wearable sensors. The approach outlined here tackles several challenges, including low information density in wearable sensor data, insufficiency of single wearable sensors in human activities recognition, and model's limited capacity for Question-Answering and interactive conversations. To resolve these obstacles, transfer learning and student-teacher networks are utilized to leverage knowledge from visual-language models. Additionally, an encoder-decoder neural network model is devised to jointly process language and sensor data for conversational purposes. Furthermore, Large Language Models are also utilized to enable interactive capabilities. The model showcases the ability to identify human activities and engage in Q\&A dialogues using various wearable sensor modalities. It performs comparably to or better than existing visual-language models in both captioning and conversational tasks. To our knowledge, this represents the first model capable of conversing about wearable sensor data, offering an innovative approach to daily activity tracking that addresses privacy and field-of-view limitations associated with current vision-based solutions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。