arXiv:2511.11811cs.HCcs.SY2025-11被引 2

打造可穿戴隐私保护设备,本地完成语音视觉推理。

Lessons Learned from Developing a Privacy-Preserving Multimodal Wearable for Local Voice-and-Vision Inference

  • 耳戴设备本地运行量化多模态模型,无云端依赖。
  • 在30克轻量设计下实现唤醒词触发与低延迟响应。
  • 适合关注隐私与实时交互的嵌入式AI研究者。

许多多模态可穿戴设备的应用需要持续感知与高强度计算,但用户因隐私担忧而拒绝使用。本文分享了构建一款佩戴于耳朵上的语音与视觉可穿戴设备的经验,该设备利用配对智能手机作为可信边缘节点,在本地完成AI推理。我们介绍了软硬件协同设计,包括在30克重量内集成摄像头、麦克风和扬声器的挑战,支持唤醒词触发的数据采集,并在离线状态下运行量化后的视觉-语言模型与大语言模型。通过多次迭代原型,我们识别出功耗管理、连接性、延迟及社会接受度等关键设计障碍。初步评估表明,在消费级移动硬件上实现完全本地的多模态推理是可行的,且具备交互式延迟。最后,我们总结了面向嵌入式AI系统开发的设计经验,强调在日常场景中平衡隐私、响应速度与可用性的重要性。

原文摘要 · Abstract (English)

Many promising applications of multimodal wearables require continuous sensing and heavy computation, yet users reject such devices due to privacy concerns. This paper shares our experiences building an ear-mounted voice-and-vision wearable that performs local AI inference using a paired smartphone as a trusted personal edge. We describe the hardware-software co-design of this privacy-preserving system, including challenges in integrating a camera, microphone, and speaker within a 30-gram form factor, enabling wake word-triggered capture, and running quantized vision-language and large-language models entirely offline. Through iterative prototyping, we identify key design hurdles in power budgeting, connectivity, latency, and social acceptability. Our initial evaluation shows that fully local multimodal inference is feasible on commodity mobile hardware with interactive latency. We conclude with design lessons for researchers developing embedded AI systems that balance privacy, responsiveness, and usability in everyday settings.

可穿戴隐私多模态边缘计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。