arXiv:2508.01915cs.CVcs.ET2025-08中稿 · ISMAR 2025 as a TV…被引 6

用声音触发拍照,让智能眼镜省电又记事。

EgoTrigger: Toward Audio-Driven Image Capture for Human Memory Enhancement in All-Day Energy-Efficient Smart Glasses

  • 听声音决定何时拍照片,只在关键时刻启动摄像头。
  • 平均减少54%照片数量,大幅降低能耗。
  • 适合想长期佩戴眼镜记生活细节的人群。

全天候智能眼镜有望成为持续上下文感知的平台,为日常生活提供前所未有的帮助。然而,在保持连续感知的同时集成多模态人工智能代理以增强人类记忆,对全天使用带来了重大能效挑战。实现这一平衡需要智能、情境感知的传感器管理。我们的方法EgoTrigger利用麦克风中的音频线索,选择性激活功耗较高的摄像头,从而在保持较高记忆辅助价值的同时实现高效感知。EgoTrigger采用轻量级音频模型YAMNet和自定义分类头,基于手-物体交互(HOI)音频线索(如抽屉打开声、药瓶开启声)触发图像捕获。除了在QA-Ego4D数据集上评估外,我们还引入并评估了人类记忆增强问答(HME-QA)数据集。该数据集包含340个来自完整Ego4D视频的人类标注的第一人称问答对,经筛选确保包含音频,聚焦于对情境理解与记忆至关重要的HOI时刻。结果表明,EgoTrigger平均可减少54%的帧数,显著节省高功耗传感组件(如摄像头)及下游操作(如无线传输)的能耗,同时在情景记忆任务数据集上达到相当的表现。我们认为这种情境感知的触发策略为实现全天候、功能完备的节能智能眼镜提供了有前景的方向,支持帮助用户回忆钥匙位置或日常活动信息(如服药情况)等应用。

原文摘要 · Abstract (English)

All-day smart glasses are likely to emerge as platforms capable of continuous contextual sensing, uniquely positioning them for unprecedented assistance in our daily lives. Integrating the multi-modal AI agents required for human memory enhancement while performing continuous sensing, however, presents a major energy efficiency challenge for all-day usage. Achieving this balance requires intelligent, context-aware sensor management. Our approach, EgoTrigger, leverages audio cues from the microphone to selectively activate power-intensive cameras, enabling efficient sensing while preserving substantial utility for human memory enhancement. EgoTrigger uses a lightweight audio model (YAMNet) and a custom classification head to trigger image capture from hand-object interaction (HOI) audio cues, such as the sound of a drawer opening or a medication bottle being opened. In addition to evaluating on the QA-Ego4D dataset, we introduce and evaluate on the Human Memory Enhancement Question-Answer (HME-QA) dataset. Our dataset contains 340 human-annotated first-person QA pairs from full-length Ego4D videos that were curated to ensure that they contained audio, focusing on HOI moments critical for contextual understanding and memory. Our results show EgoTrigger can use 54% fewer frames on average, significantly saving energy in both power-hungry sensing components (e.g., cameras) and downstream operations (e.g., wireless transmission), while achieving comparable performance on datasets for an episodic memory task. We believe this context-aware triggering strategy represents a promising direction for enabling energy-efficient, functional smart glasses capable of all-day use -- supporting applications like helping users recall where they placed their keys or information about their routine activities (e.g., taking medications).

智能眼镜音频触发节能感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。