智能眼镜实时感知并执行任务,让双手自由的AI助手。
VisionClaw: Always-On AI Agents through Smart Glasses
- 通过智能眼镜持续感知环境,语音触发任务执行
- 实验显示任务完成更快,交互负担更低
- 适合需要双手自由的场景,如办公、出行
我们提出VisionClaw,一种运行在Meta Ray-Ban智能眼镜上的始终在线可穿戴AI代理,结合实时第一人称感知与代理式任务执行。VisionClaw持续感知真实世界上下文,支持通过语音直接启动和委派任务,例如将实物物品添加至亚马逊购物车、从纸质文档生成笔记、在途中获取会议摘要、从海报创建日程事件或控制物联网设备。我们在实验室(N=12)和长期部署(N=5)中评估了该系统。结果表明,感知与执行的集成使任务完成速度更快,交互开销更低,优于非始终在线及非代理基线。此外,部署发现用户行为发生变化:任务在日常活动中主动触发,执行逐渐由代理代为完成。这揭示了一种新范式:感知与行动持续耦合,支持情境化、免手持交互。
原文摘要 · Abstract (English)
We present VisionClaw, an always-on wearable AI agent that integrates live egocentric perception with agentic task execution. Running on Meta Ray-Ban smart glasses, VisionClaw continuously perceives real-world context and enables in-situ, speech-driven action initiation and delegation via OpenClaw AI agents. Therefore, users can directly execute tasks through the smart glasses, such as adding real-world objects to an Amazon cart, generating notes from physical documents, receiving meeting briefings on the go, creating events from posters, or controlling IoT devices. We evaluate VisionClaw through a controlled laboratory study (N=12) and a longitudinal deployment study (N=5). Results show that integrating perception and execution enables faster task completion and reduces interaction overhead compared to non-always-on and non-agent baselines. Beyond performance gains, deployment findings reveal a shift in interaction: tasks are initiated opportunistically during ongoing activities, and execution is increasingly delegated rather than manually controlled. These results suggest a new paradigm for wearable AI agents, where perception and action are continuously coupled to support situated, hands-free interaction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。