让智能眼镜通过网页原生AI助手实现无屏、无手操作的日常辅助。
Egocentric Co-Pilot: Web-Native Smart-Glasses Agents for Assistive Egocentric AI
- 用大模型协调感知、推理与网络工具,实现第一视角长时间任务决策。
- 在Egolife和HD-EPIC数据集上达到顶尖的视角问答性能。
- 适合视障、认知负荷高或需双手自由的人群使用。
如果获取网络信息无需屏幕、稳定桌面甚至双手会怎样?对于穿梭于密集城市、低视力或面临认知过载的人群,智能眼镜结合AI代理可将网络变为日常生活的持续辅助层。我们提出Egocentric Co-Pilot,一种运行在智能眼镜上的网页原生神经符号框架,利用大语言模型(LLM)调度感知、推理与网络工具箱。一个第一人称推理核心结合时序思维链与分层上下文压缩,支持对连续第一视角视频的长程问题回答与决策支持,远超单个模型上下文窗口限制。此外,轻量级多模态意图层将嘈杂语音与凝视转化为结构化指令。我们还实现了基于云原生WebRTC的流式通信管道,集成语音、视频与控制消息至统一通道,用于眼镜与浏览器交互;同时部署本地WebSocket基线,揭示了本地推理与云端卸载在延迟、移动性与资源消耗间的权衡。在Egolife和HD-EPIC上的实验表明性能达到竞争水平或领先,人类在环测试显示任务完成率与用户满意度优于主流商业方案。这些结果表明,联网的第一人称协作者可成为更普及、情境感知辅助的实际路径。通过以网页原生通信原语为基础,模块化且可审计的工具使用,Egocentric Co-Pilot为教育、无障碍与社会包容提供了可落地的上下文感知协作者蓝图。
原文摘要 · Abstract (English)
What if accessing the web did not require a screen, a stable desk, or even free hands? For people navigating crowded cities, living with low vision, or experiencing cognitive overload, smart glasses coupled with AI agents could turn the web into an always-on assistive layer over daily life. We present Egocentric Co-Pilot, a web-native neuro-symbolic framework that runs on smart glasses and uses a Large Language Model (LLM) to orchestrate a toolbox of perception, reasoning, and web tools. An egocentric reasoning core combines Temporal Chain-of-Thought with Hierarchical Context Compression to support long-horizon question answering and decision support over continuous first-person video, far beyond a single model's context window. Additionally, a lightweight multimodal intent layer maps noisy speech and gaze into structured commands. We further implement and evaluate a cloud-native WebRTC pipeline integrating streaming speech, video, and control messages into a unified channel for smart glasses and browsers. In parallel, we deploy an on-premise WebSocket baseline, exposing concrete trade-offs between local inference and cloud offloading in terms of latency, mobility, and resource use. Experiments on Egolife and HD-EPIC demonstrate competitive or state-of-the-art egocentric QA performance, and a human-in-the-loop study on smart glasses shows higher task completion and user satisfaction than leading commercial baselines. Taken together, these results indicate that web-connected egocentric co-pilots can be a practical path toward more accessible, context-aware assistance in everyday life. By grounding operation in web-native communication primitives and modular, auditable tool use, Egocentric Co-Pilot offers a concrete blueprint for assistive, always-on web agents that support education, accessibility, and social inclusion for people who may benefit most from contextual, egocentric AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。