首人智能平台需打通感知-行动闭环,这篇综述首次系统构建评估框架。
From Seeing to Acting: Smart Glasses as First-Person Intelligence Platforms

- 提出从捕捉到行动的完整闭环评估框架
- 建立L0-L5能力分级体系,覆盖感知到交互全过程
- 适合研究智能眼镜系统集成与可信部署的学者
智能眼镜正从视觉采集设备演变为第一人称智能平台,连接人类感知、持续上下文与数字或物理动作。其贴身视角与佩戴者视觉、听觉、运动及手物交互同步,但受限于能效、散热、隐私和反馈等约束。尽管增强现实、第一人称视觉、多模态模型、人机交互和具身智能快速发展,该领域文献仍分散于不同设备、任务与评测标准。核心挑战不在于单一模型能否识别、回答、记忆或行动,而在于整个系统能否维持可靠、时间有效、可纠正且可管控的感知-状态-交互-行动闭环。本文是首个以统一框架系统研究智能眼镜的综述,形式化第一人称数据流与受限任务效用,沿八个可验证硬件维度刻画设备,围绕七项相互依赖的基础能力组织文献,并提出L0-L5分级框架,涵盖采集、反应式感知、上下文辅助、持久状态、受控行动与具身耦合。在九个应用场景中,关联任务与数据集、系统、产品、利益相关方、失败后果及证据缺口。进一步提出九维部署框架、条件驱动评估协议与从受控测量到纵向实地验证与审计的证据阶梯。这些工具使智能眼镜更可比、可部署、可重复评估,同时勾勒出可信第一人称智能的发展路线图。
原文摘要 · Abstract (English)
Smart glasses are evolving from capture and display accessories into first-person intelligence platforms that connect human perception, persistent context, and digital or physical action. Their on-body viewpoint aligns with the wearer's vision, audition, motion, and hand-object interaction, but must operate under tight energy, thermal, privacy, and feedback constraints. Despite rapid progress in augmented reality, egocentric vision, multimodal models, human-computer interaction, and embodied intelligence, the literature remains fragmented across devices, tasks, and benchmarks. \textit{The key challenge is not whether a model can recognize, answer, remember, or act in isolation, but whether a complete system can sustain a reliable, temporally valid, correctable, and governable perception-state-interaction-action loop.} This survey is \textit{the \textbf{first} to systematically study smart glasses through such a unified framework}. We formalize first-person data flow and constrained task utility, characterize devices along eight verifiable hardware capability axes, organize the literature around seven interdependent foundational capabilities, and introduce an L0-L5 framework spanning capture, reactive perception, contextual assistance, persistent state, governed action, and embodied coupling. Across nine application scenes, we connect tasks with datasets, systems, products, stakeholders, failure consequences, and evidence gaps. We further present a nine-dimensional deployment framework, a claim-conditioned evaluation protocol, and an evidence ladder from controlled measurement to longitudinal field validation and audit. Together, these elements make smart glasses more comparable, deployable, and reproducibly evaluated, while outlining a roadmap toward trustworthy first-person intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。