将教程视频转化为可穿戴设备辅助工具,帮视障者更准确完成任务。
Vid2Coach: Transforming How-To Videos into Task Assistants
- 基于视频生成带步骤细节和完成标准的无障碍指令
- 通过智能眼镜监测进度,减少58.5%操作错误
- 融合非视觉替代方案,适合视障用户日常使用
人们常通过视频学习烹饪、健身和手工艺等技能,但视障或低视力(BLV)人群难以跟随,因依赖视觉比对。我们观察了视觉康复治疗师(VRTs)指导BLV用户的实践,发现其提供主动与响应式支持,包括详细描述、非视觉替代方案及进度反馈。为此提出Vid2Coach系统,将教程视频转化为基于可穿戴摄像头的助手,提供可访问指令与混合主动性反馈。系统从视频中生成包含示范细节和每步完成标准的无障碍指令;利用检索增强生成技术,从专为视障者设计的资源中提取相关非视觉替代方案;并通过嵌入商用智能眼镜的摄像头监控用户进展,提供情境感知指令、主动反馈及问答服务。8名视障参与者使用Vid2Coach完成烹饪任务时,错误率比传统方式降低58.5%,且普遍希望在日常生活中持续使用。该系统展示了人工智能视觉辅助能强化而非取代非视觉专业能力的机会。
原文摘要 · Abstract (English)
People use videos to learn new recipes, exercises, and crafts. Such videos remain difficult for blind and low vision (BLV) people to follow as they rely on visual comparison. Our observations of visual rehabilitation therapists (VRTs) guiding BLV people to follow how-to videos revealed that VRTs provide both proactive and responsive support including detailed descriptions, non-visual workarounds, and progress feedback. We propose Vid2Coach, a system that transforms how-to videos into wearable camera-based assistants that provide accessible instructions and mixed-initiative feedback. From the video, Vid2Coach generates accessible instructions by augmenting narrated instructions with demonstration details and completion criteria for each step. It then uses retrieval-augmented-generation to extract relevant non-visual workarounds from BLV-specific resources. Vid2Coach then monitors user progress with a camera embedded in commercial smart glasses to provide context-aware instructions, proactive feedback, and answers to user questions. BLV participants (N=8) using Vid2Coach completed cooking tasks with 58.5\% fewer errors than when using their typical workflow and wanted to use Vid2Coach in their daily lives. Vid2Coach demonstrates an opportunity for AI visual assistance that strengthens rather than replaces non-visual expertise.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。