构建可复现的人类错误视频数据集,提升动作流程监控的鲁棒性。
How to Correctly Make Mistakes: A Framework for Constructing and Benchmarking Mistake Aware Egocentric Procedural Videos

- 基于心理学设计误差注入策略,分阶段生成自然合理的操作失误。
- 在17个任务中注入102处错误并生成27次修复动作,覆盖真实场景变化。
- 提供九维评估标准,适合研究动作识别与纠错系统的团队使用。
可靠的程序化视频监控需要暴露于自然发生的操作错误及其后续恢复行为。在第一人称视角视频中,错误常被手部部分遮挡,通过细微物体状态变化体现,而现有程序数据集对错误与纠正过程的记录有限且不一致。本文提出PIE-V(心理启发式错误注入视频框架),通过在原始步骤视频中加入可控、符合人类行为逻辑的偏差,构建并评测具备错误感知能力的第一人称程序视频。PIE-V结合心理学驱动的误差规划器(依据任务阶段与语义步骤负载)、恢复行为建模的纠正规划器、用于级联一致性重写的LLM写作者,以及验证流程连贯性的LLM评判器。对于视频片段编辑,采用文本引导生成替换片段并拼接至原片,保持视觉合理性。在17项任务和50个Ego-Exo4D场景中,共注入102处错误,生成27次恢复修正。为基准测试,提出统一分类体系与人工评分标准,涵盖九项指标:步骤级与流程级质量,包括合理性、流程逻辑性(含标注者置信度)、状态变化一致性及图文对齐性。利用该协议,审计多个现有资源,并在相同标准下对比PIE-V与自由形式的LLM生成基线。整体框架与评估体系支持第一人称程序视频中错误检测与修复的后处理验证。
原文摘要 · Abstract (English)
Reliable procedural monitoring in video requires exposure to naturally occurring human errors and the recoveries that follow. In egocentric recordings, mistakes are often partially occluded by hands and revealed through subtle object state changes, while existing procedural datasets provide limited and inconsistent mistake and correction traces. We present PIE-V (Psychologically Inspired Error injection for Videos), a framework for constructing and benchmarking mistake-aware egocentric procedural videos by augmenting clean keystep procedures with controlled, human-plausible deviations. PIE-V combines a psychology-informed error planner conditioned on procedure phase and semantic step load, a correction planner that models recovery behavior, an LLM writer that performs cascade-consistent rewrites, and an LLM judge that validates procedural coherence and repairs failures. For video segment edits, PIE-V synthesizes replacement clips with text-guided video generation and stitches them into the episode to preserve visual plausibility. Applied to 17 tasks and 50 Ego-Exo4D scenarios, PIE-V injects 102 mistakes and generates 27 recovery corrections. For benchmarking, we introduce a unified taxonomy and a human rubric with nine metrics that cover step-level and procedure-level quality, including plausibility, procedure logic with annotator confidence, state change coherence, and grounding between text and video. Using this protocol, we audit several existing resources and compare PIE-V against a freeform LLM generation baseline under the same criteria. Together, the framework and rubric support post-completion verification for egocentric procedural mistake detection and correction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。