arXiv:2508.00299cs.CVcs.AI2025-08ICCV被引 1

通过动作序列控制,实现多视角行车场景中行人视频的精准编辑。

Controllable Pedestrian Video Editing for Multi-View Driving Scenarios via Motion Sequence

  • 融合视频修复与人体动作控制,跨视角统一编辑行人区域。
  • 支持插入、替换、移除行人,保持时空连贯与视角一致性。
  • 适用于自动驾驶数据增强与危险场景仿真,提升模型鲁棒性。

自动驾驶系统中的行人检测模型常因训练数据中危险行人场景不足而缺乏鲁棒性。为此,我们提出一种基于动作序列控制的多视角行车场景行人视频可控编辑框架,结合视频修复与人体运动控制技术。首先在多相机视角中识别行人感兴趣区域,以固定比例扩展检测框,并将区域重采样拼接为统一画布,同时保持跨视角空间关系。随后应用二值掩码指定可编辑区域,利用姿态序列控制条件引导行人编辑,实现插入、替换和移除等灵活功能。大量实验表明,该框架能生成高质量、视觉真实且时空连贯、跨视角一致的行人视频,验证了其在多视角行人视频生成中的鲁棒性与通用性,具有广泛的数据增强与场景仿真应用潜力。

原文摘要 · Abstract (English)

Pedestrian detection models in autonomous driving systems often lack robustness due to insufficient representation of dangerous pedestrian scenarios in training datasets. To address this limitation, we present a novel framework for controllable pedestrian video editing in multi-view driving scenarios by integrating video inpainting and human motion control techniques. Our approach begins by identifying pedestrian regions of interest across multiple camera views, expanding detection bounding boxes with a fixed ratio, and resizing and stitching these regions into a unified canvas while preserving cross-view spatial relationships. A binary mask is then applied to designate the editable area, within which pedestrian editing is guided by pose sequence control conditions. This enables flexible editing functionalities, including pedestrian insertion, replacement, and removal. Extensive experiments demonstrate that our framework achieves high-quality pedestrian editing with strong visual realism, spatiotemporal coherence, and cross-view consistency. These results establish the proposed method as a robust and versatile solution for multi-view pedestrian video generation, with broad potential for applications in data augmentation and scenario simulation in autonomous driving.

视频编辑自动驾驶多视角

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。