arXiv:2601.14973cs.ROcs.AI2026-01中稿 · HRI, Late Breaking…被引 2

用视觉直接生成搜救无人机路径,靠近目标人并安全送医

HumanDiffusion: A Vision-Based Diffusion Trajectory Planner with Human-Conditioned Goals for Search and Rescue UAV

  • 基于图像的扩散模型直接生成避障飞行轨迹
  • 像素空间预测误差仅0.02,真实场景任务成功率80%
  • 无需地图或复杂计算,适合紧急救援场景

紧急场景中可靠的人机协作需要自主系统具备检测人类、推断导航目标并在动态环境中安全运行的能力。本文提出HumanDiffusion,一种轻量级图像条件扩散规划器,可直接从RGB图像生成面向人类的导航轨迹。系统结合YOLO-11人体检测与扩散驱动轨迹生成,使四旋翼无人机在无先验地图或高计算开销规划管道的情况下,接近目标人员并递送医疗援助。轨迹在像素空间内预测,确保运动平滑且始终与人类保持安全距离。我们在仿真和真实室内模拟灾害场景中评估了HumanDiffusion。在300个样本的测试集中,模型在像素空间轨迹重建上的均方误差为0.02。真实实验表明,在部分遮挡条件下,事故响应与搜寻定位任务的整体任务成功率达到80%。结果表明,以人为条件的扩散规划为时间敏感的辅助场景中的无人机自主导航提供了实用且鲁棒的解决方案。

原文摘要 · Abstract (English)

Reliable human--robot collaboration in emergency scenarios requires autonomous systems that can detect humans, infer navigation goals, and operate safely in dynamic environments. This paper presents HumanDiffusion, a lightweight image-conditioned diffusion planner that generates human-aware navigation trajectories directly from RGB imagery. The system combines YOLO-11 based human detection with diffusion-driven trajectory generation, enabling a quadrotor to approach a target person and deliver medical assistance without relying on prior maps or computationally intensive planning pipelines. Trajectories are predicted in pixel space, ensuring smooth motion and a consistent safety margin around humans. We evaluate HumanDiffusion in simulation and real-world indoor mock-disaster scenarios. On a 300-sample test set, the model achieves a mean squared error of 0.02 in pixel-space trajectory reconstruction. Real-world experiments demonstrate an overall mission success rate of 80% across accident-response and search-and-locate tasks with partial occlusions. These results indicate that human-conditioned diffusion planning offers a practical and robust solution for human-aware UAV navigation in time-critical assistance settings.

无人机导航扩散模型搜救机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。