用无人机和可穿戴设备实现战场伤员远程分诊,提升救援效率与安全性。
ATRACT: A Trustworthy Robotic Autonomous system to support Casualty Triage

- 融合无人机视频与可穿戴传感器数据,多模态分析伤员状态。
- 动作识别准确率达85.7%,轻量CNN模型表现接近强预训练模型。
- 适合战场救援、应急响应人员,解决前线医疗人员高风险问题。
当无人机日益与敌对行动关联时,我们将其重新用于人道主义与生命救援场景。然而,将搜救无人机应用于战场分诊仍面临巨大挑战:系统需在极端不确定性、通行受限与高个人风险下可靠运行。鉴于冲突区伤员撤离愈发脆弱,本文提出ATRACT(A Trustworthy Robotic Autonomous system to support Casualty Triage),一种新型人机协同决策支持系统,以在创伤后关键时期实现早期战场分诊。ATRACT结合无人机捕获的视频与可穿戴传感器输入,开展多模态学习,用于评估伤员状态,克服现有系统局限。无人机视频捕捉细微行为线索,如姿态、姿势;可穿戴设备提供心率、呼吸频率与活动等生理信号。通过双模态融合,ATRACT在无法即时接触伤员时为医务人员提供判断依据。为缓解受伤动作数据真实感差距,设计条件变分自编码器进行数据增强。在自建无人机采集数据集上的实验表明,所提流程在动作分类上达到85.7%准确率;轻量级CNN视觉编码器性能优于或媲美更强的预训练视频主干网络。结果表明,ATRACT是实现对抗环境中远程分诊的实用进展,多模态感知、人工监督与可信决策支持可提升伤员优先级判定,并降低前线医务人员暴露风险。
原文摘要 · Abstract (English)
At a time when drones are increasingly associated with hostile operations, we re-purpose them for humanitarian and life-saving applications. However, adapting search and rescue drones for battlefield triage remains extremely challenging; the technology must perform reliably to support frontline medics who are forced to operate under extreme uncertainty, restricted access, and significant personal risk. Due to growing vulnerabilities of casualty evacuation in conflicting zones, this paper presents ATRACT (A Trustworthy Robotic Autonomous system to support Casualty Triage), a novel human-in-the-loop decision support system to enable early battlefield triage during the critical post-trauma period. ATRACT integrates drone-captured video with wearable sensor input for multi-modal learning to support casualty-state assessment, thereby addressing the limitations of existing systems. Drone video captures fine-grained behavioural cues, such as pose, posture, while body-worn sensors provide complementary physiological signals, including heart rate, breathing rate, and movement. By combining two modalities, ATRACT provides evidence to support the early judgement of medics when direct access to the casualty is delayed, risky, or restricted. To mitigate the data realism gap pertaining to injured actions, a conditional variational autoencoder is devised for data augmentation. Experimental results on our drone captured dataset show that proposed pipeline achieves 85.7% accuracy for action classification; while our lightweight CNN visual encoder remains competitive with stronger pre-trained video backbones. Overall, the results support ATRACT as a practically meaningful step towards remote triage in contested environments, where multi-modal sensing, human oversight and trustworthy decision support can improve casualty prioritisation, and lessen the exposure of frontline medics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。