arXiv:2505.19889cs.CV2025-05被引 1

构建跨域跌倒检测数据集,提升真实场景下跌倒识别能力

OmniFall: From Staged Through Synthetic to Wild, A Unified Multi-Domain Dataset for Robust Fall Detection

  • 整合三类数据:模拟、合成和真实事故视频,统一标注体系
  • 合成数据训练模型在真实跌倒事件中表现更优,尤其区分躺卧状态
  • 释放完整数据集与标注,推动非受控环境跌倒检测研究

视觉跌倒检测模型通常在小规模模拟数据上训练,其真实应用效果不明确,因数据多样性不足且评估标准不一。本文提出OmniFall,一个包含15,000段视频(80小时)的统一基准数据集,采用单一体系的16类细粒度帧级标注。涵盖三个领域:OF-Staged整合八个模拟数据集,支持跨主体与跨视角划分;OF-Synthetic新增12,000段视频(17小时),可控地覆盖多个人口统计与环境变量;OF-In-the-Wild提供仅用于测试的真实事故视频集。我们评估了微调模型及更大规模的零样本多模态大模型。在真实跌倒事件中两者表现相当,但在临床关键的“倒地”状态识别上差异明显:零样本模型仍混淆倒地与平躺,而基于合成数据中显式倒地场景微调的模型显著更优。我们公开统一标注、合成数据与真实测试集,以促进复杂环境中跌倒及倒地状态检测器的发展。

原文摘要 · Abstract (English)

Visual fall detection models are usually trained on small, staged datasets. Their real-world utility remains unclear; such data lacks diversity and evaluation protocols differ from paper to paper. We propose OmniFall, a unified benchmark of 15k videos (80 hours) with frame-level annotations in a single 16-class taxonomy. It spans three domains: OF-Staged unifies eight staged datasets with cross-subject and cross-view splits; OF-Synthetic adds 12k videos (17 h) with controlled demographic and environmental diversity; and OF-In-the-Wild provides a test-only set of genuine accident videos. We evaluate fine-tuned models as well as much larger zero-shot multimodal LLMs. On in-the-wild fall events, both do comparably well. The clinically critical fallen state is where they part: zero-shot models keep confusing fallen with lying, whereas models fine-tuned on synthetic data with explicit fallen-state scenes do substantially better. We release the unified annotations, the synthetic data, and the in-the-wild test set to foster the development of fall and fallen-state detectors for uncontrolled environments. Dataset: https://hf.co/datasets/simplexsigil2/omnifall

跌倒检测多模态数据集真实场景

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。