构建分层人体动作识别基准,实现从原子动作到高层意图的合成与评估。
Compositional Benchmark Synthesis for Hierarchical Human Action Recognition

- 从单标签动作数据生成四层语义层级的合成片段
- 生成15,002个符合语义规则的连续动作序列
- 通过逻辑解耦设计防止模型仅学习生成器规则
跨抽象层次的人体行为识别(从原子动作到长时序意图)需要沿语义层级标注的数据。现有大规模数据集仅提供原子级标签且无时间组合,而真实复合活动数据集则局限于浅层、窄域、固定层级。本文提出一个基准生成与评估框架,从扁平单标签动作语料中合成包含动作、活动、低层意图(LLIs)、高层意图(HLIs)四层的分层意图基准,同时保留真实预提取的动作特征。在主体一致性约束下,通过转移模型组装事件,并使用覆盖感知采样器将主体使用吉尼系数从0.566降至0.248。合成基准带来循环监督风险:若生成规则与评估规则一致,模型可能仅通过恢复生成器而非真正推理。本研究通过设计实现生成规则与一阶逻辑评估规则分离以解决该问题。实例化生成15,002个事件。四个来自不同模型家族的基线表明,所有模型均存在0.13至0.17的宏F1组成式留白,即使表现最佳的图感知模型也未能填补,说明此为基准结构性特征而非模型缺陷。无逻辑基线仍违反留出语义规则,高于其内在数据率;破坏顺序的控制实验使宏F1在种子变化范围内波动,作为生成器一致性检验。已公开本体、转移模型与生成器,支持基准重生成与扩展。
原文摘要 · Abstract (English)
Recognizing human behavior across levels of abstraction, from atomic actions to long-horizon intentions, requires data annotated along a semantic hierarchy. Large corpora provide isolated, atomically labeled clips without temporal composition, whereas recorded composite-activity corpora offer shallow, domain-narrow, fixedhierarchies. A benchmark-generation and evaluation frameworkis proposed that synthesizes a four-level hierarchical-intention benchmark, spanning actions, activities, low-level intentions (LLIs), and high-level intentions (HLIs), from a flat single-label action corpus while retaining real pre-extracted features at the action level. Episodes are assembled by a transition model under a subject-consistency constraint, and a coverage-aware sampler reduces the subject usage Gini from 0.566 to 0.248. Synthesizing such a benchmark raises a circular-supervision risk that recorded datasets avoid: if the rules generating the episodes also govern the evaluation, models can succeed by recovering the generator rather than through genuine reasoning. Validity is addressed by design, holding sequence-generation rules disjoint from the first-order-logic rules used at evaluation. The instantiation yields 15,002 episodes. Four reference baselines from different model families characterize difficulty, not as recognition methods. A compositional held-out gap of 0.13 to 0.17 macro-F1 appears across all baselines, including a graph-aware model that recognizes best yet does not close the gap, indicating a structural property of the benchmark rather than a model artifact. A logic-free baseline still violates the held-out semantic rules above their intrinsic data rate, and the order-destroying control changes macro-F1 within seed variation, serving as a generator-consistency check. Theontology, transition model, and generator are released so the benchmark can beregenerated and extended.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。