arXiv:2602.06765cs.SD2026-02中稿 · ICASSP 2026被引 1

构建长时音频活动理解新数据集与统一模型

Hierarchical Activity Recognition and Captioning from Long-Form Audio

  • 提出多层级标注的长时音频数据集MultiAct
  • 实现活动/子活动/事件三级联合识别与描述生成
  • 适合研究长时序建模与多粒度语音理解的学者

真实世界中的复杂音频活动持续时间长且具有层次结构,但现有研究多聚焦短片段与孤立事件。为此,我们提出MultiAct数据集与基准,涵盖长时间厨房录音,包含活动、子活动和事件三个语义层级的标注,并配有细粒度描述与高层总结。进一步提出统一的分层模型,可联合完成分类、检测、序列预测与多分辨率字幕生成。在MultiAct上的实验建立强基线,揭示了建模长时音频中层次性与组合结构的关键挑战。未来工作可探索更适配复杂长程关系的建模方法。

原文摘要 · Abstract (English)

Complex activities in real-world audio unfold over extended durations and exhibit hierarchical structure, yet most prior work focuses on short clips and isolated events. To bridge this gap, we introduce MultiAct, a new dataset and benchmark for multi-level structured understanding of human activities from long-form audio. MultiAct comprises long-duration kitchen recordings annotated at three semantic levels (activities, sub-activities and events) and paired with fine-grained captions and high-level summaries. We further propose a unified hierarchical model that jointly performs classification, detection, sequence prediction and multi-resolution captioning. Experiments on MultiAct establish strong baselines and reveal key challenges in modelling hierarchical and compositional structure of long-form audio. A promising direction for future work is the exploration of methods better suited to capturing the complex, long-range relationships in long-form audio.

音频理解层次建模长时序多粒度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。