用大模型实现开放词汇的人类互动识别,突破传统标签限制。
OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models
- 结合大语言模型生成任意交互的文本描述,支持未见动作识别。
- 构建统一大规模人机互动数据集,涵盖多样真实场景交互。
- 适合需要灵活理解复杂互动的监控与智能视觉系统使用。
理解人与人之间的互动,尤其是在公共安全监控等场景中,对维护安全至关重要。传统活动识别系统受限于固定词汇、预定义标签和僵化的交互类别,常依赖编排好的视频,忽略并发互动群体,难以适应真实世界中多样化且不可预测的交互。本文提出一种开放词汇的人类互动识别框架(OV-HHIR),利用大语言模型在开放世界场景中生成已见与未见人类互动的开放式文本描述,不受限于固定词汇。此外,我们通过标准化并整合现有公开的人类互动数据集,构建了一个全面且大规模的基准数据集。大量实验表明,该方法优于传统固定词汇分类系统及现有跨模态语言模型,在视频理解任务中表现更优,为监控及其他领域更智能、更灵活的视觉理解系统奠定基础。
原文摘要 · Abstract (English)
Understanding human-to-human interactions, especially in contexts like public security surveillance, is critical for monitoring and maintaining safety. Traditional activity recognition systems are limited by fixed vocabularies, predefined labels, and rigid interaction categories that often rely on choreographed videos and overlook concurrent interactive groups. These limitations make such systems less adaptable to real-world scenarios, where interactions are diverse and unpredictable. In this paper, we propose an open vocabulary human-to-human interaction recognition (OV-HHIR) framework that leverages large language models to generate open-ended textual descriptions of both seen and unseen human interactions in open-world settings without being confined to a fixed vocabulary. Additionally, we create a comprehensive, large-scale human-to-human interaction dataset by standardizing and combining existing public human interaction datasets into a unified benchmark. Extensive experiments demonstrate that our method outperforms traditional fixed-vocabulary classification systems and existing cross-modal language models for video understanding, setting the stage for more intelligent and adaptable visual understanding systems in surveillance and beyond.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。