统一建模人物交互检测与生成,提升理解能力
UniHOI: Unified Human-Object Interaction Understanding via Unified Token Space
- 用统一标记空间联合建模检测与生成任务
- 长尾检测准确率提升4.9%,开放词汇生成指标提高42.0%
- 适合需要多任务交互理解的研究者与应用开发
在人物交互(HOI)领域,检测与生成是两个长期分离的任务,限制了全面交互理解的发展。为此,我们提出UniHOI,通过统一标记空间联合建模HOI检测与生成,有效促进知识共享并增强泛化能力。具体而言,引入对称的交互感知注意力模块和统一的半监督学习范式,即使在标注有限的情况下也能实现图像与交互语义间的高效双向映射。大量实验表明,UniHOI在HOI检测与生成任务上均达到领先性能。尤其在长尾场景下,检测准确率提升4.9%;在开放词汇生成任务中,交互相关指标提升42.0%。
原文摘要 · Abstract (English)
In the field of human-object interaction (HOI), detection and generation are two dual tasks that have traditionally been addressed separately, hindering the development of comprehensive interaction understanding. To address this, we propose UniHOI, which jointly models HOI detection and generation via a unified token space, thereby effectively promoting knowledge sharing and enhancing generalization. Specifically, we introduce a symmetric interaction-aware attention module and a unified semi-supervised learning paradigm, enabling effective bidirectional mapping between images and interaction semantics even under limited annotations. Extensive experiments demonstrate that UniHOI achieves state-of-the-art performance in both HOI detection and generation. Specifically, UniHOI improves accuracy by 4.9% on long-tailed HOI detection and boosts interaction metrics by 42.0% on open-vocabulary generation tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。