提出可灵活处理不同提示、帧数和实例的群体活动识别新框架
PromptGAR: Flexible Promptive Group Activity Recognition
- 将多种视觉提示统一为点提示,实现跨模态输入融合
- 在完整与部分提示下均表现优异,支持无需重训练的灵活输入
- 适合真实场景中提示不全、对象变化频繁的群体活动分析
我们提出 PromptGAR,一种新型群体活动识别(GAR)框架,兼具输入灵活性与高识别精度。现有方法因依赖完整提示标注、固定帧数与实例数、缺乏演员一致性而难以应用于真实场景。PromptGAR 是首个无需重训练即可适应不同提示、帧数和实例数量的 GAR 模型。通过将边界框、骨骼关键点、实例身份等多样视觉提示统一为点提示,并利用交叉更新的分类与提示令牌提升性能。为保证长时间活动中的演员一致性,引入相对实例注意力机制,直接编码实例身份。大量实验表明,PromptGAR 在完整提示与部分提示输入下均表现良好,验证了其在真实应用中对输入灵活性与泛化能力的有效性。
原文摘要 · Abstract (English)
We present PromptGAR, a novel framework for Group Activity Recognition (GAR) that offering both input flexibility and high recognition accuracy. The existing approaches suffer from limited real-world applicability due to their reliance on full prompt annotations, fixed number of frames and instances, and the lack of actor consistency. To bridge the gap, we proposed PromptGAR, which is the first GAR model to provide input flexibility across prompts, frames, and instances without the need for retraining. We leverage diverse visual prompts, like bounding boxes, skeletal keypoints, and instance identities, by unifying them as point prompts. A recognition decoder then cross-updates class and prompt tokens for enhanced performance. To ensure actor consistency for extended activity durations, we also introduce a relative instance attention mechanism that directly encodes instance identities. Comprehensive evaluations demonstrate that PromptGAR achieves competitive performances both on full prompts and partial prompt inputs, establishing its effectiveness on input flexibility and generalization ability for real-world applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。