发现动作识别模型过度依赖背景,提出有效缓解方法。
Seeing Beyond the Scene: Analyzing and Mitigating Background Bias in Action Recognition
- 分析三类模型发现均严重依赖背景信息
- 使用人体分割输入可降低3.78%背景偏差
- 优化提示词能引导模型关注人体动作,提升9.85%
人类动作识别模型常依赖背景线索而非人体运动与姿态进行判断,这种现象称为背景偏差。本文系统分析了分类模型、对比文本-图像预训练模型及视频大语言模型(VLLM)中的背景偏差,发现所有模型均表现出强烈的背景推理倾向。针对分类模型,提出缓解策略,通过引入人体分割输入,使背景偏差降低3.78%。此外,探索了人工与自动提示调优在VLLM中的应用,证明提示设计可将预测导向以人体为中心的推理,提升效果达9.85%。
原文摘要 · Abstract (English)
Human action recognition models often rely on background cues rather than human movement and pose to make predictions, a behavior known as background bias. We present a systematic analysis of background bias across classification models, contrastive text-image pretrained models, and Video Large Language Models (VLLM) and find that all exhibit a strong tendency to default to background reasoning. Next, we propose mitigation strategies for classification models and show that incorporating segmented human input effectively decreases background bias by 3.78%. Finally, we explore manual and automated prompt tuning for VLLMs, demonstrating that prompt design can steer predictions towards human-focused reasoning by 9.85%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。