构建首个基于多模态大模型的社交场景重要人物标注数据集
MIP-GAF: A MLLM-annotated Benchmark for Most Important Person Localization and Group Context Understanding
- 用多模态大模型自动标注真实场景中人物重要性
- 新数据集上现有算法性能大幅下降超过30%
- 适合研究社会情境理解与鲁棒性感知的学者
估计社交场景中的最重要人物(MIP)是一项挑战性任务,主要源于情境复杂性和标注数据稀缺。此外,MIP判断的因果关系主观且多样。为此,我们通过标注大规模真实场景数据集,捕捉人类对图像中‘最重要人物’的感知。论文详细描述了基于多模态大语言模型(MLLM)的数据标注策略及数据质量分析。进一步利用前沿MIP定位方法对该数据集进行全面基准测试,结果显示性能相比已有数据集显著下降(平均下降超30%),表明现有算法在真实场景下仍需更强鲁棒性。我们认为该数据集将推动下一代社会情境理解方法的发展。代码与数据已开源于https://github.com/surbhimadan92/MIP-GAF。
原文摘要 · Abstract (English)
Estimating the Most Important Person (MIP) in any social event setup is a challenging problem mainly due to contextual complexity and scarcity of labeled data. Moreover, the causality aspects of MIP estimation are quite subjective and diverse. To this end, we aim to address the problem by annotating a large-scale `in-the-wild' dataset for identifying human perceptions about the `Most Important Person (MIP)' in an image. The paper provides a thorough description of our proposed Multimodal Large Language Model (MLLM) based data annotation strategy, and a thorough data quality analysis. Further, we perform a comprehensive benchmarking of the proposed dataset utilizing state-of-the-art MIP localization methods, indicating a significant drop in performance compared to existing datasets. The performance drop shows that the existing MIP localization algorithms must be more robust with respect to `in-the-wild' situations. We believe the proposed dataset will play a vital role in building the next-generation social situation understanding methods. The code and data is available at https://github.com/surbhimadan92/MIP-GAF.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。