arXiv:2502.20490cs.CVcs.AI2025-02ACL被引 15

构建首个基于第一视角视频的社交规范评测基准,评估视觉语言模型的常识理解能力。

EgoNormia: Benchmarking Physical Social Norm Understanding

  • 从第一视角视频中自动生成带语境的多选题,覆盖七大类社会规范。
  • 顶尖视觉语言模型在该评测上最高仅得54%,安全与隐私类表现尤其薄弱。
  • 提出检索增强生成方法,可有效提升模型对社会规范的理解能力。

人类行为受社会规范制约,但针对物理或社会情境中的规范推理监督数据极为稀缺。为此,我们构建了EGONORMIA(ε),包含1853道(其中200道经验证)基于第一视角人际互动视频的多选题,用于评估和改进视觉语言模型(VLMs)的规范理解能力。该数据集涵盖安全、隐私、空间距离、礼貌、合作、协调/主动性及沟通清晰度七类规范。为规模化构建,我们提出一种新型管道,从原始第一视角视频生成具语境的多选题。实验表明,当前最先进VLM在EGONORMIA上最高仅达54%,在验证集上为65%,且各类规范表现不均,提示其在真实应用中存在安全与隐私风险。我们进一步探索改进方法,发现使用EGONORMIA的检索增强生成(RAG)能有效提升模型规范推理能力。

原文摘要 · Abstract (English)

Human activity is moderated by norms; however, supervision for normative reasoning is sparse, particularly where norms are physically- or socially-grounded. We thus present EGONORMIA $\|ε\|$, comprising 1,853 (200 for EGONORMIA-verified) multiple choice questions (MCQs) grounded within egocentric videos of human interactions, enabling the evaluation and improvement of normative reasoning in vision-language models (VLMs). EGONORMIA spans seven norm categories: safety, privacy, proxemics, politeness, cooperation, coordination/proactivity, and communication/legibility. To compile this dataset at scale, we propose a novel pipeline to generate grounded MCQs from raw egocentric video. Our work demonstrates that current state-of-the-art VLMs lack robust grounded norm understanding, scoring a maximum of 54% on EGONORMIA and 65% on EGONORMIA-verified, with performance across norm categories indicating significant risks of safety and privacy when VLMs are used in real-world agents. We additionally explore methods for improving normative understanding, demonstrating that a naive retrieval-based generation (RAG) method using EGONORMIA can enhance normative reasoning in VLMs.

视觉语言模型社会规范第一视角视频评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。