arXiv:2606.04806cs.CVcs.AI2026-06

评测视觉模型能否基于真实场景合理推理并解释下一步动作。

NoRA: Evaluating Grounded Reasonableness in Visual First-person Normative Action Reasoning

论文配图:NoRA: Evaluating Grounded Reasonableness in Visual First-person Normative Action Reasoning
图 1 · 摘自论文原文
  • 构建视频基准,要求模型自动生成动作并用事实-理由-行动链解释
  • 12个多模态系统在1420段视频上测试,仅少数能完整构建合理动作空间
  • 适合关注具身智能可解释性与社会规范推理的研究者

大语言模型和自主系统越来越多地部署在社交环境中,规范能力对安全适当行为至关重要。但现有方法或仅评估文本中的规范判断,或将其简化为从固定候选动作中选择。我们认为两者均不足:现实中代理不会被给出选项菜单,而需从零识别合理动作,并基于可见事实提供可检查的理由。为此,我们提出NoRA,一个视觉第一人称视频基准,要求模型生成候选下一步动作,并通过显式的事实-理由-行动支持图进行解释。该基准包含1,420个标注视频片段,包括HumanGold-190和LLMSilver-1230两部分。每项评估包含动作对齐、事实接地性和支持绑定,综合为单一可解释合理性得分。我们在直接、刻意和结构化提示三种模式下测试12个多模态系统,发现当前视觉语言模型虽常能恢复合理动作和相关场景事实,但持续难以构建完整的合理动作空间,也无法将所选动作正确关联到局部支持。NoRA使这一差距可度量,推动评估问题从‘能否选动作’转向‘能否为合理动作提供正确可见理由’。

原文摘要 · Abstract (English)

LLMs and agentic systems are increasingly deployed in social environments, making normative competence critical for safe and appropriate behavior. However, existing approaches either assess normative judgment in text alone or reduce it to choosing among a fixed set of candidate actions. We argue both are insufficient. In practice, agents are never handed a menu of options; they must identify a reasonable action from scratch, grounded in visible facts and supported by inspectable reasons. We introduce NoRA, a visual first-person video benchmark that requires models to generate candidate next actions and justify each through an explicit fact-reason-action support graph. The benchmark comprises 1,420 annotated video clips, including HumanGold-190 and LLMSilver-1230 splits. Each instance is evaluated through action alignment, factual grounding, and support binding, aggregated into a single grounded reasonableness score. We benchmark 12 multimodal systems under direct, deliberate, and structured prompting regimes, finding that current VLMs frequently recover plausible actions and relevant scene facts, but consistently struggle to construct the full reasonable action space and bind selected actions to the correct local support. NoRA makes this gap measurable, shifting the evaluation question from whether a model can pick an action to whether it can justify an appropriate action for the right visible reasons.

视觉推理规范判断可解释性多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。