arXiv:2509.21651cs.AI2025-09被引 11

测试大模型对物理危险的感知与干预能力,发现其表现堪忧且可改进。

Can AI Perceive Physical Danger and Intervene?

  • 用真实伤情故事生成逼真图像视频,持续评估机器人安全感知。
  • 主流模型对重物、热饮等危险判断错误率超40%,需干预时常不响应。
  • 通过后训练让模型显式推理安全规则,决策过程可解释且性能最优。

当人工智能与物理世界交互(如机器人或辅助系统)时,会面临数字AI之外的直接物理安全风险。本文研究当前主流基础模型对常识性物理安全的理解能力,例如物体过重无法搬运、热咖啡不应递给孩子等。贡献有三:首先,提出一种可扩展的连续物理安全评测方法,基于真实伤情案例和操作安全约束,利用先进生成模型构建从安全到危险状态的逼真图像与视频;其次,全面评估主流基础模型在风险识别、安全推理及干预触发方面的表现,揭示其在关键任务中部署的不足;最后,设计一种后训练范式,使模型通过系统指令显式推理具身安全约束,生成可解释的思考轨迹,在约束满足评估中达到当前最优表现。评测数据集已开源:https://asimov-benchmark.github.io/v2

原文摘要 · Abstract (English)

When AI interacts with the physical world -- as a robot or an assistive agent -- new safety challenges emerge beyond those of purely ``digital AI". In such interactions, the potential for physical harm is direct and immediate. How well do state-of-the-art foundation models understand common-sense facts about physical safety, e.g. that a box may be too heavy to lift, or that a hot cup of coffee should not be handed to a child? In this paper, our contributions are three-fold: first, we develop a highly scalable approach to continuous physical safety benchmarking of Embodied AI systems, grounded in real-world injury narratives and operational safety constraints. To probe multi-modal safety understanding, we turn these narratives and constraints into photorealistic images and videos capturing transitions from safe to unsafe states, using advanced generative models. Secondly, we comprehensively analyze the ability of major foundation models to perceive risks, reason about safety, and trigger interventions; this yields multi-faceted insights into their deployment readiness for safety-critical agentic applications. Finally, we develop a post-training paradigm to teach models to explicitly reason about embodiment-specific safety constraints provided through system instructions. The resulting models generate thinking traces that make safety reasoning interpretable and transparent, achieving state of the art performance in constraint satisfaction evaluations. The benchmark is released at https://asimov-benchmark.github.io/v2

具身智能安全评测可解释性生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。