arXiv:2604.19638cs.AIcs.CL2026-04中稿 · ACL

评测大模型在厨房场景中主动规避安全风险的能力

SafetyALFRED: Evaluating Safety-Conscious Planning of Multimodal Large Language Models

论文配图:SafetyALFRED: Evaluating Safety-Conscious Planning of Multimodal Large Language Models
图 1 · 摘自论文原文
  • 在ALFRED基础上增加六类真实厨房隐患,构建具身安全评估基准
  • 模型识别隐患准确率高,但实际避险成功率普遍低于30%
  • 强调需从问答评测转向具身规划的主动安全评估

多模态大语言模型日益被用作交互环境中的自主智能体,但其主动应对安全风险的能力仍不足。我们提出SafetyALFRED,基于具身智能体基准ALFRED,新增六类真实厨房安全隐患。现有安全评估多依赖非具身问答(QA)任务进行隐患识别,而我们评估了来自Qwen、Gemma和Gemini系列的十一款先进模型,不仅考察隐患识别能力,更关注其通过具身规划实现主动风险缓解的表现。实验结果揭示显著能力鸿沟:模型在问答任务中识别准确率较高,但在实际避险任务中的平均成功率为30%以下。研究证明,仅靠静态问答评测无法有效衡量物理安全性能,亟需转向以纠正行为为核心的具身化评估范式。代码与数据集已开源,地址为https://github.com/sled-group/SafetyALFRED.git。

原文摘要 · Abstract (English)

Multimodal Large Language Models are increasingly adopted as autonomous agents in interactive environments, yet their ability to proactively address safety hazards remains insufficient. We introduce SafetyALFRED, built upon the embodied agent benchmark ALFRED, augmented with six categories of real-world kitchen hazards. While existing safety evaluations focus on hazard recognition through disembodied question answering (QA) settings, we evaluate eleven state-of-the-art models from the Qwen, Gemma, and Gemini families on not only hazard recognition, but also active risk mitigation through embodied planning. Our experimental results reveal a significant alignment gap: while models can accurately recognize hazards in QA settings, average mitigation success rates for these hazards are low in comparison. Our findings demonstrate that static evaluations through QA are insufficient for physical safety, thus we advocate for a paradigm shift toward benchmarks that prioritize corrective actions in embodied contexts. We open-source our code and dataset under https://github.com/sled-group/SafetyALFRED.git

具身智能安全评估大模型多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。