arXiv:2608.29489cs.CRcs.CV2026-08

构建环境风险识别评测基准,提升大模型在安全认证中的可信判断能力

SpatialTrust: A Benchmark for Environmental Risk Recognition in Secure Authentication

论文配图:SpatialTrust: A Benchmark for Environmental Risk Recognition in Secure Authentication
图 1 · 摘自论文原文
  • 设计多维度问答评测集,评估模型对环境风险的感知与解释能力
  • 发现现有模型对间接风险理解不足,整体表现仅36.78%至41.12%
  • 提出结构化审计流程,有效提升模型在安全认证场景的可靠性

视觉环境风险识别在安全认证中至关重要,用户周围环境可能暴露敏感信息或带来安全隐患。然而,现有多模态大语言模型(MLLMs)评估很少关注其在空间化认证场景中可靠识别、定位和解释风险的能力。本文提出SpatialTrust,一个用于评估安全认证中环境风险识别的问答基准。该基准评估五项互补能力:敏感因素检测、直接因素识别、间接因素识别、直接因素解释与间接因素解释。我们评估了多种专有与开源MLLMs,发现当前模型表现有限,尤其在理解与解释间接风险方面存在明显短板,表明空间风险意识仍是MLLMs的挑战。此外,我们引入SpatialTrustGuard——一种结构化的问答与审计流水线,将Qwen3-VL-30B-A3B-Instruct的整体性能从36.78%提升至41.12%。研究结果凸显了专用评测基准与结构化推理方法对提升MLLMs在安全认证中可信度的必要性。

原文摘要 · Abstract (English)

Visual environmental risk recognition plays an important role in secure authentication, where a user's surroundings may reveal sensitive information or introduce potential security risks. However, existing evaluations of multimodal large language models (MLLMs) rarely examine whether models can reliably recognize, localize, and explain such risks in spatially grounded authentication scenarios. We present SpatialTrust, a question-answering benchmark for evaluating environmental risk recognition in secure authentication. SpatialTrust assesses five complementary abilities: sensitive factor detection, direct factor identification, indirect factor identification, direct factor explanation, and indirect factor explanation. We evaluate both proprietary and open-source MLLMs and find that current models show limited performance, especially in understanding and explaining indirect risks, indicating that spatial risk awareness remains a challenging capability for MLLMs. In addition, we introduce SpatialTrustGuard, a structured QA-and-audit pipeline that improves Qwen3-VL-30B-A3B-Instruct from 36.78% to 41.12% overall. Our findings highlight the need for dedicated benchmarks and structured inference methods to improve the trustworthiness of MLLMs in secure authentication.

安全认证多模态模型风险识别评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。