arXiv:2602.20976cs.CLcs.CY2026-02

测试大模型能否提前预警生态风险,发现其在限长、跨语言时预警能力下降。

Evaluating Proactive Risk Awareness of Large Language Models

  • 构建蝴蝶数据集,模拟日常行为可能引发的隐性生态影响。
  • 五款主流大模型在限长响应下主动预警率显著下降。
  • 跨语言和物种保护场景存在持续盲区,适合关注AI生态责任的研究者。

随着大语言模型(LLMs)越来越多地融入日常决策,其安全责任应从应对明确有害意图扩展到预见潜在但严重后果的风险。本文提出一种主动风险意识评估框架,用于衡量LLMs能否在损害发生前预见潜在危害并提供预警。我们构建了蝴蝶数据集(Butterfly dataset),在环境与生态领域实例化该框架,包含1,094个查询,模拟普通求解行为可能引发的潜在生态影响。通过对五款广泛使用的LLMs进行实验,分析响应长度、语言和模态的影响。结果表明,在响应长度受限、跨语言情境下,主动风险意识出现显著且一致的下降,且在(多模态)物种保护方面存在持续盲点。这些发现揭示了当前安全对齐与真实世界生态责任之间的关键差距,强调在部署中引入主动防护机制的必要性。

原文摘要 · Abstract (English)

As large language models (LLMs) are increasingly embedded in everyday decision-making, their safety responsibilities extend beyond reacting to explicit harmful intent toward anticipating unintended but consequential risks. In this work, we introduce a proactive risk awareness evaluation framework that measures whether LLMs can anticipate potential harms and provide warnings before damage occurs. We construct the Butterfly dataset to instantiate this framework in the environmental and ecological domain. It contains 1,094 queries that simulate ordinary solution-seeking activities whose responses may induce latent ecological impact. Through experiments across five widely used LLMs, we analyze the effects of response length, languages, and modality. Experimental results reveal consistent, significant declines in proactive awareness under length-restricted responses, cross-lingual similarities, and persistent blind spots in (multimodal) species protection. These findings highlight a critical gap between current safety alignment and the requirements of real-world ecological responsibility, underscoring the need for proactive safeguards in LLM deployment.

大模型安全生态风险主动预警

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。