arXiv:2601.15679cs.AI2026-01

多国联合推进智能体评估方法,聚焦信息泄露与网络攻击风险。

Improving Methodologies for Agentic Evaluations Across Domains: Leakage of Sensitive Information, Fraud and Cybersecurity Threats

  • 联合多国开展智能体测试,分设信息泄露与网络安全两条主线。
  • 使用公开与闭源模型在多个基准上测试,重点关注评估方法而非结果。
  • 为全球智能体安全评估提供标准化框架,适合政策制定者参考。

随着自主AI系统快速发展,其在真实世界交互中缺乏有效监管,带来新风险。目前智能体测试仍处于早期阶段。为此,国际先进AI测评科学网络(INAAIMS)汇聚新加坡、日本、澳大利亚、加拿大、欧盟委员会、法国、肯尼亚、韩国及英国代表,开展第三次联合测试,延续2024年11月和2025年2月的前期成果。目标是完善高级AI系统测试的最佳实践。测试分为两个方向:由新加坡人工智能与信息安全研究所(AISI)主导的通用风险(含敏感信息泄露与欺诈);由英国AISI主导的网络安全。评估覆盖开放与闭源权重模型,任务来自多个公开智能体基准。由于智能体测试尚处萌芽,重点在于识别评估过程中的方法论问题,而非分析模型性能。该合作标志着智能体评估科学的重要进展。

原文摘要 · Abstract (English)

The rapid rise of autonomous AI systems and advancements in agent capabilities are introducing new risks due to reduced oversight of real-world interactions. Yet agent testing remains nascent and is still a developing science. As AI agents begin to be deployed globally, it is important that they handle different languages and cultures accurately and securely. To address this, participants from The International Network for Advanced AI Measurement, Evaluation and Science, including representatives from Singapore, Japan, Australia, Canada, the European Commission, France, Kenya, South Korea, and the United Kingdom have come together to align approaches to agentic evaluations. This is the third exercise, building on insights from two earlier joint testing exercises conducted by the Network in November 2024 and February 2025. The objective is to further refine best practices for testing advanced AI systems. The exercise was split into two strands: (1) common risks, including leakage of sensitive information and fraud, led by Singapore AISI; and (2) cybersecurity, led by UK AISI. A mix of open and closed-weight models were evaluated against tasks from various public agentic benchmarks. Given the nascency of agentic testing, our primary focus was on understanding methodological issues in conducting such tests, rather than examining test results or model capabilities. This collaboration marks an important step forward as participants work together to advance the science of agentic evaluations.

智能体评估信息安全国际合作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。