arXiv:2507.02986cs.CL2025-07被引 2

GAF-Guard让大模型能自动识别风险并按场景监控,保障使用安全。

GAF-Guard: An Agentic Framework for Risk Management and Governance in Large Language Models

  • 用智能代理动态监测大模型在具体场景中的风险
  • 可针对不同用户需求定制风险检测与报告机制
  • 适合需要高安全性的大模型应用开发者

随着大语言模型在各领域的广泛应用,其部署需严格监控以避免意外负面后果并确保稳健性。模型还需符合人类价值观,如防止有害内容和确保负责任使用。当前的自动化监控系统多聚焦于模型自身问题(如幻觉),较少考虑具体应用场景和用户偏好。本文提出GAF-Guard,一种以用户、使用场景和模型为核心的新型智能体框架,用于检测和监控基于大模型的应用部署风险。该框架通过自主代理识别风险、调用检测工具,并在特定场景中实现持续监控与报告,提升AI安全性与用户预期满足度。代码已开源:https://github.com/IBM/risk-atlas-nexus-demos/tree/main/gaf-guard。

原文摘要 · Abstract (English)

As Large Language Models (LLMs) continue to be increasingly applied across various domains, their widespread adoption necessitates rigorous monitoring to prevent unintended negative consequences and ensure robustness. Furthermore, LLMs must be designed to align with human values, like preventing harmful content and ensuring responsible usage. The current automated systems and solutions for monitoring LLMs in production are primarily centered on LLM-specific concerns like hallucination etc, with little consideration given to the requirements of specific use-cases and user preferences. This paper introduces GAF-Guard, a novel agentic framework for LLM governance that places the user, the use-case, and the model itself at the center. The framework is designed to detect and monitor risks associated with the deployment of LLM based applications. The approach models autonomous agents that identify risks, activate risk detection tools, within specific use-cases and facilitate continuous monitoring and reporting to enhance AI safety, and user expectations. The code is available at https://github.com/IBM/risk-atlas-nexus-demos/tree/main/gaf-guard.

大模型治理风险监控智能体框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。