提出黑盒框架HalMit,通过探索泛化边界检测大模型幻觉。
Towards Mitigation of Hallucination for LLM-empowered Agents: Progressive Generalization Bound Exploration and Watchdog Monitor
- 基于概率分形采样生成大量查询,探测智能体的泛化边界。
- 在不依赖模型内部结构前提下,显著提升幻觉检测准确率。
- 适合需保障大模型系统可信性的实际部署场景。
得益于大语言模型(LLMs),智能体已成为与开放环境交互、推动人工智能部署的流行范式。然而,由LLMs产生的幻觉——输出与事实不符——带来了重大挑战,削弱了智能体的可信度。唯有有效缓解幻觉,智能体才能在真实世界中安全应用。因此,幻觉的检测与缓解对确保智能体可靠性至关重要。遗憾的是,现有方法或依赖白盒访问LLM,或无法准确识别幻觉。为此,本文提出HalMit,一种新型黑盒看门狗框架,通过建模LLM赋能智能体的泛化边界,实现无需了解模型内部结构的幻觉检测。具体而言,提出概率分形采样技术,以并行方式生成足够多查询,触发异常响应,高效识别目标智能体的泛化边界。实验表明,HalMit在幻觉监控方面显著优于现有方法。其黑盒特性与卓越性能使其成为提升LLM驱动系统可靠性的有前景解决方案。
原文摘要 · Abstract (English)
Empowered by large language models (LLMs), intelligent agents have become a popular paradigm for interacting with open environments to facilitate AI deployment. However, hallucinations generated by LLMs-where outputs are inconsistent with facts-pose a significant challenge, undermining the credibility of intelligent agents. Only if hallucinations can be mitigated, the intelligent agents can be used in real-world without any catastrophic risk. Therefore, effective detection and mitigation of hallucinations are crucial to ensure the dependability of agents. Unfortunately, the related approaches either depend on white-box access to LLMs or fail to accurately identify hallucinations. To address the challenge posed by hallucinations of intelligent agents, we present HalMit, a novel black-box watchdog framework that models the generalization bound of LLM-empowered agents and thus detect hallucinations without requiring internal knowledge of the LLM's architecture. Specifically, a probabilistic fractal sampling technique is proposed to generate a sufficient number of queries to trigger the incredible responses in parallel, efficiently identifying the generalization bound of the target agent. Experimental evaluations demonstrate that HalMit significantly outperforms existing approaches in hallucination monitoring. Its black-box nature and superior performance make HalMit a promising solution for enhancing the dependability of LLM-powered systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。