SIRIN统一检测大模型在检索与记忆增强下的虚假回答,支持实时分析和多方法对比。
SIRIN: A Unified Toolkit for Detecting Contextual Hallucinations in Retrieval-Augmented and Memory-Grounded LLM Systems

- 整合三种检测方法:表征探针、不确定性估计与判别式验证
- 支持响应级与片段级检测,覆盖白盒与黑盒场景
- 适合研究者与开发者用于提升模型可信度与长期记忆系统可靠性
SIRIN(语义不一致识别与检查枢纽)是一个统一的工具包及交互式网页界面,用于检测检索增强型、代理型及基于记忆的大语言模型系统中的上下文幻觉(即流畅但无证据支持的回应)。SIRIN融合了三种检测范式(表征探针、不确定性估计、判别式验证)以及生成前查询可回答性这一互补任务,集成于同一界面、配置系统与评估流程中,支持响应级与片段级的检查,适用于白盒与黑盒设置。其网页界面可对用户提供的上下文-查询-回答三元组进行实时分析,提供幻觉评分、未支持片段高亮及多检测器并行比对,并采用轻量级插件设计,便于新增检测器。我们展示了SIRIN在幻觉检测、查询可回答性判断以及作为长期记忆系统的忠实性闸门的应用。源代码已公开于https://github.com/sb-ai-lab/SIRIN。
原文摘要 · Abstract (English)
SIRIN (Semantic Inconsistency Recognition and Inspection Nexus) is a unified toolkit and interactive web UI for detecting contextual hallucinations (fluent, plausible responses unsupported by the provided evidence) in retrieval-augmented, agentic, and memory-grounded LLM systems. SIRIN unifies three detector paradigms (representation probing, uncertainty estimation, and judge-style verification) and the complementary task of pre-generation query answerability under one interface, configuration system, and evaluation pipeline, supporting response- and span-level inspection in both white-box and black-box settings. The web UI enables live analysis of user-supplied context-query-answer triples through hallucination scores, unsupported-span highlighting, and side-by-side detector comparison, with a lightweight plug-in design for adding new detectors. We demonstrate SIRIN on hallucination detection, query answerability, and as a faithfulness gate within long-term memory systems. The source code is publicly available at https://github.com/sb-ai-lab/SIRIN.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。