arXiv:2410.08764cs.CL2024-10被引 7

评估法律问答生成结果的可信度,提升AI在法律场景的可靠性。

Measuring the Groundedness of Legal Question-Answering Systems

  • 构建法律领域专用的置信度评估数据集,测试生成回答与原文的匹配度。
  • 最佳方法在分类任务中达到0.8的宏F1值,有效识别无依据的回答。
  • 兼顾检测速度,适合集成到真实法律AI系统中作为校验环节。

在法律等高风险领域,生成式AI系统的准确性与可信度至关重要。本文提出一个全面的基准测试框架,用于评估AI生成回答的置信度(groundedness),旨在显著提升其可靠性。实验涵盖基于相似度的度量方法与自然语言推理模型,判断回答是否合理依赖于上下文。同时探索了多种大语言模型提示策略,以增强对无依据回答的检测能力。通过新构建的置信度分类语料库进行验证,该数据集专为法律查询及检索增强提示下的回答设计,聚焦生成内容与原始资料的一致性。结果显示,最佳方法在置信度分类任务中获得0.8的宏F1得分。此外,还评估了各方法的延迟表现,以判断其在实际应用中的可行性,该环节常用于触发人工复核或自动重生成。本研究证明了多种检测方法在提升法律领域生成式AI可信度方面的潜力。

原文摘要 · Abstract (English)

In high-stakes domains like legal question-answering, the accuracy and trustworthiness of generative AI systems are of paramount importance. This work presents a comprehensive benchmark of various methods to assess the groundedness of AI-generated responses, aiming to significantly enhance their reliability. Our experiments include similarity-based metrics and natural language inference models to evaluate whether responses are well-founded in the given contexts. We also explore different prompting strategies for large language models to improve the detection of ungrounded responses. We validated the effectiveness of these methods using a newly created grounding classification corpus, designed specifically for legal queries and corresponding responses from retrieval-augmented prompting, focusing on their alignment with source material. Our results indicate potential in groundedness classification of generated responses, with the best method achieving a macro-F1 score of 0.8. Additionally, we evaluated the methods in terms of their latency to determine their suitability for real-world applications, as this step typically follows the generation process. This capability is essential for processes that may trigger additional manual verification or automated response regeneration. In summary, this study demonstrates the potential of various detection methods to improve the trustworthiness of generative AI in legal settings.

法律AI置信度评估生成式AI评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。