测试大模型在法律咨询中识别信息缺失的能力
InsufficiencyBench: Evaluating LLM legal advice on underspecified user queries

- 构建首个针对查询不完整性的法律AI评测基准
- 模型对缺失要素识别准确率不足46%,平均召回仅44%
- 适合法律AI安全评估与司法合规研究者使用
法律AI系统日益用于解答法律问题,但现有评测假设问题已完整。现实中用户常遗漏关键事实,而这些事实直接影响法律结论。我们提出InsufficiencyBench,首个聚焦查询侧信息不足的法律评测基准:评估模型是否能识别查询缺乏法律关键信息、判断缺什么、并避免过早下结论。我们定义了三大结构失效模式下的八类典型缺失要素,并构建202个评测条目(58个基础问题,144个缺陷变体),覆盖六个法律领域和24个美国司法管辖区,由执业律师标注。评估十个前沿模型发现,无一模型在缺失要素识别上的F2超过0.46,中位召回率仅为0.44。模型或盲目兜底,或在虚构前提下沉默回应。无一模型既能准确识别缺陷问题,又能合理回应完整问题。
原文摘要 · Abstract (English)
Legal AI systems are increasingly used to answer legal questions, yet existing benchmarks assume queries arrive fully specified. In practice, users omit facts that materially determine the legal outcome. We introduce InsufficiencyBench, the first legal benchmark targeting query-side insufficiency: whether a model recognizes when a query lacks legally material information, identifies what is missing, and refrains from premature conclusions. We formalize a taxonomy of eight canonical missing-element categories across three structural failure modes---switch, gating, and fatal prerequisite--- and construct 202 benchmark items (58 base queries, 144 deficient variants) spanning six legal domains and 24 US jurisdictions and annotated by practising attorneys. Evaluating ten frontier models, we find that no model exceeds F2 = 0.46 on missing-element identification and that the median recall is 0.44. Models either hedge indiscriminately or answer silently under fabricated presumptions. No model both identifies and qualifies responses to deficient queries while directly addressing complete ones.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。