arXiv:2601.03553cs.CLcs.AI2026-01AAAI被引 4

为警用大模型设计评估框架,发现商用模型在事实推理上表现不佳。

Evaluating LLMs for Police Decision-Making: A Framework Based on Police Action Scenarios

  • 构建覆盖执法全流程的警察行动场景评估框架
  • 基于8000+官方文件创建新数据集,验证模型在事实推荐上能力不足
  • 适合关注警务AI安全与评估的研究者和政策制定者

大型语言模型在警务中的应用日益增多,但缺乏针对警务场景的评估框架。尽管模型回答未必违法,但未经验证的使用仍可能导致非法逮捕和证据收集不当等问题。为此,本文提出PAS(警察行动场景)框架,系统覆盖评估全过程。基于该框架,我们从超过8000份官方文档构建了新的问答数据集,并通过统计分析与警务专家判断验证了关键评估指标。实验表明,商用大模型在本研究所设的警务任务中表现较差,尤其在提供基于事实的建议方面存在明显短板。研究强调建立可扩展的评估体系对保障警务人工智能可靠性至关重要。相关数据与提示模板已公开。

原文摘要 · Abstract (English)

The use of Large Language Models (LLMs) in police operations is growing, yet an evaluation framework tailored to police operations remains absent. While LLM's responses may not always be legally incorrect, their unverified use still can lead to severe issues such as unlawful arrests and improper evidence collection. To address this, we propose PAS (Police Action Scenarios), a systematic framework covering the entire evaluation process. Applying this framework, we constructed a novel QA dataset from over 8,000 official documents and established key metrics validated through statistical analysis with police expert judgements. Experimental results show that commercial LLMs struggle with our new police-related tasks, particularly in providing fact-based recommendations. This study highlights the necessity of an expandable evaluation framework to ensure reliable AI-driven police operations. We release our data and prompt template.

大模型评估警务AI事实推理数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。