arXiv:2511.04705cs.CLcs.AI2025-11

首个面向政务双语政策任务的多维评测基准,提升大模型在政府场景下的可靠性。

POLIS-Bench: Towards Multi-Dimensional Evaluation of LLMs for Bilingual Policy Tasks in Governmental Scenarios

  • 构建最新双语政策语料库,扩大评估样本规模
  • 设计三类场景化任务,全面测试模型理解与应用能力
  • 创新双指标评估框架,兼顾内容对齐与任务合规性

我们提出POLIS-Bench,首个专为政府双语政策场景下大语言模型设计的系统性评估基准。相较于现有基准,其三大创新为:(i) 构建大规模、时效性强的双语政策语料库,显著扩大有效评估样本量,确保与当前治理实践相关;(ii) 设计三类场景化任务——条款检索与解释、解决方案生成、合规性判断,全面探测模型理解与应用能力;(iii) 提出结合语义相似度与准确率的双指标评估框架,精确衡量内容一致性与任务要求符合度。对超过10个前沿大模型的大规模评估显示,推理型模型在跨任务中表现更稳定、准确,凸显合规任务挑战性。基于该基准,我们成功微调轻量开源模型,POLIS系列模型在多项政策子任务上达到或超越主流闭源基线,成本显著降低,为政府场景下可靠、合规、低成本部署提供可行路径。

原文摘要 · Abstract (English)

We introduce POLIS-Bench, the first rigorous, systematic evaluation suite designed for LLMs operating in governmental bilingual policy scenarios. Compared to existing benchmarks, POLIS-Bench introduces three major advancements. (i) Up-to-date Bilingual Corpus: We construct an extensive, up-to-date policy corpus that significantly scales the effective assessment sample size, ensuring relevance to current governance practice. (ii) Scenario-Grounded Task Design: We distill three specialized, scenario-grounded tasks -- Clause Retrieval & Interpretation, Solution Generation, and the Compliance Judgmen--to comprehensively probe model understanding and application. (iii) Dual-Metric Evaluation Framework: We establish a novel dual-metric evaluation framework combining semantic similarity with accuracy rate to precisely measure both content alignment and task requirement adherence. A large-scale evaluation of over 10 state-of-the-art LLMs on POLIS-Bench reveals a clear performance hierarchy where reasoning models maintain superior cross-task stability and accuracy, highlighting the difficulty of compliance tasks. Furthermore, leveraging our benchmark, we successfully fine-tune a lightweight open-source model. The resulting POLIS series models achieves parity with, or surpasses, strong proprietary baselines on multiple policy subtasks at a significantly reduced cost, providing a cost-effective and compliant path for robust real-world governmental deployment.

大模型评测政务AI双语处理合规性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。