arXiv:2512.14554cs.CLcs.AI2025-12

首个面向越南法律推理的AI评估基准,专为测试大模型法律理解力设计。

VLegal-Bench: Cognitively Grounded Benchmark for Vietnamese Legal Reasoning of Large Language Models

  • 基于认知分类理论构建多层级法律任务
  • 含10450个专家标注的真实法律样本
  • 适合研究越南法律AI或开发智能法律助手者

大语言模型在法律领域的应用迅速发展,但越南法律体系复杂、层级分明且频繁修订,给模型法律知识理解与应用能力的评估带来挑战。为此,本文提出首个系统性评估越南法律任务的基准——VLegal-Bench。该基准基于布卢姆认知分类理论,设计涵盖法律理解多个层次的任务,模拟真实法律辅助工作流程,包括通用问答、检索增强生成、多步推理及情景化问题解决。所有10,450个样本均由法律专家通过严格标注与交叉验证流程生成,确保内容源自权威法律文件并反映实际应用场景。该基准提供标准化、透明且认知驱动的评估框架,为提升越南法律AI系统的可靠性、可解释性与伦理对齐性奠定基础。公开访问页面已上线:https://vilegalbench.cmcai.vn/。

原文摘要 · Abstract (English)

The rapid advancement of large language models (LLMs) has enabled new possibilities for applying artificial intelligence within the legal domain. Nonetheless, the complexity, hierarchical organization, and frequent revisions of Vietnamese legislation pose considerable challenges for evaluating how well these models interpret and utilize legal knowledge. To address this gap, the Vietnamese Legal Benchmark (VLegal-Bench) is introduced, the first comprehensive benchmark designed to systematically assess LLMs on Vietnamese legal tasks. Informed by Bloom's cognitive taxonomy, VLegal-Bench encompasses multiple levels of legal understanding through tasks designed to reflect practical usage scenarios. The benchmark comprises 10,450 samples generated through a rigorous annotation pipeline, where legal experts label and cross-validate each instance using our annotation system to ensure every sample is grounded in authoritative legal documents and mirrors real-world legal assistant workflows, including general legal questions and answers, retrieval-augmented generation, multi-step reasoning, and scenario-based problem solving tailored to Vietnamese law. By providing a standardized, transparent, and cognitively informed evaluation framework, VLegal-Bench establishes a solid foundation for assessing LLM performance in Vietnamese legal contexts and supports the development of more reliable, interpretable, and ethically aligned AI-assisted legal systems. To facilitate access and reproducibility, we provide a public landing page for this benchmark at https://vilegalbench.cmcai.vn/.

法律AI大模型评估越南语认知分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。