arXiv:2604.16270cs.CLcs.AI2026-04中稿 · the FISU Joint Con…

评测大模型在越南法律文本上的真实能力,发现读得顺不等于懂法律。

From Benchmarking to Reasoning: A Dual-Aspect, Large-Scale Evaluation of LLMs on Vietnamese Legal Text

论文配图:From Benchmarking to Reasoning: A Dual-Aspect, Large-Scale Evaluation of LLMs on Vietnamese Legal Text
图 1 · 摘自论文原文
  • 构建双维度评估框架,测准确率、可读性与一致性
  • 揭示模型在细节法律准确性上存在严重推理缺陷
  • 适合法律AI研究者和司法智能化开发者参考

越南法律文本的复杂性阻碍公众获得司法正义。尽管大语言模型在法律文本简化方面前景广阔,但评估其真实能力需超越表面指标。本文提出一个双维度大规模评估框架:首先,在准确率、可读性和一致性三个维度上对四种先进大模型(GPT-4o、Claude 3 Opus、Gemini 1.5 Pro、Grok-1)建立性能基准;其次,基于60篇精选越南法律条文的专家验证错误分类体系,开展大规模错误分析。结果揭示关键权衡:Grok-1在可读性和一致性上表现优异,但牺牲了细粒度法律准确性;Claude 3 Opus虽有高准确率,却隐藏大量细微但关键的推理错误。错误分析指出‘错误举例’和‘误解读’是最常见失败类型,表明当前大模型的核心挑战并非摘要,而是受控的法律推理。通过量化基准与质性深度分析结合,本研究为法律应用中的大模型评估提供了全面且可操作的方案。

原文摘要 · Abstract (English)

The complexity of Vietnam's legal texts presents a significant barrier to public access to justice. While Large Language Models offer a promising solution for legal text simplification, evaluating their true capabilities requires a multifaceted approach that goes beyond surface-level metrics. This paper introduces a comprehensive dual-aspect evaluation framework to address this need. First, we establish a performance benchmark for four state-of-the-art large language models (GPT-4o, Claude 3 Opus, Gemini 1.5 Pro, and Grok-1) across three key dimensions: Accuracy, Readability, and Consistency. Second, to understand the "why" behind these performance scores, we conduct a large-scale error analysis on a curated dataset of 60 complex Vietnamese legal articles, using a novel, expert-validated error typology. Our results reveal a crucial trade-off: models like Grok-1 excel in Readability and Consistency but compromise on fine-grained legal Accuracy, while models like Claude 3 Opus achieve high Accuracy scores that mask a significant number of subtle but critical reasoning errors. The error analysis pinpoints \textit{Incorrect Example} and \textit{Misinterpretation} as the most prevalent failures, confirming that the primary challenge for current LLMs is not summarization but controlled, accurate legal reasoning. By integrating a quantitative benchmark with a qualitative deep dive, our work provides a holistic and actionable assessment of LLMs for legal applications.

法律AI大模型评测越南语推理能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。