测试大模型对税务罚单合法性的判断能力,发现其推理仍有限。
Taxation Perspectives from Large Language Models: A Case Study on Additional Tax Penalties
- 构建300例韩国土税判例数据集,含选择题与论述题
- 大模型在复杂案例中准确率不足,尤其难把握法律推理关键环节
- 适合研究AI法律推理、税务AI应用的学者与从业者参考
大语言模型在税务领域的表现如何?尽管已有大量法律领域研究,但专门针对税务的研究仍很稀缺,且现有数据集或过于简化,或未开源。为填补这一空白,我们提出PLAT——一个用于评估大模型预测额外税罚合法性能力的新基准。PLAT包含300个样本:100个二选一题、100个多选题、100道论述题,全部来自100个韩国法院判例。该数据集旨在检验模型对税法的理解力及处理需综合推理的复杂案件的能力。系统性实验表明:(1)模型基础能力有限,尤其在涉及矛盾问题、需结合纳税人具体情况时;(2)即使是o3等具备推理增强能力的先进模型,在IRAC框架中的‘AC’(分析与结论)阶段仍表现不佳。数据集已公开于Hugging Face。
原文摘要 · Abstract (English)
How capable are large language models (LLMs) in the domain of taxation? Although numerous studies have explored the legal domain, research dedicated to taxation remains scarce. Moreover, the datasets used in these studies are either simplified, failing to reflect the real-world complexities, or not released as open-source. To address this gap, we introduce PLAT, a new benchmark designed to assess the ability of LLMs to predict the legitimacy of additional tax penalties. PLAT comprises 300 examples: (1) 100 binary-choice questions, (2) 100 multiple-choice questions, and (3) 100 essay-type questions, all derived from 100 Korean court precedents. PLAT is constructed to evaluate not only LLMs' understanding of tax law but also their performance in legal cases that require complex reasoning beyond straightforward application of statutes. Our systematic experiments with multiple LLMs reveal that (1) their baseline capabilities are limited, especially in cases involving conflicting issues that require a comprehensive understanding (not only of the statutes but also of the taxpayer's circumstances), and (2) LLMs struggle particularly with the "AC" stages of "IRAC" even for advanced reasoning models like o3, which actively employ inference-time scaling. The dataset is publicly available at: https://huggingface.co/collections/sma1-rmarud/plat-predicting-the-legitimacy-of-punitive-additional-tax
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。