专用于德国税法分析的大型语言模型,精准处理法律条文引用与结构化推理。
SteuerLLM: Local specialized large language model for German tax law analysis
- 基于真实考试题生成合成数据,用检索增强管道训练领域专用模型
- 280亿参数模型在税法任务上超越多数更大通用模型,关键在领域适配而非参数量
- 开源完整数据集与模型,适合法律AI研究者及税务数字化从业者使用
大语言模型在通用推理和语言理解方面表现强劲,但在受严格规则、精确术语和法律约束结构支配的领域中性能下降。税法即为典型挑战:正确答案需精确引用法律条文、结构化法律论证,并在严格评分体系下保持数值准确。我们算法生成了首个源自真实德国大学税法考试的开放基准测试集SteuerEx,包含115道专家验证的考题,覆盖六个核心税法领域及多个学术层次,并采用逐条部分得分评估框架,贴近真实考试实践。我们进一步提出SteuerLLM,一种基于大规模合成数据训练的德国税法领域适配大模型,该数据由真实考试材料通过受控检索增强管道生成。SteuerLLM(28B参数)在多数任务中持续优于同规模通用指令调优模型,甚至超过部分更大系统,表明领域特定数据与架构适配对真实法律推理任务的表现比参数规模更具决定性。所有基准数据、训练数据、模型权重及评估代码均公开发布,以支持可复现的领域专用法律人工智能研究。SteuerLLM在线演示可访问:https://steuerllm.i5.ai.fau.de。
原文摘要 · Abstract (English)
Large language models (LLMs) demonstrate strong general reasoning and language understanding, yet their performance degrades in domains governed by strict formal rules, precise terminology, and legally binding structure. Tax law exemplifies these challenges, as correct answers require exact statutory citation, structured legal argumentation, and numerical accuracy under rigid grading schemes. We algorithmically generate SteuerEx, the first open benchmark derived from authentic German university tax law examinations. SteuerEx comprises 115 expert-validated examination questions spanning six core tax law domains and multiple academic levels, and employs a statement-level, partial-credit evaluation framework that closely mirrors real examination practice. We further present SteuerLLM, a domain-adapted LLM for German tax law trained on a large-scale synthetic dataset generated from authentic examination material using a controlled retrieval-augmented pipeline. SteuerLLM (28B parameters) consistently outperforms general-purpose instruction-tuned models of comparable size and, in several cases, substantially larger systems, demonstrating that domain-specific data and architectural adaptation are more decisive than parameter scale for performance on realistic legal reasoning tasks. All benchmark data, training datasets, model weights, and evaluation code are released openly to support reproducible research in domain-specific legal artificial intelligence. A web-based demo of SteuerLLM is available at https://steuerllm.i5.ai.fau.de.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。