首个面向中文税务实践的可扩展评估基准,填补真实场景评测空白。
TaxPraBen: A Scalable Benchmark for Structured Evaluation of LLMs in Chinese Real-World Tax Practice

- 构建10项传统任务+3类真实税务场景的综合评估体系
- 19个模型测试显示闭源大模型表现最优,中文模型普遍优于多语言模型
- 适合关注AI在法律财税领域落地的研究者与开发者
尽管大型语言模型在通用领域表现优异,但在高度专业化、知识密集且受法律监管的中文税务领域仍存在明显短板。现有税务评测多聚焦孤立NLP任务,忽视真实业务能力。为此,我们提出TaxPraBen,首个专用于中文税务实践的基准。它融合10项传统应用任务及3项开创性真实场景——税务风险防范、税务稽查分析、税务策略规划,源自14个数据集,共7.3K实例。TaxPraBen采用“结构化解析-字段对齐提取-数值与文本匹配”的可扩展评估范式,实现端到端税务实践评估,并具备向其他领域扩展的能力。基于布卢姆分类法评估19个LLMs,结果表明显著性能差异:所有闭源大参数模型表现优异,中文模型如Qwen2.5普遍优于多语言模型,而经少量税务数据微调的YaYi2模型仅获有限提升。TaxPraBen为推进LLMs在实际应用中的评估提供关键资源。
原文摘要 · Abstract (English)
While Large Language Models (LLMs) excel in various general domains, they exhibit notable gaps in the highly specialized, knowledge-intensive, and legally regulated Chinese tax domain. Consequently, while tax-related benchmarks are gaining attention, many focus on isolated NLP tasks, neglecting real-world practical capabilities. To address this issue, we introduce TaxPraBen, the first dedicated benchmark for Chinese taxation practice. It combines 10 traditional application tasks, along with 3 pioneering real-world scenarios: tax risk prevention, tax inspection analysis, and tax strategy planning, sourced from 14 datasets totaling 7.3K instances. TaxPraBen features a scalable structured evaluation paradigm designed through process of "structured parsing-field alignment extraction-numerical and textual matching", enabling end-to-end tax practice assessment while being extensible to other domains. We evaluate 19 LLMs based on Bloom's taxonomy. The results indicate significant performance disparities: all closed-source large-parameter LLMs excel, and Chinese LLMs like Qwen2.5 generally exceed multilingual LLMs, while the YaYi2 LLM, fine-tuned with some tax data, shows only limited improvement. TaxPraBen serves as a vital resource for advancing evaluations of LLMs in practical applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。