测试大模型算税能力,发现现有效果不足三分之一。
TaxCalcBench: Evaluating Frontier Models on the Tax Calculation Task
- 构建税务计算基准测试,评估模型理解文本并精准算税的能力。
- 顶尖模型在简化样本上仅能正确计算不到三分之一的联邦个税申报表。
- 模型普遍存在误用税率表、计算错误和资格判断失误问题。
AI能帮你报税吗?目前还不能。计算美国个人所得税需要理解大量英文文本,并据此精确计算结果。我们提出TaxCalcBench,一个用于评估模型在给定必要信息时计算个人所得税申报表能力的基准测试。实验表明,即使在简化样本集上,当前最先进模型也仅能正确完成不到三分之一的联邦个税申报表计算。分析显示,模型普遍存在误用税率表、计算错误及资格判定失误等问题。研究结果表明,要将大语言模型应用于个人所得税计算,仍需构建额外基础设施。
原文摘要 · Abstract (English)
Can AI file your taxes? Not yet. Calculating US personal income taxes is a task that requires building an understanding of vast amounts of English text and using that knowledge to carefully compute results. We propose TaxCalcBench, a benchmark for determining models' abilities to calculate personal income tax returns given all of the necessary information. Our experiment shows that state-of-the-art models succeed in calculating less than a third of federal income tax returns even on this simplified sample set. Our analysis concludes that models consistently misuse tax tables, make errors in tax calculation, and incorrectly determine eligibility. Our findings point to the need for additional infrastructure to apply LLMs to the personal income tax calculation task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。