arXiv:2508.21051cs.CLcs.AI2025-08AAAI被引 7

用符号推理+大模型结合,让税务计算更准更可信。

Language Models and Logic Programs for Trustworthy Tax Reasoning

  • 将自然语言税则转为逻辑程序,再由大模型辅助推理
  • 在SARA数据集上准确率显著提升,部署成本低于270美元/人
  • 适合需要高可信度税务系统的开发者和政策制定者

美国国税局数据显示,普通美国人纳税需花费270美元和13小时。全球范围内的税务申报涉及复杂规则交叉与数值计算,出错可能带来高昂罚款,现有大语言模型(LLMs)难以满足准确性与可审计性要求。本文提出一种融合LLMs与符号求解器的方案,用于计算税务义务。在挑战性较强的法定推理评估(SARA)数据集上评估了多种变体,并引入基于真实罚金的系统部署成本估算方法。研究发现,预先将文本规则转化为形式化逻辑程序,并结合智能检索的形式化案例示例,能显著提升性能,使系统成本远低于实际平均值。结果表明,语义解析技术在法定推理中有效,且神经符号架构具有良好的经济可行性,有助于提升可靠税务协助的可及性。

原文摘要 · Abstract (English)

According to the United States Internal Revenue Service, ``the average American spends $\$270$ and 13 hours filing their taxes''. Even beyond the U.S., tax filing requires complex reasoning, combining application of overlapping rules with numerical calculations. Because errors can incur costly penalties, any automated system must deliver high accuracy and auditability, making modern large language models (LLMs) poorly suited for this task. We propose an approach that integrates LLMs with a symbolic solver to calculate tax obligations. We evaluate variants of this system on the challenging StAtutory Reasoning Assessment (SARA) dataset, and include a novel method for estimating the cost of deploying such a system based on real-world penalties for tax errors. We further show how combining up-front translation of plain-text rules into formal logic programs, combined with intelligently retrieved exemplars for formal case representations, can dramatically improve performance on this task and reduce costs to well below real-world averages. Our results demonstrate the effectiveness of applying semantic parsing methods to statutory reasoning, and show promising economic feasibility of neuro-symbolic architectures for increasing access to reliable tax assistance.

税务计算神经符号逻辑推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。