用智能体动态调用法规,提升法律推理准确率
DAR: Deontic Reasoning with Agentic Harnesses

- 让模型按需调用法规文本,动态构建推理路径
- 在DeonticBench硬集上显著提升推理正确率
- 适合法律、合规等需要精准规则执行的场景
德性推理是根据明确规则和政策对具体案例作出判断的任务,例如依据税法计算税负或裁定移民上诉结果。基于大模型的德性推理面临关键挑战:规则集可能冗长且相互引用,模型难以定位特定推理步骤所需规则。本文提出德性智能体推理(DAR),一种模型按需与法规交互的智能体推理框架。我们在DeonticBench的多个硬子集上评估了DAR在不同智能体设定下的表现。结果显示,智能体框架能推动德性推理任务的边界,但提升并非均匀:弱模型在数值任务上表现反而下降,且消耗大量令牌。
原文摘要 · Abstract (English)
Deontic reasoning is the task of answering questions by applying explicit rules and policies to case-specific facts, for example computing tax liability under a statute or determining the outcome of an immigration appeal. A key technical challenge for LLM-based deontic reasoning is that the relevant ruleset can be long and cross-referenced, so models may still fail to locate the rules needed for a particular reasoning step. We introduce Deontic Agentic Reasoning (DAR), an agentic reasoning setup in which the model interacts with the statutes on demand. We evaluate DAR under multiple harnesses on hard subsets of DeonticBench. Across these settings, we find that agentic harnesses can push the frontier on deontic reasoning tasks, but improvements are not uniform: weaker models often degrade on numerical tasks while consuming far more tokens.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。