用外部状态控制器让大模型在法庭质询中持续推进,避免卡壳。
Enforcing Monotonic Progress in Legal Cross-Examination: Preventing Long-Horizon Stagnation in LLM-Based Inquiry
- 设计神经符号架构,通过外部控制器强制推进关键信息单元
- 实测三起命案质询任务完成度超97%,冗余接近零
- 适合需严格流程控制的法律、医疗等高风险场景
大型语言模型虽具备出色的语言流畅性,但在明确程序约束下的长周期任务中难以可靠完成。在法庭质询场景中,纯概率生成常维持行为连贯性,却无法确保程序进展。我们将其失败归因于程序停滞,并提出软有限状态机(Soft-FSM),通过外部确定性状态控制器,对累积的关键信息单元(KIUs)强制单调推进。在三个真实台湾刑事杀人案件上的实验表明,基线方法完成度低于40%,而Soft-FSM始终超过97%且冗余近乎为零。结果表明,在此类领域,仅依赖大模型的涌现行为无法保证任务完成,必须通过显式且可验证的外部状态控制来实现可靠执行。
原文摘要 · Abstract (English)
Large language models (LLMs) exhibit impressive linguistic fluency but struggle to reliably complete long-horizon tasks under explicit procedural constraints. In legal cross-examination, purely proba-bilistic generation often maintains behavioral coherence while failing to ensure procedural advancement. We characterize this failure as procedural stagnation and propose Soft-FSM, a neuro-symbolic architecture that enforces monotonic progress over accumulated Key Information Units (KIUs) via an external deterministic state controller. Experiments on three real-world Taiwanese criminal homicide cases show that baseline methods collapse below 40% completeness, while Soft-FSM consistently achieves over 97% with near-zero redundancy. These results suggest that, in such domains, reliable task completion cannot be guaranteed by emergent LLM behavior alone, and can be reliably enforced through explicit and verifiable external state control.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。