让大模型在企业流程中更准更省还可审计。
FRAMES: Guarded and Dual-Objective Skill Evolution for Agents in Policy-Governed Enterprise Workflows

- 用共识变异+帕累托筛选进化技能,自动优化
- 准确率提升同时推理成本不变,且无回退
- 适合需要合规与可追溯的生产级AI应用
大语言模型代理正越来越多地应用于受政策约束的企业工作流,如文档审计,要求规则一致应用、每个值都有依据,并保持可审计性。提升这些代理面临挑战:操作反馈稀疏且无标签,修改一条规则可能引发其他情况回退,且需在不增加推理成本的前提下提高准确率。我们提出FRAMES,一种闭环框架:从现有资产冷启动可部署技能,通过基于共识的变异、准确率与成本的帕累托选择,以及反回退保证进行演化,全程保持可审计性。部署于内部生产系统后,FRAMES在准确率-成本权衡上优于所有基线,且在tau-bench上复现了相同成果。
原文摘要 · Abstract (English)
LLM agents increasingly run policy-bound enterprise workflows such as document auditing, where they must apply rules consistently, ground every value, and stay auditable. Improving these agents is hard: operational feedback is sparse and unlabeled, edits to one rule can regress unrelated cases, and accuracy must improve without inflating inference cost or losing auditability. We present FRAMES, a closed-loop framework that cold-starts deployable skills from existing assets and then evolves them through consensus-based mutation, Pareto selection over accuracy and cost, and an anti-regression guarantee, all while preserving auditability. Deployed on our internal production system, FRAMES attains the best accuracy-cost trade-off among baselines, with the same gains reproduced on tau-bench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。