构建可自我进化法律代理框架,让AI像律所一样从经验中学习。
Parthenon Law: A Self-Evolving Legal-Agent Framework
- 分模型、工具、技能等模块设计可审计的法律代理架构
- 12510次实验显示强模型仍难一次性完成法律任务
- 通过失败反馈自动优化技能与知识,不改模型权重
随着智能体能力提升,法律领域大模型代理有望将文书繁重的工作转化为可审查的成果——但可靠部署面临三大障碍:缺乏对当前最强模型与调用组合在端到端法律任务上的大规模实证;缺乏专为法律垂直领域设计的代理架构,仅有通用调用框架;在事实、法规和截止日期不断变化的环境中,缺乏系统基于自身结果进行学习的机制。本文逐一应对。基于 Harvey LAB 的大规模实证研究(12,510 条代理轨迹)表明,即使前沿代理在单次通过中仍难以完成任务:尽管逐项准确率随模型增强而提升,但严格意义上的任务完成率停滞不前。随后提出 extsc{Parthenon}——一种自演化法律代理框架,将模型、调用、代理角色、法律知识、确定性工具和程序技能拆解为可审计的表征面,支持溯源追踪、时间与数值锚定、交付物合规与问题闭环。最后引入反泄露学习环,将评分失败转化为与任务无关的技能、工具与知识修订,使系统能像律所一样通过每件案件的经验持续优化,而无需修改模型权重。大规模实证分析显示, extsc{Parthenon} 显著提升了现有最先进模型与调用组合在法律任务上的表现。
原文摘要 · Abstract (English)
As agents grow more capable, legal-domain LLM agents promise to turn document-heavy matters into reviewable work products -- yet reliable deployment faces three obstacles: no large-scale evidence on how today's strongest model-and-harness combinations behave on end-to-end legal matters; no agent architecture adapted to the legal vertical, only general-purpose harnesses; and, in a setting that keeps shifting with new facts, authorities, and deadlines, no mechanism for systems to learn from their own outcomes. We address each. A large-scale empirical study on Harvey LAB -- $12{,}510$ agent trajectories -- shows that even frontier agents remain far from completing matters in a single pass: per-criterion accuracy climbs with stronger models while strict matter completion stalls. We then introduce \textsc{Parthenon}, a self-evolving legal-agent framework that factors Model, Harness, Agent roles, legal Knowledge, deterministic Tools, and procedural Skills into auditable surfaces for source traceability, date and number grounding, deliverable compliance, and issue closure. Finally, an anti-leakage learning loop converts scored failures into task-agnostic edits to skills, tools, and knowledge, letting the system improve with experience -- as a firm refines its checklists and playbooks after each matter -- without touching model weights. Across our large-scale empirical analysis, \textsc{Parthenon} substantially improves the performance of state-of-the-art models and harnesses on legal-matter tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。