用AI代理自动优化文档处理配置,效率提升10倍以上。
IDP AutoOpt: Agent-Driven Optimization of Document Processing Pipeline Configurations

- 基于大模型的自主代理闭环优化文档处理配置
- 提取任务准确率90.2%,成本降低4.6倍,耗时从周级降至两小时
- 适合需要快速部署企业级AI系统的团队
我们提出IDP AutoOpt,一个自主的LLM代理,用于发现智能文档处理(IDP)流水线的高性能配置。当前联合调优提示词、模型、OCR设置和模式需领域专家投入20至80+人时/文档类型,且难以扩展。IDP AutoOpt采用闭环流程:在小规模标注集上评分,诊断字段级错误,生成针对性修改并重新评估,由人工编写的领域技能引导,体现生产经验。在医疗、营销情报和金融领域的抽取、分类与分包任务中,其表现匹配或超越人类专家,在同等或更低成本下达成更高精度(抽取基准上90.2%对比81.6%,每页成本降低4.6倍),将配置时间从数周缩短至两小时内。我们进一步表明,代理能力存在不可逾越的阈值,结构化领域技能优于原始代码访问,后者若无组织反而降低性能。还分享了上下文管理与方差控制的实践经验。该方法仅需可配置流水线、评分函数和少量标注数据,可扩展至RAG及多智能体工作流等其他企业AI系统。
原文摘要 · Abstract (English)
We present IDP AutoOpt, an autonomous LLM agent that discovers high-performing configurations for intelligent document processing (IDP) pipelines. Tuning IDP prompts, models, OCR settings, and schemas jointly currently costs domain specialists 20 to 80+ person-hours per document type and does not scale as enterprises add document classes. IDP AutoOpt runs a closed loop: it scores a configuration on a small labeled set, diagnoses field-level errors, generates targeted edits, and re-evaluates, guided by human-authored domain skills that encode production expertise. Across extraction, classification, and packet-splitting tasks deployed in healthcare, marketing-intelligence, and financial-services settings, IDP AutoOpt matches or exceeds human-expert accuracy at equal or lower cost (on an extraction benchmark, 90.2% vs 81.6% at 4.6 x lower per-page cost), cutting configuration time from weeks to under two hours. We further show that agent LLM capability has a hard threshold below which optimization fails, and that curated domain skills outperform raw source-code access, which can degrade performance when provided without structure. We also share practical lessons on context management and variance mitigation. Requiring only a configurable pipeline, a scoring function, and a small labeled set, the approach extends beyond IDP to other enterprise AI systems, such as RAG and multi-agent workflows, where configuration bottlenecks deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。