分阶段混合任务提升跨领域操作流程理解能力
FM SO.P: A Progressive Task Mixture Framework with Automatic Evaluation for Cross-Domain SOP Understanding
- 分三阶段逐步训练:术语消歧、流程顺序理解、场景图推理
- 7B模型达34.3%通过率,仅用10%参数媲美72B大模型
- 自动多智能体评估适配不同领域规则与测试标准
标准操作流程(SOP)对企事业单位运营至关重要,但现有语言模型在SOP理解与跨领域泛化上表现不佳。原因在于联合训练无法区分术语精准性、流程顺序性与约束推理等核心能力。本文提出FM SO.P框架,通过两项创新解决该问题:第一,设计渐进式任务混合机制,分三阶段依次引入概念消歧(术语精准)、动作序列理解(流程正确)、场景感知图推理(条件逻辑),并累积数据训练;第二,构建自动多智能体评估系统,由三个智能体协同生成评分标准、分层测试集和评分打分,可自适应不同领域需求(如驾照办理的时间约束、银行合规要求)。在涵盖银行、驾照中心、医疗、市场、大学、图书馆、酒店共七个领域的SOPBench基准上,使用32B模型达到48.3%通过率,7B开源模型达34.3%,与Qwen-2.5-72B-Instruct基线(34.4%)相当,但仅需其十分之一的参数量。
原文摘要 · Abstract (English)
Standard Operating Procedures (SOPs) are critical for enterprise operations, yet existing language models struggle with SOP understanding and cross-domain generalization. Current methods fail because joint training cannot differentiate between reasoning capabilities that SOP requires: terminology precision, sequential ordering, and constraint reasoning. We propose FM SO.P, solving these challenges through two novelties. First, we introduce progressive task mixtures that build capabilities by stages across three task types with cumulative data: concept disambiguation for terminology precision, action sequence understanding for procedural correctness, and scenario-aware graph reasoning for conditional logic. Second, we propose an automatic multi-agent evaluation system consisting of three agents that adaptively generate rubrics, stratified test sets, and rubric scoring, adapting to domains (e.g., temporal constraints for DMV, regulatory compliance for banking). Evaluated on SOPBench across seven domains (Bank, DMV, Healthcare, Market, University, Library, Hotel), FM SO.P achieves 48.3\% pass rate with our 32B model and 34.3\% with our opensource 7B model, matching Qwen-2.5-72B-Instruct baseline (34.4\%) with 10x fewer parameters.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。