arXiv:2605.16297cs.CYcs.AI2026-05

为金融IT流程中的任务分级,判断大模型能否可靠替代人工

Task-Level AI Readiness Assessment for Business Process Management:The T-IPO Model and LARA Matrix in Financial-Services IT Operations

论文配图:Task-Level AI Readiness Assessment for Business Process Management:The T-IPO Model and LARA Matrix in Financial-Services IT Operations
图 1 · 摘自论文原文
  • 将任务拆解为八元组表示,用五维评分矩阵评估大模型适配度
  • 任务自动完成率从L1级95%降至L3级40%,合规性权重达1.5倍
  • 适合企业流程自动化团队与AI落地决策者参考

企业工作流中哪些任务可由大语言模型代理可靠处理?当前多数流程建模框架仍停留在活动层面,但单个活动可能包含难度迥异的工作。本文在金融服务业IT场景下提出两个设计工具:T-IPO将每个任务表示为八元组,LARA(LLM Agent Readiness Assessment)是五维度评分量表,用于评估任务对代理替代的准备度。合规敏感度权重设为1.5倍,经三轮德尔菲研究与AHP验证。量表分为四级(L1-L4),并设有保底规则:即使其他得分低,合规负荷最高的任务也不得低于L3级。两项工具嵌入我们提出的PARTIS方法论,并映射至BWW本体。在127个任务上评估,评分者间一致性达Fleiss' κ=0.80;三家机构复现结果κ=0.73。与活动级评估对比显示,任务级评估预测效用可能更优。120个任务实例的试点部署表明,自动完成率从L1级的95%单调下降至L3级的约40%。探索性因子分析揭示双因素结构:任务就绪度由认知执行复杂度与治理合规强度共同决定。最后提出校准程序LARA-TCA,以适应不断演进的大模型能力。

原文摘要 · Abstract (English)

Which tasks inside an enterprise workflow can a large-language-model agent reliably handle, and under what conditions? Most business process modeling frameworks still answer this at the activity level, even though a single activity can bundle work of radically different difficulty. This paper takes the analysis a step smaller. We describe two design artifacts developed in a financial-services IT setting: T-IPO, which represents each task as an eight-element tuple, and LARA (LLM Agent Readiness Assessment), a five-dimension rubric that scores a task's readiness for agent substitution. Compliance Sensitivity carries $1.5\times$ weight, a value we fixed through a three-round Delphi study and cross-checked with AHP. The rubric produces four levels, L1 to L4, and applies a floor rule so that a task with maximum compliance load cannot be classified below L3 no matter what the other scores say. Both artifacts sit inside a larger methodology (PARTIS) that we map onto BWW ontology in Section 3. We evaluate the instruments across 127 tasks. Inter-rater agreement reaches Fleiss' $κ= 0.80$; a replication at three further institutions returns $κ= 0.73$. A controlled comparison against activity-level assessment suggests, though does not prove, an improvement in predictive utility at the task level. Pilot deployment of 120 task instances confirms that auto-completion decays monotonically from $95\%$ at L1 through about $70\%$ at L2 to about $40\%$ at L3. Exploratory factor analysis points to a two-factor structure: task readiness seems to be determined jointly by cognitive-execution complexity and governance-compliance intensity. We close with a recalibration procedure (LARA-TCA) so the rubric can keep pace with evolving LLM capabilities.

流程自动化LLM评估金融IT

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。