arXiv:2606.01513cs.DCcs.AI2026-06被引 1

用合规得分筛选生成结果,提升支付争议文档生成的准确与效率。

Compliance-Scored Best-of-N Guardrail Orchestration for Multimodal Document Generation in Payments Dispute Defense

论文配图:Compliance-Scored Best-of-N Guardrail Orchestration for Multimodal Document Generation in Payments Dispute Defense
图 1 · 摘自论文原文
  • 并行生成多份文档,按合规得分优先选择最优输出。
  • 20秒内完成5次尝试,合规率高达91%。
  • 适合需要高合规性与低延迟的企业级文档生成场景。

高风险企业文档生成(如金融争议叙述、合规通知、审计摘要)需保证结构正确性、政策合规性及大规模低延迟运行。此前系统将PII脱敏、内容审核、格式校验等步骤分散处理,导致逻辑碎片化、响应慢、运维成本高。本文提出一种文本与图像输入的合规护航编排层,将多候选生成与显式合规评分结合,实现早期退出。该框架可配置并行生成头,对候选输出按加权护航规则(包括PII检测、内容审核、结构约束、领域规则)打分,返回最优结果并附带选择元数据。运营数据显示:20秒内完成5次尝试,合规率达91%。针对支付争议辩护摘要,通过聚合操作读数分析而非随机A/B测试,变量组整体胜率高于对照组(301/659 对比 536/1548),提升11.0个百分点(95%置信区间[6.6, 15.5],p < 0.001);调整后未收货案例提升7.5个百分点(95%置信区间[0.2, 15.7],p = 0.045)。欺诈与本地证据排名差异方向为正但不显著。还报告了770条生成证据的评审员校准结果与70例OCR样本,并完整披露请求接口、评分逻辑、伪代码及可复现边界。

原文摘要 · Abstract (English)

High-stakes enterprise document generation, including financial dispute narratives, compliance notices, and audit summaries, demands schema correctness, policy compliance, and low-latency operation at scale. Prior to a unified guardrail layer, production systems often stitched together separate PII redaction, content moderation, and format validation steps, leading to fragmented logic, slower request paths, and higher operational cost. We present a guardrail orchestration layer for text and image inputs that couples multi-candidate generation with an explicit compliance score used for early exit. The framework runs configurable parallel generation heads, scores candidates against weighted guardrails including PII detection, content moderation, schema constraints, and domain rules, and returns the best-scoring output with selection metadata. The available operational readout reports 5 attempts within 20 seconds and 91 percent compliance. For payments dispute defense summaries, we analyze aggregate operational scenario readouts rather than a randomized A/B test. Variable cohorts show higher count win rates than controls overall, 301/659 versus 536/1548, corresponding to +11.0 percentage points with 95 percent confidence interval [6.6, 15.5] and p < 0.001, and for adjusted item-not-received cases, +7.5 percentage points with 95 percent confidence interval [0.2, 15.7] and p = 0.045. Fraud and local evidence-ranking deltas are directionally positive but not statistically significant from the aggregate count data. We also report reviewer-calibrated Responsible-AI evidence-quality signals from 770 generated-evidence reviews and a 70-case OCR slice, and document the reproducibility boundary through the request interface, scoring logic, pseudocode, and operational evidence boundary.

文档生成合规评分支付争议多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。