通过双循环反馈优化提示,提升关键文档生成的准确性与可审计性。
Hierarchical Online Prompt Mutation with Dual-Loop Feedback for Guardrailed Evidence Document Generation: A Production-Evaluation Case Study

- 分层在线提示变异框架,动态调整提示策略以适应真实业务场景。
- 相比静态提示,生成准确率提升11.0个百分点,金额加权准确率提高19.1个百分点。
- 适合需要高可靠性、可追溯性的生产级文本生成系统开发者参考。
高风险生产级文档生成系统要求语言模型具备自适应性、证据支撑性和可审计性。本文提出HOPM框架,在真实市场争议证据工作流中进行评估。HOPM将提示视为在线策略:家族/版本路由选择提示,确定性护栏将失败归因于可变提示词类别,人类评审与自动裁判的双重反馈同时更新路由与变异优先级。主证据来自一次匹配的生产-评估消融实验:七种变体在相同600个案例上测试,对比静态提示、人工迭代、仅强化学习路由、仅变异适应、仅人工反馈、仅自动裁判反馈及完整双循环HOPM。完整HOPM使计数胜率从34.7%提升至45.7%(+11.0个百分点;配对McNemar检验p=1.31e-11),金额加权胜率从22.3%升至41.4%(+19.1个百分点;95%配对自举置信区间[10.3, 28.9]个百分点)。平均李克特评分从3.18升至4.40,问题标记率从15.3%降至5.2%。支持性评审资料包含770条生成文本评审、318份标注评审导出、10案例/61评分校准子集、70案例/350评分OCR基准;这些资料用于校准评分标准、护栏、标题风险与OCR风险解读,而非替代生产消融。论文提供控制设置、样本量、置信区间、配对检验、提示词类别、伪代码、架构图、评分表、护栏分类体系及构造示例,确保评估结构可复现且不暴露专有证据。
原文摘要 · Abstract (English)
High-stakes production document-generation systems require language models to be adaptive, evidence-grounded, and auditable. We present HOPM, a hierarchical online prompt mutation framework evaluated on a real marketplace dispute-evidence workflow. HOPM treats prompts as online policies: a family/version router selects a prompt, deterministic guardrails attribute failures to mutable prompt-token categories, and dual feedback from human review and an automated judge updates both routing and mutation priorities. The primary evidence is an observed matched production-evaluation ablation: seven variants are evaluated on the same 600 cases each, enabling component comparisons against static prompting, manual iteration, bandit-only routing, mutation-only adaptation, human-only feedback, auto-judge-only feedback, and full dual-loop HOPM. Full HOPM improves count win rate over a static control from 34.7% to 45.7% (+11.0 pp; paired McNemar p = 1.31e-11) and amount-weighted win rate from 22.3% to 41.4% (+19.1 pp; 95% paired bootstrap CI [10.3, 28.9] pp). It also increases mean Likert quality from 3.18 to 4.40 and reduces issue-flag rate from 15.3% to 5.2%. Supporting review artifacts cover 770 generated-text reviews, 318 labeled reviewer exports, a 10-case/61-rating calibration slice, and a 70-case/350-rating OCR benchmark; these artifacts calibrate rubric, guardrail, title-risk, and OCR-risk interpretation rather than substituting for the production ablation. The paper includes control setup, sample sizes, confidence intervals, paired tests, prompt-token categories, pseudocode, schema, rubric, guardrail taxonomy, and a constructed example so the evaluation structure can be reproduced without exposing proprietary evidence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。