解决广告问答中幻觉网址问题,提升生成准确性和安全性。
Towards Faithful Industrial RAG: A Reinforced Co-adaptation Framework for Advertising QA
- 构建图结构检索与多维奖励强化学习协同优化
- 幻觉率降低72%,网址幻觉减少92.7%
- 适合工业级高可靠性问答系统研发者参考
工业广告问答是高风险任务,幻觉内容(尤其是虚构链接)可能导致财务损失、合规违规和法律风险。尽管检索增强生成(RAG)被广泛应用,但在生产部署中仍面临挑战,因工业知识具有强关联性、频繁更新且与生成目标对齐不足。本文提出一种强化协同优化框架,通过两个组件联合优化检索与生成:(1) 图感知检索(GraphRAG),在高引用知识子图上建模实体-关系结构,实现多跳、领域特定证据选择;(2) 基于组相对策略优化(GRPO)的证据约束强化学习,采用多维度奖励(包括忠实性、风格合规、安全性和链接有效性)。在内部广告问答数据集上的实验显示,在专家评估的准确性、完整性与安全性维度均有持续提升,幻觉率降低72%。两周线上A/B测试表明,点赞率提升28.6%,点踩率下降46.2%,链接幻觉减少92.7%。该系统已上线运行超半年,服务数百万次问答交互。
原文摘要 · Abstract (English)
Industrial advertising question answering (QA) is a high-stakes task in which hallucinated content, particularly fabricated URLs, can lead to financial loss, compliance violations, and legal risk. Although Retrieval-Augmented Generation (RAG) is widely adopted, deploying it in production remains challenging because industrial knowledge is inherently relational, frequently updated, and insufficiently aligned with generation objectives. We propose a reinforced co-adaptation framework that jointly optimizes retrieval and generation through two components: (1) Graph-aware Retrieval (GraphRAG), which models entity-relation structure over a high-citation knowledge subgraph for multi-hop, domain-specific evidence selection; and (2) evidence-constrained reinforcement learning via Group Relative Policy Optimization (GRPO) with multi-dimensional rewards covering faithfulness, style compliance, safety, and URL validity. Experiments on an internal advertising QA dataset show consistent gains across expert-judged dimensions including accuracy, completeness, and safety, while reducing the hallucination rate by 72\%. A two-week online A/B test demonstrates a 28.6\% increase in like rate, a 46.2\% decrease in dislike rate, and a 92.7\% reduction in URL hallucination. The system has been running in production for over half a year and has served millions of QA interactions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。