arXiv:2603.18911cs.CLcs.AI2026-03

让中英文双语对话模型零幻觉,靠逐阶段训练和引用溯源。

Progressive Training for Explainable Citation-Grounded Dialogue: Reducing Hallucination to Zero in English-Hindi LLMs

  • 分四步渐进训练,逐步加入引用生成与多语言支持。
  • 编码器-解码器模型在第二阶段后幻觉率降至0.0%。
  • 小模型经训练可媲美大模型,适合资源有限场景。

知识驱动对话系统通过外部知识源生成信息丰富、上下文相关的回复,但现有方法多集中于英语,缺乏显式引用机制验证事实,且决策过程透明度低。本文提出XKD-Dial,一种面向中英双语的四阶段渐进式训练框架:(1) 多语言适配,(2) 基于引用的英文对话SFT,(3) 双语对话SFT,(4) 基于引用感知奖励的GRPO对齐。评估六种模型(编码器-解码器:250M-3B;解码器仅:1B-7B),在各阶段均进行测试。关键贡献包括:(i) 三种后处理可解释性分析——交叉注意力对齐、积分梯度归因与遮挡因果溯源——系统揭示引用行为如何被学习;(ii) 引用接地式SFT使编码器-解码器模型从第二阶段起幻觉率降至0.0%;(iii) 渐进式流程避免灾难性遗忘并提升印地语能力;(iv) 小模型经SFT后在英文表现上可媲美大模型;(v) GRPO在结构化引用任务上仅提供边际改进。评估涵盖六项自动指标(BLEU、ROUGE、BERTScore、FactScore、Citation-F1、幻觉率)。

原文摘要 · Abstract (English)

Knowledge-grounded dialogue systems aim to generate informative, contextually relevant responses by conditioning on external knowledge sources. However, most existing approaches focus exclusively on English, lack explicit citation mechanisms for verifying factual claims, and offer limited transparency into model decision-making. We present XKD-Dial, a progressive four-stage training pipeline for explainable, knowledge-grounded dialogue generation in a bilingual (English-Hindi) setting, comprising: (1) multilingual adaptation, (2) English dialogue SFT with citation grounding, (3) bilingual dialogue SFT, and (4) GRPO alignment with citation-aware rewards. We evaluate six models spanning encoder-decoder (250M-3B) and decoder-only (1B-7B) architectures at every pipeline stage. Our key contributions are: (i) three post-hoc explainability analyses - cross-attention alignment, Integrated Gradients attribution, and occlusion-based causal grounding - applied systematically across the training trajectory to reveal how citation behaviour is learned, not only whether it is learned; (ii) citation-grounded SFT reduces hallucination to 0.0% for encoder-decoder models from Stage 2 onward; (iii) the progressive pipeline prevents catastrophic forgetting while improving Hindi capabilities; (iv) smaller models match larger models on English after SFT; and (v) GRPO provides marginal improvement over well-designed SFT for structured citation tasks. We evaluate across six automatic metrics (BLEU, ROUGE, BERTScore, FactScore, Citation-F1, and hallucination rate).

对话系统双语模型零幻觉可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。