arXiv:2505.08167cs.CLcs.AI2025-05被引 1

双向思维链+奖励机制提升非遗问答准确率

Fusing Bidirectional Chains of Thought and Reward Mechanisms A Method for Enhancing Question-Answering Capabilities of Large Language Models for Chinese Intangible Cultural Heritage

  • 用正反双向推理激活模型隐含知识
  • 在非遗问答任务中显著提升准确率与生成质量
  • 方法可迁移至金融、百科等多领域

大规模语言模型的快速发展为领域专用模型提供了重要支持,但使用非物质文化遗产(ICH)数据微调时,仍面临偏见、错误知识继承和灾难性遗忘等问题。为此,我们提出一种融合双向思维链与奖励机制的新训练方法,基于专用于非遗领域的 ICH-Qwen 模型。该方法不仅支持正向推理,还通过反向提问与反向推理激活模型潜藏知识,提升答案准确性。同时,在训练中引入奖励机制,通过结构与内容评估并采用不同权重优化决策过程。对比实验表明,该方法在 ICH-Qwen 上优于零样本、逐步推理、知识蒸馏和问题增强方法,在准确率、Bleu-4 和 Rouge-L 指标上均表现更优。消融实验验证了双向思维链与奖励机制的协同有效性。泛化实验显示,该方法在金融、Wikidata、StrategyQA 等多个领域数据集和先进模型上均带来性能提升,证明其跨领域适应能力,为未来多领域模型训练提供有效路径。

原文摘要 · Abstract (English)

The rapid development of large language models (LLMs) has provided significant support and opportunities for the advancement of domain-specific LLMs. However, fine-tuning these large models using Intangible Cultural Heritage (ICH) data inevitably faces challenges such as bias, incorrect knowledge inheritance, and catastrophic forgetting. To address these issues, we propose a novel training method that integrates a bidirectional chains of thought and a reward mechanism. This method is built upon ICH-Qwen, a large language model specifically designed for the field of intangible cultural heritage. The proposed method enables the model to not only perform forward reasoning but also enhances the accuracy of the generated answers by utilizing reverse questioning and reverse reasoning to activate the model's latent knowledge. Additionally, a reward mechanism is introduced during training to optimize the decision-making process. This mechanism improves the quality of the model's outputs through structural and content evaluations with different weighting schemes. We conduct comparative experiments on ICH-Qwen, with results demonstrating that our method outperforms 0-shot, step-by-step reasoning, knowledge distillation, and question augmentation methods in terms of accuracy, Bleu-4, and Rouge-L scores on the question-answering task. Furthermore, the paper highlights the effectiveness of combining the bidirectional chains of thought and reward mechanism through ablation experiments. In addition, a series of generalizability experiments are conducted, with results showing that the proposed method yields improvements on various domain-specific datasets and advanced models in areas such as Finance, Wikidata, and StrategyQA. This demonstrates that the method is adaptable to multiple domains and provides a valuable approach for model training in future applications across diverse fields.

非遗问答双向推理奖励机制领域适配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。