arXiv:2504.16832cs.CL2025-04被引 3

绿心模型提升越南语逻辑推理能力,解决语言混杂与事实错误问题。

GreenMind: A Next-Generation Vietnamese Large Language Model for Structured and Logical Reasoning

  • 基于分组相对策略优化微调,结合高质量合成推理数据。
  • 在VLSP 2023上超越已有方法,语言一致性显著提升。
  • 适合需要精准越南语推理的AI应用开发者使用。

链式思维(CoT)是一种有效的大型语言模型推理方法,需在生成最终答案前完成中间推理步骤。本文提出受分组相对策略优化启发的越南语推理模型GreenMind-Medium-14B-R1,利用高质量越南语合成推理数据集,并设计两种奖励函数以克服该技术的主要缺陷:(i) 语言混杂问题——在采样过程中显式检测偏倚语言字符的存在;(ii) 使用Sentence Transformer模型确保生成推理内容的事实正确性,避免影响最终输出。在VLSP 2023挑战赛的越南语数据集上的实验表明,该模型优于先前工作,提升了响应的语言一致性。此外,我们在SeaExam——一个多语言选择题数据集上扩展评估,结果表明其推理方法相较少样本提示技术更具有效性。

原文摘要 · Abstract (English)

Chain-of-Thought (CoT) is a robust approach for tackling LLM tasks that require intermediate reasoning steps prior to generating a final answer. In this paper, we present GreenMind-Medium-14B-R1, the Vietnamese reasoning model inspired by the finetuning strategy based on Group Relative Policy Optimization. We also leverage a high-quality Vietnamese synthesized reasoning dataset and design two reward functions to tackle the main limitations of this technique: (i) language mixing, where we explicitly detect the presence of biased language characters during the process of sampling tokens, and (ii) we leverage Sentence Transformer-based models to ensure that the generated reasoning content maintains factual correctness and does not distort the final output. Experimental results on the Vietnamese dataset from the VLSP 2023 Challenge demonstrate that our model outperforms prior works and enhances linguistic consistency in its responses. Furthermore, we extend our evaluation to SeaExam-a multilingual multiple-choice dataset, showing the effectiveness of our reasoning method compared to few-shot prompting techniques.

越南语逻辑推理链式思维语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。