arXiv:2511.06680cs.CL2025-11中稿 · LREC 2026被引 1

用迭代反馈提升大模型韩语方言翻译准确性

Steering LLMs toward Korean Local Speech: Iterative Refinement Framework for Faithful Dialect Translation

  • 通过翻译-验证-反馈循环,引导大模型生成地道方言
  • 新指标DFS和TDR有效区分真假方言转换成功案例
  • 适合关注方言生成与评估的NLP研究者

标准语到方言的机器翻译仍面临大语言模型固有的方言差距及n-gram度量固有的评估扭曲问题,后者倾向于奖励源文本复制而非真实方言转换。本文提出方言精炼(DIA-REFINE)框架,通过外部方言分类器驱动大模型进行翻译、验证与反馈的迭代循环,实现目标方言输出的精准引导。为克服n-gram度量缺陷,引入方言保真度分数(DFS)量化语言偏移,目标方言比率(TDR)衡量方言转换成功率。在零样本与上下文学习基线上的实验表明,DIA-REFINE持续提升方言保真度。新指标能区分‘假成功’(高n-gram分但非真实方言)与‘真尝试’(虽低n-gram分但真实尝试)案例。观察还发现模型响应程度不一,且融入上下文示例可进一步改善方言表达翻译效果。本工作建立了一套目标导向、包容性更强的方言翻译框架,提供严谨评估与模型性能洞察。

原文摘要 · Abstract (English)

Standard-to-dialect machine translation remains challenging due to a persistent dialect gap in large language models and evaluation distortions inherent in n-gram metrics, which favor source copying over authentic dialect translation. In this paper, we propose the dialect refinement (DIA-REFINE) framework, which guides LLMs toward faithful target dialect outputs through an iterative loop of translation, verification, and feedback using external dialect classifiers. To address the limitations of n-gram-based metrics, we introduce the dialect fidelity score (DFS) to quantify linguistic shift and the target dialect ratio (TDR) to measure the success of dialect translation. Experiments on Korean dialects across zero-shot and in-context learning baselines demonstrate that DIA-REFINE consistently enhances dialect fidelity. The proposed metrics distinguish between False Success cases, where high n-gram scores obscure failures in dialectal translation, and True Attempt cases, where genuine attempts at dialectal translation yield low n-gram scores. We also observed that models exhibit varying degrees of responsiveness to the framework, and that integrating in-context examples further improves the translation of dialectal expressions. Our work establishes a robust framework for goal-directed, inclusive dialect translation, providing both rigorous evaluation and critical insights into model performance.

方言翻译大模型评估指标

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。