arXiv:2604.05302cs.CL2026-04ACL

用强化学习实现多语言文本简化,无需平行语料库。

Right at My Level: A Unified Multilingual Framework for Proficiency-Aware Text Simplification

论文配图:Right at My Level: A Unified Multilingual Framework for Proficiency-Aware Text Simplification
图 1 · 摘自论文原文
  • 基于强化学习构建统一框架,无需平行语料库即可实现多语言简化。
  • 在4种语言上训练40亿参数模型,提升目标水平词汇覆盖率。
  • 适合外语教学与个性化内容生成场景,支持多种语言水平标准。

文本简化通过提供可理解输入支持第二语言学习,符合输入假说。然而,构建个性化平行语料库成本高,现有大模型阅读控制方法依赖预标注语料,主要针对英语。我们提出Re-RIGHT,一种无需平行语料监督的自适应多语言文本简化统一强化学习框架。实验表明,基于提示的词汇简化在较低水平及非英语语言上表现不佳,即使使用GPT-5.2和Gemini 2.5等先进模型。为此,我们在英语、日语、韩语和中文中收集了4.3万条词汇级数据,训练了一个40亿参数的策略模型。该模型融合词汇覆盖、语义保留与连贯性三个奖励模块。相比更强基线模型,Re-RIGHT在保持原意与流畅性的前提下,显著提升了目标语言水平下的词汇覆盖率。

原文摘要 · Abstract (English)

Text simplification supports second language (L2) learning by providing comprehensible input, consistent with the Input Hypothesis. However, constructing personalized parallel corpora is costly, while existing large language model (LLM)-based readability control methods rely on pre-labeled sentence corpora and primarily target English. We propose Re-RIGHT, a unified reinforcement learning framework for adaptive multilingual text simplification without parallel corpus supervision. We first show that prompting-based lexical simplification at target proficiency levels (CEFR, JLPT, TOPIK, and HSK) performs poorly at easier levels and for non-English languages, even with state-of-the-art LLMs such as GPT-5.2 and Gemini 2.5. To address this, we collect 43K vocabulary-level data across four languages (English, Japanese, Korean, and Chinese) and train a compact 4B policy model using Re-RIGHT, which integrates three reward modules: vocabulary coverage, semantic preservation, and coherence. Compared to the stronger LLM baselines, Re-RIGHT achieves higher lexical coverage at target proficiency levels while maintaining original meaning and fluency.

文本简化多语言强化学习语言学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。