arXiv:2609.04823cs.CLcs.AI2026-09

用强化学习提升大模型对加泰罗尼亚语的文本简化能力

Reinforcement Learning for improving Large Language Models' Catalan text simplification capabilities

论文配图:Reinforcement Learning for improving Large Language Models' Catalan text simplification capabilities
图 1 · 摘自论文原文
  • 设计新奖励函数,结合SARI与惩罚项指导模型简化风格
  • 在ASSET数据集上微调后,加泰罗尼亚语简化效果显著提升
  • 适合关注低资源语言文本简化与RL应用的研究者

尽管自动文本简化(ATS)对可及性至关重要,但其进展未能跟上自然语言处理技术的快速演进。本文研究将强化学习(RL)应用于大语言模型(LLMs),以提升低资源语言的文本简化质量。提出一种新型奖励函数,通过组相对策略优化(GRPO)引导模型实现目标简化风格,该函数融合SARI指标与特定惩罚项。在ASSET数据集上对IberianLLM-7B-Instruct进行微调后,模型在两个精心构建的加泰罗尼亚语基准测试中性能提升,同时有效抑制了先前观察到的负面行为。通过将ASSET翻译为加泰罗尼亚语和西班牙语并分别微调模型,探索跨语言迁移,但这些方法在域外基准上未表现出显著提升。

原文摘要 · Abstract (English)

Although automatic text simplification (ATS) is critical for accessibility, its progress has not matched the rapid evolution of broader natural language processing techniques. This paper investigates the application of reinforcement learning (RL) to improve the quality of ATS for low-resource languages using Large Language Models (LLMs). The paper introduces a novel reward function, designed to guide LLMs toward a targeted simplification style with Group Relative Policy Optimization (GRPO), that combines the SARI metric with specific penalty components. The effectiveness of GRPO with this reward function is motivated and demonstrated by post-training IberianLLM-7B-Instruct on the ASSET dataset. After post-training on the English ASSET, the model's ATS performance improves on two curated Catalan benchmarks while also successfully suppressing previously observed negative behaviors. Cross-lingual transfer learning is explored by translating ASSET into Catalan and Spanish and post-training the model on each version, but these fail to show a significant improvement on the out-of-domain benchmark.

文本简化强化学习低资源语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。