用强化学习让大模型翻译精准控时,兼顾语义和字幕时限
HOMURA: Taming the Sand-Glass for Time-Constrained LLM Translation via Reinforcement Learning
- 通过强化学习动态调节翻译长度,平衡语义与时间约束
- 在砂漏基准上实现95%的时长达标率,且不牺牲语义质量
- 适合字幕、配音等严格限时场景,对语言密度敏感
大型语言模型在多语言翻译中取得显著进展,但存在跨语言冗余偏差,难以满足字幕、配音等严格时长限制任务。现有提示工程方法难以调和语义保真度与时间可行性之间的矛盾。为此,我们首先提出「砂漏(Sand-Glass)」基准,专门评估音节数级时长约束下的翻译表现。进一步提出HOMURA框架,基于强化学习显式优化语义保留与时间合规性的权衡。采用带KL正则化的目标函数与新颖的动态音节比例奖励机制,有效“驯服”输出长度。实验表明,该方法显著优于多个强基线模型,在保持语义充分性的同时,精确控制输出时长并尊重语言密度层级。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have achieved remarkable strides in multilingual translation but are hindered by a systemic cross-lingual verbosity bias, rendering them unsuitable for strict time-constrained tasks like subtitling and dubbing. Current prompt-engineering approaches struggle to resolve this conflict between semantic fidelity and rigid temporal feasibility. To bridge this gap, we first introduce Sand-Glass, a benchmark specifically designed to evaluate translation under syllable-level duration constraints. Furthermore, we propose HOMURA, a reinforcement learning framework that explicitly optimizes the trade-off between semantic preservation and temporal compliance. By employing a KL-regularized objective with a novel dynamic syllable-ratio reward, HOMURA effectively "tames" the output length. Experimental results demonstrate that our method significantly outperforms strong LLM baselines, achieving precise length control that respects linguistic density hierarchies without compromising semantic adequacy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。