用智能筛选小数据,实现低资源语言高质翻译与低碳训练。
SAGE: Sustainable Agent-Guided Expert-tuning for Culturally Attuned Translation in Low-Resource Southeast Asia
- 用强化学习代理自动筛选高质量文化相关语料,替代海量噪声数据。
- 在7种东南亚低资源语言上超越现有模型,数据量减少97.1%,耗能降95.2%。
- 适合关注可持续性、低资源语言翻译的开发者与政策制定者。
全球网络的包容愿景因严重语言鸿沟受阻,尤其在东南亚低资源地区。尽管大语言模型(LLMs)为翻译提供可能,但在数据匮乏场景下面临双重挑战:高质量、文化相关数据稀缺,以及在大规模噪声网络语料上训练带来的高昂能耗。为调和数字包容与环境可持续性之间的矛盾,我们提出可持续代理引导专家微调(SAGE)。该框架开创性地采用能源感知范式,优先选择“优质数据”而非“大数据”。不依赖碳密集型的大规模训练,SAGE利用强化学习(RL)代理,通过组相对策略优化(GRPO)进行优化,自主构建紧凑训练集。代理基于小型专家构建的社区对话集生成语义奖励信号,过滤噪声与文化偏差。随后使用低秩适配(LoRA)高效微调开源LLM。我们将SAGE应用于英语与东南亚七种低资源语言间的翻译任务,在BLEU-4和COMET-22指标上均达到新最优表现,有效捕捉本地语言特征。关键的是,相比全数据基线,SAGE仅使用3.0%的数据量,训练能耗降低95.2%,同时性能更优。SAGE为全球南方数字鸿沟提供了一条可扩展、负责任的解决路径。
原文摘要 · Abstract (English)
The vision of an inclusive World Wide Web is impeded by a severe linguistic divide, particularly for communities in low-resource regions of Southeast Asia. While large language models (LLMs) offer a potential solution for translation, their deployment in data-poor contexts faces a dual challenge: the scarcity of high-quality, culturally relevant data and the prohibitive energy costs of training on massive, noisy web corpora. To resolve the tension between digital inclusion and environmental sustainability, we introduce Sustainable Agent-Guided Expert-tuning (SAGE). This framework pioneers an energy-aware paradigm that prioritizes the "right data" over "big data". Instead of carbon-intensive training on unfiltered datasets, SAGE employs a reinforcement learning (RL) agent, optimized via Group Relative Policy Optimization (GRPO), to autonomously curate a compact training set. The agent utilizes a semantic reward signal derived from a small, expert-constructed set of community dialogues to filter out noise and cultural misalignment. We then efficiently fine-tune open-source LLMs on this curated data using Low-Rank Adaptation (LoRA). We applied SAGE to translation tasks between English and seven low-resource languages (LRLs) in Southeast Asia. Our approach establishes new state-of-the-art performance on BLEU-4 and COMET-22 metrics, effectively capturing local linguistic nuances. Crucially, SAGE surpasses baselines trained on full datasets while reducing data usage by 97.1% and training energy consumption by 95.2%. By delivering high-performance models with a minimal environmental footprint, SAGE offers a scalable and responsible pathway to bridge the digital divide in the Global South.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。