用语法解析器替代人工标注,提升北欧低资源语言模型的语法质量。
SAGA: Score-Weighted Adaptive Generation Alignment for Low-Resource Nordic Language Models

- 用依赖解析器生成偏好对,替代人工标注进行模型优化。
- 丹麦语法准确率从69.0%提升至93.8%,冰岛语在独立评测中提升4.5个百分点。
- 适合无大量人工标注但有高质量语法解析器的低资源语言研究者。
偏好优化已证明能有效提升大语言模型性能,但通常依赖昂贵的人工偏好标注。将此类方法扩展到形态丰富且数据稀缺的低资源语言仍具挑战性。本文提出SAGA(Score-weighted Adaptive Generation Alignment),一种基于解析器引导的偏好优化框架,用依赖解析器的判断代替人工标签。SAGA将解析器输出转化为delta-DPO的偏好对,结合解析质量与词汇多样性构建复合奖励,通过奖励差阈值过滤低信息偏好对,并监控奖励滥用以确保监督可靠性。在丹麦语、冰岛语和挪威语(Bokmål)上使用GPT-SW3-1.3B模型测试,SAGA无需人工偏好标注即显著提升语法质量:丹麦语解析成功率从69.0%升至93.8%,冰岛语在独立Stanza评测中平均提升3.3个百分点(三轮均值),母语者在80%的成对比较中更偏好SAGA输出,挪威语提升28个百分点。结果表明,在具备高质量依赖解析器的前提下,解析器生成的监督可作为人工偏好标注的有效替代方案。
原文摘要 · Abstract (English)
Preference optimisation has proven effective for improving large language models but typically relies on costly human preference annotations. Extending these methods to morphologically rich, low-resource languages remains challenging because such annotations are scarce. We present SAGA (Score-weighted Adaptive Generation Alignment), a parser-guided preference optimisation framework that replaces human labels with dependency-parser supervision. SAGA converts parser judgements into preference pairs for delta-DPO, combines parser quality with lexical diversity in a composite reward, filters low-information pairs using a reward-gap criterion, and monitors reward hacking to maintain reliable supervision. Across Danish, Icelandic, and Norwegian Bokmål using GPT-SW3-1.3B, SAGA consistently improves grammatical quality without requiring human preference labels. Danish parse success increases from 69.0% to 93.8%, Icelandic achieves a +4.5 percentage-point improvement on an independent Stanza evaluation (three-run mean +3.3 percentage points) while native speakers prefer SAGA outputs in 80% of pairwise comparisons, and Norwegian Bokmål improves by +28 percentage points. These results demonstrate that parser-derived supervision is a practical alternative to human preference annotation for grammatical alignment in low-resource languages where high-quality dependency parsers are available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。