arXiv:2504.18180cs.CLcs.AI2025-04被引 1

用人类偏好训练提升模型生成冰岛语法律摘要的准确性。

Aligning Language Models for Icelandic Legal Text Summarization

  • 采用人类反馈强化学习与直接偏好优化改进模型。
  • 偏好训练使法律术语准确率提升,但语言流畅性未明显改善。
  • 适合关注法律文本生成与人机对齐的研究者。

将语言模型应用于法律领域有望提升处理大量文书的工作效率。然而,法律文本的专业术语、细微语义和正式风格带来显著挑战。本研究探讨基于偏好的训练方法(包括从人类反馈中强化学习与直接偏好优化)是否能提升模型生成符合冰岛语法律规范和用户偏好的摘要能力。通过对比使用偏好训练与传统监督学习微调的模型,结果表明偏好训练可提高生成摘要的法律准确性,但对冰岛语整体语言质量改善不显著。自动化指标与人工评估结果存在差异,凸显在法律领域开发语言模型时进行定性评价的重要性。

原文摘要 · Abstract (English)

The integration of language models in the legal domain holds considerable promise for streamlining processes and improving efficiency in managing extensive workloads. However, the specialized terminology, nuanced language, and formal style of legal texts can present substantial challenges. This study examines whether preference-based training techniques, specifically Reinforcement Learning from Human Feedback and Direct Preference Optimization, can enhance models' performance in generating Icelandic legal summaries that align with domain-specific language standards and user preferences. We compare models fine-tuned with preference training to those using conventional supervised learning. Results indicate that preference training improves the legal accuracy of generated summaries over standard fine-tuning but does not significantly enhance the overall quality of Icelandic language usage. Discrepancies between automated metrics and human evaluations further underscore the importance of qualitative assessment in developing language models for the legal domain.

法律文本偏好训练冰岛语模型对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。