arXiv:2502.17664cs.CL2025-02

用轻量模型提升低资源语言生成的忠实度。

Towards Typologically Aware Rescoring to Mitigate Unfaithfulness in Lower-Resource Languages

  • 用小规模BERT模型重评分大模型输出,提升忠实性。
  • 在越、波、格三语中准确率达88.33%,无需微调。
  • 方法对多任务有效,适合关注语言多样性研究者。

多语言大模型在资源匮乏语言中更易生成不忠实内容(Guerreiro et al., 2023 - arXiv:2303.16104),可能因这些类型多样语言在训练数据中代表性不足。为缓解此问题,我们提出使用计算轻量的辅助模型对大模型输出进行重评分。作为可行性证明,我们发现仅用不到700MB数据从头预训练的单语4层BERT模型,在未微调情况下,可对三种谱系无关且形态复杂度不同的语言——越南语、波兰语和格鲁吉亚语——的摘要忠实性识别达到平均88.33%准确率。相同超参数配置在三个其他任务上也表现良好,表明该方法在重评分场景外具泛化潜力。为支持类型学感知的模型选择,我们还研究了形态复杂度与正则化、模型深度及训练目标的交互关系,最终表明:形态复杂语言更受益于丢弃(dropout),而跨语言下游性能最优来自浅层架构以及标准BERT训练目标。

原文摘要 · Abstract (English)

Multilingual large language models (LLMs) are known to more frequently generate non-faithful output in resource-constrained languages (Guerreiro et al., 2023 - arXiv:2303.16104), potentially because these typologically diverse languages are underrepresented in their training data. To mitigate unfaithfulness in such settings, we propose using computationally light auxiliary models to rescore the outputs of larger architectures. As proof of the feasibility of such an approach, we show that monolingual 4-layer BERT models pretrained from scratch on less than 700 MB of data without fine-tuning are able to identify faithful summaries with a mean accuracy of 88.33% in three genetically unrelated languages that differ in their morphological complexity - Vietnamese, Polish and Georgian. The same hyperparameter combination moreover generalises well to three other tasks, suggesting applications for rescoring beyond improving faithfulness. In order to inform typologically aware model selection, we also investigate how morphological complexity interacts with regularisation, model depth and training objectives, ultimately demonstrating that morphologically complex languages are more likely to benefit from dropout, while across languages downstream performance is enhanced most by shallow architectures as well as training using the standard BERT objectives.

语言模型忠实性低资源重评分

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。