arXiv:2412.15453cs.CLcs.AI2024-12被引 8

用偏好优化提升多语言反仇恨言论生成质量

Northeastern Uni at Multilingual Counterspeech Generation: Enhancing Counter Speech Generation with LLM Alignment through Direct Preference Optimization

  • 通过直接偏好优化对齐大模型,使回复更符合人类偏好
  • 在多语言测试中表现优于传统微调方法,且可扩展至多种语言
  • 适合研究反仇恨内容生成与跨语言AI对齐的学者使用

自动生成反仇恨言论(CS)是应对仇恨言论的关键策略,能提供建设性、有依据的回应。然而,现有方法在跨语言场景下难以生成高质量、有影响力且可扩展的回应。本文提出一种新方法,通过监督微调(SFT)和直接偏好优化(DPO)对大语言模型进行对齐,使输出更符合人类偏好,具备上下文适切性和语言适应性。同时引入知识增强机制,提升生成内容的事实准确性和相关性。实验表明,经DPO对齐的模型在多个反仇恨言论基准上显著优于SFT基线,并能有效扩展至多种语言。模型训练在英文环境下完成,但相同模型用于评估巴斯克语、意大利语和西班牙语等语言的性能,验证了其跨语言泛化能力。

原文摘要 · Abstract (English)

The automatic generation of counter-speech (CS) is a critical strategy for addressing hate speech by providing constructive and informed responses. However, existing methods often fail to generate high-quality, impactful, and scalable CS, particularly across diverse linguistic contexts. In this paper, we propose a novel methodology to enhance CS generation by aligning Large Language Models (LLMs) using Supervised Fine-Tuning (SFT) and Direct Preference Optimization (DPO). Our approach leverages DPO to align LLM outputs with human preferences, ensuring contextually appropriate and linguistically adaptable responses. Additionally, we incorporate knowledge grounding to enhance the factual accuracy and relevance of generated CS. Experimental results demonstrate that DPO-aligned models significantly outperform SFT baselines on CS benchmarks while scaling effectively to multiple languages. These findings highlight the potential of preference-based alignment techniques to advance CS generation across varied linguistic settings. The model supervision and alignment is done in English and the same model is used for reporting metrics across other languages like Basque, Italian, and Spanish.

反仇恨生成大模型对齐多语言偏好优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。