多个大模型辩论提升文化适应性,小模型也能达到大模型水平。
Multiple LLM Agents Debate for Equitable Cultural Alignment
- 用多个大模型辩论方式优化文化判断,比单模型更准。
- 7-90亿参数的小模型表现媲美270亿参数的大模型。
- 适合关注跨文化公平与多模型协作的研究者。
大型语言模型(LLMs)需适应多元文化背景,以服务全球不同群体。以往研究多采用单模型、单轮次方法,本文提出利用多个大模型的互补优势,提升文化适应能力。我们设计了多智能体辩论框架:两个基于大模型的代理就文化情境进行辩论,并协同达成最终决策。提出两种变体:一种为代理仅进行辩论,另一种在回合中动态选择自我反思或辩论。我们在7个开源大模型(及21种组合)上,使用包含75个国家社会礼节规范的NormAd-ETI基准进行评估。实验表明,辩论机制在整体准确率和文化群体公平性上均优于单模型基线。尤其显著的是,7-9B参数的小模型在该框架下表现可匹敌27B参数的大模型。
原文摘要 · Abstract (English)
Large Language Models (LLMs) need to adapt their predictions to diverse cultural contexts to benefit diverse communities across the world. While previous efforts have focused on single-LLM, single-turn approaches, we propose to exploit the complementary strengths of multiple LLMs to promote cultural adaptability. We introduce a Multi-Agent Debate framework, where two LLM-based agents debate over a cultural scenario and collaboratively reach a final decision. We propose two variants: one where either LLM agents exclusively debate and another where they dynamically choose between self-reflection and debate during their turns. We evaluate these approaches on 7 open-weight LLMs (and 21 LLM combinations) using the NormAd-ETI benchmark for social etiquette norms in 75 countries. Experiments show that debate improves both overall accuracy and cultural group parity over single-LLM baselines. Notably, multi-agent debate enables relatively small LLMs (7-9B) to achieve accuracies comparable to that of a much larger model (27B parameters).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。