法国政府打造开源平台,收集法语人类偏好数据以提升非英语大模型表现。
compar:IA: The French Government's LLM arena to collect French-language human prompts and preference data
- 通过盲式成对比较界面,收集真实用户对法语模型的自由提问与判断。
- 截至2026年2月7日,已积累超60万条法语提示与25万条偏好投票数据。
- 数据开放可复用,适合关注多语言大模型训练与人机交互的研究者。
大型语言模型在非英语语言中常表现出性能下降、文化契合度低和安全鲁棒性差的问题,部分原因在于预训练数据和人类偏好对齐数据以英语为主。基于强化学习的人类反馈(RLHF)与直接偏好优化(DPO)等训练方法依赖人类偏好数据,但多数非英语语言仍缺乏此类数据且多不公开。为此,我们推出 compar:IA,一个由法国政府内部开发的开源数字公共服务平台,旨在从主要法语使用者群体中收集大规模人类偏好数据。该平台采用盲式成对比较界面,捕捉多样语言模型下的自由形式提示与用户判断,同时保持低参与门槛与隐私保护的自动过滤机制。截至2026年2月7日,已累计收集超过60万条自由文本提示和25万条偏好投票,其中约89%为法语内容。我们发布了三个互补数据集——对话、投票与反应,均采用开放许可,并展示初步分析结果,包括法语模型排行榜与用户交互模式。除法语场景外,compar:IA正向国际公共产品演进,为多语言模型训练、评估及人机交互研究提供可复用基础设施。
原文摘要 · Abstract (English)
Large Language Models (LLMs) often show reduced performance, cultural alignment, and safety robustness in non-English languages, partly because English dominates both pre-training data and human preference alignment datasets. Training methods like Reinforcement Learning from Human Feedback (RLHF) and Direct Preference Optimization (DPO) require human preference data, which remains scarce and largely non-public for many languages beyond English. To address this gap, we introduce compar:IA, an open-source digital public service developed inside the French government and designed to collect large-scale human preference data from a predominantly French-speaking general audience. The platform uses a blind pairwise comparison interface to capture unconstrained, real-world prompts and user judgments across a diverse set of language models, while maintaining low participation friction and privacy-preserving automated filtering. As of 2026-02-07, compar:IA has collected over 600,000 free-form prompts and 250,000 preference votes, with approximately 89% of the data in French. We release three complementary datasets -- conversations, votes, and reactions -- under open licenses, and present initial analyses, including a French-language model leaderboard and user interaction patterns. Beyond the French context, compar:IA is evolving toward an international digital public good, offering reusable infrastructure for multilingual model training, evaluation, and the study of human-AI interaction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。