arXiv:2507.22159cs.CLcs.AI2025-07中稿 · IJCNLP-AACL 2025被引 1

首个印尼语多领域偏好数据集,用于评估大模型生成文本质量。

IndoPref: A Multi-Domain Pairwise Preference Dataset for Indonesian

  • 全由母语者撰写,覆盖10个领域,确保语言文化真实。
  • 包含522个提示和4099组人工标注的偏好比较结果。
  • 适合研究多语言大模型或关注印尼语应用的开发者。

超过2亿人使用印尼语,但在基于偏好的大语言模型(LLM)研究中仍严重不足。现有多数多语言数据集依赖英语翻译,常缺乏文化与语言真实性。为此,我们提出IndoPref,首个完全由母语者撰写的多领域印尼语偏好数据集,旨在评估LLM生成文本的自然度与质量。数据集包含522个提示,生成4,099组人类标注的成对偏好,来自五个指令微调的LLM之间的比较。所有标注均为印尼语原生撰写,具有高一致性,经Krippendorff's alpha测量。基准涵盖10个不同类别,使从业者可精准识别LLM的优劣势。

原文摘要 · Abstract (English)

Over 200 million people speak Indonesian, yet the language remains significantly underrepresented in preference-based research for large language models (LLMs). Most existing multilingual datasets are derived from English translations, often resulting in content that lacks cultural and linguistic authenticity. To address this gap, we introduce IndoPref, the first fully human-authored and multi-domain Indonesian preference dataset designed to evaluate the naturalness and quality of LLM-generated text. The dataset contains 522 prompts and yields 4,099 human-annotated pairwise preferences from comparisons across five instruction-tuned LLMs. All annotations are natively written in Indonesian with strong inter-annotator agreement, measured by Krippendorff's alpha. Our benchmark spans 10 diverse categories, enabling practitioners to identify LLMs' fine-grained strengths and weaknesses.

多语言偏好数据集印尼语大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。