首个日语大模型可控生成评测基准,解决多语言控制评估难题
LCTG Bench: LLM Controlled Text Generation Benchmark
- 构建统一框架评估大模型在日语场景下的可控生成能力
- 对比9个模型发现日语专用模型显著优于通用多语言模型
- 为日语业务应用选型提供可量化的可控性参考标准
大型语言模型(LLMs)的兴起带来了更丰富、高质量的机器生成文本,但其强大的表达能力使基于特定业务指令控制输出变得困难。现有评测基准存在两大问题:(1) 主要覆盖英语、中文等主流语言,忽视日语等低资源语言;(2) 采用任务特定评估指标,缺乏跨场景统一的可控性选型框架。为此,本文提出LCTG Bench,首个面向日语的大型语言模型可控生成评测基准。该基准提供统一评估框架,支持用户根据实际使用场景选择最合适的可控模型。通过对GPT-4等九种日语专用及多语言模型的评估,揭示了当前日语大模型在可控性方面的现状与挑战,并明确指出多语言模型与日语专用模型之间存在显著差距。
原文摘要 · Abstract (English)
The rise of large language models (LLMs) has led to more diverse and higher-quality machine-generated text. However, their high expressive power makes it difficult to control outputs based on specific business instructions. In response, benchmarks focusing on the controllability of LLMs have been developed, but several issues remain: (1) They primarily cover major languages like English and Chinese, neglecting low-resource languages like Japanese; (2) Current benchmarks employ task-specific evaluation metrics, lacking a unified framework for selecting models based on controllability across different use cases. To address these challenges, this research introduces LCTG Bench, the first Japanese benchmark for evaluating the controllability of LLMs. LCTG Bench provides a unified framework for assessing control performance, enabling users to select the most suitable model for their use cases based on controllability. By evaluating nine diverse Japanese-specific and multilingual LLMs like GPT-4, we highlight the current state and challenges of controllability in Japanese LLMs and reveal the significant gap between multilingual models and Japanese-specific models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。