arXiv:2503.21500cs.CL2025-03ACL被引 3

首个匈牙利语大模型评测基准,专评语言特性与生成能力

OpenHuEval: Evaluating Large Language Model on Hungarian Specifics

  • 基于真实网络查询构建多维度评测数据集
  • 涵盖5个任务3953题,覆盖8类匈牙利语特异性
  • 适用于非英语大模型研究与匈牙利语优化

我们提出OpenHuEval,首个聚焦匈牙利语及语言特性的大模型评测基准。该基准源于多源匈牙利语材料,遵循最新评估设计原则:采用真实用户查询、强调生成能力评估,并引入大模型作为评判者以提升评估多维性与准确性。最终包含8个匈牙利语特异性维度,5项任务,共3953道题目。我们评估了主流大模型(含传统LMM与大型推理模型),结果表明针对匈牙利语的专项评估与优化至关重要。同时,基于OpenHuEval构建了分析推理模型思维过程的框架,揭示了非英语语言中大模型的内在模式与机制。相关数据集将公开于https://github.com/opendatalab/OpenHuEval。

原文摘要 · Abstract (English)

We introduce OpenHuEval, the first benchmark for LLMs focusing on the Hungarian language and specifics. OpenHuEval is constructed from a vast collection of Hungarian-specific materials sourced from multiple origins. In the construction, we incorporated the latest design principles for evaluating LLMs, such as using real user queries from the internet, emphasizing the assessment of LLMs' generative capabilities, and employing LLM-as-judge to enhance the multidimensionality and accuracy of evaluations. Ultimately, OpenHuEval encompasses eight Hungarian-specific dimensions, featuring five tasks and 3953 questions. Consequently, OpenHuEval provides the comprehensive, in-depth, and scientifically accurate assessment of LLM performance in the context of the Hungarian language and its specifics. We evaluated current mainstream LLMs, including both traditional LLMs and recently developed Large Reasoning Models. The results demonstrate the significant necessity for evaluation and model optimization tailored to the Hungarian language and specifics. We also established the framework for analyzing the thinking processes of LRMs with OpenHuEval, revealing intrinsic patterns and mechanisms of these models in non-English languages, with Hungarian serving as a representative example. We will release OpenHuEval at https://github.com/opendatalab/OpenHuEval .

大模型评测匈牙利语生成能力非英语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。