arXiv:2503.15837cs.CLcs.AI2025-03被引 6

首个兼顾古文理解与生成的综合评测基准,揭示大模型在古文创作上的短板。

Fùxì: A Benchmark for Evaluating Language Models on Ancient Chinese Text Understanding and Generation

  • 构建21项任务覆盖理解与生成,含诗词、对联等新题型
  • 模型在理解任务表现良好,生成任务准确率不足40%
  • 引入文化真实性与格式规范双评估机制,适合古籍数字化研究者

古汉语处理因独特的语言特征、复杂的结构约束和丰富的文化背景,对大语言模型提出特殊挑战。现有评测多聚焦选择题式理解,却缺乏对古典中文生成能力的评估。我们提出Fùxì,一个涵盖21个多样化任务的综合性基准,全面评估理解与生成能力。其核心贡献包括:(1)平衡覆盖理解与生成任务,新增诗词创作、对联补全等创新题型;(2)设计专用生成评估指标,结合规则验证与微调的语言模型评分器;(3)建立系统性评估框架,同时考量语言准确性和文化真实性。对主流LLMs的广泛测试显示,模型在理解任务表现优异,但在生成任务上差距显著,尤其在需深度文化知识与经典格式遵循的任务中表现不佳。研究揭示了当前古文处理的局限,并为未来模型优化提供方向。该基准、评估工具与基线结果已公开,助力该领域研究。

原文摘要 · Abstract (English)

Ancient Chinese text processing presents unique challenges for large language models (LLMs) due to its distinct linguistic features, complex structural constraints, and rich cultural context. While existing benchmarks have primarily focused on evaluating comprehension through multiple-choice questions, there remains a critical gap in assessing models' generative capabilities in classical Chinese. We introduce Fùxì, a comprehensive benchmark that evaluates both understanding and generation capabilities across 21 diverse tasks. Our benchmark distinguishes itself through three key contributions: (1) balanced coverage of both comprehension and generation tasks, including novel tasks like poetry composition and couplet completion, (2) specialized evaluation metrics designed specifically for classical Chinese text generation, combining rule-based verification with fine-tuned LLM evaluators, and (3) a systematic assessment framework that considers both linguistic accuracy and cultural authenticity. Through extensive evaluation of state-of-the-art LLMs, we reveal significant performance gaps between understanding and generation tasks, with models achieving promising results in comprehension but struggling considerably in generation tasks, particularly those requiring deep cultural knowledge and adherence to classical formats. Our findings highlight the current limitations in ancient Chinese text processing and provide insights for future model development. The benchmark, evaluation toolkit, and baseline results are publicly available to facilitate research in this domain.

古汉语文本生成评测基准大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。