用韵律诗测试大模型的规则理解能力,提出首个系统性评估框架。
MetricalARGS: A Taxonomy for Studying Metrical Poetry with LLMs
- 构建四维度评估框架:分析、检索、生成、支持,专攻韵律诗任务。
- 以泰卢固语为例,展示如何用该框架评估模型对严格音节规则的遵守。
- 适合研究语言模型在复杂文学规则下的推理能力,推动诗歌理解发展。
以往的自然语言处理研究多聚焦于诗歌自动生成与摘要。许多语言拥有成熟的韵律诗传统,其对音节和音素模式有严格约束,为探测大语言模型(LLMs)的深层推理与语言理解能力提供了契机。本文提出MetricalARGS,首个面向跨语言韵律诗的NLP任务分类体系,涵盖分析、检索、生成和支持四个维度。我们探讨这些任务与现有NLP任务的关系,讨论数据集与评估指标问题。以泰卢固语为例,说明该分类体系的实际应用。MetricalARGS揭示了通过韵律诗视角理解当前大模型能力与局限性的广阔前景。
原文摘要 · Abstract (English)
Prior NLP work studying poetry has focused primarily on automatic poem generation and summarization. Many languages have well-studied traditions of poetic meter which enforce constraints on a poem in terms of syllable and phoneme patterns. Such advanced literary forms offer opportunities for probing deeper reasoning and language understanding in Large Language Models (LLMs) and their ability to follow strict pre-requisites and rules. In this paper, we introduce MetricalARGS, the first taxonomy of poetry-related NLP tasks designed to evaluate LLMs on metrical poetry across four dimensions: Analysis, Retrieval, Generation, and Support. We discuss how these tasks relate to existing NLP tasks, addressing questions around datasets and evaluation metrics. Taking Telugu as our example language, we illustrate how the taxonomy can be used in practice. MetricalARGS highlights the broader possibilities for understanding the capabilities and limitations of today's LLMs through the lens of metrical poetry.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。