arXiv:2409.19691cs.CL2024-09EMNLP被引 4

构建中文修辞理解与生成的综合数据集,助力写作能力提升

CERD: A Comprehensive Chinese Rhetoric Dataset for Rhetorical Understanding and Generation in Essays

  • 构建涵盖4类宏观与23类细粒度修辞的中文数据集
  • 大模型在多数任务中表现最佳,多任务联合微调效果更优
  • 适合语言生成、写作辅助与中文NLP研究者使用

现有修辞理解与生成数据集多聚焦单一粗粒度或细粒度类别,忽视不同修辞手法间的关联性,将其视为独立任务。本文提出中文作文修辞数据集CERD,包含4种常见粗粒度类别(隐喻、拟人、夸张、排比)和23种跨形式与内容层面的细粒度类别。CERD为人工标注的综合性中文修辞数据集,涵盖五个相互关联的子任务。不同于以往工作,该数据集支持多种修辞的理解、成分识别及在给定条件下的修辞句生成,有助于提升作者写作水平与语言运用能力。通过大量实验验证了CERD中多个任务间的关联性,并为未来修辞研究建立基准。实验结果表明,大语言模型在多数任务中表现最佳,多任务联合微调进一步提升性能。

原文摘要 · Abstract (English)

Existing rhetorical understanding and generation datasets or corpora primarily focus on single coarse-grained categories or fine-grained categories, neglecting the common interrelations between different rhetorical devices by treating them as independent sub-tasks. In this paper, we propose the Chinese Essay Rhetoric Dataset (CERD), consisting of 4 commonly used coarse-grained categories including metaphor, personification, hyperbole and parallelism and 23 fine-grained categories across both form and content levels. CERD is a manually annotated and comprehensive Chinese rhetoric dataset with five interrelated sub-tasks. Unlike previous work, our dataset aids in understanding various rhetorical devices, recognizing corresponding rhetorical components, and generating rhetorical sentences under given conditions, thereby improving the author's writing proficiency and language usage skills. Extensive experiments are conducted to demonstrate the interrelations between multiple tasks in CERD, as well as to establish a benchmark for future research on rhetoric. The experimental results indicate that Large Language Models achieve the best performance across most tasks, and jointly fine-tuning with multiple tasks further enhances performance.

修辞分析中文NLP文本生成数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。