构建土耳其语引文意图数据集并实现91.3%准确率分类
A Large-Scale Dataset and Citation Intent Classification in Turkish with LLMs
- 用DSPy框架自动优化提示,提升大模型分类性能
- 集成多模型输出,用XGBoost实现91.3%准确率
- 开源土耳其语引文数据集,助力语言研究
理解引文的定性意图对全面评估学术研究至关重要,但对像土耳其语这样的黏着语而言极具挑战。本文提出一种系统方法并构建首个公开的土耳其语引文意图数据集,通过定制标注工具完成。评估标准上下文学习(ICL)发现其性能受人工设计提示不一致影响。为此,我们引入基于DSPy的可编程分类流水线,自动优化提示。最终采用堆叠泛化集成策略,以XGBoost作为元模型融合多个优化模型输出,实现91.3%的准确率,达到当前最优水平。本研究为土耳其语NLP社区及更广泛的学术界提供基础数据集与稳健分类框架,推动未来定性引文研究发展。
原文摘要 · Abstract (English)
Understanding the qualitative intent of citations is essential for a comprehensive assessment of academic research, a task that poses unique challenges for agglutinative languages like Turkish. This paper introduces a systematic methodology and a foundational dataset to address this problem. We first present a new, publicly available dataset of Turkish citation intents, created with a purpose-built annotation tool. We then evaluate the performance of standard In-Context Learning (ICL) with Large Language Models (LLMs), demonstrating that its effectiveness is limited by inconsistent results caused by manually designed prompts. To address this core limitation, we introduce a programmable classification pipeline built on the DSPy framework, which automates prompt optimization systematically. For final classification, we employ a stacked generalization ensemble to aggregate outputs from multiple optimized models, ensuring stable and reliable predictions. This ensemble, with an XGBoost meta-model, achieves a state-of-the-art accuracy of 91.3\%. Ultimately, this study provides the Turkish NLP community and the broader academic circles with a foundational dataset and a robust classification framework paving the way for future qualitative citation studies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。