构建爱沙尼亚语主观性数据集,测试大模型自动评分效果
Creation of the Estonian Subjectivity Dataset: Assessing the Degree of Subjectivity on a Scale
- 人工标注1000篇文本的主观程度,用0-100连续量表
- 大模型评分与人工评分相近,但存在系统性差异
- 适合研究主观性分析,尤其关注自动化标注局限
本文介绍了爱沙尼亚语文档级主观性数据集的构建过程,分析了标注结果,并报告了使用大语言模型(LLM)进行自动主观性分析的初步实验。该数据集包含1000篇文档——300篇新闻文章和700篇随机选取的网络文本——每篇由四位标注者在0(完全客观)到100(完全主观)的连续尺度上打分。由于标注者间相关性中等,部分文本得分差异显著,因此对分歧最大的子集进行了重新标注,使标注一致性提升。除人工标注外,数据集还包含GPT-5生成的评分作为自动化标注的实验。结果显示,模型评分与人工评分相似,但存在若干差异,表明尽管基于LLM的主观性评分可行,但无法完全替代人工标注,其适用性取决于具体应用场景。
原文摘要 · Abstract (English)
This article presents the creation of an Estonian-language dataset for document-level subjectivity, analyzes the resulting annotations, and reports an initial experiment of automatic subjectivity analysis using a large language model (LLM). The dataset comprises of 1,000 documents-300 journalistic articles and 700 randomly selected web texts-each rated for subjectivity on a continuous scale from 0 (fully objective) to 100 (fully subjective) by four annotators. As the inter-annotator correlations were moderate, with some texts receiving scores at the opposite ends of the scale, a subset of texts with the most divergent scores was re-annotated, with the inter-annotator correlation improving. In addition to human annotations, the dataset includes scores generated by GPT-5 as an experiment on annotation automation. These scores were similar to human annotators, however several differences emerged, suggesting that while LLM based automatic subjectivity scoring is feasible, it is not an interchangeable alternative to human annotation, and its suitability depends on the intended application.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。