大模型能高效辅助专家标注事件,但无法独立替代人工。
Large Language Models Are Effective Human Annotation Assistants, But Not Good Independent Annotators
- 用大模型辅助筛选、合并文档并提取事件变量
- 专家与大模型协作时标注一致性提升37%以上
- 适合需要快速处理海量文本的新闻监测场景
事件标注对于识别市场变化、追踪突发新闻和理解社会趋势至关重要。尽管专家标注是黄金标准,但人工标注成本高且效率低。不同于仅关注单一语境的信息抽取实验,本文评估了一个完整工作流:剔除无关文档、合并同一事件的文档,并进行事件标注。尽管基于大模型的自动化标注优于传统TF-IDF方法或事件集整理(Event Set Curation),但其可靠性仍不及人类专家。然而,在事件集整理中引入大模型可显著降低专家在变量标注上的时间与认知负担。当大模型协助提取事件变量时,专家更倾向于采纳这些结果,其一致性高于完全自动化的标注方式。
原文摘要 · Abstract (English)
Event annotation is important for identifying market changes, monitoring breaking news, and understanding sociological trends. Although expert annotators set the gold standards, human coding is expensive and inefficient. Unlike information extraction experiments that focus on single contexts, we evaluate a holistic workflow that removes irrelevant documents, merges documents about the same event, and annotates the events. Although LLM-based automated annotations are better than traditional TF-IDF-based methods or Event Set Curation, they are still not reliable annotators compared to human experts. However, adding LLMs to assist experts for Event Set Curation can reduce the time and mental effort required for Variable Annotation. When using LLMs to extract event variables to assist expert annotators, they agree more with the extracted variables than fully automated LLMs for annotation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。