arXiv:2508.07827cs.CL2025-08中稿 · COLM被引 13

测试大模型能否像专家一样准确标注专业文本,发现效果有限。

Evaluating Large Language Models as Expert Annotators

  • 用多智能体讨论模拟专家协作,让模型互相辩论后打标签。
  • 单独使用大模型时,思维链等技巧提升微弱甚至有害。
  • 专家领域标注中,模型易固执己见,难被说服。

文本标注通常耗时费力。尽管大语言模型在通用自然语言处理任务中表现出替代人类标注者的能力,但在需要专业知识的领域(如金融、生物医学、法律)中的有效性仍不明确。本文评估了顶级大模型及多智能体方法在这三个高度专业化领域的表现。提出一种多智能体讨论框架,模拟人类标注员通过交流彼此的标注和理由后形成最终标签。同时引入推理模型(如o3-mini)进行对比。实验结果表明:(1)采用推理时技术(如思维链CoT、自洽性)的单个模型仅带来微弱或负面性能提升,与以往文献结论相反;(2)多数情况下,推理模型与非推理模型相比无显著差异,说明长思维链对专业领域标注帮助有限;(3)多智能体环境中出现特定行为,例如Claude 3.7 Sonnet在开启思考模式后极少修改初始标注,即使其他智能体给出正确答案或有效推理。

原文摘要 · Abstract (English)

Textual data annotation, the process of labeling or tagging text with relevant information, is typically costly, time-consuming, and labor-intensive. While large language models (LLMs) have demonstrated their potential as direct alternatives to human annotators for general domains natural language processing (NLP) tasks, their effectiveness on annotation tasks in domains requiring expert knowledge remains underexplored. In this paper, we investigate: whether top-performing LLMs, which might be perceived as having expert-level proficiency in academic and professional benchmarks, can serve as direct alternatives to human expert annotators? To this end, we evaluate both individual LLMs and multi-agent approaches across three highly specialized domains: finance, biomedicine, and law. Specifically, we propose a multi-agent discussion framework to simulate a group of human annotators, where LLMs are tasked to engage in discussions by considering others' annotations and justifications before finalizing their labels. Additionally, we incorporate reasoning models (e.g., o3-mini) to enable a more comprehensive comparison. Our empirical results reveal that: (1) Individual LLMs equipped with inference-time techniques (e.g., chain-of-thought (CoT), self-consistency) show only marginal or even negative performance gains, contrary to prior literature suggesting their broad effectiveness. (2) Overall, reasoning models do not demonstrate statistically significant improvements over non-reasoning models in most settings. This suggests that extended long CoT provides relatively limited benefits for data annotation in specialized domains. (3) Certain model behaviors emerge in the multi-agent discussion environment. For instance, Claude 3.7 Sonnet with thinking rarely changes its initial annotations, even when other agents provide correct annotations or valid reasoning.

大模型专家标注多智能体专业领域

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。