arXiv:2602.05493cs.CLcs.AI2026-02

LinguistAgent用多模型协作自动标注语言数据,提升研究效率。

LinguistAgent Technical Report: A Reflective Multi-Model Platform for Automated Linguistic Annotation

  • 采用反思式多模型架构,模拟同行评审流程。
  • 在隐喻识别任务中实现与人工标注相当的准确率(F1=0.82)。
  • 适合语言学、人文学科研究者快速验证标注方案。

数据标注仍是人文学科和社科领域的重大瓶颈,尤其在隐喻识别等复杂语言任务中。尽管大语言模型(LLMs)展现出潜力,但其理论能力与实际应用之间仍存在显著差距。本文提出LinguistAgent,一个集成且用户友好的平台,利用反思式多模型架构实现语言标注自动化。该平台包含标注器(Annotator)和可选的评审器(Reviewer),模拟同行评审过程。支持三种主要范式对比实验:提示工程(零样本/少样本/思维链)、检索增强生成,以及微调。通过复现一篇已发表研究中的隐喻识别任务,LinguistAgent实现了基于实时标记级别的评估(F1=0.82,Cohen's kappa=0.76),并与人工黄金标准进行比对。相关应用与代码已开源至https://github.com/Bingru-Li/LinguistAgent。

原文摘要 · Abstract (English)

Data annotation remains a significant bottleneck in the field of humanities and social sciences, particularly for complex linguistic tasks such as metaphor identification. While Large Language Models (LLMs) show promise, a significant gap remains between the theoretical capability of LLMs and their practical utility for researchers. This paper introduces LinguistAgent, an integrated, user-friendly platform that leverages a reflective multi-model architecture to automate linguistic annotation. The platform comprises an Annotator and an optional Reviewer to simulate a peer-review process. This platform supports comparative experiments across three main paradigms: Prompt Engineering (Zero-shot/Few-shot/Chain-of-thought), Retrieval-Augmented Generation, and Fine-tuning. We demonstrate LinguistAgent's efficacy by replicating the task of metaphor identification from a published study, which provides real-time token-level evaluation (F1 and Cohen's kappa) against human gold standards. The application and codes are released on https://github.com/Bingru-Li/LinguistAgent.

语言标注大模型应用隐喻识别自动化工具

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。