arXiv:2602.19467cs.CYcs.AI2026-02

测试低成本大模型能否替代人工编码,发现顶尖模型准确率超97%。

Can Large Language Models Replace Human Coders? Introducing ContentBench

  • 构建可版本化更新的基准测试集,评估大模型在文本编码任务中的表现。
  • 顶级低预算模型对学术社交内容分类达97%-99%准确率,成本仅数美元处理五万条。
  • 适合关注自动化内容分析、模型评估与伦理治理的研究者使用。

低成本大语言模型能否替代人工完成仍支撑实证内容分析的解释性编码工作?本文提出 ContentBench,一个公开的基准测试套件,用于追踪低预算 LLM 在相同解释性编码任务中达成的一致性水平及其成本。该套件采用可版本化的赛道设计,鼓励研究者贡献新数据集。报告首项结果:ContentBench-ResearchTalk v1.0 包含1,000条模拟社交媒体风格的学术研究帖子,标注为五大类——赞美、批评、讽刺、提问和程序性评论。参考标签仅在三个前沿推理模型(GPT-5、Gemini 2.5 Pro、Claude Opus 4.1)一致同意时生成,所有最终标签均由作者进行质量审计。在59个被评估模型中,表现最佳的低预算 LLM 与评审标签达成约97%-99%的一致性,远超 GPT-3.5 Turbo。部分顶尖模型仅需数美元即可完成50,000条文本的编码,使大规模解释性编码从人力瓶颈转向验证、报告与治理问题。然而,本地运行的小型开源模型在讽刺类内容上仍表现不佳(如 Llama 3.2 3B 在高难度讽刺样本上仅4%一致率)。ContentBench 已发布数据、文档及交互式测验平台(contentbench.github.io),支持长期可比评估并欢迎社区扩展。

原文摘要 · Abstract (English)

Can low-cost large language models (LLMs) take over the interpretive coding work that still anchors much of empirical content analysis? This paper introduces ContentBench, a public benchmark suite that helps answer this replacement question by tracking how much agreement low-cost LLMs achieve and what they cost on the same interpretive coding tasks. The suite uses versioned tracks that invite researchers to contribute new benchmark datasets. I report results from the first track, ContentBench-ResearchTalk v1.0: 1,000 synthetic, social-media-style posts about academic research labeled into five categories spanning praise, critique, sarcasm, questions, and procedural remarks. Reference labels are assigned only when three state-of-the-art reasoning models (GPT-5, Gemini 2.5 Pro, and Claude Opus 4.1) agree unanimously, and all final labels are checked by the author as a quality-control audit. Among the 59 evaluated models, the best low-cost LLMs reach roughly 97-99% agreement with these jury labels, far above GPT-3.5 Turbo, the model behind early ChatGPT and the initial wave of LLM-based text annotation. Several top models can code 50,000 posts for only a few dollars, pushing large-scale interpretive coding from a labor bottleneck toward questions of validation, reporting, and governance. At the same time, small open-weight models that run locally still struggle on sarcasm-heavy items (for example, Llama 3.2 3B reaches only 4% agreement on hard-sarcasm). ContentBench is released with data, documentation, and an interactive quiz at contentbench.github.io to support comparable evaluations over time and to invite community extensions.

大模型评测内容分析自动化编码开源模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。