arXiv:2505.14918cs.CLcs.LG2025-05被引 2

提出评估大模型分类一致性框架,帮机构选对模型、用对数据量。

Reliable Decision Support with LLMs: A Framework for Evaluating Consistency in Binary Text Classification Applications

  • 基于心理测量学设计评估框架,量化模型内部与跨模型一致性。
  • 14个模型在1350篇财经新闻上测试,90%-98%样本达成完全一致。
  • 小模型反超大模型,但全模型无法预测真实市场走势。

本研究提出一种评估大语言模型(LLM)二分类任务一致性的框架,弥补现有可靠性评估方法的缺失。借鉴心理测量学原理,确定样本量需求,开发无效响应检测指标,并评估模型内与跨模型一致性。案例研究分析14个LLM(包括claude-3-7-sonnet、gpt-4o、deepseek-r1、gemma3、llama3.2、phi4和command-r-plus)在1,350篇财经新闻上的金融情绪分类表现,每模型执行五次重复实验。所有模型均展现高内部一致性,在90%-98%样本上达成完全一致,且高价与低价同系列模型差异微小。与StockNewsAPI标签对比,模型准确率达0.76-0.88;其中gemma3:1B、llama3.2:3B和claude-3-5-haiku等小型模型表现优于大型模型。然而所有模型在预测实际市场波动时均仅达随机水平,表明问题本质在于任务定义而非模型能力。该框架为模型选择、样本量规划与可靠性评估提供系统指导,助力组织优化资源分配。

原文摘要 · Abstract (English)

This study introduces a framework for evaluating consistency in large language model (LLM) binary text classification, addressing the lack of established reliability assessment methods. Adapting psychometric principles, we determine sample size requirements, develop metrics for invalid responses, and evaluate intra- and inter-rater reliability. Our case study examines financial news sentiment classification across 14 LLMs (including claude-3-7-sonnet, gpt-4o, deepseek-r1, gemma3, llama3.2, phi4, and command-r-plus), with five replicates per model on 1,350 articles. Models demonstrated high intra-rater consistency, achieving perfect agreement on 90-98% of examples, with minimal differences between expensive and economical models from the same families. When validated against StockNewsAPI labels, models achieved strong performance (accuracy 0.76-0.88), with smaller models like gemma3:1B, llama3.2:3B, and claude-3-5-haiku outperforming larger counterparts. All models performed at chance when predicting actual market movements, indicating task constraints rather than model limitations. Our framework provides systematic guidance for LLM selection, sample size planning, and reliability assessment, enabling organizations to optimize resources for classification tasks.

大模型评估分类一致性金融文本心理测量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。