用大模型评估论点质量,通过两两比较实现可靠排序。
Argument Quality Assessment with Large Language Models: A Pairwise Bradley-Terry Approach
- 基于两两比较的伯莱德-特里模型,让大模型判断论点优劣。
- 70B参数模型与人类判断相关性达中等水平(κ=0.493)。
- 结果稳定,适合需要可复现论点评估的研究者使用。
大语言模型在推理与判断任务中表现突出,但论点质量评估仍需严谨方法。我们测试了12个不同规模与架构的开源大模型,在零样本、少样本及思维链提示下,模拟人类对论点在逻辑、修辞和辩证三个维度的两两比较,并利用伯莱德-特里模型推导出论点的潜在强度得分与排名。结果显示,大模型与人类判断具有中等但有限的相关性,其中Llama-70B表现最佳,达到中等一致性(Cohen's κ = 0.493),其伯莱德-特里得分与人工标注的相关系数在肯德尔、皮尔逊和斯皮尔曼中为0.327–0.477。其他模型虽在大小与家族上差异明显,但整体表现互补,对质量维度有部分理解。模型预测在多次运行中稳定,仅不足7.75%案例出现标签分歧,剩余差异通过多数投票与少样本提示处理。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have demonstrated remarkable capabilities in tasks related to reasoning and judgment. However, assessing the quality of arguments requires a rigorous evaluation. We investigate the extent to which LLMs can effectively perform this task. We tested 12 open-weight LLMs of different sizes and families under zero-shot, few-shot, and chain-of-thought to approximate human pairwise comparisons of argument quality across three dimensions--logical, rhetorical, and dialectic--and used these comparisons in a Bradley-Terry model to infer latent strength scores and derive a ranking of arguments. Our insights show that LLMs have promising but moderate correlation with human judgment, with Llama-70B obtaining the strongest alignment, reaching moderate Cohen's $κ$ = 0.493 and moderate correlations with Bradley-Terry scores derived from these annotations (Kendall, Pearson, and Spearman: 0.327-0.477). Other LLMs exhibit weak, moderate, or high alignment with Llama-70B while achieving comparable results against human judgment, suggesting partial but complementary understanding of underlying quality dimensions despite differences in model size and family. Moreover, LLM predictions are stable across trial runs, with fewer than 7.75% of cases yielding different labels. Remaining variability is handled via majority voting and few-shot prompting for large-size models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。