测试大模型能否像人类一样判断辩论中论点的强弱。
Can LLMs Judge Debates? Evaluating Non-Linear Reasoning via Argumentation Theory Semantics
- 用辩论对话数据让模型排序论点,模拟正式论证结构。
- 模型对论点评分与理论标准相关性中等,长文本下表现下降。
- 先进提示策略能缓解论点位置和长度带来的偏差。
大型语言模型(LLMs)在线性推理任务中表现出色,但在自然辩论这类非线性结构上的表现仍缺乏研究,而辩论最适合用论点图表示。本文评估了LLMs是否能通过计算论证理论(CAT)中的量化辩论(QuAD)语义来近似结构化推理。具体而言,QuAD根据论点间的攻击与支持关系赋予其可接受性分数。基于两个NoDE数据集中的对话式辩论,模型被要求在不访问底层图结构的情况下对论点进行排序。我们测试了多种LLM在高级指令策略(如思维链和上下文学习)下的表现。尽管模型与QuAD排名存在一定一致性,但输入过长或论述流被打断时性能显著下降。先进提示策略有助于减轻由论点长度和位置引起的偏差。研究结果揭示了大模型在建模形式化论证语义方面的潜力与局限,也推动了未来图感知推理的研究。
原文摘要 · Abstract (English)
Large Language Models (LLMs) excel at linear reasoning tasks but remain underexplored on non-linear structures such as those found in natural debates, which are best expressed as argument graphs. We evaluate whether LLMs can approximate structured reasoning from Computational Argumentation Theory (CAT). Specifically, we use Quantitative Argumentation Debate (QuAD) semantics, which assigns acceptability scores to arguments based on their attack and support relations. Given only dialogue-formatted debates from two NoDE datasets, models are prompted to rank arguments without access to the underlying graph. We test several LLMs under advanced instruction strategies, including Chain-of-Thought and In-Context Learning. While models show moderate alignment with QuAD rankings, performance degrades with longer inputs or disrupted discourse flow. Advanced prompting helps mitigate these effects by reducing biases related to argument length and position. Our findings highlight both the promise and limitations of LLMs in modeling formal argumentation semantics and motivate future work on graph-aware reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。