arXiv:2411.05639cs.CL2024-11被引 5

评测4个开源大模型在论辩挖掘任务中的表现

Assessing Open-Source Large Language Models on Argumentation Mining Subtasks

  • 对比Mistral 7B等4个开源模型在零样本和少样本下的论辩能力
  • 在三个数据集上测试论辩单元分类与关系分类效果
  • 为未来开放模型的论辩研究提供评估基准

本文评估了四个开源大语言模型(Mistral 7B、Mixtral 8x7B、Llama2 7B 和 Llama3 8B)在论辩挖掘(AM)任务中的表现。实验基于三个语料库:说服性作文(PE)、论辩微文本(AMT)第1部分和第2部分,涵盖两个子任务:论辩话语单元分类(ADUC)和论辩关系分类(ARC)。研究在零样本与少样本场景下进行,旨在评估这些模型在论辩理解方面的能力,为未来基于开源大模型的计算论辩研究提供参考。

原文摘要 · Abstract (English)

We explore the capability of four open-sourcelarge language models (LLMs) in argumentation mining (AM). We conduct experiments on three different corpora; persuasive essays(PE), argumentative microtexts (AMT) Part 1 and Part 2, based on two argumentation mining sub-tasks: (i) argumentative discourse units classifications (ADUC), and (ii) argumentative relation classification (ARC). This work aims to assess the argumentation capability of open-source LLMs, including Mistral 7B, Mixtral8x7B, LlamA2 7B and LlamA3 8B in both, zero-shot and few-shot scenarios. Our analysis contributes to further assessing computational argumentation with open-source LLMs in future research efforts.

论辩挖掘大模型评测开源模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。