首个波斯语可解释立场检测数据集,助力理解社交媒体观点
FarExStance: Explainable Stance Detection for Farsi
- 构建波斯语可解释立场数据集,含论点、立场与证据提取
- 微调的RoBERTa和少量样本Claude-3.5-Sonnet在立场识别上表现最佳
- 少量样本Claude-3.5-Sonnet生成解释最自然,适合需要可信推理的场景
我们提出了FarExStance,一个用于波斯语可解释立场检测的新数据集。每个实例包含一个论点、文章或社交媒体帖子对该论点的立场标签,以及支持该标签的抽取式证据。我们在新数据集上比较了微调的多语言RoBERTa模型与多个大语言模型在零样本、少样本及参数高效微调设置下的表现。在立场检测任务中,表现最好的模型是微调后的RoBERTa模型、经参数高效微调的LLM Aya-23-8B,以及少样本的Claude-3.5-Sonnet。在解释质量方面,自动评估显示少样本GPT-4o生成的解释最连贯;人工评估则表明少样本Claude-3.5-Sonnet获得最高整体解释评分(OES)。微调后的Aya-32-8B模型生成的解释与参考答案最接近。
原文摘要 · Abstract (English)
We introduce FarExStance, a new dataset for explainable stance detection in Farsi. Each instance in this dataset contains a claim, the stance of an article or social media post towards that claim, and an extractive explanation which provides evidence for the stance label. We compare the performance of a fine-tuned multilingual RoBERTa model to several large language models in zero-shot, few-shot, and parameter-efficient fine-tuned settings on our new dataset. On stance detection, the most accurate models are the fine-tuned RoBERTa model, the LLM Aya-23-8B which has been fine-tuned using parameter-efficient fine-tuning, and few-shot Claude-3.5-Sonnet. Regarding the quality of the explanations, our automatic evaluation metrics indicate that few-shot GPT-4o generates the most coherent explanations, while our human evaluation reveals that the best Overall Explanation Score (OES) belongs to few-shot Claude-3.5-Sonnet. The fine-tuned Aya-32-8B model produced explanations most closely aligned with the reference explanations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。