评测视觉语言模型在亚里士多德说服三要素上的表现
Evaluating VLMs on Multimodal Aristotelian Persuasion Tasks
- 基于亚里士多德说服模型设计图文推理任务
- 通义千问系列在逻辑/情感检测上表现优异,通义千问2在伦理检测中竞争力强
- 适合关注多模态推理与人类认知机制的科研人员
视觉语言模型(VLMs)在多项任务中表现出色,但在复杂任务上的评估仍不充分。受亚里士多德说服模型启发,我们构建了以逻辑(Logos)、修辞(Ethos)、情感(Pathos)为核心的三要素评估框架,并基于ImageArg数据集开展测评。结果显示,通义千问系列模型整体表现提升,其中Qwen3在逻辑和情感识别任务中表现突出,而Qwen2在更复杂的伦理判断任务中也展现出良好性能。研究代码已开源,以推动该方向的持续探索。
原文摘要 · Abstract (English)
Vision Language Models (VLMs) have demonstrated exceptional performance across various tasks. However, they have not yet been thoroughly evaluated on more complex tasks. The Persuasion Model, conceived by Aristotle, resembles a triangle shape, which highlights its inherent challenges related to personal biases. To assess the progress of VLMs on these complex tasks, we use the ImageArg datasets, focusing on the Logos, Ethos, and Pathos detection tasks. Our findings indicate that models from the Qwen family achieve improved F1 scores, with Qwen3 performing exceptionally well on the Logos and Pathos tasks, while Qwen2 exhibits competitive performance on the more complex Ethos detection task. We release the code to foster research in this direction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。