arXiv:2412.07019cs.CLcs.CY2024-12被引 4

用大模型评估阴谋论影响,发现多步推理更准但模型有偏见。

Assessing the Impact of Conspiracy Theories Using Large Language Models

  • 设计多步推理策略,让大模型像人一样分析阴谋论证据。
  • 多步推理评估准确率高,但模型对前置阴谋论更敏感。
  • 适合研究舆情、危机应对与模型偏见的学者或从业者。

衡量阴谋论(CTs)的相对影响对于危机时期优先响应和资源分配至关重要。然而,评估其对公众的实际影响面临独特挑战,不仅需要获取特定于阴谋论的知识,还需涵盖社会、心理和文化等多维信息。近期大语言模型(LLMs)的发展表明其在此场景中具有潜力,因其具备大规模训练语料中的丰富知识,并能支持复杂推理。本文构建了包含主流阴谋论的人工标注影响数据集,借鉴人类评估流程,设计定制化策略以利用大模型进行类人化的阴谋论影响评估。通过严格实验发现,采用多步推理分析更多相关证据的评估模式可产生准确结果;但多数大模型存在显著偏差,例如在提示中较早出现的阴谋论会被赋予更高影响评分,且对情绪化、冗长的阴谋论生成的评估准确性更低。

原文摘要 · Abstract (English)

Measuring the relative impact of CTs is important for prioritizing responses and allocating resources effectively, especially during crises. However, assessing the actual impact of CTs on the public poses unique challenges. It requires not only the collection of CT-specific knowledge but also diverse information from social, psychological, and cultural dimensions. Recent advancements in large language models (LLMs) suggest their potential utility in this context, not only due to their extensive knowledge from large training corpora but also because they can be harnessed for complex reasoning. In this work, we develop datasets of popular CTs with human-annotated impacts. Borrowing insights from human impact assessment processes, we then design tailored strategies to leverage LLMs for performing human-like CT impact assessments. Through rigorous experiments, we textit{discover that an impact assessment mode using multi-step reasoning to analyze more CT-related evidence critically produces accurate results; and most LLMs demonstrate strong bias, such as assigning higher impacts to CTs presented earlier in the prompt, while generating less accurate impact assessments for emotionally charged and verbose CTs.

大模型阴谋论偏见评估多步推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。