arXiv:2508.10769cs.AIcs.MM2025-08

构建首个多模态AI内容人类反应数据集,揭示图文不一致最易暴露虚假内容。

Modeling Human Responses to Multimodal AI Content

  • 建立包含15万条帖子的MhAIM数据集,分析人类对多模态AI内容的反应机制。
  • 发现图文不一致时,人更易识别出AI生成内容,准确率显著提升。
  • 提出T-Lens系统,让大模型基于预测人类反应来增强交互与可解释性。

随着AI生成内容日益普及,虚假信息风险随之上升。以往研究多聚焦于内容真实性验证,但对其如何影响人类认知与行为了解甚少。在金融、股市等场景中,预测用户反应(如新闻是否传播)的重要性甚至超过事实核查。为此,本文采用以人为中心的方法,构建MhAIM数据集,包含154,552条在线帖子(其中111,153条为AI生成),支持大规模分析人类对多模态AI内容的响应。人类实验表明,当文本与视觉存在不一致时,人们更易识别出AI生成内容。为此,我们提出可信度、影响力、开放性三项新指标,量化用户对内容的判断与互动。进一步设计了基于LLM的T-Lens系统,其核心为HR-MCP(Human Response Model Context Protocol),基于标准化的模型上下文协议(MCP),可无缝集成任意大模型。该系统能预测人类对多模态信息的反应,从而提升模型的可解释性与交互能力。本研究提供实证洞察与实用工具,助力大模型具备人类感知能力,揭示AI、认知与信息接收之间的复杂互动,提出可操作的应对策略以缓解AI驱动的虚假信息风险。

原文摘要 · Abstract (English)

As AI-generated content becomes widespread, so does the risk of misinformation. While prior research has primarily focused on identifying whether content is authentic, much less is known about how such content influences human perception and behavior. In domains like trading or the stock market, predicting how people react (e.g., whether a news post will go viral), can be more critical than verifying its factual accuracy. To address this, we take a human-centered approach and introduce the MhAIM Dataset, which contains 154,552 online posts (111,153 of them AI-generated), enabling large-scale analysis of how people respond to AI-generated content. Our human study reveals that people are better at identifying AI content when posts include both text and visuals, particularly when inconsistencies exist between the two. We propose three new metrics: trustworthiness, impact, and openness, to quantify how users judge and engage with online content. We present T-Lens, an LLM-based agent system designed to answer user queries by incorporating predicted human responses to multimodal information. At its core is HR-MCP (Human Response Model Context Protocol), built on the standardized Model Context Protocol (MCP), enabling seamless integration with any LLM. This integration allows T-Lens to better align with human reactions, enhancing both interpretability and interaction capabilities. Our work provides empirical insights and practical tools to equip LLMs with human-awareness capabilities. By highlighting the complex interplay among AI, human cognition, and information reception, our findings suggest actionable strategies for mitigating the risks of AI-driven misinformation.

多模态人类行为大模型信息风险

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。