arXiv:2606.22738cs.IR2026-06

构建可模拟用户信任与验证行为的智能体,评估AI生成内容的可信度。

PA-User: Simulating Trust and Verification under AI-Generated Content

  • 引入信任预算与动态信念更新机制,模拟用户验证行为。
  • 在HC3数据集上将信任校准误差降至0.162,显著优于无信任机制方案。
  • 适用于研究AI内容可信度、信息检索系统评估的学者与工程师。

当前在线信息用户普遍认为部分内容由AI生成或修改,混合情形(如人类文本经语言模型重写、AI筛选内容伪装成编辑、基于检索的实时回答)难以辨别,用户需持续投入成本验证真伪。现有信息检索用户模拟器无法建模此现象。本文提出PA-User用户模拟器,包含三项新组件:验证耗能预算(会话间恢复)、对各来源类别的事实性独立贝塔信念(按来源领域划分)、以及依赖当前信任、资源与领域风险的接受/验证/丢弃决策规则。框架具备两项验证属性:信任后验收敛于真实事实性(面效度),各组件影响可通过消融实验分离(结构效度)。在含85,449对人工与ChatGPT答案的HC3数据集上,带信任组件的模型信任校准误差为0.162,远低于无信任机制的0.356。相比始终接受的基线,高风险后悔值从0.171降至0.122(相对减少29%),验证率达34.5%,是无预算消融版本的一半。单机制消融可单独诊断各组件作用。

原文摘要 · Abstract (English)

Most users of online information now assume that some of what they read has been written, edited, or selected by an AI model. Hybrid cases are the hardest to tell apart: human prose rewritten by a language model, AI-curated lists presented as editorial, retrieval-augmented answers composed on the fly from human sources. Users cannot reliably distinguish these cases, and the ongoing cost of checking what is genuine has become part of how they search. Current user simulators in information retrieval do not model this. We propose PA-User, a user simulator with three new components: a detection-effort budget that is spent on verification and recovers between sessions; a trust component that holds a separate Beta belief over the factuality of each source class (domain by provenance) and updates from observed outcomes; and a decision rule that picks accept, verify, or discard for each result, conditional on current trust, current effort, and per-domain stakes. We state two verification-and-validation (V\&V) properties of the framework. The trust posterior converges to the true class factuality (face validity). Each component's contribution to any observable can be isolated by ablation (structural validity). On the HC3 corpus (85,449 paired human and ChatGPT answers in five domains), PA-User reaches a trust-calibration error of $0.162$, against $0.356$ for any configuration without the trust component. PA-User reduces high-stakes regret from $0.171$ to $0.122$ ($29\%$ relative) against an always-accept ablation, and verifies $34.5\%$ of results, half the rate of an ablation with no effort budget. Each single-mechanism ablation isolates one component, which makes the framework individually diagnosable.

用户模拟可信度评估AI内容

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。