arXiv:2606.18268cs.SIcs.AI2026-06

用多智能体模拟真实用户评价,提升社区事实核查效率。

Towards Multi-Agent-Simulation-Based Community Note Evaluation

论文配图:Towards Multi-Agent-Simulation-Based Community Note Evaluation
图 1 · 摘自论文原文
  • 构建多角色智能体模拟不同用户评分行为,生成可解释判断。
  • 在250万条社区笔记上达到84.7%准确率,优于现有方法。
  • 适合研究社会媒体可信度评估与自动化审核系统的人参考。

基于社区共识的事实核查正快速扩展于社交媒体平台,但人工贡献者带来的跨共识审核存在延迟高、覆盖率低的问题。为此,我们首先构建了ComRate数据集,包含250万条社区笔记及超过2.09亿条评分,数据源自$\\$X$。随后提出MultiCom框架,一种基于角色引导的多智能体评分系统,通过在矩阵分解的评分者空间中聚类贡献者,并驱动角色化智能体依据官方评分标准生成结构化评估。这些智能体输出包含置信度、一致性信号和理由等可解释判断。采用交叉验证校准的聚合算法整合原始投票与诊断性理由信号,实现可靠预测。大量实验表明,MultiCom在评测集上平均准确率达84.7%(平衡准确率68.3%,宏平均F1为60.1%),显著优于基线方法。

原文摘要 · Abstract (English)

Community-based fact-checking that relies on cross-consensus is expanding rapidly on social media platforms. However, the delay and low-ratio of cross-consensus community fact-checks rated by human contributors remains a significant challenge. To address this, we first created ComRate, a large-scale dataset comprising 2.5 million community notes and over 209 million ratings sourced from $\mathbb{X}$. We then propose MultiCom, a persona-guided multi-agent rating framework for community note evaluation. MultiCom simulates diverse rater population by clustering contributors in a matrix-factorized rater space and prompting persona agents to generate structured assessments based on the official community notes rating schema. These agents output structured and explainable judgments, such as confidence, agreement signals and reasons. An out-of-fold calibrated aggregation algorithm combines features such as raw votes and diagnostic reason signals for reliable prediction. Extensive evaluations demonstrate that MultiCom outperforms alternative methods, achieving an average accuracy of 84.7% (balanced accuracy 68.3%, macro-F1 60.1%) on the evaluation set.

社区核查多智能体可解释性评分预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。