arXiv:2606.11692cs.CYcs.AI2026-06

用智能体模拟器评估协商投票中的观点覆盖问题,提升决策代表性。

Evaluation of Alternative-Based Information Systems for Deliberative Polling using an Agentic Simulator

  • 构建基于大模型的智能体辩论模拟器,动态生成股东观点与论证链。
  • 实验显示推荐数量和论证密度显著影响观点覆盖度,最大覆盖率达87%。
  • 提出反向PageRank加权法,有效抵御恶意观点淹没攻击,适合政策评估场景。

协商投票通过在投票前向参与者呈现广泛论据来改善集体决策,但确保每位选民接触到代表性论据集合(覆盖问题)仍是挑战,尤其在大规模或存在策略性动机的选民中。本文引入一种基于大语言模型的智能体双极论证模拟器(ABAS),将投票过程形式化为六元组⟨Jend, Jopp, Ratt, Renh, VA, VR⟩,包含支持与反对理由、攻击与增强关系以及股东与关系权重。模拟器运行N个自主股东智能体,其潜在观点按[-1,1]分布,依次投票、选择或创作理由,并可提交论证图链接。系统采用可观测支持量排名现有理由的推荐机制,以覆盖率(即每个股东收到的K条推荐中,来自语料库理由标签集的比例)评估解决方案对NP难的包络理由问题的有效性。实验分析了创造力率(pown)、推荐规模(K)、论证密度(plinks)和种群规模(N)对覆盖率与语料多样性的影响。在无法进行Sybil攻击但仅关系图可被操控的认证选民环境中,测试评分机制抗策略攻击能力:标签洪水攻击导致覆盖率崩溃,而通过反向PageRank规则实现的作者计数加权显著优于均匀权重,具备更强鲁棒性。

原文摘要 · Abstract (English)

Deliberative polling promises to improve collective decision-making by exposing shareholders to a broad range of arguments before they vote. Yet ensuring that every voter encounters a representative sample of the reason space, the coverage problem, remains an open challenge, particularly at scale and in adversarial or strategically motivated electorates. This paper introduces a way of evaluating solutions using the LLM-based Agentic Bipolar Argumentation Simulator, grounded in a framework which formalises a poll as a six-tuple <Jend, Jopp, Ratt, Renh, VA, VR> of endorsing and opposing justifications, attack and enhance relations, and shareholder- and relation-weights. ABAS simulates N autonomous shareholder agents, each assigned a latent opinion according to desired distributions in [-1, 1], who sequentially vote, choose or author justifications, and optionally submit argumentation-graph links. The simulator implements recommendations that rank existing justifications by their observable endorsement mass. It evaluates the mechanism's success by coverage, namely the fraction of the corpus reason-tag set represented in the K recommendations presented to each shareholder, as a solution to the NP-hard Subsuming Justification Problem. Reported experiments characterise how creativity rate (pown), recommendation size (K), argumentation density (plinks), and population size (N) affect coverage and corpus diversity. In an authenticated electorate where Sybil attacks are impossible and only the relation graph is gameable, we stress-test the scoring with coordinated strategic voting attacks: a tag-flood attack collapses coverage, while author-count relation weighting through a reversed-PageRank rule resists the flood markedly better than uniform weights.

协商投票智能体模拟观点覆盖对抗攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。