提升去中心化大模型推理的奖励公平性,抵御恶意评分干扰。
Adaptive and Robust Cost-Aware Proof of Quality for Decentralized LLM Inference Networks
- 引入自适应信任加权共识,动态调整评价节点权重。
- 在多种攻击下,鲁棒聚合使评分与真实质量更一致。
- 适合构建抗作弊的分布式大模型推理激励系统。
去中心化大语言模型推理网络需要轻量级机制,在异构延迟和成本条件下奖励高质量输出。Proof of Quality 通过采样评估节点对候选输出打分,并聚合得分形成共识信号以决定奖励。然而,评估节点异质性及恶意评分操纵会扭曲共识,导致奖励虚高,削弱开放参与下的激励一致性。本文扩展了成本感知的 Proof of Quality 机制,加入抗敌对的共识生成方法。研究了中位数、截尾均值等鲁棒聚合规则,以及基于偏差信号动态更新评价节点权重的自适应信任加权共识。通过问答与摘要任务,利用真实答案作为离线分析代理,量化评估者可靠性,发现其差异显著,且存在任务依赖的错配,甚至反转相关性。进一步在四种对抗策略(噪声注入、抬高、破坏、间歇操纵)下,测试不同恶意比例与采样规模组合的表现。结果表明,鲁棒聚合显著提升共识与真实质量代理的一致性,降低对噪声和策略性攻击的敏感度。同时揭示采样带来的权衡:更大采样集虽降低评估者奖励并增加收益方差,但推理奖励相对稳定。这些发现支持将鲁棒共识作为成本感知 Proof of Quality 的默认组件,并为在对抗风险与资源约束下选择采样参数提供实用指导。
原文摘要 · Abstract (English)
Decentralized large language model inference networks require lightweight mechanisms to reward high quality outputs under heterogeneous latency and cost. Proof of Quality provides scalable verification by sampling evaluator nodes that score candidate outputs, then aggregating their scores into a consensus signal that determines rewards. However, evaluator heterogeneity and malicious score manipulation can distort consensus and inflate payouts, which weakens incentive alignment in open participation settings. This paper extends a cost-aware Proof of Quality mechanism by adding adversary-resilient consensus formation. We study robust aggregation rules, including median and trimmed mean, and an adaptive trust-weighted consensus that updates evaluator weights from deviation signals. Using question answering and summarization workloads with a ground truth proxy for offline analysis, we quantify evaluator reliability and show strong variance across evaluators, including task-dependent misalignment that can invert correlations. We then evaluate robustness under four adversarial strategies, including noise injection, boosting, sabotage, and intermittent manipulation, across a sweep of malicious ratios and evaluator sample sizes. Our results show that robust aggregation improves consensus alignment with the ground truth proxy and reduces sensitivity to noisy and strategic attacks compared with simple averaging. We further characterize the operational trade-off introduced by evaluator sampling, where larger evaluator sets reduce evaluator rewards and increase payoff variance while inference rewards remain relatively stable in our configuration. These findings motivate robust consensus as a default component for cost-aware Proof of Quality and provide practical guidance for selecting evaluator sampling parameters under adversarial risk and resource constraints.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。