arXiv:2505.23842cs.CLecon.GN2025-05被引 9

用谢泼德值公平分配LLM摘要中文档的贡献,保障内容创作者收益。

Fair Document Valuation in LLM Summaries via Shapley Values

  • 基于语义聚类的谢泼德值近似方法,提升大规模计算效率
  • 在亚马逊评论数据上显著优于蒙特卡洛等传统近似方法
  • 不依赖具体模型或评估方式,适用于多种摘要场景

大型语言模型(LLMs)正日益用于搜索引擎和AI助手,直接提供多源内容的摘要答案,却遮蔽了原始内容创作者的贡献,威胁内容生态的可持续性。本文将此问题定义为公平文档估值与补偿,并提出基于谢泼德值的框架。由于精确计算成本过高,我们开发了聚类谢泼德值(Cluster Shapley),通过LLM嵌入对语义相似文档进行分组,在集群层面计算谢泼德值,并给出近似误差与收益归属误差的理论边界。在亚马逊产品评论数据上,现有近似方法(如蒙特卡洛采样、Kernel SHAP)在LLM场景下表现不佳,而聚类谢泼德值显著提升了效率-精度权衡。简单的分配启发式(如均分或基于相关性的分配)虽计算便宜,但结果极不公平。该方法对具体使用的LLM、摘要过程及评估方式均无感,具有广泛适用性。

原文摘要 · Abstract (English)

Large Language Models (LLMs) increasingly power search engines and AI assistants that retrieve and summarize content from many sources. By serving answers directly, these systems obscure the original content creators' contributions, threatening the compensation that sustains a healthy content ecosystem. We frame this as a problem of fair document valuation and compensation, and propose a framework based on the Shapley value. Because exact Shapley computation is prohibitively expensive at scale, we develop Cluster Shapley, an approximation that groups semantically similar documents via LLM embeddings and computes Shapley values at the cluster level, with formal bounds on both the approximation error and the induced revenue-attribution error. On Amazon product review data, off-the-shelf approximations such as Monte Carlo sampling and Kernel SHAP perform suboptimally in LLM settings, whereas Cluster Shapley substantially improves the efficiency--accuracy frontier. Simple attribution heuristics (e.g., equal or relevance-based allocation), though computationally cheap, yield highly unfair outcomes. Our approach is agnostic to the exact LLM used, the summarization process used, and the evaluation procedure, which makes it broadly applicable to a variety of summarization settings.

公平性谢泼德值文档估值LLM应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。