用户推理过程可作个人数据,社区可据此蒸馏出更契合自身需求的模型。
Conscious Data Contribution via Community-Driven Chain-of-Thought Distillation
- 将用户推理中间步骤视为个人数据,支持数据可携权
- 社区聚合知识蒸馏出性能更优、目标对齐的新模型
- 适合关注隐私保护与集体智能的AI研究者
当前大模型训练依赖大规模数据,催生了如大语言模型对话机器人等新应用,也引发数据隐私与用户选择权问题。本文聚焦于使用思维链(CoT)推理的大语言模型,分析其计算过程中产生的中间文本是否构成用户个人数据。基于数据可携性法律解读,我们主张这些中间产物应被视为用户数据。在此基础上,结合有意识数据贡献框架,提出社区可整合共享知识,蒸馏出更契合自身目标的替代模型。实证验证了该方法有效性,并研究了社区多样性、推理粒度与规模对蒸馏性能的影响。
原文摘要 · Abstract (English)
The current era of AI development places a heavy emphasis on training large models on increasingly scaled-up datasets. This paradigm has catalyzed entirely new product categories, such as LLM chatbots, while also raising concerns about data privacy and consumer choice. In this paper, we consider questions of data portability and user autonomy in the context of LLMs that "reason" using chain-of-thought (CoT) traces, computing intermediate text artifacts from user input before producing a final output. We first interpret recent data privacy and portability law to argue that these intermediate computations qualify as users' personal data. Then, building on the existing framework of Conscious Data Contribution, we show how communities who receive low utility from an available model can aggregate and distill their shared knowledge into an alternate model better aligned with their goals. We verify this approach empirically and investigate the effects of community diversity, reasoning granularity, and community size on distillation performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。