arXiv:2507.03059cs.CYcs.AI2025-07

用已故学者的电脑数据训练AI,复刻其写作风格与专业知识。

AI-Based Reconstruction from Inherited Personal Data: Analysis, Feasibility, and Prospects

  • 用一百万字文本微调GPT-4,还原研究者文风与专业能力。
  • 包含元数据和非文本数据可进一步提升复刻精度。
  • 适合关注数字遗产、AI辅助科研的人士参考。

本文探讨通过分析已故研究人员个人电脑中的数据,训练人工智能(AI)以创建其“电子副本”的可行性。典型研究者电脑中约有100万字的文本数据,包括论文、邮件和草稿,足以对GPT-4等先进预训练模型进行微调,高保真还原研究者的写作风格、领域知识与修辞特征。研究还指出,引入非文本数据与文件元数据可进一步丰富AI对研究者的表征。该理念可扩展至活人与电子副本的交流、多个电子副本间的协作,以及组织级电子副本的构建与互联,以优化信息获取与战略决策。伦理问题如所有权与安全性被强调为实现过程中需重点关注的事项。结果表明,基于AI的智力遗产保存与增强具有广阔前景。

原文摘要 · Abstract (English)

This article explores the feasibility of creating an "electronic copy" of a deceased researcher by training artificial intelligence (AI) on the data stored in their personal computers. By analyzing typical data volumes on inherited researcher computers, including textual files such as articles, emails, and drafts, it is estimated that approximately one million words are available for AI training. This volume is sufficient for fine-tuning advanced pre-trained models like GPT-4 to replicate a researcher's writing style, domain expertise, and rhetorical voice with high fidelity. The study also discusses the potential enhancements from including non-textual data and file metadata to enrich the AI's representation of the researcher. Extensions of the concept include communication between living researchers and their electronic copies, collaboration among individual electronic copies, as well as the creation and interconnection of organizational electronic copies to optimize information access and strategic decision-making. Ethical considerations such as ownership and security of these electronic copies are highlighted as critical for responsible implementation. The findings suggest promising opportunities for AI-driven preservation and augmentation of intellectual legacy.

AI复刻数字遗产文本生成研究助手

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。