用搜索记录模拟个人知识,小模型微调后表现超人
Individual Text Corpora Predict User-Specific Knowledge: Benchmarks of Individualized Knowledge Simulation

- 用用户搜索历史构建文本语料,微调小模型预测知识
- 微调后模型在公开题上超过真人和德国常模
- 适合研究个性化知识建模与模型校准的学者
本研究检验了从搜索历史中提取的个体文本语料(ICs)是否可用于模拟个体知识。我们收集了316名成人的ICs,他们回答了36道多项选择题,并比较了多个大语言模型(LLMs)的表现,仅Qwen3-1.7B在任务中可行。通过低秩适应(LoRA)进行特定任务微调后,该模型在公开题目上的表现超越了参与者及德国常模样本。然而,在非公开问题上,模型表现仍低于参与者,提示公开题可能存在训练数据污染。将ICs融入检索增强生成以预测个体回答时,模型与参与者匹配准确率显著高于随机水平,表明存在可检测的个体知识信号。但模型对正确答案的概率估计偏低,校准性差。知识空白预测表现不佳,但当语料超过五百万词元时有所改善。我们提出基于熵的评估基准,作为个性化知识模拟的校准指标。
原文摘要 · Abstract (English)
This study examines whether individual text corpora (ICs) from search histories can be used to simulate individual knowledge. We collected ICs from 316 adults, who answered 36 multiple-choice knowledge items, and compared several large language models (LLMs) on this task, of which only Qwen3-1.7B proved viable. After task-specific fine-tuning via Low-Rank Adaptation (LoRA), Qwen3-1.7B outperformed both participants and a representative German norm sample on publicly available items. On non-public questions, however, the LLM performed worse than our participants, suggesting possible training data contamination for the public questions. When integrating ICs into retrieval-augmented generation to predict individual responses, LLM-participant Match accuracies significantly exceeded chance, which demonstrates a detectable individual knowledge signal. The probabilities assigned to the participants' answers were, however, low and far below the probability of correct answers, indicating poor calibration toward individual response patterns. Knowledge-gap prediction was sub-optimal, though it improved for corpora exceeding five million tokens. We discuss our entropy based evaluation benchmarks as calibration indices for individualized knowledge simulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。