arXiv:2605.22971cs.CLcs.HC2026-05

用聊天记录估算员工专业能力,大模型效果参差不齐

Can AI Guess What You Know? Performance Comparison of Large Language Models for Human Domain Knowledge Estimation From Communication Logs

论文配图:Can AI Guess What You Know? Performance Comparison of Large Language Models for Human Domain Knowledge Estimation From Communication Logs
图 1 · 摘自论文原文
  • 用对话日志零样本推断个人专业领域知识
  • Gemini 2.5 Flash误差最低(MAE 21.13%)
  • 文本量多不等于推断准,适合组织知识管理

员工常难以判断'谁懂什么',导致组织效率损失。本文研究大型语言模型(LLMs)能否从长期Slack日志中直接推断个体领域知识。基于43名用户的27,188条消息,评估了七种模型(含Gemini、Claude及GPT系列),将零样本估计结果与27名参与者的自评技能评分对比。Gemini 2.5 Flash表现最佳,平均绝对误差(MAE)为21.13%,而GPT系列模型误差显著更高。值得注意的是,推断准确率与消息数量相关性很弱,表明单纯增加文本量无法保证更好推断。研究证实了自动化专家地图的可行性与当前局限,强调需考虑隐私保护及更结构化的知识表示。

原文摘要 · Abstract (English)

Employees often struggle to identify ``who knows what,'' leading to organizational productivity losses. We investigate whether Large Language Models (LLMs) can infer individual domain knowledge directly from long-term Slack logs. Analyzing 27,188 messages from 43 users, we evaluated seven models (including Gemini, Claude, and GPT families) by comparing their zero-shot estimates against self-reported skill ratings from 27 participants. Gemini 2.5 Flash achieved the lowest error (MAE 21.13%), while GPT models showed significantly larger discrepancies. Notably, estimation accuracy depended only weakly on message volume, indicating that more text alone does not guarantee better inference. These findings demonstrate the feasibility and current limits of automated expertise mapping, highlighting the need for privacy-preserving deployments and richer, structure-aware representations of human knowledge.

知识图谱大模型应用职场智能隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。