arXiv:2603.16120cs.CL2026-03ACL被引 2

让AI研究助手真正懂用户:用真人反馈改进个性化科研工具

Language Models Don't Know What You Want: Evaluating Personalization in Deep Research Needs Real Users

  • 构建个性化科研代理MySQA,根据用户兴趣生成定制化报告
  • 真人测试发现9种机器无法识别的个性化错误,远超自动评估结果
  • 强调真实用户参与是衡量个性化进展的关键,不能只靠大模型打分

深度科研(DR)系统帮助研究人员应对论文数量激增的问题。这类工具可合成文献回答研究问题,但缺乏对用户的理解。为此,我们提出MyScholarQA(MySQA),一个个性化DR代理:1)推断用户的科研兴趣画像;2)针对用户输入查询提出个性化操作建议;3)生成符合用户认可操作的多章节报告。我们首先采用NLP标准流程测试:构建基于合成用户和大模型裁判的基准,结果显示MySQA在引用指标和个性化动作遵循方面优于基线。然而,我们怀疑该流程未能涵盖个性化DR的全部价值维度,因此通过在线版本对真实用户进行访谈以揭示深层问题。我们发现了九类大模型裁判无法检测的个性化缺陷,并基于定性反馈总结出未来设计的启示。总体而言,我们主张一个核心观点:真实用户参与是实现个性化进步的关键支柱,而当前易于使用的LLM裁判可能使NLP领域忽视这一本质。

原文摘要 · Abstract (English)

Deep Research (DR) systems help researchers cope with ballooning publishing counts. Such tools synthesize scientific papers to answer research queries, but lack understanding of their users. We address this with MyScholarQA (MySQA), a personalized DR agent that: 1) infers a profile with a user's research interests; 2) proposes personalized actions for a user's input query; and 3) writes a multi-section report for the query that follows user-approved actions. We first test MySQA with NLP's standard protocol: we build a benchmark with synthetic users and LLM judges, where MySQA beats baselines in citation metrics and personalized action-following. However, we suspect this process does not cover all aspects of personalized DR users value, so we interview users in an online version of MySQA to unmask them. We reveal nine nuanced errors of personalized DR undetectable by our LLM judges, and we study qualitative feedback to form lessons for future DR design. In all, we argue for a pillar of personalization that easy-to-use LLM judges can lead NLP to overlook: real progress in personalization is only possible with real users.

个性化科研工具真实用户大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。