arXiv:2512.04343cs.IR2025-12

个性化让AI回答更精准,但现有评估方法会误判其优点。

The Personalization Paradox: Semantic Loss vs. Reasoning Gains in Agentic AI Q&A

  • 用12个真实学生问题对比10种配置,测试个性化对AI问答的影响。
  • 个性化提升推理与事实准确度,但降低语义相似度评分。
  • 揭示当前评估体系缺陷,建议用多维度指标衡量个性化效果。

本文以AIVisor这一用于学生咨询的智能体检索增强大模型为研究对象,通过12个刻意强调词汇精确性的真实咨询问题,对比了十种个性化与非个性化系统配置。采用线性混合效应模型分析了词汇(BLEU、ROUGE-L)、语义(METEOR、BERTScore)和事实依据(RAGAS)三类指标。结果表明:个性化显著提升推理质量与事实准确性,但导致语义相似度下降,原因并非答案质量差,而是现有评估指标对个性化表达偏差过度惩罚。这暴露了当前大模型评估体系在用户特定响应上的结构性缺陷。全集成个性化配置在整体表现上最优,说明当使用多维指标时,个性化可真正提升系统效能。研究证明个性化带来的是指标依赖性变化而非全面改进,为构建更透明、稳健的个性化智能体提供了方法论基础。

原文摘要 · Abstract (English)

AIVisor, an agentic retrieval-augmented LLM for student advising, was used to examine how personalization affects system performance across multiple evaluation dimensions. Using twelve authentic advising questions intentionally designed to stress lexical precision, we compared ten personalized and non-personalized system configurations and analyzed outcomes with a Linear Mixed-Effects Model across lexical (BLEU, ROUGE-L), semantic (METEOR, BERTScore), and grounding (RAGAS) metrics. Results showed a consistent trade-off: personalization reliably improved reasoning quality and grounding, yet introduced a significant negative interaction on semantic similarity, driven not by poorer answers but by the limits of current metrics, which penalize meaningful personalized deviations from generic reference texts. This reveals a structural flaw in prevailing LLM evaluation methods, which are ill-suited for assessing user-specific responses. The fully integrated personalized configuration produced the highest overall gains, suggesting that personalization can enhance system effectiveness when evaluated with appropriate multidimensional metrics. Overall, the study demonstrates that personalization produces metric-dependent shifts rather than uniform improvements and provides a methodological foundation for more transparent and robust personalization in agentic AI.

个性化评估方法大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。