测试大模型能否根据用户偏好个性化摘要,发现多数模型实际做不到。
Are Large Language Models In-Context Personalized Summarizers? Get an iCOPERNICUS Test Done!
- 设计新测试框架iCOPERNICUS,检验模型是否利用阅读史和用户差异做个性化摘要。
- 17个主流模型中15个在更丰富提示下性能下降,降幅1.6%至3.6%。
- 适合关注模型个性化能力、评测真实个性化的研究者与开发者。
大型语言模型在基于上下文学习(ICL)的摘要任务上表现优异,但摘要的突出性应匹配用户的特定偏好历史。因此,需要可靠的上下文个性化学习(ICPL)能力。为评估任意大模型是否具备ICPL,需能识别用户画像差异。近期研究首次提出度量个性化程度的指标EGISES,衡量模型对用户差异的响应能力。然而,该方法无法检验模型是否同时利用三种提示线索:(i) 示例摘要,(ii) 用户阅读历史,(iii) 用户画像对比。为此,我们提出iCOPERNICUS框架,以EGISES为对比基准,系统审视大模型的摘要个性化学习能力。作为案例研究,我们基于报告的ICL性能评估17个前沿大模型,发现其中15个在使用更丰富提示时,其个性化能力下降(最小降幅1.6%,最大降幅3.6%),表明其缺乏真正的上下文个性化学习能力。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have succeeded considerably in In-Context-Learning (ICL) based summarization. However, saliency is subject to the users' specific preference histories. Hence, we need reliable In-Context Personalization Learning (ICPL) capabilities within such LLMs. For any arbitrary LLM to exhibit ICPL, it needs to have the ability to discern contrast in user profiles. A recent study proposed a measure for degree-of-personalization called EGISES for the first time. EGISES measures a model's responsiveness to user profile differences. However, it cannot test if a model utilizes all three types of cues provided in ICPL prompts: (i) example summaries, (ii) user's reading histories, and (iii) contrast in user profiles. To address this, we propose the iCOPERNICUS framework, a novel In-COntext PERsonalization learNIng sCrUtiny of Summarization capability in LLMs that uses EGISES as a comparative measure. As a case-study, we evaluate 17 state-of-the-art LLMs based on their reported ICL performances and observe that 15 models' ICPL degrades (min: 1.6%; max: 3.6%) when probed with richer prompts, thereby showing lack of true ICPL.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。