为个性化智能体设计动态评估框架,提升推荐可信度与适应性。
Dynamic Evaluation Framework for Personalized and Trustworthy Agents: A Multi-Session Approach to Preference Adaptability
- 用模拟用户画像和结构化访谈收集偏好,构建可迭代评估体系。
- 基于LLM驱动的动态仿真,实现对推荐策略的持续测试与优化。
- 适合研究可信智能体、个性化推荐系统的学者与开发者。
生成式AI的进展推动了个性化智能体的研究热潮。随着个性化程度提高,用户对智能体决策与行动的信任需求也随之上升。然而,现有评估方法陈旧且无法捕捉用户交互的动态演化特性。本文提出一种全新的概念性评估框架,通过为用户建模独特属性与偏好,让智能体通过结构化访谈与模拟用户互动,获取偏好并提供定制化推荐。这些推荐由大语言模型驱动的仿真系统进行动态评估,支持自适应与迭代优化。该灵活框架适用于多种智能体与应用场景,全面评估主动、个性化、可信推荐策略的性能。
原文摘要 · Abstract (English)
Recent advancements in generative AI have significantly increased interest in personalized agents. With increased personalization, there is also a greater need for being able to trust decision-making and action taking capabilities of these agents. However, the evaluation methods for these agents remain outdated and inadequate, often failing to capture the dynamic and evolving nature of user interactions. In this conceptual article, we argue for a paradigm shift in evaluating personalized and adaptive agents. We propose a comprehensive novel framework that models user personas with unique attributes and preferences. In this framework, agents interact with these simulated users through structured interviews to gather their preferences and offer customized recommendations. These recommendations are then assessed dynamically using simulations driven by Large Language Models (LLMs), enabling an adaptive and iterative evaluation process. Our flexible framework is designed to support a variety of agents and applications, ensuring a comprehensive and versatile evaluation of recommendation strategies that focus on proactive, personalized, and trustworthy aspects.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。