构建首个个性化问答核心意图识别基准,填补系统理解用户优先需求的空白
IPQA: A Benchmark for Core Intent Identification in Personalized Question Answering
- 基于用户答案选择行为推断其优先意图,结合满意理论建模决策过程
- 在多领域数据集上测试发现主流模型对历史信息中的核心意图识别率低
- 适合研究个性化对话、用户意图建模及智能助手可解释性的学者使用
意图识别是生成个性化问答恰当响应的基础。然而,现有基准仅评估响应质量或检索性能,未直接衡量意图识别能力。这一缺口至关重要,因为若无法理解用户优先关注的意图,系统就无法满足其信息需求。为此,我们提出“核心意图”概念:用户为满足信息需求而优先选择的答案意图。为评估该能力,我们构建了IPQA——个性化问答中核心意图识别的基准。由于用户不明确表达优先意图,我们通过分析答案选择中的可观测行为模式,依据满意理论(satisficing theory)推断核心意图。数据集通过系统性筛选、大模型标注与自动化验证结合人工校验构建,覆盖多个领域。实验表明,当前先进语言模型在个性化场景下难以识别核心意图,且随着问题复杂度提升,性能持续下降。代码与数据集将公开,以推动该方向研究。
原文摘要 · Abstract (English)
Intent identification serves as the foundation for generating appropriate responses in personalized question answering (PQA). However, existing benchmarks evaluate only response quality or retrieval performance without directly measuring intent identification capabilities. This gap is critical because without understanding which intents users prioritize, systems cannot generate responses satisfying individual information needs. To address this, we introduce the concept of core intents: intents users prioritize when selecting answers to satisfy their information needs. To evaluate these core intents, we propose IPQA, a benchmark for core Intent identification in Personalized Question Answering. Since users do not explicitly state their prioritized intents, we derive core intents from observable behavior patterns in answer selection, grounded in satisficing theory where users choose answers meeting their acceptance thresholds. We construct a dataset with various domains through systematic filtering, LLM-based annotation, and rigorous quality control combining automated verification with human validation. Experimental evaluations across state-of-the-art language models reveal that current systems struggle with core intent identification in personalized contexts. Models fail to identify core intents from user histories, with performance degrading as question complexity increases. The code and dataset will be made publicly available to facilitate future research in this direction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。