用逻辑框架让大模型为老人生成有目的的个性化故事,减少幻觉。
A Reflective Storytelling Agent for Older Adults: Integrating Argumentation Schemes and Argument Mining in LLM-Based Personalised Narratives

- 结合知识图谱与论证理论,引导和检查故事生成过程。
- 超三分之二的故事被认出有个人意义,一半体现论证逻辑。
- 适合关注老年人健康叙事、人机交互设计的研究者。
本研究探讨基于知识驱动的大语言模型(LLM)讲故事能否支持老年人与数字伴侣进行有意义的叙事互动。针对LLM存在的幻觉和透明度不足问题,提出一种反思性叙事代理,整合知识图谱、用户建模、论证理论与论证挖掘技术,以指导并检验叙事生成。研究分两阶段:第一阶段通过11位领域专家参与的设计工作坊开展形成性评估,迭代优化系统;最终系统生成基于结构化用户模型的故事,反映促进健康的活动与动机。第二阶段由55名老年人评估四种提示下、两种创意水平的个性化叙事,评价其目的感、实用性、文化相关性和不一致性。系统同时计算幻觉风险指标。参与者在约三分之二的故事中识别出个人相关的目的,其中近半数体现论证逻辑。文化可识别性显著影响使用意愿,轻微不一致在故事仍可理解且具个人关联时通常被接受。高幻觉风险指标的故事更常被视作不一致,而高论证质量指标则与更高清晰度和意义评分相关。研究证明,论证挖掘可作为健康导向的老年人故事生成中,形式化信号与人类评价之间的反思性检验机制。
原文摘要 · Abstract (English)
This work investigates whether knowledge-driven large language model (LLM)-based storytelling can support purposeful narrative interaction with a digital companion for older adults. To address known limitations of LLMs, including hallucinations and limited transparency, we present a reflective storytelling agent integrating knowledge graphs, user modelling, argumentation theory, and argument mining to guide and inspect narrative generation. The study consisted of two phases. Phase I employed participatory design involving 11 domain experts in a formative evaluation that informed iterative refinement. The resulting system generates narratives grounded in structured user models representing health-promoting activities and motivations. Phase II involved 55 older adults evaluating persona-based narratives across four prompts and two creativity levels. Participants assessed perceived purpose, usefulness, cultural relatability, and inconsistencies. The system additionally computed hallucination-risk indicators to evaluate generated narratives. Participants recognised personally relevant purposes in roughly two thirds of narratives, while argument-based purposes were identified in around half of these cases. Cultural recognisability strongly influenced willingness to use the functionality, whereas minor inconsistencies were often tolerated when narratives remained understandable and personally relevant. Narratives with higher hallucination-risk indicators were more often perceived as inconsistent, while higher argument-quality indicators tended to co-occur with higher clarity and meaningfulness ratings. Overall, the study positions argument mining as a reflective inspection mechanism for comparing formal grounding signals with human evaluations in health-oriented LLM storytelling for older adults.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。