自动生成带记忆的对话数据,提升大模型长短期记忆能力。
AgenticAI-DialogGen: Topic-Guided Conversation Generation for Fine-Tuning and Evaluating Short- and Long-Term Memories of LLMs

- 用智能体框架自动构建人物设定与话题引导的对话
- 生成的对话数据让模型在记忆问答任务上表现更优
- 适合研究记忆机制或需要高质量对话数据的研究者
大语言模型虽能处理长对话上下文,但缺乏同时包含短时与长时对话历史的数据集,导致记忆能力的微调与评估困难。现有对话数据集或缺乏记忆锚点,或忽视话题连贯性,或依赖昂贵的人工标注。为此,我们提出 AgenticAI-DialogGen——一个无需人工干预的模块化智能体框架,通过大模型代理从非结构化对话中提取知识图谱、识别话题、构建说话人角色,并模拟话题引导的对话。此外,引入问答模块生成基于短时与长时对话历史的记忆锚定问题对。我们还构建了新数据集 TopicGuidedChat (TGC),其中长时记忆以说话人专属知识图谱形式编码,短时记忆则为新生成的话题引导对话。评估表明,AgenticAI-DialogGen 生成的对话质量更高,且在 TGC 数据集上微调的模型在记忆问答任务上表现显著提升。
原文摘要 · Abstract (English)
Recent advancements in Large Language Models (LLMs) have improved their ability to process extended conversational contexts, yet fine-tuning and evaluating short- and long-term memories remain difficult due to the absence of datasets that encode both short- and long-term conversational history. Existing conversational datasets lack memory grounding, overlook topic continuity, or rely on costly human annotation. To address these gaps, we introduce AgenticAI-DialogGen, a modular agent-based framework that generates persona-grounded and topic-guided conversations without human supervision. The framework uses LLM agents to extract knowledge graphs, identify topics, build speaker personas, and simulate topic-guided conversations from unstructured conversations. A QA module generates memory-grounded Question Answer (QA) pairs drawn from short- and long-term conversational histories. We also generated a new dataset entitled, TopicGuidedChat (TGC), where long-term memory is encoded as speaker-specific knowledge graphs and short-term memory as newly generated topic-guided conversations. Evaluations depict that AgenticAI-DialogGen yields higher conversational quality and LLMs fine-tuned on TGC dataset achieve improved performance on memory-grounded QA tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。