通过少量语音数据构建精简个性化词汇表,提升人机对话中的语言一致性。
Towards Stable and Personalised Profiles for Lexical Alignment in Spoken Human-Agent Dialogue
- 基于10分钟语音数据,为每类词性设计5-10个核心词项
- 小而紧凑的词汇表在召回率与覆盖率上表现最优
- 适合追求高效、稳定人机对齐的对话系统开发者
词汇对齐指对话双方逐渐使用相似词汇,有助于沟通成功。然而,当前对话代理中对此研究仍不足,尤其在大语言模型背景下。本研究借鉴个性化策略,探索构建稳定、个性化的词汇特征档案作为对齐基础。通过调整用于构建的语音转录数据量及每词性类别中的词项数量,评估档案随时间的表现,使用召回率、覆盖率和余弦相似度指标。结果表明,仅需10分钟转录语音,包含每类词性各5项(形容词、连词)或10项(副词、名词、代词、动词),即可实现性能与数据效率的最佳平衡。研究为对话代理实现词汇对齐提供了实用路径。
原文摘要 · Abstract (English)
Lexical alignment, where speakers start to use similar words across conversation, is known to contribute to successful communication. However, its implementation in conversational agents remains underexplored, particularly considering the recent advancements in large language models (LLMs). As a first step towards enabling lexical alignment in human-agent dialogue, this study draws on strategies for personalising conversational agents and investigates the construction of stable, personalised lexical profiles as a basis for lexical alignment. Specifically, we varied the amounts of transcribed spoken data used for construction as well as the number of items included in the profiles per part-of-speech (POS) category and evaluated profile performance across time using recall, coverage, and cosine similarity metrics. It was shown that smaller and more compact profiles, created after 10 min of transcribed speech containing 5 items for adjectives, 5 items for conjunctions, and 10 items for adverbs, nouns, pronouns, and verbs each, offered the best balance in both performance and data efficiency. In conclusion, this study offers practical insights into constructing stable, personalised lexical profiles, taking into account minimal data requirements, serving as a foundational step toward lexical alignment strategies in conversational agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。