arXiv:2503.17936cs.CLcs.AI2025-03被引 4

研究大模型问答中不完整与模糊问题如何影响多轮交互需求

An Empirical Study of the Role of Incompleteness and Ambiguity in Interactions with Large Language Models

  • 用神经符号框架分析人机交互中的问题不完整性和模糊性
  • 发现不完整或模糊问题需多轮交互,且交互越长问题越清晰
  • 提出可量化评估问答质量的新指标,适合对话系统研究者

自然语言作为人机交互媒介,在大语言模型(LLMs)出现后迎来重大变革。如今人们常将LLM视作现代“预言家”,可随时提问。与古希腊德尔斐神谕不同,使用LLM无需单次问答;也不同于皮提亚女祭司,其回答可通过补充上下文改进。本文旨在研究在何种情况下需要多轮交互才能成功解答问题,或判断问题无法回答。我们提出一种神经符号框架,建模人与LLM之间的交互过程,定义问题的不完整性和模糊性为可从对话消息中推导出的属性。基于基准数据集的结果表明,答案正确性取决于问题是否表现出不完整或模糊特征。实验显示,不完整或模糊问题比例高的数据集通常需要多轮交互;增加交互长度能有效降低这些问题的显著性。结果还表明,所提出的不完整性和模糊性度量可作为刻画问答交互的重要工具。

原文摘要 · Abstract (English)

Natural language as a medium for human-computer interaction has long been anticipated, has been undergoing a sea-change with the advent of Large Language Models (LLMs) with startling capacities for processing and generating language. Many of us now treat LLMs as modern-day oracles, asking it almost any kind of question. Unlike its Delphic predecessor, consulting an LLM does not have to be a single-turn activity (ask a question, receive an answer, leave); and -- also unlike the Pythia -- it is widely acknowledged that answers from LLMs can be improved with additional context. In this paper, we aim to study when we need multi-turn interactions with LLMs to successfully get a question answered; or conclude that a question is unanswerable. We present a neural symbolic framework that models the interactions between human and LLM agents. Through the proposed framework, we define incompleteness and ambiguity in the questions as properties deducible from the messages exchanged in the interaction, and provide results from benchmark problems, in which the answer-correctness is shown to depend on whether or not questions demonstrate the presence of incompleteness or ambiguity (according to the properties we identify). Our results show multi-turn interactions are usually required for datasets which have a high proportion of incompleteness or ambiguous questions; and that that increasing interaction length has the effect of reducing incompleteness or ambiguity. The results also suggest that our measures of incompleteness and ambiguity can be useful tools for characterising interactions with an LLM on question-answeringproblems

大模型问答系统交互设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。