arXiv:2608.24920cs.CLcs.AI2026-08中稿 · the AI in Measurem…

不同大模型对同一对话的回复语义差异大,影响评估一致性。

Semantic Variability of Replies Across LLMs: Implications for Designing Conversation-Based Assessment

论文配图:Semantic Variability of Replies Across LLMs: Implications for Designing Conversation-Based Assessment
图 1 · 摘自论文原文
  • 用真实对话数据测试多模型回复语义一致性
  • 有上下文时模型间回复相似度下降30%以上
  • 适合教育评估系统设计者参考

本研究考察当底层大模型改变时,LLM生成的回复是否保持语义一致。基于真实协作对话消息,在有无前序聊天历史两种条件下,对比了不同LLM生成回复的语义相似性。结果表明,模型选择和对话上下文均会影响回复相似性及与人类回复的一致性。这说明仅靠提示工程和上下文无法保证跨模型回复的一致性,凸显在大模型快速迭代背景下,需构建能维持稳定、可比回复的基础设施与设计策略。

原文摘要 · Abstract (English)

This study examines whether LLM-generated replies remain semantically consistent when the underlying LLM changes. Using messages from real collaborative conversations, we compared the semantic similarity of generated replies across LLMs under two conditions: with and without preceding chat history. Results show that model choice and conversational context both affect response similarity and alignment with human replies. These findings indicate that prompting and conversational context alone may not be sufficient to preserve response consistency across LLMs, highlighting the need for infrastructure and design strategies that can maintain stable and comparable responses amid the rapid and continuous evolution of LLMs.

大模型评估对话一致性LLM设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。