arXiv:2607.22658cs.AIcs.SD2026-07中稿 · Interspeech 2026被引 1

构建对话中人际立场评估基准,测试语音大模型的判断能力

StanceBench: A Benchmark for Audio LLM-Based Interpersonal Stance Evaluation from Speech

论文配图:StanceBench: A Benchmark for Audio LLM-Based Interpersonal Stance Evaluation from Speech
图 1 · 摘自论文原文
  • 用角色提示极点定义9类人际立场维度
  • 发现共情与礼貌最易识别,诚实最难且受提示顺序影响大
  • 适合研究对话智能、人机交互与语音大模型评估的学者

语音对话模型日益依赖语调和互动细节传递社交意图,但相关评估基准仍有限。我们提出StanceBench,一个用于测量对话中人际立场并评估语音大模型作为自动评判者的能力的基准。基于Seamless Interaction语料库,StanceBench(1)通过角色提示极点定义9种立场维度,(2)统一单说话人与互动式评估标准,(3)报告了大模型作为评判者的鲁棒性、偏见及立场推断表现。在各类立场中,共情与礼貌最容易被识别;温暖与自信中等可分,存在正向偏差/不对称;诚实最难,提示顺序影响显著,需跨轮次证据支持。注意力可区分但与人类判断弱相关。互动立场更依赖上下文,存在阈值差距与高方差,尤其在冲突调节方面。

原文摘要 · Abstract (English)

Speech-to-speech dialogue models increasingly depend on prosody and interactional nuance to convey social intent, yet benchmarks for these cues remain limited. We introduce StanceBench, a benchmark for measuring interpersonal stance in conversational speech and evaluating audio-capable LLMs as automated judges. Using the Seamless Interaction corpus, StanceBench (1) specifies 9 stance dimensions via role-prompt poles, (2) standardizes single-speaker and interaction-based evaluations, and (3) reports LLM-as-a-judge robustness, bias, and stance inference. Across evaluated stances, empathy and politeness are the easiest. Warmth and assertiveness are moderately separable with positivity skew/asymmetry. Honesty is the hardest and shows high prompt order bias, consistent with needing cross-turn evidence. Attentiveness is separable but aligns weakly with humans. Interaction stances are more context-sensitive, with threshold gaps and high variance, especially conflict regulation.

语音大模型立场识别人机交互评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。