arXiv:2606.25552cs.CLcs.AI2026-06

解决语音理解中多意图预测不一致问题,提升准确性。

SFL-MTSC: Leveraging Semantic Frame-Level Multi-Task Self-Consistency for Robust Multi-Intent Spoken Language Understanding

  • 在语义框架层进行多任务自一致性校验,分解并重组预测结果。
  • 零样本测试下槽位F1与整体准确率显著提升,意图准确率稳定。
  • 适合需要高鲁棒性的多意图语音理解场景,如智能助手。

基于提示的语音语言理解(SLU)使用大语言模型时,常因解码随机性导致多意图场景下的意图-槽位结构不一致。为此,我们提出语义框架级多任务自一致性(SFL-MTSC)框架,在语义框架层面操作。不同于输出层多数投票,SFL-MTSC将预测分解为特定意图的框架,通过领域-意图分组和槽位层级聚类,并采用路径支持评分评估聚类可靠性。可靠框架被保留并重新整合形成最终预测。在MAC-SLU基准数据集的零样本实验中,相比单路径推理,槽位F1和整体准确率均有提升,而意图准确率在大多数设置下保持稳定。

原文摘要 · Abstract (English)

Prompt-based spoken language understanding (SLU) with large language models (LLMs) often suffers from inconsistent intent--slot structures due to decoding stochasticity, particularly in multi-intent scenarios. In view of this, we propose Semantic Frame-Level Multi-Task Self-Consistency (SFL-MTSC), a novel structured aggregation framework operating at the semantic frame level. Instead of output-level majority voting, SFL-MTSC decomposes predictions into intent-specific frames, applies domain--intent grouping and slot-level clustering, and evaluates cluster reliability using path support scoring. Reliable frames are retained and re-integrated to form the final prediction. Zero-shot experiments on the MAC-SLU benchmark dataset show improved slot F1 and overall accuracy over single-path inference, while intent accuracy remains largely stable across most settings.

语音理解多意图自一致性大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。