arXiv:2506.22679cs.CL2025-06中稿 · Interspeech 2025被引 1

用大模型分析太空任务中团队对话的细微行为,效果因模型类型而异。

Assessing the feasibility of Large Language Models for detecting micro-behaviors in team interactions during space missions

  • 对比了零样本、微调等五种方法检测对话中的细微行为
  • 解码器类模型Llama-3.1在二分类上达68%宏F1分数
  • 适合高风险场景下仅基于文本的团队沟通分析

我们探索大语言模型(LLMs)在模拟太空任务中通过对话转录文本检测团队互动中的细微行为的可行性。研究采用编码器类序列分类模型(如RoBERTa、DistilBERT)进行零样本分类、微调及改写增强微调,并使用解码器类因果语言建模模型进行少样本文本生成,以预测每个对话片段对应的微行为。结果表明,编码器类模型即使经过加权微调,仍难以识别少数类微行为,尤其是抑制性言语。相反,指令微调后的解码器模型Llama-3.1表现更优,最佳模型在三分类任务中实现44%宏F1,在二分类任务中达到68%。该研究对开发面向高风险环境中团队通信动态分析的语音技术具有启示意义,尤其适用于仅能获取文本数据的场景。

原文摘要 · Abstract (English)

We explore the feasibility of large language models (LLMs) in detecting subtle expressions of micro-behaviors in team conversations using transcripts collected during simulated space missions. Specifically, we examine zero-shot classification, fine-tuning, and paraphrase-augmented fine-tuning with encoder-only sequence classification LLMs, as well as few-shot text generation with decoder-only causal language modeling LLMs, to predict the micro-behavior associated with each conversational turn (i.e., dialogue). Our findings indicate that encoder-only LLMs, such as RoBERTa and DistilBERT, struggled to detect underrepresented micro-behaviors, particularly discouraging speech, even with weighted fine-tuning. In contrast, the instruction fine-tuned version of Llama-3.1, a decoder-only LLM, demonstrated superior performance, with the best models achieving macro F1-scores of 44% for 3-way classification and 68% for binary classification. These results have implications for the development of speech technologies aimed at analyzing team communication dynamics and enhancing training interventions in high-stakes environments such as space missions, particularly in scenarios where text is the only accessible data.

大模型行为检测团队沟通太空任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。