将冗长字幕翻译压缩为简洁表达,提升跨语言内容理解效率
SUMART: SUMmARizing Translation from Wordy to Concise Expression
- 用大模型自动压缩冗长字幕,生成简洁版翻译数据用于训练
- 在真实场景中实现更紧凑的外语言语字幕输出,提升可读性
- 适合快速获取演讲、会议、播客等长内容要点的用户
我们提出 SUMART,一种将冗长字幕翻译压缩为简洁表达的方法。该方法适用于跨语言对话(如字幕翻译)或观看外语音频配字幕的场景。当说话人表述冗长时,系统利用本地大语言模型对字幕进行压缩,并将压缩数据存入数据库用于微调。随后,使用原始ASR结果与压缩翻译结果的配对数据,微调翻译模型以生成更简洁的实用翻译。实际应用中,该模型可生成精炼字幕。此外,我们开发了基于增强现实空间的字幕翻译对话应用。初步定性调研表明,用户普遍认可其在快速获取信息场景中的价值,尤其适用于演讲、讲座、播客及会议问答等需要高效处理大量信息的场合。
原文摘要 · Abstract (English)
We propose SUMART, a method for summarizing and compressing the volume of verbose subtitle translations. SUMART is designed for understanding translated captions (e.g., interlingual conversations via subtitle translation or when watching movies in foreign language audio and translated captions). SUMART is intended for users who want a big-picture and fast understanding of the conversation, audio, video content, and speech in a foreign language. During the training data collection, when a speaker makes a verbose statement, SUMART employs a large language model on-site to compress the volume of subtitles. This compressed data is then stored in a database for fine-tuning purposes. Later, SUMART uses data pairs from those non-compressed ASR results and compressed translated results for fine-tuning the translation model to generate more concise translations for practical uses. In practical applications, SUMART utilizes this trained model to produce concise translation results. Furthermore, as a practical application, we developed an application that allows conversations using subtitle translation in augmented reality spaces. As a pilot study, we conducted qualitative surveys using a SUMART prototype and a survey on the summarization model for SUMART. We envision the most effective use case of this system is where users need to consume a lot of information quickly (e.g., Speech, lectures, podcasts, Q&A in conferences).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。