FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs

[Paper] [Code] [Modelscope:SenseVoice CosyVoice] [HuggingFace: SenseVoice CosyVoice]


Tongyi SpeechTeam

Alibaba Group

Abstract: This report introduces FunAudioLLM, a framework designed to enhance natural voice interactions between humans and large language models (LLMs). At its core are two innovative models: SenseVoice for high-precision multilingual speech recognition, emotion recognition, and audio event detection; and CosyVoice for natural speech generation with multi-language, timbre, and emotion control. SenseVoice delivers exceptionally low latency and supports over 50 languages, while CosyVoice excels in multi-lingual voice generation, zero-shot voice generation, cross-lingual voice cloning, and instruction-following capabilities. The models related to SenseVoice and CosyVoice have been open-sourced on Modelscope and Huggingface, along with the corresponding training, inference, and fine-tuning codes released on GitHub. By integrating these models with LLMs, FunAudioLLM enables applications such as speech translation, emotional voice chat, interactive podcasts, and expressive audiobook narration, thereby pushing the boundaries of voice interaction technology.

Contents

Speech-to-Speech Translation

By integrating SenseVoice, LLMs, and CosyVoice, we can effortlessly perform speech-to-speech translation (S2ST). Note that the original recordings are highlighted in bold.

ZH EN JP Yue KO

对,所以说你现在的话,这个账单的话,你既然说能处理,那你就想办法处理掉。

Yes, that's why I'm saying, regarding the bill you're currently discussing, if you say you can handle it, then find a way to take care of it.

そう、だから今あなたが言っていること、この請求書について、あなたが処理できると言ったのなら、何とかして処理してください。

对,所以话你而家讲嘅,呢张账单嘅话,你既然话得掂,噉你就要想办法搞掂佢。

맞아, 그래서 네가 지금 말하는 것, 이 계산서에 대해서, 네가 처리할 수 있다고 했다면, 그렇다면 방법을 찾아서 처리해야 해.

在那之后,完全收购那家公司。因此,保持管理层的一致性,利益与即将加入家族的资产保持一致,这就是我们有时不买下全部的原因。

And then later on, fully acquiring that company. So keeping management in line, interest in line with the asset that's coming into the family is a reason why sometimes we don't buy the whole thing.

その後、その会社を完全に買収する。だから、経営陣を一列に並べ、家族に入る資産との利益を一致させることが、私たちが全てを買わない理由の一つです。

之后,就完全收购咗嗰间公司。所以,保持管理层同兴趣与即将加入家族嘅资产保持一致,系我们有时唔买晒全部嘅原因。

그리고 나중에, 그 회사를 완전히 인수하게 됩니다. 그래서 경영진을 일치시키고, 가족에 들어오는 자산과의 이익을 일치시키는 것이 우리가 가끔 전부를 사지 않는 이유입니다.

只是雨滴有什么麻烦的?这还没有打雷呢!

Rainfall alone does not constitute a storm. Thunder is required.

雨なんて大したことありません。まだ雷も鳴っていないのですから。

只系雨滴有咩麻烦嘅?连打雷都未有呢!

고작 빗방울로 호들갑은, 아직 번개도 치지 않았는데.

早点回家吧,到家了如果可以的话给我回个信息。

Come home early, and if possible, let me know when you're back.

早く家に帰って、家に着いたら可能なら私にメッセージを返してください。

早啲翻屋企吧,到屋企可以嘅话俾我回个信息。

일찍 집에 돌아오세요. 집에 도착하면 가능하다면 저에게 연락해 주세요.

虽然这种时候已经错过了云先生的戏,但夜晚的璃月仍有许多好去处。

By this time of night, we will have already missed Yun's play. But LeeYew is full of interesting places to explore, even at night.

もう雲殿の芝居に間に合わないが、夜の璃月はまだまだいいところがたくさんあるぞ。

虽然呢个时候已经错过咗云先生嘅戏,但夜晚嘅璃月仲有好多好去处。

어디가서 눈을 피하지. 난 괜찮은데, 넌 감기 걸릴지도 모르니.

Emotional VoiceChat

By integrating SenseVoice, LLMs, and CosyVoice, we can develop an Emotional Voice Chat application. In the following example, both the User and Assistant content are synthesized by CosyVoice.

--
Audio

User: 早上好

Assistant: 生成风格: warm tone, slow speaking rate, and friendly emotion.
播报内容: 早上好呀,今天的精神状态如何呀?

User: <|HAPPY|>今天特别开心

Assistant: 生成风格: high pitch, fast speaking rate, and happy emotion.
播报内容: [breath]哎哟,听到你这么开心[breath]我也很高兴呢,发生什么好事啦?

User: 我们要发布一个新的语音大模型啦<|Laughter|>

Assistant: 生成风格: normal tone, medium speaking rate, with a hint of amusement.
播报内容: [laughter]太棒了,听说你们要发布新的语音大模型,真是个令人兴奋的消息啊!

User: 悲伤的说一句光头强台词吧

Assistant: 生成风格: low pitch, slow speaking rate, and sad emotion.
播报内容: "唉---看来这次又要空手而归了---"

User: 真棒

Assistant: 生成风格: positive tone, moderate speaking rate, and encouraging emotion.
播报内容: 那真是太棒啦,继续保持这种势头哦!