Voicebox：大规模多语言通用语音合成模型

邀
请
朋
友
一
起
学

主讲人：徐煒甯 | Meta研究科学家、MIT博士

开课时间

2023.09.20 20:00
课程时长

65分钟
学习人数

5851人次学习

立即学习

添加客服获取课件

立即学习

Voicebox：大规模多语言通用语音合成模型

Large-scale generative models such as GPT and DALL-E have revolutionized natural language processing and computer vision research. These models not only generate high fidelity text or image outputs, but are also generalists which can solve tasks not explicitly taught. In contrast, speech generative models are still primitive in terms of scale and task generalization.

In this talk, I will present Voicebox, the most versatile text-guided generative model for speech at scale. Voicebox is a non-autoregressive flow-matching model trained to infill speech, given audio context and text, trained on over 50K hours of speech that are neither filtered nor enhanced. Similar to GPT, Voicebox can perform many different tasks through in-context learning. Voicebox can be used for mono or cross-lingual zero-shot text-to-speech synthesis, noise removal, content editing, style conversion, and diverse sample generation. In particular, Voicebox outperforms the state-of-the-art zero-shot TTS model VALL-E on both intelligibility and audio similarity while being up to 20 times faster.