WeSpeaker
A comprehensive speaker embedding learning toolkit supporting industrial-scale training, research, and production deployment.
View on GitHubRepresentative open-source systems, datasets, and invited talks. The complete and most current publication record is maintained on Google Scholar.
More than 60 papers in leading speech conferences and journals. Visit Google Scholar for the latest papers and citation information.
Toolkits and benchmarks spanning speaker representation, target speaker extraction, music generation, and speech language model evaluation.
A comprehensive speaker embedding learning toolkit supporting industrial-scale training, research, and production deployment.
View on GitHubThe first open-source target speaker extraction toolkit, designed for scalable and generalizable personalized speech extraction.
Fast diffusion-based rhythmic music generation with an emphasis on efficient, high-quality synthesis.
View on GitHubAutoregressive diffusion-based music generation for high-quality, high-fidelity, and diverse song synthesis.
View on GitHubA real-world, conversation-centric benchmark for target speaker extraction in natural multi-speaker interactions.
Project websiteA multi-tier, multi-speaker, multilingual, multi-scenario, and multi-task benchmark for large speech language models.
Project website