Research Outputs

Publications & Projects

Representative open-source systems, datasets, and invited talks. The complete and most current publication record is maintained on Google Scholar.

Complete Publication List

More than 60 papers in leading speech conferences and journals. Visit Google Scholar for the latest papers and citation information.

Google Scholar

Open-Source Projects & Benchmarks

Toolkits and benchmarks spanning speaker representation, target speaker extraction, music generation, and speech language model evaluation.

WeSpeaker

A comprehensive speaker embedding learning toolkit supporting industrial-scale training, research, and production deployment.

View on GitHub

WeSep

The first open-source target speaker extraction toolkit, designed for scalable and generalizable personalized speech extraction.

DiffRhythm

Fast diffusion-based rhythmic music generation with an emphasis on efficient, high-quality synthesis.

View on GitHub

SongBloom

Autoregressive diffusion-based music generation for high-quality, high-fidelity, and diverse song synthesis.

View on GitHub

Real-T

A real-world, conversation-centric benchmark for target speaker extraction in natural multi-speaker interactions.

Project website

MSU-Bench

A multi-tier, multi-speaker, multilingual, multi-scenario, and multi-task benchmark for large speech language models.

Project website

Invited Talks & Tutorials

Honors & Awards

  • 2024 · Best Paper Award, ISCSLP 2024
  • 2024 · Best Student Paper Award, ISCSLP 2024
  • 2019 · Ranked 1st in both tracks, VoxSRC 2019
  • 2019 · Ranked 1st in all four tracks, DIHARD 2019
  • 2018 · IEEE Ganesh N. Ramaswamy Memorial Student Grant Award