PRLab @ Nanjing University

Speech and Audio Intelligence Lab

SAIL

Pattern Recognition Laboratory
School of Intelligence Science and Technology
Nanjing University

以声音启航,以智能致远。Sail with Sound. Navigate with Intelligence.
Nanjing University Suzhou Campus, Suzhou, China

About SAIL

The Speech and Audio Intelligence Lab (SAIL) is part of PRLab (模式识别实验室), School of Intelligence Science and Technology, Nanjing University (Suzhou Campus).

Led by Dr. Shuai Wang, a Tenure-Track Associate Professor, the group develops intelligent technologies for speech, audio, and music, with an emphasis on impactful research and real-world applications.

We are actively recruiting.Openings for 2027 Fall Ph.D. and Master's students, as well as year-round research assistants.
View openings

Research Directions

Our research spans speaker-centric speech intelligence, speech and audio generation, speech language models, embodied auditory intelligence, and speech technologies for health.

Speech, Audio & Music Generation

Speech synthesis, voice conversion, controllable audio generation, and high-quality music generation using autoregressive and diffusion-based models.

Speech SynthesisVoice ConversionFoley Audio GenerationSong Generation

Speech Language Models

Understanding and evaluating end-to-end speech language models, with an emphasis on speaker awareness, multi-speaker interaction, full-duplex dialogue, and speech agents.

Speech AgentsFull-Duplex Dialogue Systems

Embodied Auditory Intelligence

Auditory perception for embodied agents, spanning spatial listening, audio-visual scene understanding, and on-device audio processing in distributed environments.

Spatial Audio ProcessingAudio-Visual UnderstandingDistributed / Edge Audio Processing

Speech Processing for Health

Speech, audio, and neural-signal modeling for health assessment, clinical speech restoration, neurological rehabilitation, and assistive communication.

EEG-Audio LearningClinical Speech RestorationDepression Detection

Recent Highlights

2026.08
2027 Fall recruitment is open for Ph.D. students, Master's students, and research assistants.
2025.10
Delivered the tutorial Deep Speaker Representation Learning at NCMMSC 2025 and APSIPA 2025.
2025.08
Presented invited talks on rethinking speaker modeling across speech applications and the Real-T dataset at Interspeech 2025.
2024
Received the Best Paper Award and Best Student Paper Award at ISCSLP 2024.
2019
Ranked first in both tracks of VoxSRC 2019 and all four tracks of DIHARD 2019.

Join Us

We welcome students who are interested in this field and hope to make a real impact in speech and audio intelligence. / 欢迎对本方向感兴趣、希望在这一领域大展身手的同学联系申请。

Current openings include 2027 Fall Ph.D. and Master's positions, undergraduate internships, and year-round Research Assistant positions. Research assistants may work from Nanjing University Suzhou Campus, Shenzhen Loop Area Institute, CUHK-Shenzhen, or remotely.

Graduate Students

  • 2027 Fall Ph.D. and Master's intake
  • Computer science, electronic engineering, mathematics, or related backgrounds
  • Strong curiosity and self-motivation
  • Good English reading and writing

Research Assistants

  • Year-round openings for undergraduate and graduate students
  • Hands-on frontier research projects
  • Flexible on-site or remote arrangements
  • Joint supervision opportunities with Prof. Haizhou Li

What We Value

  • Solid mathematical foundations
  • Python, C++, and practical deep-learning skills
  • Ability to turn questions into experiments
  • Initiative, rigor, and collaboration

How to Apply / 申请方式

Email shuaiwang@nju.edu.cn with your CV, transcripts, and a short statement of research interests.

Suggested subject: Application Type – Name – University – Year. NJU students are also welcome to visit Room 536, West Wing, Nanyong Building.