Listening demo · speech enhancement

Training-Free Combination of Discriminative and Generative Speech Enhancement via Posterior Sampling

Listening room

One tab per model pair, 10 clips each. Clips are selected trade-off cases: on every clip the discriminative model wins at least one metric and the generative model wins at least one other (PESQ / ESTOI / SI-SDR / WER sub+del against the clean reference) — the situation posterior sampling is built for. Speech is GigaSpeech (audiobook / podcast / YouTube), mixed with test-set noises at −5 dB. Playback keeps its position when you switch systems, so you can A/B the same instant of speech across models. Audio is 16 kHz, Opus-compressed.

0.0 / 0.0 s

space play/pause  ·  switch system in place  ·  previous / next clip