Listening demo · speech enhancement
One tab per model pair, 10 clips each. Clips are selected trade-off cases: on every clip the discriminative model wins at least one metric and the generative model wins at least one other (PESQ / ESTOI / SI-SDR / WER sub+del against the clean reference) — the situation posterior sampling is built for. Speech is GigaSpeech (audiobook / podcast / YouTube), mixed with test-set noises at −5 dB. Playback keeps its position when you switch systems, so you can A/B the same instant of speech across models. Audio is 16 kHz, Opus-compressed.
space play/pause · ↑↓ switch system in place · ←→ previous / next clip