π§ Audio Samples β Clean | Noise | Noisy
About This Tool
This application is inspired by NVIDIA NeMo Speech Data Explorer, a Dash-based tool for interactive exploration and analysis of ASR datasets.
What this tool shows
- π Dataset-level WER/CER comparison across 22 Indian languages
- π Audio playback: clean speech, isolated noise, and noisy mix
- π WER degradation ranking and per-sample impact breakdown
- ποΈ Sortable / filterable evaluation table
- π Correlation between clean performance and noise robustness
Model & setup
- Model: ai4bharat/indic-conformer-600m-multilingual
- Dataset: IndicVoices (valid split, 22 languages)
- Noise: CAIMAN-ASR-BackgroundNoise, normalized to RMS = 0.1
- SNR: 5 dB (active speech RMS-based mixing)
- Decoding: CTC
How noise is isolated
The βNoise (isolated)β track is computed as noise = noisy β clean. Because the
mixing is linear addition (noisy = speech + noise_scaled), subtracting the clean
audio recovers the exact noise component. It is peak-normalised to 80% for easier listening.
Static build: metrics are computed over the full evaluation set (77,566 utterances), while audio playback covers the first 10 samples per language. Hindi has no clean audio in the source dataset, so its clean and noise tracks are unavailable.