Selected Topics for Music and Acoustic Engineering โ 2024/2025
Stefano Polimeno, Riccardo Corร , Riccardo Moschen
This project explores music source separation, the task of isolating individual instruments (vocals, drums, bass, other) from a mixture audio signal. We investigate both traditional signal processing approaches and deep learning models, with a focus on U-Net-based architectures enhanced with:
- BiLSTM bottlenecks for temporal modeling
- Attention-gated skip connections
- Instance Normalization for improved training stability
Our models are trained and evaluated on the MUSDB18 dataset, the standard benchmark for music source separation.
- โ Implementation of classical separation methods: Wiener filtering, Binary Masking, HPSS, NMF, Spectral Subtraction
- โ Design of a BiLSTM-enhanced U-Net with attention and normalization strategies
- โ
Evaluation using standard separation metrics:
- SDR (Signal-to-Distortion Ratio)
- SIR (Signal-to-Interference Ratio)
- SAR (Signal-to-Artifacts Ratio)
- โ Comprehensive analysis of strengths, limitations, and trade-offs between traditional methods and deep neural models
We use the MUSDB18 dataset:
- 150 full-length stereo tracks (~10 hours of audio)
- 4 isolated stems: vocals, drums, bass, other
- Provided at 44.1 kHz, split into training (100) and test (50) tracks
- Baselines: Spectral masking, HPSS, NMF, spectral subtraction
- Deep U-Net: Encoderโdecoder with skip connections
- Enhancements:
- BiLSTM at bottleneck
- Attention-gated skip connections
- InstanceNorm + dropout for regularization
- Output: Multi-source mask prediction (vocals, drums, bass, other)
- SDR: 3.22 ยฑ 2.91 dB โ moderate separation quality
- SIR: 6.42 ยฑ 6.33 dB โ reasonable suppression of interfering sources
- SAR: 0.31 ยฑ 4.65 dB โ artifacts introduced by recurrent layers
Key Insight:
BiLSTM enhances temporal modeling but introduces artifacts, highlighting the trade-off between temporal understanding and signal fidelity.
- Reduce processing artifacts via alternative temporal modeling strategies
- Improve phase consistency in reconstruction
- Explore genre-adaptive models and attention refinements
- Ronneberger et al. (2015) โ U-Net
- Jansson et al. (2017) โ Singing voice separation with deep U-Nets
- Rafii et al. (2017) โ MUSDB18 dataset
- Dรฉfossez et al. (2019) โ Demucs
- Driedger & Mรผller (2014) โ HPSS