Skip to content

stepoliste/Music-Source-Separation

Folders and files

NameName
Last commit message
Last commit date

Latest commit

ย 

History

3 Commits
ย 
ย 
ย 
ย 
ย 
ย 

Repository files navigation

๐ŸŽถ Music Source Separation using Deep U-Net Architectures

Selected Topics for Music and Acoustic Engineering โ€” 2024/2025
Stefano Polimeno, Riccardo Corร , Riccardo Moschen


๐Ÿ“Œ Overview

This project explores music source separation, the task of isolating individual instruments (vocals, drums, bass, other) from a mixture audio signal. We investigate both traditional signal processing approaches and deep learning models, with a focus on U-Net-based architectures enhanced with:

  • BiLSTM bottlenecks for temporal modeling
  • Attention-gated skip connections
  • Instance Normalization for improved training stability

Our models are trained and evaluated on the MUSDB18 dataset, the standard benchmark for music source separation.


โšก Key Contributions

  • โœ… Implementation of classical separation methods: Wiener filtering, Binary Masking, HPSS, NMF, Spectral Subtraction
  • โœ… Design of a BiLSTM-enhanced U-Net with attention and normalization strategies
  • โœ… Evaluation using standard separation metrics:
    • SDR (Signal-to-Distortion Ratio)
    • SIR (Signal-to-Interference Ratio)
    • SAR (Signal-to-Artifacts Ratio)
  • โœ… Comprehensive analysis of strengths, limitations, and trade-offs between traditional methods and deep neural models

๐Ÿ“‚ Dataset

We use the MUSDB18 dataset:

  • 150 full-length stereo tracks (~10 hours of audio)
  • 4 isolated stems: vocals, drums, bass, other
  • Provided at 44.1 kHz, split into training (100) and test (50) tracks

๐Ÿ—๏ธ Methodology

  1. Baselines: Spectral masking, HPSS, NMF, spectral subtraction
  2. Deep U-Net: Encoderโ€“decoder with skip connections
  3. Enhancements:
    • BiLSTM at bottleneck
    • Attention-gated skip connections
    • InstanceNorm + dropout for regularization
  4. Output: Multi-source mask prediction (vocals, drums, bass, other)

๐Ÿ“Š Results

  • SDR: 3.22 ยฑ 2.91 dB โ†’ moderate separation quality
  • SIR: 6.42 ยฑ 6.33 dB โ†’ reasonable suppression of interfering sources
  • SAR: 0.31 ยฑ 4.65 dB โ†’ artifacts introduced by recurrent layers

Key Insight:
BiLSTM enhances temporal modeling but introduces artifacts, highlighting the trade-off between temporal understanding and signal fidelity.


๐Ÿ”ฎ Future Work

  • Reduce processing artifacts via alternative temporal modeling strategies
  • Improve phase consistency in reconstruction
  • Explore genre-adaptive models and attention refinements

๐Ÿ“š References

  • Ronneberger et al. (2015) โ€” U-Net
  • Jansson et al. (2017) โ€” Singing voice separation with deep U-Nets
  • Rafii et al. (2017) โ€” MUSDB18 dataset
  • Dรฉfossez et al. (2019) โ€” Demucs
  • Driedger & Mรผller (2014) โ€” HPSS

About

Deep U-Net architectures for music source separation, with BiLSTM bottlenecks, attention-gated skip connections, and evaluation on the MUSDB18 dataset.

Resources

Stars

Watchers

Forks

Releases

Packages

Contributors

Languages