Unsupervised Deep Clustering for Source Separation: Direct Learning from Mixtures Using Spatial Information

Efthymios Tzinis, Shrikant Venkataramani, Paris Smaragdis

Research output: Chapter in Book/Report/Conference proceedingConference contribution

Abstract

We present a monophonic source separation system that is trained by only observing mixtures with no ground truth separation information. We use a deep clustering approach which trains on multichannel mixtures and learns to project spectrogram bins to source clusters that correlate with various spatial features. We show that using such a training process we can obtain separation performance that is as good as making use of ground truth separation information. Once trained, this system is capable of performing sound separation on monophonic inputs, despite having learned how to do so using multi-channel recordings.

Original languageEnglish (US)
Title of host publication2019 IEEE International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2019 - Proceedings
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages81-85
Number of pages5
ISBN (Electronic)9781479981311
DOIs
StatePublished - May 2019
Externally publishedYes
Event44th IEEE International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2019 - Brighton, United Kingdom
Duration: May 12 2019May 17 2019

Publication series

NameICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings
Volume2019-May
ISSN (Print)1520-6149

Conference

Conference44th IEEE International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2019
CountryUnited Kingdom
CityBrighton
Period5/12/195/17/19

Fingerprint

Source separation
Bins
Acoustic waves

Keywords

  • Deep clustering
  • source separation
  • unsupervised learning

ASJC Scopus subject areas

  • Software
  • Signal Processing
  • Electrical and Electronic Engineering

Cite this

Tzinis, E., Venkataramani, S., & Smaragdis, P. (2019). Unsupervised Deep Clustering for Source Separation: Direct Learning from Mixtures Using Spatial Information. In 2019 IEEE International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2019 - Proceedings (pp. 81-85). [8683201] (ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings; Vol. 2019-May). Institute of Electrical and Electronics Engineers Inc.. https://doi.org/10.1109/ICASSP.2019.8683201

Unsupervised Deep Clustering for Source Separation : Direct Learning from Mixtures Using Spatial Information. / Tzinis, Efthymios; Venkataramani, Shrikant; Smaragdis, Paris.

2019 IEEE International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2019 - Proceedings. Institute of Electrical and Electronics Engineers Inc., 2019. p. 81-85 8683201 (ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings; Vol. 2019-May).

Research output: Chapter in Book/Report/Conference proceedingConference contribution

Tzinis, E, Venkataramani, S & Smaragdis, P 2019, Unsupervised Deep Clustering for Source Separation: Direct Learning from Mixtures Using Spatial Information. in 2019 IEEE International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2019 - Proceedings., 8683201, ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings, vol. 2019-May, Institute of Electrical and Electronics Engineers Inc., pp. 81-85, 44th IEEE International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2019, Brighton, United Kingdom, 5/12/19. https://doi.org/10.1109/ICASSP.2019.8683201
Tzinis E, Venkataramani S, Smaragdis P. Unsupervised Deep Clustering for Source Separation: Direct Learning from Mixtures Using Spatial Information. In 2019 IEEE International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2019 - Proceedings. Institute of Electrical and Electronics Engineers Inc. 2019. p. 81-85. 8683201. (ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings). https://doi.org/10.1109/ICASSP.2019.8683201
Tzinis, Efthymios ; Venkataramani, Shrikant ; Smaragdis, Paris. / Unsupervised Deep Clustering for Source Separation : Direct Learning from Mixtures Using Spatial Information. 2019 IEEE International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2019 - Proceedings. Institute of Electrical and Electronics Engineers Inc., 2019. pp. 81-85 (ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings).
@inproceedings{624e4dff1a75457d9520b01f556c68ae,
title = "Unsupervised Deep Clustering for Source Separation: Direct Learning from Mixtures Using Spatial Information",
abstract = "We present a monophonic source separation system that is trained by only observing mixtures with no ground truth separation information. We use a deep clustering approach which trains on multichannel mixtures and learns to project spectrogram bins to source clusters that correlate with various spatial features. We show that using such a training process we can obtain separation performance that is as good as making use of ground truth separation information. Once trained, this system is capable of performing sound separation on monophonic inputs, despite having learned how to do so using multi-channel recordings.",
keywords = "Deep clustering, source separation, unsupervised learning",
author = "Efthymios Tzinis and Shrikant Venkataramani and Paris Smaragdis",
year = "2019",
month = "5",
doi = "10.1109/ICASSP.2019.8683201",
language = "English (US)",
series = "ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings",
publisher = "Institute of Electrical and Electronics Engineers Inc.",
pages = "81--85",
booktitle = "2019 IEEE International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2019 - Proceedings",
address = "United States",

}

TY - GEN

T1 - Unsupervised Deep Clustering for Source Separation

T2 - Direct Learning from Mixtures Using Spatial Information

AU - Tzinis, Efthymios

AU - Venkataramani, Shrikant

AU - Smaragdis, Paris

PY - 2019/5

Y1 - 2019/5

N2 - We present a monophonic source separation system that is trained by only observing mixtures with no ground truth separation information. We use a deep clustering approach which trains on multichannel mixtures and learns to project spectrogram bins to source clusters that correlate with various spatial features. We show that using such a training process we can obtain separation performance that is as good as making use of ground truth separation information. Once trained, this system is capable of performing sound separation on monophonic inputs, despite having learned how to do so using multi-channel recordings.

AB - We present a monophonic source separation system that is trained by only observing mixtures with no ground truth separation information. We use a deep clustering approach which trains on multichannel mixtures and learns to project spectrogram bins to source clusters that correlate with various spatial features. We show that using such a training process we can obtain separation performance that is as good as making use of ground truth separation information. Once trained, this system is capable of performing sound separation on monophonic inputs, despite having learned how to do so using multi-channel recordings.

KW - Deep clustering

KW - source separation

KW - unsupervised learning

UR - http://www.scopus.com/inward/record.url?scp=85068960173&partnerID=8YFLogxK

UR - http://www.scopus.com/inward/citedby.url?scp=85068960173&partnerID=8YFLogxK

U2 - 10.1109/ICASSP.2019.8683201

DO - 10.1109/ICASSP.2019.8683201

M3 - Conference contribution

AN - SCOPUS:85068960173

T3 - ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings

SP - 81

EP - 85

BT - 2019 IEEE International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2019 - Proceedings

PB - Institute of Electrical and Electronics Engineers Inc.

ER -