Point Cloud Audio Processing

Krishna Subramani, Paris Smaragdis

Research output: Chapter in Book/Report/Conference proceedingConference contribution

Abstract

Most audio processing pipelines involve transformations that act on fixed-dimensional input representations of audio. For example, when using the Short Time Fourier Transform (STFT) the DFT size specifies a fixed dimension for the input representation. As a consequence, most audio machine learning models are designed to process fixed-size vector inputs which often prohibits the repurposing of learned models on audio with different sampling rates or alternative representations. We note, however, that the intrinsic spectral information in the audio signal is invariant to the choice of the input representation or the sampling rate. Motivated by this, we introduce a novel way of processing audio signals by treating them as a collection of points in feature space, and we use point cloud machine learning models that give us invariance to the choice of representation parameters, such as DFT size or the sampling rate. Additionally, we observe that these methods result in smaller models, and allow us to significantly subsample the input representation with minimal effects to a trained model performance.

Original languageEnglish (US)
Title of host publication2021 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics, WASPAA 2021
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages31-35
Number of pages5
ISBN (Electronic)9781665448703
DOIs
StatePublished - 2021
Event2021 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics, WASPAA 2021 - New Paltz, United States
Duration: Oct 17 2021Oct 20 2021

Publication series

NameIEEE Workshop on Applications of Signal Processing to Audio and Acoustics
Volume2021-October
ISSN (Print)1931-1168
ISSN (Electronic)1947-1629

Conference

Conference2021 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics, WASPAA 2021
Country/TerritoryUnited States
CityNew Paltz
Period10/17/2110/20/21

Keywords

  • Point Clouds
  • Sample Rate Invariance
  • Transformers

ASJC Scopus subject areas

  • Electrical and Electronic Engineering
  • Computer Science Applications

Fingerprint

Dive into the research topics of 'Point Cloud Audio Processing'. Together they form a unique fingerprint.

Cite this