Evaluating and Interpreting Pooling Techniques in Spectrogram-Based Audio Analysis Using Diverse Metrics

Abstract

Audio analysis is a rapidly advancing field that spans various domains, including speech, music, and environmental sound data. Using spectrograms with Convolutional Neural Networks (CNNs) enables the visualization and extraction of critical audio features by combining time-frequency representations with deep learning. Pooling plays a crucial role in this process, as it reduces dimensionality while retaining essential information. However, existing evaluations of pooling methods primarily emphasize downstream task performance, such as classification accuracy, often overlooking their effectiveness in preserving critical signal features. To address this gap, we use 17 distinct metrics, categorized into four domains, to comprehensively assess various pooling operations. Furthermore, we explore the underex-amined relationship between specific pooling techniques and their impact on feature retention across diverse audio applications. Our analysis encompasses spectrograms from three audio domains (speech, music, and environmental sound), identifying their key characteristics, and grouping them accordingly. Using this setup, we evaluate the performance of 12 pooling methods across these applications. By investigating the features critical to each task and evaluating how well different pooling techniques preserve them, we give insights into their suitability for specific applications. This work aims to guide researchers in selecting the most appropriate pooling strategies for their applications, enabling more granular evaluations, improving explainability, and thereby advancing the precision and efficiency of audio analysis pipelines.

Details

Subject

Background noise;
Speech;
Music;
Performance evaluation;
Machine learning;
Spectrograms;
Artificial neural networks;
Accuracy;
Deep learning;
Computer science;
Fourier transforms;
Larynx;
Neural networks;
Signal processing;
Sound

Identifier / keyword

Audio data analysis; pooling; deep learning; dimensionality reduction; spectrograms

Title

Evaluating and Interpreting Pooling Techniques in Spectrogram-Based Audio Analysis Using Diverse Metrics

Author

PDF

Publication title

International Journal of Advanced Computer Science and Applications; West Yorkshire

Volume

Issue

Number of pages

Publication year

2025

Publication date

2025

Publisher

Science and Information (SAI) Organization Limited

Place of publication

West Yorkshire

Country of publication

United Kingdom

Publication subject

Computers, Sciences: Comprehensive Works

ISSN

2158107X

e-ISSN

21565570

Source type

Scholarly Journal

Language of publication

English

Document type

Journal Article

DOI

https://doi.org/10.14569/IJACSA.2025.0160795

ProQuest document ID

3240918339

Document URL

https://www.proquest.com/scholarly-journals/evaluating-interpreting-pooling-techniques/docview/3240918339/se-2?accountid=208611

© 2025. This work is licensed under http://creativecommons.org/licenses/by/4.0/ (the “License”). Notwithstanding the ProQuest Terms and Conditions, you may use this content in accordance with the terms of the License.

Last updated

2025-08-29

Database

ProQuest One Academic

Evaluating and Interpreting Pooling Techniques in Spectrogram-Based Audio Analysis Using Diverse Metrics

Content area

Abstract

Details