A Comprehensive Review of Speech Signal Processing, Voice Activity Detection, and Automatic Speech Recognition: Techniques, Challenges, and Future Directions
DOI:
https://doi.org/10.18486/ijcsnt/14.3.017Keywords:
Speech Signal Processing, Voice Activity De-tection, Automatic Speech Recognition, MFCC, LPC, Hidden Markov Models, Deep Learning, LSTM, Transformer, End-to-End ASR, Noise Robustness, Feature Extraction, Self-Supervised LearningAbstract
Speech signal processing has become one of the most critical enablers of modern human–computer interaction, under-pinning technologies ranging from voice assistants and telephony to clinical diagnostics and security systems. This paper presents a comprehensive and systematic review of the field, spanning three tightly coupled domains: (i) fundamental speech signal analysis and preprocessing, (ii) voice activity detection (VAD), and (iii) automatic speech recognition (ASR). We begin by examining the physical and mathematical foundations of speech production and perception, followed by a detailed treatment of classical and contemporary feature extraction methods, including Mel-Frequency Cepstral Coefficients (MFCC), Linear Predictive Coding (LPC), Perceptual Linear Prediction (PLP), filter-bank features, and learned representations produced by deep neural networks. For VAD, we critically compare energy-based, statisti-cal, model-driven, and end-to-end deep learning approaches with respect to accuracy, latency, and robustness under diverse noise conditions. In the ASR domain, we trace the evolution from hid-den Markov model (HMM) systems through hybrid HMM–deep-neural-network (DNN) architectures to fully neural sequence-to-sequence models, including connectionist temporal classifica-tion (CTC) and attention-based encoder–decoder frameworks. Benchmark performance metrics across standardised datasets (TIMIT, LibriSpeech, CHiME, VoxCeleb) are consolidated and discussed. We further identify persistent open challenges—environmental noise, speaker variability, low-resource languages, and real-time edge deployment—and outline promising research directions, including self-supervised pre-training, multimodal fusion, federated learning, and neuromorphic processing. The survey is intended to serve as a unified reference for researchers and practitioners working at the intersection of signal processing, machine learning, and human–computer interaction.
References
H. Dudley, “Remaking Speech,” Journal of the Acoustical Society of America, vol. 11, no. 2, pp. 169–177, 1939. DOI: https://doi.org/10.1121/1.1916020
K. H. Davis, R. Biddulph, and S. Balashek, “Automatic Recognition of Spoken Digits,” Journal of the Acoustical Society of America, vol. 24, no. 6, pp. 637–642, 1952. DOI: https://doi.org/10.1121/1.1906946
Grand View Research, “Voice and Speech Recognition Market Size Report,” 2023. [Online]. Available: https://www.grandviewresearch.com.
G. Fant, Acoustic Theory of Speech Production. The Hague, Netherlands: Mouton, 1960.
D. O’Shaughnessy, Speech Communication: Human and Machine. Reading, MA, USA: Addison-Wesley, 1987.
F. J. Harris, “On the Use of Windows for Harmonic Analysis with the Discrete Fourier Transform,” Proceedings of the IEEE, vol. 66, no. 1, pp. 51–83, Jan. 1978. DOI: https://doi.org/10.1109/PROC.1978.10837
L. R. Rabiner and R. W. Schafer, Digital Processing of Speech Signals. Englewood Cliffs, NJ, USA: Prentice-Hall, 1978.
S. Boll, “Suppression of Acoustic Noise in Speech Using Spectral Subtraction,” IEEE Transactions on Acoustics, Speech, and Signal Processing, vol. 27, no. 2, pp. 113–120, Apr. 1979. DOI: https://doi.org/10.1109/TASSP.1979.1163209
B. S. Atal and S. L. Hanauer, “Speech Analysis and Synthesis by Linear Prediction of the Speech Wave,” Journal of the Acoustical Society of America, vol. 50, no. 2, pp. 637–655, 1971. DOI: https://doi.org/10.1121/1.1912679
S. B. Davis and P. Mermelstein, “Comparison of Parametric Representations for Monosyllabic Word Recognition in Continuously Spoken Sentences,” IEEE Transactions on Acoustics, Speech, and Signal Processing, vol. 28, no. 4, pp. 357–366, Aug. 1980. DOI: https://doi.org/10.1109/TASSP.1980.1163420
H. Hermansky, “Perceptual Linear Predictive (PLP) Analysis of Speech,” Journal of the Acoustical Society of America, vol. 87, no. 4, pp. 1738–1752, Apr. 1990. DOI: https://doi.org/10.1121/1.399423
O. Abdel-Hamid, A. Mohamed, H. Jiang, L. Deng, G. Penn, and D. Yu, “Convolutional Neural Networks for Speech Recognition,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 22, no. 10, pp. 1533–1545, Oct. 2014. DOI: https://doi.org/10.1109/TASLP.2014.2339736
S. Young, et al., The HTK Book, Version 3.4. Cambridge, U.K.: Cambridge University Engineering Department, 2006.
L. R. Rabiner and M. R. Sambur, “An Algorithm for Determining the Endpoints of Isolated Utterances,” Bell System Technical Journal, vol. 54, no. 2, pp. 297–315, Feb. 1975. DOI: https://doi.org/10.1002/j.1538-7305.1975.tb02840.x
J. Sohn, N. S. Kim, and W. Sung, “A Statistical Model-Based Voice Activity Detection,” IEEE Signal Processing Letters, vol. 6, no. 1, pp. 1–3, Jan. 1999. DOI: https://doi.org/10.1109/97.736233
D. A. Reynolds and R. C. Rose, “Robust Text-Independent Speaker Identification Using Gaussian Mixture Speaker Models,” IEEE Transactions on Speech and Audio Processing, vol. 3, no. 1, pp. 72–83, Jan. 1995. DOI: https://doi.org/10.1109/89.365379
ITU-T, “Recommendation G.729 Annex B: A Silence Compression Scheme for G.729 Optimized for Terminals Conforming to Recommendation V.70,” 1996.
I. Tashev, S. Mirsamadi, and A. Acero, “Voice Activity Detector Using Neural Network,” Proceedings of the IEEE International Workshop on Machine Learning for Signal Processing (MLSP), Santander, Spain, pp. 1–6, 2012.
X. Zhang and D. Wang, “Boosting Contextual Information for Deep Neural Network Based Voice Activity Detection,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 21, no. 11, pp. 2237–2240, Nov. 2013.
N. Ryant, M. Liberman, and J. Yuan, “Speech Activity Detection on YouTube Using Deep Neural Networks,” Proceedings of Interspeech, Lyon, France, pp. 728–731, 2013. DOI: https://doi.org/10.21437/Interspeech.2013-203
F. Eyben, F. Weninger, S. Squartini, and B. Schuller, “Real-Life Voice Activity Detection with LSTM Recurrent Neural Networks and an Application to Hollywood Movies,” Proceedings of ICASSP, Vancouver, Canada, pp. 483–487, 2013. DOI: https://doi.org/10.1109/ICASSP.2013.6637694
T. Hughes and K. Mierle, “Recurrent Neural Networks for Voice Activity Detection,” Proceedings of ICASSP, Vancouver, Canada, pp. 7378–7382, 2013. DOI: https://doi.org/10.1109/ICASSP.2013.6639096
Silero Team, “Silero VAD: Pre-Trained Enterprise-Grade Voice Activity Detector,” GitHub, 2021. [Online]. Available: https://github.com/snakers4/silero-vad.
H. Bredin, et al., “Pyannote.audio: Neural Building Blocks for Speaker Diarization,” Proceedings of ICASSP, Barcelona, Spain, pp. 7124–7128, 2020. DOI: https://doi.org/10.1109/ICASSP40776.2020.9052974
O. Ghahabi, J. Hernando, and J. A. Trujillo, “Unsupervised Noise Robust Voice Activity Detection for Meeting Recordings,” Computer Speech & Language, vol. 47, pp. 1–14, Jan. 2018. DOI: https://doi.org/10.1016/j.csl.2017.06.007
G. E. Hinton, et al., “Deep Neural Networks for Acoustic Modeling in Speech Recognition,” IEEE Signal Processing Magazine, vol. 29, no. 6, pp. 82–97, Nov. 2012. DOI: https://doi.org/10.1109/MSP.2012.2205597
D. Povey, et al., “The Kaldi Speech Recognition Toolkit,” Proceedings of IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), Hawaii, USA, 2011.
A. Graves, A. Mohamed, and G. Hinton, “Deep Recurrent Neural Networks for Acoustic Modelling,” Proceedings of ICASSP, Vancouver, Canada, pp. 6645–6649, 2013. DOI: https://doi.org/10.1109/ICASSP.2013.6638947
A. Hannun, et al., “Deep Speech: Scaling Up End-to-End Speech Recognition,” arXiv preprint arXiv:1412.5567, 2014.
W. Chan, N. Jaitly, Q. V. Le, and O. Vinyals, “Listen, Attend and Spell: A Neural Network for Large Vocabulary Conversational Speech Recognition,” Proceedings of ICASSP, Shanghai, China, pp. 4960–4964, 2016. DOI: https://doi.org/10.1109/ICASSP.2016.7472621
A. Vaswani, et al., “Attention Is All You Need,” Proceedings of NeurIPS, vol. 30, pp. 5998–6008, 2017.
L. Dong, S. Xu, and B. Xu, “Speech-Transformer: A No-Recurrence Sequence-to-Sequence Model for Speech Recognition,” Proceedings of ICASSP, Calgary, Canada, pp. 5884–5888, 2018. DOI: https://doi.org/10.1109/ICASSP.2018.8462506
A. Gulati, et al., “Conformer: Convolution-Augmented Transformer for Speech Recognition,” Proceedings of Interspeech, Shanghai, China, pp. 5036–5040, 2020. DOI: https://doi.org/10.21437/Interspeech.2020-3015
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations,” Proceedings of NeurIPS, vol. 33, pp. 12449–12460, 2020.
W.-N. Hsu, B. Bolte, Y.-H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, “HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 29, pp. 3451–3460, 2021. DOI: https://doi.org/10.1109/TASLP.2021.3122291
A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust Speech Recognition via Large-Scale Weak Supervision,” Proceedings of the International Conference on Machine Learning (ICML), pp. 28492–28518, 2023.
C. Gulcehre, et al., “On Using Monolingual Corpora in Neural Machine Translation,” arXiv preprint arXiv:1503.03535, 2015.
Y. Zhang, et al., “Very Deep Convolutional Networks for End-to-End Speech Recognition,” Proceedings of ICASSP, New Orleans, USA, pp. 4845–4849, 2017. DOI: https://doi.org/10.1109/ICASSP.2017.7953077
X. Hao, X. Su, R. Horaud, and X. Li, “FullSubNet: A Full-Band and Sub-Band Fusion Model for Real-Time Single-Channel Speech Enhancement,” Proceedings of ICASSP, Toronto, Canada, pp. 6633–6637, 2021. DOI: https://doi.org/10.1109/ICASSP39728.2021.9414177
S. Pascual, A. Bonafonte, and J. Serra, “SEGAN: Speech Enhancement Generative Adversarial Network,” Proceedings of Interspeech, Stockholm, Sweden, pp. 3642–3646, 2017. DOI: https://doi.org/10.21437/Interspeech.2017-1428
J.-M. Valin, “A Hybrid DSP/Deep Learning Approach to Real-Time Full-Band Speech Enhancement,” Proceedings of MMSP, Vancouver, Canada, pp. 1–5, 2018. DOI: https://doi.org/10.1109/MMSP.2018.8547084
N. Carlini and D. Wagner, “Audio Adversarial Examples: Targeted Attacks on Speech-to-Text,” Proceedings of IEEE Security & Privacy Workshops, San Francisco, USA, pp. 1–7, 2018. DOI: https://doi.org/10.1109/SPW.2018.00009
H. B. McMahan, E. Moore, D. Ramage, S. Hampson, and B. A. y Arcas, “Communication-Efficient Learning of Deep Networks from Decentralized Data,” Proceedings of AISTATS, Fort Lauderdale, USA, pp. 1273–1282, 2017.
P. K. Rubenstein, et al., “AudioPaLM: A Large Language Model That Can Speak and Listen,” arXiv preprint arXiv:2306.12925, 2023.
C. Tang, et al., “SALMONN: Towards Generic Hearing Abilities for Large Language Models,” Proceedings of the International Conference on Learning Representations (ICLR), Vienna, Austria, 2024.
Downloads
Published
Issue
Section
License
Copyright (c) 2025 Sapna Rajput, Stuti Pandey, Tanushka Mishra

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.