A Review of Dynamic Time Warping in Speech Technology for Inclusive Digital Development

[featured_image]
  • Version
  • Download 0
  • File Size 366.98 KB
  • File Count 1
  • Create Date September 22, 2026
  • Last Updated September 22, 2026

A Review of Dynamic Time Warping in Speech Technology for Inclusive Digital Development

ABSTRACT

Dynamic Time Warping (DTW) remains an important technique in the temporal alignment of speech signals and continues to play an important role in the application of Automatic Speech Recognition (ASR), pronunciation evaluation, and other speech technologies. This comprehensive review provides an extensive overview of the contemporary developments in DTW, ranging from the classical, Multiview, and optimization techniques, and evaluates their application in speech recognition, speech disorder recognition, and sign language retrieval and would enhance increased access to education, healthcare, and other public services among the underserved populations. The comparison with advanced recognition techniques, namely Wav2Vec 2.0, Whisper, Conformer Transducer, DeepSpeech, Kaldi, and Recurrent Neural Network Transducer (RNN-T), It shows DTW's strength in low resource and small vocabulary speech recognition. The main advantages of DTW are the reduced data requirements, the low computational complexity, and the high interpretability facilitated by the alignment paths. The limitations of the datasets and the potential offered by the hybrid DTW and machine learning techniques are also discussed.

Keywords: Dynamic Time Warping (DTW), Automatic Speech Recognition (ASR), Pronunciation Error Detection, Assistive Speech Systems, Mel-Frequency Cepstral Coefficients (MFCCs)

SHARE