Abstract
Dysarthria, a motor speech disorder, impairs the muscles involved in speech production, leading to challenges in articulation, pronunciation, and overall communication. This results in slow, slurred speech that is difficult to understand. Augmentative and Alternative Communication (AAC) aids integrated with speech recognition technology offer a promising solution for individuals with dysarthria. However, Automatic Speech Recognition (ASR) systems trained on typical speech data often struggle to recognize dysarthric speech due to its unique speech patterns and limited training data. To address these challenges, a hybrid Transformer-CTC model has been proposed for improving ASR performance on dysarthric speech. The Transformer architecture employs a self-attention mechanism that models complex dependencies between speech features, enabling it to identify and emphasize important patterns even when training data is limited. This ability is particularly crucial for dysarthric speech, where speech signals often exhibit high variability. On the other hand, Connectionist Temporal Classification (CTC) acts as an effective transcription layer. It aligns speech features with character sequences without requiring precise input-output alignment, making it well-suited for handling the inconsistencies and distortions present in dysarthric speech. The integration of these components creates a powerful architecture capable of learning nuanced speech patterns and delivering accurate transcriptions for dysarthric speech. The model was trained using the UA speech corpus, containing 13 hours of speech from 15 speakers with varying dysarthria levels. The proposed hybrid system achieves an impressive Word Recognition Accuracy (WRA) of 89%, demonstrating its effectiveness in accurately transcribing dysarthric speech. This innovative approach significantly advances the development of ASR technologies tailored to diverse and variable speech patterns, ultimately enhancing communication for individuals with speech disorders.
| Original language | American English |
|---|---|
| Pages (from-to) | 82479-82491 |
| Number of pages | 13 |
| Journal | IEEE Access |
| Volume | 13 |
| DOIs | |
| State | Published - May 8 2025 |
ASJC Scopus Subject Areas
- General Computer Science
- General Materials Science
- General Engineering
Keywords
- CTC
- Dysarthria
- Transformer
- assistive technology
- deep learning
- speech recognition
Cite this
- APA
- Standard
- Harvard
- Vancouver
- Author
- BIBTEX
- RIS