Speech Recognition and Synthesis
Speech recognition and synthesis research investigates how machines can accurately convert spoken language into text and generate natural-sounding speech from written input, drawing on statistical models and, increasingly, deep neural networks. Early systems relied on Hidden Markov Models to represent the sequential structure of speech, but modern approaches—including end-to-end architectures and sequence-to-sequence models—learn these representations directly from large amounts of audio data, dramatically improving accuracy across diverse speakers and acoustic conditions. Researchers are now pushing toward more robust speaker verification, better handling of noisy or low-resource languages, and tighter integration of acoustic and language models so that systems behave reliably outside controlled settings. Open challenges include making recognition fair across dialects and accents, reducing the data requirements for training high-quality models, and building synthesis systems whose output remains indistinguishable from human speech under careful scrutiny.
- Works
- 99,063
- Total citations
- 1,002,745
- Keywords
- Deep Neural NetworksAcoustic ModelingSpeaker VerificationConvolutional Neural NetworksEnd-to-End Speech RecognitionHidden Markov Models
Top papers in Speech Recognition and Synthesis
Ordered by total citation count.
- AI-Assisted Pipeline for Dynamic Generation of Trustworthy Health Supplement Content at Scale↗ 45,820OA
- Learning Phrase Representations using RNN Encoder–Decoder for Statistical Machine Translation↗ 24,756OA
- A tutorial on hidden Markov models and selected applications in speech recognition↗ 22,833
- Efficient Estimation of Word Representations in Vector Space↗ 18,162OA
- Conditional Random Fields: Probabilistic Models for Segmenting and Labeling Sequence Data↗ 13,013OA
- Efficient Estimation of Word Representations in Vector Space↗ 11,738OA
- Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling↗ 10,811OA
- Deep Neural Networks for Acoustic Modeling in Speech Recognition: The Shared Views of Four Research Groups↗ 10,360
- Bidirectional recurrent neural networks↗ 10,151
- Speech recognition with deep recurrent neural networks↗ 8,872
- Fundamentals of speech recognition↗ 7,704
- LSTM: A Search Space Odyssey↗ 6,925OA
Active researchers
Top authors in this area, ranked by h-index.