Physical SciencesComputer ScienceArtificial Intelligence

Speech Recognition and Synthesis

Speech recognition and synthesis research investigates how machines can accurately convert spoken language into text and generate natural-sounding speech from written input, drawing on statistical models and, increasingly, deep neural networks. Early systems relied on Hidden Markov Models to represent the sequential structure of speech, but modern approaches—including end-to-end architectures and sequence-to-sequence models—learn these representations directly from large amounts of audio data, dramatically improving accuracy across diverse speakers and acoustic conditions. Researchers are now pushing toward more robust speaker verification, better handling of noisy or low-resource languages, and tighter integration of acoustic and language models so that systems behave reliably outside controlled settings. Open challenges include making recognition fair across dialects and accents, reducing the data requirements for training high-quality models, and building synthesis systems whose output remains indistinguishable from human speech under careful scrutiny.

Works
99,063
Total citations
1,002,745
Keywords
Deep Neural NetworksAcoustic ModelingSpeaker VerificationConvolutional Neural NetworksEnd-to-End Speech RecognitionHidden Markov Models

Top papers in Speech Recognition and Synthesis

Ordered by total citation count.

Active researchers

Top authors in this area, ranked by h-index.

Related topics