Speech recognition is the computational process of converting spoken language into written text. Usually called automatic speech recognition (ASR), it connects audio processing with machine learning and natural language processing. Its central task is determining what was said, rather than identifying the speaker or interpreting the speaker’s intention. Recognition therefore provides a component of a spoken-language interface, not necessarily a complete language-understanding system. (web.stanford.edu)
Development and scope
Early speech recognizers operated with restricted vocabularies and controlled speaking conditions. IBM introduced its experimental Shoebox device in 1961 and exhibited it at the 1962 Seattle World’s Fair. It recognized ten spoken digits and six command words, allowing users to control an adding machine. Later research expanded from isolated words to continuous speech, in which users need not pause between words. Statistical approaches became central to this development, followed by increasingly integrated neural systems. (ibm.com)
Recognition tasks differ substantially in difficulty. Read speech, telephone conversations, spontaneous meetings, and distant-microphone recordings have different acoustic and linguistic characteristics. Vocabulary restrictions also matter: selecting among a few commands is different from transcribing unrestricted conversation. A recognition score consequently describes performance under particular conditions rather than a universal ability to recognize speech. (web.stanford.edu)
From audio to linguistic units
Speech recordings are digitized waveforms. A recognizer may transform overlapping short segments into frequency-based representations through a Fourier transform. Common inputs include logarithmic mel-filterbank energies and mel-frequency cepstral coefficients, which summarize spectral characteristics relevant to speech sounds. This feature extraction reduces the raw signal to a sequence suitable for further modeling; some neural systems instead learn representations directly from waveform samples. (web.stanford.edu)
The acoustic sequence and the written sequence have different lengths. A vowel may occupy many audio frames, while a word may contain several phonemes whose realizations vary with neighboring sounds. Systems must therefore resolve both sound identity and temporal alignment. Outputs may be phonemes, characters, words, or subword units selected through tokenization. These choices affect vocabulary coverage and the relationship between pronunciation and spelling. (web.stanford.edu)
Statistical recognition
A conventional statistical formulation seeks the word sequence (W) most probable given acoustic observations (X):
[ \hat W=\arg\max_W P(W\mid X) =\arg\max_W P(X\mid W)P(W). ]
The second expression follows from Bayes’ theorem, omitting a denominator constant across candidate sequences. An acoustic model supplies the likelihood of the audio, while a language model supplies probabilities for word sequences. A pronunciation lexicon connects words with sound units. (arxiv.org)
Traditional systems frequently use a hidden Markov model to represent progression through speech states and a Gaussian mixture model to describe acoustic observations. Hybrid systems replace the mixture-based acoustic component with neural predictions while retaining the state-based framework. Decoding searches for plausible paths, using methods related to dynamic programming, including the Viterbi algorithm. (arxiv.org)
Neural and end-to-end models
Deep learning enabled recognizers to learn more complex acoustic representations. End-to-end approaches integrate much of the mapping from audio to output symbols within a trainable neural system, reducing reliance on separately engineered components. They nevertheless differ in their alignment assumptions, decoding procedures, and use of external language models. (arxiv.org)
Connectionist temporal classification (CTC), introduced in 2006, is a loss function for learning from sequences without frame-level transcription alignments. It introduces a blank symbol and sums the probabilities of valid alignments that yield the target transcription. This allows training from paired recordings and transcripts rather than manually segmented speech. (cs.toronto.edu)
Attention-based systems use an encoder–decoder architecture to generate output symbols while consulting encoded audio through an attention mechanism. Transducer models provide another approach, combining acoustic information with preceding output symbols. Implementations may employ a recurrent neural network or Transformer architecture. Streaming configurations must balance access to future audio against response latency. (arxiv.org)
Training and multilingual recognition
Supervised learning uses recordings paired with reference transcriptions. The composition of training data influences performance across recording conditions, speaking styles, and languages. Self-supervised learning also exploits untranscribed audio. The 2020 wav2vec 2.0 study demonstrated pretraining through masked speech representations and a contrastive objective, followed by fine-tuning on transcribed speech, including experiments with very limited labeled data. (arxiv.org)
Another approach scales training with large collections of imperfectly supervised audio. The 2022 Whisper study trained multilingual, multitask models on 680,000 hours of audio and reported transfer to multiple benchmarks without dataset-specific fine-tuning. Its tasks included transcription and speech-to-English translation. Transcription preserves the source language, whereas machine translation produces another language; the two objectives remain distinct even when supported by one model. (arxiv.org)
Evaluation, limitations, and applications
The principal transcription metric is word error rate:
[ \mathrm{WER}=\frac{S+D+I}{N}, ]
where substitutions, deletions, and insertions are counted relative to (N) reference words. Alignment uses edit distance. Scores depend on reference conventions and scoring rules, so comparisons require consistent evaluation conditions. WER measures transcription differences, not their practical importance: errors involving names or negation can matter disproportionately. (nist.gov)
Accuracy can also differ across speaker groups. A 2020 study of five commercial systems found average error rates of 35% for Black speakers and 19% for white speakers in its sampled US interviews. These results describe the tested systems and recordings, not every recognizer or subsequent version. Applications include dictation, automated customer service, voice-controlled interfaces, and hands-free computer access; their usefulness depends on both recognition accuracy and the surrounding interface. (doi.org)