Computers can talk to us. 
Computers can learn to talk. 
Computers can learn to make human speech. This is called speech synthesis.
There are different ways to make these sounds. Some systems use a database of recorded speech. They glue small pieces of sound together to make words. Other systems use a model of the vocal tract. The vocal tract is the part of the body used for speaking. This can make a completely new, synthetic voice.
These tools help many people. They help those with visual impairments or reading disabilities. 
Speech synthesis is the way computers create human speech. 
A text-to-speech system usually has two main parts. The first part is called the front-end. It takes raw text and turns it into written-out words. This step is called text normalization. The front-end also gives each word a phonetic transcription. This is a way to show how a word sounds. It also marks where phrases and sentences begin.
People have tried to make machines speak for a very long time. Long ago, there were legends about "Brazen Heads" that could talk. In 1779, Christian Gottlieb Kratzenstein built models of the human vocal tract. These models could make five different long vowel sounds. Later, Wolfgang von Kempelen built a machine using bellows. 
Many famous machines and milestones helped shape this technology. In 1961, John Larry Kelly, Jr. used an IBM 704 computer to synthesize speech. The computer even sang the song "Daisy Bell."
Today, we use artificial intelligence to make voices sound even more natural. In 2016, DeepMind released WaveNet to start the field of deep learning speech synthesis. This uses models to create speech from acoustic features. In 2018, Google AI released Tacotron 2. This system used neural networks to make very natural speech.
Speech synthesis is the artificial production of human speech. This technology allows computers to turn written text into spoken words. A system used for this purpose is called a speech synthesizer. These systems can exist as software or as dedicated hardware products.
A text-to-speech (TTS) system generally consists of two distinct parts. The first part is called the front-end. The front-end performs several major tasks to prepare the text. First, it uses a process called text normalization. This converts raw text, like numbers or abbreviations, into written-out words. Next, it performs text-to-phoneme conversion. This means it assigns phonetic transcriptions to each word. The front-end also divides the text into prosodic units. These units include phrases, clauses, and sentences. This creates a symbolic linguistic representation for the next step.
The second part is the back-end, which is often called the synthesizer. The back-end takes the symbolic linguistic representation and converts it into actual sound. There are different ways a synthesizer can create this sound. Some systems use concatenation. This method glues together small pieces of recorded speech from a database. The size of the stored units changes the output. A system using phones or diphones has a large range but may lack clarity. Other systems store entire words or sentences for higher quality.
Humans have tried to build speaking machines for centuries. Long ago, legends spoke of "Brazen Heads" that could talk. In 1779, Christian Gottlieb Kratzenstein built models of the human vocal tract. These models could produce five long vowel sounds. Later, Wolfgang von Kempelen built an acoustic-mechanical speech machine. This machine used bellows to move air. It included models of the tongue and lips to produce consonants. 
Computer-based speech synthesis began in the late 1950s. In 1961, John Larry Kelly, Jr. and Louis Gerstman used an IBM 704 computer to synthesize speech. The computer sang "Daisy Bell." This famous event even inspired a scene in the film 2001: A Space Odyssey. In 1968, Noriko Umeda and others developed the first general English TTS system in Japan. By 1974, the Unix operating system included a speech utility. In 1978, Texas Instruments released the Speak & Spell toy. This toy used special speech chips based on Linear Predictive Coding (LPC).
Newer technologies have moved into the field of artificial intelligence. In 2016, DeepMind released WaveNet. This used deep learning to model raw waveforms. A modified version called Parallel WaveNet was released a year later. It was 1,000 times faster than the original. In 2018, Google AI released Tacotron 2. This system used neural networks to produce very natural speech. It required tens of hours of audio training data to work well.
Modern AI has changed how much data we need to create voices. The platform 15.ai became famous for its ability to clone voices. It can synthesize an expressive voice using only 15 seconds of training data. This is a massive reduction from the tens of hours once required. This technology has become popular in memes and content creation. However, it also led to the first instance of speech synthesis NFT fraud in 2022. 
🖼️ Images & Media (19)
+ 7 more
More to explore
✨ What else?
Related topics you might enjoy
🔬 Go deeper
More advanced topics to explore
🪜 Step back
Simpler topics to build understanding
What is Nepedia?
A free, ad-free encyclopedia for children. Every article is written at five reading levels, so the same page works for a five-year-old and a fifteen-year-old — use the level switcher above to see this one change. No account needed to read.