[QUOTE=Nava]
… formant?
M-W:
Main Entry: for·mant
Function: noun
Date: 1901
: a characteristic component of the quality of a speech sound; specifically : any of several resonance bands held to determine the phonetic quality of a vowel
…
still leaves me kind of “ei?”
Let’s see if I get this, a “formant” is any of the elements that make a particular voice be that particular voice? Tone, pitch, etc only they have another technical name? And you’re saying that being able to recognize two different operatic tenors singing the same aria is an extra formant over regular people?
(The “first, second, third” reminds me of how in Statistics stuff like the “average” and the “standard deviation” are officially called first, second… Idon’trememberthetermrightnow - people still use the old fashioned names, because they may be less exact but we don’t really need that level of exactitude)
[/QUOTE]
Heheh, I apologize; I was trying to go for a really generalized answer that skipped the theoretical stuff, but let me be pseudo-precise. (Warning! This deals with simple physics, which is not an area I specialize in. I’m also leaving out oodles and oodles of stuff for the sake of simplicity and readability. If you want a really thorough introduction, read Elements of Acoustic Phonetics by Peter Ladefoged. It’s a really neat book.):
Everyone who has replied basically has the right idea about how speech is produced. One very prevalent model in north america is the Source-Filter model, which basically claims that there are two big components at work making speech: a Source, which is vibrating air (which may or may not include buzzing introduced by the vocal folds), and a Filter, which is the post-glottal vocal tract. You know how acoustic instruments like guitars have sound boxes which amplify the sound of the string vibrating by boosting certain types of oscillation in the air? Well the Filter in the vocal tract works in almost the exact same way. Think of the throat, oral cavity, and nasal cavity like a big cavern: you make a noise, and if the cavern has the right proportions when the noise you’ve made echoes, the echo will be louder than the original noise.
The reasoning behind this amplification is fairly simple: the vocal tract essentially acts like a tube resonator, so a very good (simplified) way of looking at the properties of resonance which amplify the speech signal.
First, we’re going to look at simple harmonic resonance. The speech signal is extremely complex, but it is gudied by the basic principles.
As you’re probably aware, vibration just describes oscillation of an object (in this case, molecules of air). If I take a spring and pull on it it’ll get longer, but the moment I let go of the spring it will spring back to its resting state, and then some. In recovering from the displacement I introduced it will actually compact itself, momentarily coiling more tightly than it was originally. The configuration our spring was in before we messed with it is called its “equilibrium position”. In essence, a harmonic oscillator like a spring, or the air molecules which vibrate during speech, is simply a system which, when displaced from its equilibrium position, will experience a force to return to said equilibrium which is proportional to the displacing force. Since our spring will only have one displacing force in this simplified model - the kinetic energy introduced by the guy who first pulled on it - it is displaying simple harmonic motion: if you were to tie a pencil to the vibrating end of the spring and put a moving sheet of paper under it like a seismometer, the pencil would draw a sinusoidal wave. Anyway, since there’s friction acting on the spring and damping the oscillation it’s eventually going to stop vibrating back and forth and return to its equilibrium position.
Air in a tube resonator essentially works in the same way, and this is vital to promoting the resonance which amplifies the signal. You see, the speech signal at the source - the glottis - is actually quite quiet in most cases. However, if we picture the vocal tract as an upside-down L-shaped tube and the vibration at the source as a wave, we can readily visualize how formants work.
First off, let’s assume that we’re working with something like a vowel: the vocal folds are vibrating, air flows through the folds, through the mouth, and out past the lips. The first thing we need is the frequency of the vibration at the Source, the vocal folds. If we take the frequency of the source (F0, or the Fundamental Frequency) in hertz, it will simply be equal to the number of vibrations of the vocal folds per second. Let’s choose a nice, clean number like 100. (The numbers we’re going to develop won’t work anywhere, and would never accurately model a real-life vowel. Be forwarned :).) Now, visualize a 100hz sine wave: since we have 100 cycles per second, the wave will repeat itself 100 times. Now, as the signal travels through the vocal tract the tract will create resonances at various frequencies. What we’re interested in are the resonances which boost the signal. As it turns out, the resonances which boost the vibration at the source are the harmonics, which are odd integral multiples of the fundamental frequency. To visualize it, draw a 300hz sine wave over the 100hz wave. This wave is 3x the fundamental frequency, and as such it’s going to be the first harmonic. If you look at the drawing, you’ll notice that although the waves don’t match up in their entirety, they overlap at key points. This overlap is essentially what boosts the speech signal.
Now that we know what harmonics look like, formants are easy: a formant is just an acoustic peak in the signal, which amplifies certain key harmonics.
Wikipedia actually has a nice visual reference at File:Spectrogram -iua-.png - Wikipedia . The dark bands, which are noted with red arrows, are the formants.