Source: Laboratory of Jonathan Flombaum—Johns Hopkins University
Spoken language, a singular human achievement, relies heavily on specialized perceptu…
1. Stimuli
2. Inducing the Illusion
Language perception?in a spoken form?benefits from face-to-face interactions, as the mouth supplies good visual information for articulating specific sounds.
For instance, in an up-close and unobstructed situation, an individual can watch their friend mention going to the beach. In this case, they use visual input?observing the movement around the lips and tongue?to clearly comprehend what was said.
However, if the friend continues to talk out of sight in another room, they might be tempted to watch the muted television and therefore must solely rely on the obstructed voice to make sense of the message.
In this case, what was actually said at the tail end, pick, interfered with the silent kick and was misinterpreted as tick. This is an example of the McGurk Effect?a perceptual illusion that arises through a mismatch between sound and visual cues.
This video demonstrates how to construct the audiovisual stimuli to test the phenomenon originally discovered by McGurk and Macdonald. It also investigates how vision interacts with sound production to understand how individuals learn language at a very young age.
In this experiment, participants are asked to watch muted videos, in which a word like gain is mouthed, while a sound such as bane is played simultaneously in the background. Afterwards, they are asked to share what they heard.
To understand the outcome, how the illusion is produced, let?s first discuss how phonemes?the minimal units of speech sounds?are articulated.
For example, bane and gain share the same elements in all positions except for the first, which are the sounds /b/ and /g/.
Although words with these initial phonemes may sound similar, when /g/ is shown and /b/ is played, individuals are expected to hear a completely different third sound?/d/?instead.
The reason /d/ is heard is due to the fact that all three are basically produced in the same manner, with only a small difference in where the speaker places an obstruction in airflow, called the points of articulation, or POA.
For instance, when a /b/ sound is made, lips provide the obstruction, resulting in a labial POA, whereas for /g/, it?s referred to as palatal?in the back of the mouth. As for /d/, the POA is dental, a consequence of the tongue touching the upper teeth.
When the brain integrates the conflicting visual /g/ and auditory /b/, it concludes that the final sound must lie somewhere in the middle of POAs, thus hearing /d/ and reporting the word Dane.
In preparation for the demonstration, obtain a computer to present videos on and a smartphone with a video camera.
First position the camera so that your head fills the display. Now, record four 10-s clips, each one containing different words that should be repeated 10 times at a rate of 1 word/s. Make sure to transfer the gain and can videos to the computer for visual playback.
To conduct the experiment, sit a participant in front of the computer. Open up the video file for the word gain and turn off the audio.
On the phone, open up the video for bane. Place it behind the computer so that its screen is hidden and only the sound can be heard clearly.
Instruct the participant to watch the computer monitor and listen. Then, play both videos simultaneously.
When the clips end, ask the participant what they heard. [Participant says: "Dane"]. Repeat the procedure by playing the video of the word can on the computer and presenting the audio for pan on the phone. Once again, question the participant as to what they heard. [Participant says: "tan"].
Here, the words bane and pan were played aloud as the participant watched gain and can being mouthed. Typically, when a term with the /g/ phoneme is shown visually and paired with the sound /b/, individuals will hear /d/.
Likewise, when a word starting with /k/ is paired with the sound /p/, individuals will hear /t/.
The reason behind such auditory perception is due to the way that sounds are produced. The brain tries to resolve conflicting information from the eyes seeing labial movements?/b/ and /p/?while the ears hear palatal units?/g/ and /k/. As a result, it concludes that the sounds must lie in the middle, resulting in the perception of dental phonemes?/d/ and /t/.
Now that you are familiar with how to produce the McGurk effect, let?s look at some other ways that researchers use this perceptual phenomenon to investigate language development and cases in which the effect is altered.
Infants can even be tested on the McGurk effect as early as five months of age, when they are pre-linguistic, using an habituation-of-looking-time paradigm.
In this procedure, Rosenblum and colleagues repeatedly presented infants with a particular syllable, like va, in both the audio and visual domains before introducing mismatched phonemes in a testing phase.
Infants showed signs of habituation to va?reduced looking times?and dishabituation, noted as increased looking, when something other than va was perceived. Thus, even before infants can talk, they display similar results as adults, in which they rely on the use of visual information for language discrimination.
However, children with autism have greater difficulty exhibiting the McGurk effect as readily as controls due to their impaired ability to understand and attend to the visual facial components. This indicates fundamental differences in processing audiovisual speech, which may contribute to their difficulty with language and communication.
Lastly, patients with lesions in their left hemisphere?the side typically predominant for understanding and learning language?often use visual facial features to help during speech therapy. Interestingly, when tested on the McGurk effect, they more often reported hearing dental sounds compared to controls. Such perceptions are likely due to their higher focus on visual information.
You?ve just watched JoVE?s video on the McGurk Effect. Now you should know how to conduct this audiovisual illusion and relate phonemes to sound production. In addition, you should also have a better understanding of the interactions between vision and hearing, and how they can be affected during development and adulthood.
Thanks for watching!
View the full transcript and gain access to JoVE Science Education videos
Q1: What is the McGurk Effect and how does it demonstrate audiovisual perception?
The McGurk Effect is a perceptual illusion arising from a mismatch between sound and visual cues during speech perception. When you watch someone mouth one word while hearing a different word played simultaneously, your brain integrates the conflicting information and you perceive a third sound entirely. For example, watching 'gain' mouthed while hearing 'bane' typically results in hearing 'Dane.' This demonstrates that vision and hearing work together to interpret spoken language.
Q2: Why does the brain perceive a different phoneme when visual and auditory information conflict?
The brain resolves conflicting visual and auditory phonemes by computing a compromise based on points of articulation—where the speaker places an obstruction in airflow. When you see /g/ (palatal, back of mouth) but hear /b/ (labial, lips), your brain concludes the sound must lie between these positions, resulting in /d/ (dental, tongue on teeth). This integration mechanism reflects how perspectives on sensation and perception reveal the brain's reliance on visual information to disambiguate ambiguity in spoken language.
Q3: How can you conduct a McGurk Effect experiment with basic equipment?
Record four 10-second video clips of yourself repeating different words at one word per second. Mute one video on a computer monitor while playing audio from a different word on a hidden smartphone behind it. Ask participants what they heard. For example, play muted 'gain' video with 'bane' audio, or muted 'can' video with 'pan' audio. Participants typically report hearing the compromise phoneme rather than either the visual or auditory input alone.
Q4: What do infants reveal about audiovisual speech processing using the McGurk Effect?
Infants as young as five months old show the McGurk Effect using habituation-of-looking-time paradigms. Researchers repeatedly present infants with matched audiovisual syllables like 'va,' then introduce mismatched phonemes. Infants display habituation through reduced looking times, then dishabituation with increased looking when perceiving something different. This demonstrates that even pre-linguistic infants rely on visual information for language discrimination, similar to adults.
Q5: How does autism affect the McGurk Effect and audiovisual speech processing?
Children with autism show greater difficulty exhibiting the McGurk Effect compared to controls due to impaired ability to understand and attend to visual facial components. This indicates fundamental differences in processing audiovisual speech, which may contribute to their difficulty with language and communication. The reduced reliance on visual cues during speech perception suggests altered integration of sensory information in autism spectrum disorder.
Q6: Why do patients with left hemisphere lesions show stronger McGurk Effects during speech therapy?
Patients with left hemisphere lesions—the side typically dominant for language understanding and learning—often report hearing dental sounds more frequently than controls when tested on the McGurk Effect. This occurs because these patients rely more heavily on visual facial features to compensate for impaired auditory language processing during speech therapy. Their heightened focus on visual information results in stronger perception of the compromise phoneme.
Q7: What role does face-to-face interaction play in language perception?
Face-to-face interaction provides critical visual information about mouth and tongue movements that clarifies spoken language. The mouth can often supply better visual signals than speech supplies auditory signals, especially in close, unobstructed views. The human brain favors visual input to disambiguate inherent ambiguity in spoken language. This reliance on visual cues explains why understanding someone talking out of sight is more difficult than face-to-face communication.