来源:Jonathan Flombaum 实验室——约翰斯·霍普金斯大学
口语是人类独有的成就,其产生在很大程度上依赖于特殊的感知机制。语言感知机制的一个重要特征是,它同时依赖听觉和视觉信息。这一点合乎逻辑,因为在现代技术出现之前,人们所接触的语言大多发生在面对面的交流中。由于发出特定的语音需要精确…
1. 刺激材料
2. 诱导错觉
口语形式的语言感知得益于面对面交流,因为口部能为特定发音提供丰富的视觉信息。
例如,在近距离且无遮挡的情况下,一个人可以观察朋友提到要去海滩。此时,他们利用视觉信息——观察嘴唇和舌头周围的运动——来清晰理解所说的内容。
然而,如果朋友继续在另一个房间内说话而看不见其身影,他们可能会倾向于观看静音的电视,因此只能完全依靠被遮挡的声音来理解信息。
在此情况下,尾音实际发出的“pick”干扰了无声的踢唇动作,被误听为“tick”。这是麦格克效应(McGurk Effect)的一个实例——一种由于声音与视觉线索不匹配而产生的感知错觉。
本视频演示了如何构建视听刺激材料,以检验麦格克(McGurk)和麦克唐纳(Macdonald)最初发现的现象。同时探讨了视觉如何与声音产生相互作用,以帮助理解个体在很小年龄时如何学习语言。
在本实验中,要求参与者观看静音视频,视频中某人默念“gain”之类的单词,同时在背景中播放“bane”之类的声音。随后,要求他们说出自己所听到的内容。
为了理解这一现象的结果以及错觉是如何产生的,让我们首先讨论音素——即语音的最小单位——是如何发音的。
例如,单词“bane”和“gain”在除首位以外的所有位置上都具有相同的音素,首位音素分别为 /b/ 和 /g/。
尽管具有这些起始音素的词语听起来可能相似,但当呈现/g/的视觉信息而播放/b/的听觉信息时,个体通常会听到一个完全不同的第三种声音——/d/。
之所以听感为 /d/,是因为这三个音本质上发音方式基本相同,唯一的细微差别在于说话者对气流阻碍位置的不同,即发音部位(POA)的差异。
例如,发 /b/ 音时,双唇形成阻碍,产生双唇发音部位(labial POA);而发 /g/ 音时,发音部位位于口腔后部,称为软腭音(palatal);发 /d/ 音时,发音部位为齿音(dental),这是由于舌尖接触上齿所致。
当大脑整合了视觉上的 /g/ 和听觉上的 /b/ 这一矛盾信息时,会推断最终的声音应位于发音部位的中间位置,因此听觉上感知为 /d/,并报告听到的单词为 Dane。
演示前,准备一台用于播放视频的计算机以及一部带摄像功能的智能手机。
首先调整摄像头位置,使您的头部充满显示画面。然后录制四个10秒的视频片段,每个片段包含不同的词语,每个词语以每秒1个词的频率重复10次。请确保将增益和原始视频传输到计算机,以便进行视觉回放。
进行实验时,让受试者坐在计算机前。打开“gain”一词的视频文件,并关闭音频。
在手机上打开bane的视频,将其放置在电脑后方,使手机屏幕被遮挡,仅能清晰听到声音。
指导参与者观看电脑显示器并聆听。然后,同时播放两个视频。
当视频片段结束后,询问参与者听到了什么。[参与者回答:“Dane”]。接着重复该步骤:在计算机上播放单词“can”的视频,同时通过手机播放“pan”的音频。再次询问参与者听到了什么。[参与者回答:“tan”]。
在此实验中,当参与者观看“gain”和“can”的口型时,播放了“bane”和“pan”的发音。通常情况下,当呈现含有/g/音位的词语视觉刺激,并同时配以/b/的声音时,个体往往会感知为/d/音。
同样,当一个以/k/开头的词与/p/音组合时,人们会听成/t/。
这种听觉感知背后的原因与声音的产生方式有关。当眼睛看到双唇音的口型 /b/ 和 /p/,而耳朵听到的是硬腭音 /g/ 和 /k/ 时,大脑会试图解决这种视觉与听觉之间的信息冲突。最终,大脑判断声音应介于两者之间,从而产生了齿音 /d/ 和 /t/ 的感知。
现在您已经熟悉了如何产生麦克古克效应,接下来让我们看看研究人员如何利用这一知觉现象来探究语言发展,以及该效应发生改变的一些情况。
甚至可以在婴儿五个月大、尚未发展语言能力时,利用注视时间习惯化范式来测试其麦格克效应。
在此实验过程中,Rosenblum 及其同事在测试阶段引入不匹配的音素之前,先在听觉和视觉通道中反复向婴儿呈现某个特定音节(如 va)。
婴儿表现出对“va”刺激的习惯化迹象——注视时间减少——以及当感知到不同于“va”的刺激时出现的去习惯化,表现为注视时间增加。因此,即使在婴儿还不会说话之前,他们就已表现出与成人相似的结果,即依赖视觉信息进行语言辨别。
然而,由于自闭症儿童在理解和关注视觉面部成分方面存在障碍,他们表现出麦格克效应的能力较对照组更差。这表明他们在视听言语处理方面存在根本性差异,这可能与其在语言和沟通方面的困难有关。
最后,左半球(通常在语言理解和学习方面占优势的一侧)有病灶的患者在言语治疗过程中常利用面部视觉特征来辅助。有趣的是,在进行麦格克效应测试时,与对照组相比,他们更常报告听到了齿音。这种感知很可能是由于他们更加关注视觉信息所致。
您刚刚观看了JoVE关于麦格克效应的视频。现在,您应该了解如何进行这种视听错觉实验,并将音素与声音产生联系起来。此外,您还应更好地理解视觉与听觉之间的相互作用,以及这些相互作用在发育期和成年期如何受到影响。
感谢观看!
View the full transcript and gain access to JoVE Science Education videos
Q1: What is the McGurk Effect and how does it demonstrate audiovisual perception?
The McGurk Effect is a perceptual illusion arising from a mismatch between sound and visual cues during speech perception. When you watch someone mouth one word while hearing a different word played simultaneously, your brain integrates the conflicting information and you perceive a third sound entirely. For example, watching 'gain' mouthed while hearing 'bane' typically results in hearing 'Dane.' This demonstrates that vision and hearing work together to interpret spoken language.
Q2: Why does the brain perceive a different phoneme when visual and auditory information conflict?
The brain resolves conflicting visual and auditory phonemes by computing a compromise based on points of articulation—where the speaker places an obstruction in airflow. When you see /g/ (palatal, back of mouth) but hear /b/ (labial, lips), your brain concludes the sound must lie between these positions, resulting in /d/ (dental, tongue on teeth). This integration mechanism reflects how perspectives on sensation and perception reveal the brain's reliance on visual information to disambiguate ambiguity in spoken language.
Q3: How can you conduct a McGurk Effect experiment with basic equipment?
Record four 10-second video clips of yourself repeating different words at one word per second. Mute one video on a computer monitor while playing audio from a different word on a hidden smartphone behind it. Ask participants what they heard. For example, play muted 'gain' video with 'bane' audio, or muted 'can' video with 'pan' audio. Participants typically report hearing the compromise phoneme rather than either the visual or auditory input alone.
Q4: What do infants reveal about audiovisual speech processing using the McGurk Effect?
Infants as young as five months old show the McGurk Effect using habituation-of-looking-time paradigms. Researchers repeatedly present infants with matched audiovisual syllables like 'va,' then introduce mismatched phonemes. Infants display habituation through reduced looking times, then dishabituation with increased looking when perceiving something different. This demonstrates that even pre-linguistic infants rely on visual information for language discrimination, similar to adults.
Q5: How does autism affect the McGurk Effect and audiovisual speech processing?
Children with autism show greater difficulty exhibiting the McGurk Effect compared to controls due to impaired ability to understand and attend to visual facial components. This indicates fundamental differences in processing audiovisual speech, which may contribute to their difficulty with language and communication. The reduced reliance on visual cues during speech perception suggests altered integration of sensory information in autism spectrum disorder.
Q6: Why do patients with left hemisphere lesions show stronger McGurk Effects during speech therapy?
Patients with left hemisphere lesions—the side typically dominant for language understanding and learning—often report hearing dental sounds more frequently than controls when tested on the McGurk Effect. This occurs because these patients rely more heavily on visual facial features to compensate for impaired auditory language processing during speech therapy. Their heightened focus on visual information results in stronger perception of the compromise phoneme.
Q7: What role does face-to-face interaction play in language perception?
Face-to-face interaction provides critical visual information about mouth and tongue movements that clarifies spoken language. The mouth can often supply better visual signals than speech supplies auditory signals, especially in close, unobstructed views. The human brain favors visual input to disambiguate inherent ambiguity in spoken language. This reliance on visual cues explains why understanding someone talking out of sight is more difficult than face-to-face communication.