The translational value of primary research in human perception and cognition hinges on the extent to which the findings transfer to real-world stimuli and contexts. A long-standing question concerns how the brain processes real-world sensory inputs. Currently, knowledge of visual cognition is based almost exclusively on studies that have relied on stimuli in the form of two-dimensional (2-D) pictures, usually presented in the form of computerized images. Although image interaction is becoming increasingly common in the modern world, humans are active observers for whom the visual system has evolved to allow perception and interaction with real objects, not images1. To date, the overarching assumption in studies of human vision has been that images are equivalent to, and appropriate proxies for, real object displays. Currently, however, we know surprisingly little about whether images effectively trigger the same underlying cognitive processes as do the real objects. Therefore, it is important to determine the extent to which responses to images are like, or different from, those elicited by their real-world counterparts.
There are several important differences between real objects and images that could lead to differences in how these stimuli are processed in the brain. When we look at real objects with two eyes, each eye receives information from a slightly different horizontal vantage point. This discrepancy between the different images, known as binocular disparity, is resolved by the brain to produce a unitary sense of depth2,3. Depth cues derived from stereoscopic vision, together with other sources such as motion parallax, convey precise information to the observer about the object’s egocentric distance, location, and physical size, as well its three-dimensional (3-D) geometric shape structure4,5. Planar images of objects do not convey information about the physical size of the stimulus because only the distance to the monitor is known by the observer, not the distance to the object. While 3-D images of objects, such as stereograms, approximate more closely the visual appearance of real objects, they do not exist in 3-D space, nor do they afford genuine motor actions such as grasping with the hands6.
The practical challenges of using real object stimuli in experimental contexts
Unlike studies of the image vision in which stimulus presentation is entirely computer-controlled, working with real objects presents a range of practical challenges for the experimenter. The position, order, and timing of object presentations must be controlled manually throughout the experiment. Working with real objects (unlike images) can involve a significant time commitment due to the need to collect7,8,9 or make10 the objects, set up the stimuli prior to the experiment, and present the objects manually during the study. Moreover, in experiments that are designed to compare, directly, responses to real objects with images, it is critical to match closely the appearance of the stimuli in the different display formats8,9. Stimulus parameters, environmental conditions, as well as randomization and counterbalancing of real object and image stimuli, must all be controlled carefully to isolate causal factors and rule out alternative explanations for the observed effects.
The methods detailed below for presenting real objects (and matched images) are described in the context of a decision-making paradigm. The general approach can be extended, however, to examine whether stimulus format influences other aspects of visual cognition such as perception, memory or attention.
Are real objects processed differently to images? A case example from decision-making
The mismatch between the kinds of objects that we encounter in real-world scenarios versus those examined in laboratory experiments is especially apparent in studies of human decision-making. In most studies of dietary choice, participants are asked to make judgments about snack foods that are presented as colored 2-D images on a computer monitor 11,12,13,14. In contrast, everyday decisions about which foods to eat are usually made in the presence of real foods, such as at the supermarket or the cafeteria. Although in modern life we regularly view images of snack foods (i.e., on billboards, television screens and online platforms), the ability to detect and respond appropriately to the presence of real energy-dense foods may be adaptive from an evolutionary perspective because it facilitates growth, competitive advantage, and reproduction15,16,17.
Research outcomes in scientific studies of decision-making and dietary choice have been used to guide public health initiatives aimed at curbing rising obesity rates. Unfortunately, however, these initiatives appear to have met with little to no measurable success18,19,20,21. Obesity remains a major contributor to the global burden of a disease22 and is linked to a range of associated health problems, including coronary heart disease, dementia, Type II diabetes, certain cancers, and increased overall risk of morbidity22,23,24,25,26,27. The sharp rise in obesity and associated health conditions over recent decades28 has been linked with the availability of cheap, energy-dense foods18,29. As such, there is an intense scientific interest in understanding the underlying cognitive and neural systems that regulate everyday dietary decisions.
If there are differences in the way foods in different formats are processed in the brain, then this might provide insights into why public health approaches to combating obesity have been unsuccessful. Despite the differences between images and real-world objects, described above, surprisingly little is known about whether images of snack foods are processed similarly to their real-world counterparts. In particular, little is known about whether or not real foods are perceived to be more valuable or satiating than matched images of the same items. Classic early behavioral studies found that young children were able to delay gratification in the context of 2-D colored images of snack foods30, but not when they were confronted with real snack foods31. However, few studies have examined in adults whether the format in which a snack food is displayed influences decision-making or valuation12,32,33 and only one study to date, from our laboratory, has tested this question when stimulus parameters and environmental factors are matched across formats7. Here, we describe innovative techniques and apparatus for investigating whether decision-making in healthy human observers is influenced by the format in which the stimuli are displayed.
Our study7 was motivated by a previous experiment conducted by Bushong and colleagues12 in which college-aged students were asked to place monetary bids on a range of everyday snack foods using a Becker-DeGroot-Marschak (BDM) bidding task34. Using a between-subjects design, Bushong and colleagues12 presented the snack foods in one of three formats: text descriptors (i.e., 'Snickers bar'), 2-D colored images, or real foods. Average bids for the snacks (in dollars) were contrasted across the three participant groups. Surprisingly, students who viewed real foods were willing to pay 61% more for the items than those who viewed the same stimuli as images or text descriptors -a phenomenon the authors termed the 'real-exposure effect'12. Critically, however, participants in the text and image conditions completed the bidding task in a group setting and entered their responses via individual computer terminals; conversely, those assigned to the real food condition performed the task one-on-one with the experimenter. The appearance of the stimuli in the real and image conditions was also different. In the real food condition, the foods were presented to the observer on a silver tray, whereas in the image condition the stimuli were presented as scaled cropped images on a black background. Thus, it is possible that participant differences, environmental conditions, or stimulus-related differences, could have led to inflated bids for the real foods. Following from Bushong, et al.12, we examined whether the real foods are valued more than 2-D images of food, but critically, we used a within-subjects design in which environmental and stimulus-related factors were carefully controlled. We developed a custom-designed turntable in which the stimuli in each display format could be interleaved randomly from trial to trial. Stimulus presentation and timing were identical across the real object and image trials, thus reducing the likelihood that participants could use different strategies to perform the task in the different display conditions. Finally, we controlled carefully the appearance of the stimuli in the real object and image conditions so that the real foods and images were matched closely for apparent size, distance, viewpoint, and background. There are likely to be other procedures or mechanisms that could allow for randomizing stimulus formats across trials, but our method allows for many objects (and images) to be presented in relatively rapid interleaved succession. From a statistical standpoint, this design maximizes power to detect significant effects more so than is possible using between-subjects designs. Similarly, the effects cannot be attributed to a-priori differences in willingness-to-pay (WTP) between observers. It is, of course, the case that in within-subjects designs open the possibility for demand characteristics. However, in our study participants understood that they could 'win' a food item at the end of the experiment regardless of the display format in which it appeared in the bidding task. Participants were also informed that arbitrarily reducing bids (i.e., for the images) would reduce their chances of winning and that the best strategy for winning the desired item is to bid one’s true value34,35,36. The aim of this experiment is to compare WTP for real foods versus 2-D images using a BDM bidding task34,35.