Source: Laboratories of Gary Lewandowski, Dave Strohmetz, and Natalie Ciarocco—Monmouth University
In order to study something scientifically, a resea…
1. Define key variables.
2. Create coding categories from the operational definition of inappropriate content.
| Coding Categories | Themes and Exemplars | Count |
| Crude Behavior | Toilet humor Purposefully disgusting behaviors | |
| Rude Behavior | Disrupting others Poor Manners | |
| Language | Using curse words | |
| Verbal Aggression | Insults Yelling Name-Calling | |
| Physical Aggression | Hitting Pushing/Shoving Tripping | |
| Drug References | Verbal (suggestive statements/conversation) Nonverbal (mimicking drug use) | |
| Sexual References | Verbal (suggestive statements/conversation) Nonverbal (mimicking sexual acts) |
Table 1. Example of how to record instances of inappropriate behaviors. This log can be systematically used across raters.
3. Instruct raters to separately watch the same episode of SpongeBob SquarePants and provide coding counts.
4. Instruct raters to separately watch the same episode of Caillou and provide coding counts.
5. Compare ratings to see if the Raters came up with similar ratings for each show.
Scientific research uses precise methods to collect data, yet variability in obtaining measurements often exists.
Reliability can be assessed for any experimental measurement, and today, we?ll have a look at measurements of inappropriate behaviors in cartoons.
When viewers agree on the amount of inappropriate material within the same show?across multiple episodes?their judgments are considered highly reliable. In this case, assessments can extend across different shows because of the consistency between observers, which is referred to as inter-rater reliability.
This video demonstrates how to design and perform, as well as how to analyze and interpret, an experiment examining whether one cartoon has more inappropriate content than another.
To examine reliability and inter-rater reliability, a within-subjects design is used in this experiment. Participants are asked to watch two episodes of two different cartoons?SpongeBob SquarePants and Caillou.
Within this context of cartoon watching, the dependent variable is the number of inappropriate behaviors participants observe. These include: any crude and rude behaviors, bad language, verbal and physical aggression, and references to drugs and sexual content.
If reliability exists in the scoring of inappropriate content of a specific cartoon, participants will consistently rate that cartoon across different episodes.
Moreover, if multiple participants are in agreement with the number of inappropriate instances they count, inter-rater reliability exists.
Thus, establishing inter-rater reliability allows researchers to use the same participants to more powerfully compare data between multiple conditions.
To conduct the study, prepare four clips: two different episodes from two different cartoons, SpongeBob SquarePants and Caillou.
To allow participants to systematically identify instances of inappropriate behavior, create a coding sheet with categories, concrete examples, and space to count each occurrence.
With the participant sitting in front of the screen, hand them four coding sheets. Instruct the participant to separately watch two episodes of SpongeBob SquarePants.
As the participant watches each episode, instruct them to identify every occurrence of inappropriate behavior.
Using the same coding scheme, instruct the participant to watch and rate two episodes of Caillou.
To analyze the reliability of participants? ratings of cartoon content, compare the coding sheets between each participant across the different episodes of cartoons. Sum all of the responses on a master sheet.
Graph the total number of inappropriate behaviors for each rater across episodes and cartoons.
Note that high reliability was observed in the scoring of the two different cartoons, as SpongeBob is consistently scored higher than Caillou.
However, stronger inter-rater reliability was found in the scoring of inappropriate content in Caillou compared to SpongeBob. Reduced inter-rater reliability was more obvious in the scoring of Episode 2 of SpongeBob.
Now that you are familiar with reliability in the context of content analysis, you can apply this approach to other areas of research.?
Many psychological experiments gather information by utilizing cognitive assessments and surveys, in which reliability between each of the items must be consistent between participants.
Reliability in neurophysiological measures, such as EEG or eye tracking, is essential to conducting repeatable experiments. This reliability allows researchers to make associations between brain function and disease states across multiple subjects.
Additionally, researchers must ensure certain measurements in an experiment are consistent over time. For example, weight measurements are reliably taken to compare data before and after exercise routines.
You?ve just watched JoVE?s introduction to determining reliability in psychological experiments. Now you should have a good understanding of how to quantify a psychological construct such as inappropriate behavior, design an experiment, and finally how to evaluate reliability from the results.
Thanks for watching!?
View the full transcript and gain access to JoVE Science Education videos
Q1: What is inter-rater reliability in psychology experiments?
Inter-rater reliability occurs when multiple observers or raters agree on their measurements of the same behavior or phenomenon. When participants consistently count the same number of inappropriate instances across different episodes or shows, this demonstrates strong inter-rater reliability. Establishing inter-rater reliability allows researchers to confidently compare data between multiple conditions using the same participants.
Q2: How do you measure reliability in content analysis studies?
Reliability in content analysis is measured by comparing coding sheets between participants across different episodes or conditions. Researchers sum all responses on a master sheet and graph the total number of occurrences for each rater. High reliability is demonstrated when raters consistently score the same content similarly, such as SpongeBob consistently scoring higher than Caillou across episodes.
Q3: Why is a coding sheet important when analyzing behavioral content?
A coding sheet provides a systematic framework for identifying and counting specific behaviors. It includes concrete categories, examples, and space to record each occurrence, ensuring participants apply consistent criteria when observing. This standardization helps establish reliability by allowing multiple raters to independently assess the same content using identical definitions and measurement procedures.
Q4: What design did researchers use to examine reliability in the cartoon study?
Researchers used a within-subjects repeated-measures design where participants watched multiple episodes from two different cartoons. Each participant rated the same cartoons across different episodes, allowing researchers to assess both test-retest reliability within a cartoon and inter-rater reliability across participants. This design strengthens comparisons between conditions by using the same participants.
Q5: How does reliability apply beyond content analysis in psychology research?
Reliability is essential across multiple psychological measurement methods. Cognitive assessments and surveys require consistent item reliability between participants. Neurophysiological measures like EEG or eye tracking must be reliable to establish associations between brain function and disease states. Additionally, researchers must ensure measurements remain consistent over time, such as weight measurements taken before and after exercise interventions.
Q6: What dependent variable was measured in the cartoon content study?
The dependent variable was the number of inappropriate behaviors participants observed in each cartoon episode. Inappropriate behaviors included crude and rude actions, bad language, verbal and physical aggression, and references to drugs or sexual content. Participants used the coding sheet to systematically count and record each occurrence of these behaviors while watching the cartoons.
Q7: Why is quantifying psychological constructs challenging for researchers?
Psychological constructs like inappropriate behavior are abstract and subjective, making them difficult to measure directly. Researchers must develop operational definitions and systematic measurement tools, such as coding sheets with concrete examples, to transform abstract concepts into quantifiable data. This process requires careful design to ensure different observers can reliably identify and count the same behaviors consistently.