Study design
Before any recruitment or data collection, the study protocol was reviewed and approved by the Research Ethics Committee of the University of Malaya (approval ID: UUM20259910). All procedures involving human participants were conducted in accordance with institutional requirements and the Declaration of Helsinki. Participants provided electronic informed consent before entering the questionnaire, and no personally identifiable information was exported for analysis.
This study used a randomized between-subjects online experiment to examine risk perception and disaster information verification among Chinese university students after exposure to simulated social media posts. Following stimulus-based misinformation research12, the design manipulated predefined information cues rather than relying only on retrospective self-reported exposure. The experiment used a 2 × 2 × 2 factorial structure with three factors: source cue, information veracity, and emotional tone. The official-source + misleading-information cells represented fictional official-looking or impersonated official-style posts created solely for the experiment; they did not represent a real emergency management agency issuing false information.
The study procedure is shown in Figure 1. Participants completed eligibility screening, electronic informed consent, baseline measures, random assignment, exposure to one simulated post, post-exposure assessment, verification task, debriefing, and data-quality screening. Each participant viewed only one post to avoid learning effects across conditions. The factorial structure and randomization cells are presented in Table 1.
Participants first completed eligibility screening and electronic informed consent, followed by baseline measurement, individual-level random assignment to one of eight experimental cells, exposure to one simulated disaster-related social media post, post-exposure assessment, a task-based verification assessment, immediate debriefing, data-quality screening, and final statistical analysis. The workflow shows only procedural and screening steps and does not present empirical results.
Participants and recruitment
Participants were full-time undergraduate or postgraduate students enrolled in mainland Chinese universities. Eligibility criteria were age 18 years or older, current university enrollment, ability to read simplified Chinese, and regular use of at least one social media platform. Participants were excluded if they were younger than 18 years, not currently enrolled as university students, had joined the pilot test, submitted an incomplete questionnaire, failed the attention-check item, completed the questionnaire in an unrealistically short time, or showed patterned responses across scale items.
Recruitment was conducted through university notice groups, course announcement channels, and student association networks. The recruitment notice described the study as research on university students' evaluation of online disaster information. It did not disclose before participation that some posts contained misleading elements, because advance disclosure would change the verification task. The use of Chinese university students was justified because this population frequently encounters social media misinformation and reports uncertainty when judging online credibility13.
Sample size determination
The target sample size was set before data collection. The study aimed to retain at least 400 valid responses, with approximately 50 participants in each of the eight experimental cells. This target supported factorial comparisons, interaction testing, and sensitivity analyses after invalid responses were removed. Approximately 460 students were invited to allow for incomplete questionnaires, failed attention checks, and low-quality responses.
Given that associations between risk perception and emergency information-seeking behavior are usually modest in size14, the study retained a cell size large enough to stabilize main-effect and interaction estimates. Randomization balance across the eight cells was assessed using chi-square tests for categorical baseline variables and one-way ANOVA for continuous baseline variables before the primary models were fitted.
Power and sensitivity planning were defined before analysis. With a target retained sample of at least 400 participants and eight randomized cells, the study was adequately powered for small-to-moderate main effects in factorial models. The design was not powered to make strong confirmatory claims about three-way interactions, mediation, moderation, or logistic regression subgroup effects. Therefore, interaction, mediation, moderation, and binary verification-accuracy analyses were treated as secondary or exploratory and interpreted after the primary factorial models.
Experimental materials
Eight simulated social media posts were created for the 2 × 2 × 2 design. Each post described a plausible disaster-related situation relevant to university students, including heavy rainfall, urban flooding, temporary transport disruption, campus safety reminders, or short-term emergency preparedness. The posts used a screenshot-style layout similar to public-facing Chinese social media posts and contained a source label, headline, brief body text, time stamp, engagement indicators, and simple interface elements. No real disaster location, real victim, real emergency event, real official announcement, or identifiable public agency post was used.
The accurate versions contained internally consistent information and proportionate protective advice, such as checking official updates, avoiding flooded roads, and preparing basic necessities. The misleading versions used exaggerated threat claims, vague attribution, unsupported predictions, or unverifiable urgency statements15. The official-source condition used a fictional verified emergency-management-style label, whereas the peer-source condition used an ordinary forwarded student or peer account. The source names, locations, timestamps, engagement indicators, and interface details were fictional. Full text descriptions of all eight stimuli are provided in Table 2 so that readers can evaluate the independence of the source, veracity, and tone manipulations.
Pilot testing of stimuli
A pilot test was conducted with 30 university students who met the same eligibility criteria but were not included in the main experiment. Participants rated each post on realism, emotional intensity, source credibility, wording clarity, and perceived accuracy using 5-point scales. They also indicated whether any post appeared to refer to a real recent disaster event or a real government announcement.
Pilot testing was used to check realism and manipulation clarity rather than to make misleading posts obviously false, because misinformation may still be misjudged even under attentive processing conditions16. Stimuli were revised when wording was unclear, emotional contrast was too weak, threat wording was excessive, or content appeared too similar to a real event.
Three manipulation-check items were retained in the main questionnaire. Participants identified whether the post appeared to come from an official or non-official source, whether it appeared accurate or doubtful, and how emotionally threatening the wording appeared. Participants were not excluded only because they misjudged the veracity of a post, since misjudgment was part of the measured phenomenon. Exclusion was applied only when failed manipulation recognition occurred together with invalid response behavior, such as a failed attention check or patterned answers.
Experimental procedure
The experiment was administered through an online survey platform accessible by smartphone or computer from March 3, 2026, to March 21, 2026. After opening the link, participants viewed an electronic informed consent page describing the general purpose, voluntary participation, estimated completion time, confidentiality procedures, possible mild discomfort from disaster-related content, and the right to withdraw before submission. Participants confirmed that they were at least 18 years old and agreed to participate before entering the questionnaire.
Participants first completed baseline measures, including demographic characteristics, daily social media use, prior disaster experience, perceived information literacy, and general trust in official emergency information. The platform then randomly assigned each participant to one of the eight experimental cells using embedded individual-level randomization. Participants viewed one simulated post and could not return to previous pages after completing the post-exposure section.
Immediately after exposure, participants completed measures of perceived risk, perceived credibility, emotional response, verification intention, and sharing intention. The verification task required concrete behavioral choices rather than only intention ratings, because young users' confidence in detecting misinformation does not necessarily translate into accurate identification in realistic digital environments17. Participants selected verification actions from a fixed list before proceeding to the debriefing page.
At the end of the questionnaire, all participants received an immediate debriefing. The page stated that the post was created for research purposes, noted that some versions contained misleading elements, and reminded participants to verify disaster-related claims through official emergency management channels before sharing. Participants were also instructed not to copy, save, or forward any stimulus material outside the study.
Measures
Risk perception was measured with five items covering perceived severity, likelihood, personal relevance, campus or community impact, and need for precaution. These items reflected the established view of risk perception as a subjective judgment involving perceived consequences, uncertainty, and controllability18. Responses were recorded on a 5-point Likert scale from 1 = strongly disagree to 5 = strongly agree. Higher scores indicated higher perceived risk.
Perceived credibility was measured with four items assessing whether the post appeared believable, accurate, trustworthy, and worth relying on. Verification intention was measured with four items assessing willingness to check official channels, search for supporting evidence, compare multiple sources, and delay forwarding before confirmation. Sharing intention was measured with three items assessing willingness to repost, forward to classmates, or remind others based on the post. Fact-checking intention was treated as a separate construct because prior social media research links it to news literacy and trust judgments19.
Information literacy was measured before exposure with six items covering source evaluation, cross-checking habits, recognition of misleading headlines, awareness of emotional framing, ability to distinguish official from non-official information, and willingness to consult authoritative sources. Prior disaster experience was measured with two items on personal experience of natural disasters or emergency disruption and previous online searches for disaster-related information. General trust in official emergency information was measured with three items.
Verification behavior was scored using a task-based multiple-response rule. Participants could select any of eight actions: checking an official emergency management website, searching a verified government account, comparing the claim with recognized news outlets, checking the university emergency notice page, asking an unspecified friend, reading comments only, forwarding first and checking later, or doing nothing20. The first four actions were coded as reliable verification actions, and each received one point; the latter four actions were coded as passive, unreliable, or premature-sharing responses and received zero points. The continuous verification behavior score, therefore, ranged from 0 to 4. A binary verification accuracy variable was coded as 1 when a participant selected at least one reliable verification action and did not select forwarding before verification. Variable definitions, item numbers, coding rules, and score construction are provided in Table 3.
Data quality control
Data quality control was completed before statistical analysis. Of 460 submitted or initiated responses, 34 were excluded, and 426 were retained for analysis. The exclusion sequence was fixed in advance: age ineligibility, non-student status, incomplete questionnaire, pilot-test participation, failed attention check, unrealistically short completion time, and patterned response. Completion time was considered unrealistically short if it was less than one-third of the median completion time. Straight-line responses were removed only when they appeared together with failed attention checks or unrealistically short completion times.
Duplicate or suspicious submissions were screened using non-identifying survey indicators, including repeated response patterns and abnormal submission timing. No names, student identification numbers, phone numbers, social media account names, facial images, precise locations, IP addresses, or device identifiers were exported for analysis. Cases with missing values on primary outcome variables were excluded from the main analysis. Because the questionnaire required completion of core experimental items before submission, missingness in primary variables was expected to be minimal. The screening criteria and operational definitions are listed in Table 4.
Ethical considerations
The study used limited concealment only for the presence of misleading elements in some posts. It did not conceal the general topic, task type, voluntary nature, confidentiality procedures, or possible mild discomfort. This concealment was necessary to preserve the validity of the verification task and was followed by immediate debriefing. The disaster-related materials avoided real locations, real victims, real emergency cases, graphic images, and real government announcements because irresponsible handling of emergency misinformation can increase anxiety and interfere with appropriate public response21.
The anonymized dataset contained only research variables and randomly generated participant codes. Raw survey exports, screened analytic data, codebook files, stimulus materials, and statistical syntax were stored separately on a password-protected computer accessible only to the research team. No personally identifiable information was exported, analyzed, or shared.
Statistical analysis
Statistical analysis was conducted using IBM SPSS Statistics version 27.0 and R version 4.3. Continuous variables were summarized as means and standard deviations. Categorical variables were summarized as frequencies and percentages. Internal consistency of multi-item scales was assessed using Cronbach's alpha, with α ≥ 0.70 treated as acceptable for group-level analysis.
Before hypothesis testing, manipulation checks were examined. Independent-samples t tests compared perceived emotional intensity between neutral and high-threat conditions. Chi-square tests examined recognition of official versus peer source cues. Perceived accuracy ratings were compared between accurate and misleading posts. Randomization balance was examined across the eight experimental cells before outcome models were fitted.
The primary outcomes were risk perception, perceived credibility, verification intention, sharing intention, and the continuous task-based verification behavior score. The binary verification accuracy variable, interaction terms beyond planned two-way contrasts, mediation, moderation, and robustness models were treated as secondary or exploratory. Factorial ANOVA and general linear models were used for continuous primary outcomes. Two-way and three-way interaction terms were included to examine whether the effect of information veracity varied by source cue or emotional tone. Logistic regression was used for binary verification accuracy, with odds ratios and 95% confidence intervals reported.
Mediation and moderation models were treated as secondary analyses. Perceived credibility was tested as a mediator between source cue and verification intention. Risk perception was tested as a mediator between emotional tone and sharing intention. Information literacy was tested as a moderator of the association between perceived credibility and sharing intention. Because misinformation responses are shaped by cognitive evaluation, trust judgment, and sharing motivation rather than message accuracy alone22, these models were interpreted after the primary factorial models. Indirect effects were estimated using 5,000 bootstrap resamples.
Robustness checks were conducted in three steps. First, the primary models were repeated after excluding participants who failed any manipulation-check item. Second, models were adjusted for gender, education level, daily social media use, prior disaster experience, and baseline trust in official emergency information. Third, verification behavior was analyzed both as a continuous score and as a binary accuracy variable. Statistical significance for primary outcomes was set at p < 0.05. Secondary analyses were interpreted by effect direction, confidence intervals, and consistency with the primary models rather than by marginal p-values alone. Effect sizes were reported as partial eta squared, Cohen's d, standardized coefficients, or odds ratios according to model type.
Data management and reproducibility
The anonymized analytic dataset contained one row per participant and one column per variable. It included participant code, experimental condition, demographic variables, baseline measures, manipulation-check items, post-exposure outcomes, verification behavior indicators, composite scores, and exclusion flags. Participant codes were randomly generated and were not linked to student identity.
Composite scores were calculated by averaging valid items within each scale after confirming item direction. Higher scores consistently indicated higher levels of the measured construct. Reverse coding was applied only when required by item wording. The data-cleaning log documented all exclusions, recoding decisions, composite-score calculations, and model specifications. The reproducibility package included the anonymized analytic dataset, codebook, stimulus descriptions, screening log, and statistical syntax. Transparent documentation of stimulus design, randomization, exclusion rules, scoring procedures, and analysis syntax was maintained to support reproducibility in misinformation and risk communication research23.