This study was approved by the Biomedical Ethics Committee of Jishou University (Approval No. JSDX-2023-0034). All participants were full-time undergraduate students aged 18–25 years and provided written informed consent before participation. Permission was obtained from the participating universities. The survey data were used only for academic research and were analyzed anonymously. The study was carried out in compliance with the Declaration of Helsinki and local institutional requirements.
Participant recruitment and sampling design
A cross-sectional survey was conducted among undergraduates from three comprehensive universities in southern China between September and December 2025 using a multistage stratified cluster random sampling design. In the first stage, three comprehensive universities were selected from a list of public comprehensive universities in southern China using stratified random sampling, with strata defined by provincial administrative region to ensure geographic representation. In the second stage, within each selected university, two to three intact classes from each grade level (freshman through senior) were randomly selected as cluster sampling units using computer-generated random numbers, yielding a total of 28 classes across the three universities. All students in the selected classes were invited to participate, as shown in Figure 1. The inclusion criteria were as follows: (1) full-time undergraduate status; (2) age between 18 and 25 years; and (3) voluntary participation with written informed consent. The exclusion criteria were: (1) recent severe physical disease or contraindications to exercise, such as fracture or severe heart disease; (2) current psychiatric medication use, such as treatment for severe depression or schizophrenia; and (3) questionnaire completion time of less than 180 s or obvious logical inconsistencies, such as selecting the same option for all items. A total of 1,500 questionnaires were distributed. Of the 80 invalid questionnaires excluded during data cleaning, 24 were excluded due to completion time below 180 s, 20 due to logical inconsistencies (e.g., identical responses across all items or contradictory answers), 18 due to age outside the 18–25 year range, 11 due to excessive missing data (more than 20% of items), and 7 due to self-reported severe physical disease or current psychiatric medication use. This left 1,420 valid questionnaires for analysis, with an effective response rate of 94.67%.

Figure 1: Study design and analytical workflow. Schematic overview of participant recruitment and screening (1,500 questionnaires distributed; 80 excluded; N = 1,420 included), assessment with the PSQI and IPAQ-SF, latent profile analysis (one- to five-profile models compared using AIC, BIC, aBIC, entropy, LMRT, and BLRT), assignment to four sleep profiles, bivariate chi-square tests, multinomial logistic regression with gender as a covariate, and sensitivity analyses (full-covariate adjustment and cluster-robust standard errors). Please click here to view a larger version of this figure.
Demographic and lifestyle assessment
A self-designed questionnaire was utilized to obtain demographic and sociological characteristics, including gender (male/female), age (in years), grade (freshman/sophomore/junior/senior), place of origin (urban/rural), body mass index (BMI), smoking history, drinking history, and average daily screen time. BMI was calculated as body weight in kilograms divided by the square of height in meters (BMI = weight [kg] / height [m]2). Self-reported height and weight were screened for implausible values prior to analysis; cases with height <140 cm or >200 cm, weight <30 kg or >150 kg, or BMI <12 kg/m2 or >40 kg/m2 were flagged and verified against original records, with confirmed outliers excluded. Smoking status was categorized as never, occasional (smoking on some days but not daily), or regular (daily smoking), based on self-reported smoking behavior over the previous month. Alcohol consumption was categorized as never, occasional (drinking less than once per week), or regular (drinking at least once per week), over the previous month. Average daily screen time was assessed by asking participants to report the total number of hours per day spent on smartphones, computers, tablets, and television over the previous month, recorded as a continuous variable. The questionnaire was pilot-tested among 30 undergraduate students before formal data collection to assess item clarity and completion time; minor wording adjustments were made based on pilot feedback. Previous studies have suggested that these variables may confound the association between sleep and physical activity30,31. All instruments, software, and materials used in this study are listed in the Table of Materials.
Sleep quality assessment
The PSQI, developed by Buysse et al.9, was used to assess participants’ sleep quality over the previous month. The scale includes 19 self-rated items that form seven components: subjective sleep quality, sleep onset latency, sleep duration, habitual sleep efficiency, sleep disturbances, use of hypnotic medication, and daytime dysfunction. Each component is scored from 0 (no difficulty) to 3 (severe difficulty). The scores of these seven continuous components were used as observed indicators for LPA. The PSQI has demonstrated good reliability and validity among college student populations in China and internationally30. In the present study, Cronbach’s alpha for the scale was 0.83.
Physical activity assessment
The IPAQ-SF, developed by Craig et al.32, was used to assess participants’ physical activity over the previous seven days. The questionnaire records the frequency (days/week) and duration (minutes/day) of vigorous physical activity, moderate-intensity physical activity, and walking. Metabolic equivalent of task (MET) values were calculated according to the IPAQ scoring protocol: walking MET-min/week = 3.3 × duration × frequency; moderate-intensity MET-min/week = 4.0 × duration × frequency; and vigorous-intensity MET-min/week = 8.0 × duration × frequency. The MET values for the three activity types were summed to obtain total physical activity. Participants were classified into high PA, moderate PA, and low PA according to the IPAQ-SF scoring criteria validated in college students33. High PA was defined as vigorous activity on at least three days with a total physical activity level of at least 1,500 MET-min/week, or any combination of walking, moderate-intensity, or vigorous-intensity activities on seven or more days with a total physical activity level of at least 3,000 MET-min/week. Moderate PA was defined as vigorous activity on at least three days for at least 20 minutes per day, or moderate-intensity activity and/or walking on at least five days for at least 30 minutes per day, or any combination of activities reaching at least 600 MET-min/week. Low PA was defined as not meeting the criteria for moderate or high PA. Data cleaning followed the official IPAQ scoring protocol: participants reporting more than 16 h of total activity per day were excluded; per-category daily durations were truncated at 180 min before MET-minute calculation; and cases with logically inconsistent responses were removed during data cleaning.
Statistical analysis
First, conduct LPA using latent profile modeling software (see Table of Materials)34. Of the 1,420 participants, 33 (2.3%) had partial missing data on the PSQI, with 69 of 9,940 component-level responses missing (0.69%; i.e., 1,420 participants × seven PSQI component scores). The full-information maximum likelihood (FIML) method was used to handle these missing values under the missing-at-random assumption. Models with one to five profiles were estimated sequentially. Model fit was evaluated using the Akaike information criterion (AIC), Bayesian information criterion (BIC), sample-size adjusted BIC (aBIC), entropy, the Lo-Mendell-Rubin likelihood ratio test (LMRT), and the bootstrap likelihood ratio test (BLRT). Lower AIC, BIC, and aBIC values indicate better relative fit, whereas entropy values closer to 1 indicate higher classification accuracy. The LMRT and BLRT were used to determine whether a model with k profiles fit significantly better than a model with k − 1 profiles. The optimal profile solution was selected by jointly considering statistical fit indices, entropy, class size (no profile containing less than 5% of the sample), and substantive interpretability, following established guidelines for latent profile analysis11,35. After the optimal profile solution was selected, individuals’ most likely class membership was exported to statistical analysis software (see Table of Materials) for subsequent analyses. Pearson chi-square tests were used to examine differences in demographic characteristics and physical activity levels across sleep profiles. Bivariate chi-square tests showed that only gender and physical activity level differed significantly across the four sleep profiles (both p < 0.001), whereas age, grade, BMI, smoking, drinking, and screen time did not (all p > 0.05; see Supplementary Table 1). Therefore, only gender was included as a covariate in the multinomial logistic regression model, consistent with the principle of parsimony. A sensitivity analysis adjusting for all collected covariates yielded the same pattern and significance of results (Supplementary Table 2). Multinomial logistic regression was performed with physical activity level as the dependent variable (low PA = 0, moderate PA = 1, high PA = 2) and sleep profile as the core independent variable. Odds ratios (ORs) and 95% confidence intervals (CIs) were calculated. All tests were two-sided, and statistical significance was set at p < 0.05. Because participants were recruited through intact classes nested within universities, we conducted a sensitivity analysis using the complex-samples module of the statistical software (see Table of Materials), specifying university and class as clustering variables, to obtain design-adjusted standard errors. The results were consistent with those from the standard multinomial logistic regression (Supplementary Table 3). The de-identified raw dataset with English variable labels and a codebook is provided as Supplementary File 1.