This study applied Mendelian randomization to evaluate the potential association between age at first sexual intercourse and HIV infection risk using genetic instrumental variables to reduce confounding and strengthen causal inference.
Research Article
This study applied Mendelian randomization to evaluate the potential association between age at first sexual intercourse and HIV infection risk using genetic instrumental variables to reduce confounding and strengthen causal inference.
This study applied a two-sample Mendelian randomization framework to investigate the potential association between age at first sexual intercourse (AFS) and the risk of Human Immunodeficiency Virus (HIV) infection. Summary-level Genome-Wide Association Study (GWAS) data from European populations were analyzed, including 214,547 individuals for AFS and 357 HIV cases with 218,435 controls from the FinnGen R5 dataset. Independent single-nucleotide polymorphisms significantly associated with AFS were selected as instrumental variables following linkage disequilibrium clumping and instrument-strength assessment. Causal estimates were evaluated using inverse-variance weighting (IVW), MR-Egger regression, weighted median, weighted mode, and simple mode. Cochran’s Q test, MR-Egger intercept analysis, MR-PRESSO assessment, and leave-one-out sensitivity analyses were performed to evaluate heterogeneity, pleiotropy, and robustness. The IVW analysis suggested that genetically predicted later AFS was associated with reduced HIV infection risk (OR = 0.192, 95% CI = 0.062–0.592, P = 0.004), whereas earlier sexual debut corresponded to increased HIV susceptibility. Directionally consistent findings across multiple Mendelian randomization methods supported the stability of the observed association. However, the findings should be interpreted cautiously because the HIV outcome analysis relied on a single dataset with a limited number of HIV cases. These results support a potential association between earlier sexual debut and HIV susceptibility and demonstrate the utility of Mendelian randomization for investigating behavioral risk factors associated with infectious disease outcomes.
Human reproductive behaviors, including age at first sexual intercourse (AFS), number of sexual partners, and age at first childbirth, are closely associated with social, behavioral, and health outcomes1,2. Previous studies have demonstrated a trend toward earlier AFS among adolescents3. For example, in the United Kingdom, the median age at first heterosexual intercourse is approximately 17 years, and an increasing proportion of adolescents report sexual debut before 16 years of age2. Earlier sexual debut has also been associated with a greater likelihood of high-risk sexual behaviors, substance use, adverse mental health outcomes, and sexually transmitted infections in observational studies4.
Human Immunodeficiency Virus (HIV), the causative agent of acquired immunodeficiency syndrome (AIDS), remains a major global public health challenge5. Recent global estimates indicate that nearly 39 million individuals are living with HIV worldwide6. In China, the incidence of newly diagnosed HIV infection has continued to increase in recent years7. Several observational studies have suggested that earlier sexual debut is associated with an increased risk of HIV infection and other sexually transmitted infections, particularly among men who have sex with men (MSM)8,9. HIV-positive MSM have also been reported to experience their first same-sex intercourse at younger ages compared with bisexual men10. However, AFS should be interpreted primarily as a behavioral and social indicator reflecting patterns of sexual exposure and risk-taking behavior rather than as a direct biological determinant of HIV infection.
Although observational studies provide important epidemiological evidence, these associations may be influenced by residual confounding, socioeconomic factors, behavioral heterogeneity, and reverse causation. Mendelian randomization (MR) analysis uses genetic variants as instrumental variables to strengthen causal inference by reducing confounding and reverse-causation bias11. MR is particularly valuable for evaluating behavioral exposures associated with infectious disease outcomes because randomized controlled trials involving sexual behavior are generally not feasible or ethically appropriate. Previous investigations examining the relationship between AFS and HIV risk have been predominantly observational, and relatively few studies have applied genetic epidemiological approaches to determine whether the observed association reflects a causal relationship independent of broader behavioral or socioeconomic influences.
Access restricted. Please log in or start a trial to view this content.
All original Genome-Wide Association Study (GWAS) investigations included in this analysis obtained written informed consent from participants and received approval from their respective institutional ethics committees. Because the present study used publicly available de-identified GWAS summary statistics, additional institutional review board approval was not required. The research tools used in this protocol are listed in Table of Materials.
1. Data
Single-nucleotide polymorphism (SNP) summary statistics were obtained from publicly accessible GWAS databases for both exposure and outcome datasets. GWAS summary statistics for age at first sexual intercourse (AFS) were obtained from a 2021 GWAS meta-analysis comprising 214,547 individuals of European ancestry from the UK Biobank (UKB)12. AFS was treated as a continuous exposure variable. Participant-reported sexual behavior data collected through anonymous interviews were used, excluding individuals younger than 12 years of age. Responses regarding sexual intercourse history and age at first sexual intercourse were extracted as described in the original GWAS study12.
GWAS summary statistics for Human Immunodeficiency Virus (HIV) infection were obtained from the FinnGen consortium R5 release, which included 357 HIV cases and 218,435 controls13. AFS was designated as the exposure variable, and HIV infection was defined as the outcome variable. Only datasets involving individuals of European ancestry were included to minimize population stratification bias. The overall analytical workflow for GWAS data extraction, SNP filtering, harmonization, Mendelian randomization analyses, sensitivity testing, and output interpretation is summarized in Figure 1.
2. Study design
A two-sample Mendelian randomization (MR) framework based on random chromosomal allocation during meiosis was applied14. Genetic variants associated with AFS were used as instrumental variables to estimate the relationship between AFS and HIV infection risk. The MR analysis was conducted under three core assumptions: the selected SNPs were strongly associated with AFS, consistent with the relevance assumption; the SNPs were independent of potential confounding variables associated with HIV infection risk, consistent with the independence assumption; and the SNPs influenced HIV infection exclusively through AFS without alternative causal pathways, consistent with the exclusion restriction assumption (Figure 2)14.
Instrument strength was evaluated using genome-wide significance thresholds and F-statistics. Genome-wide significant SNPs associated with AFS were selected using stringent linkage disequilibrium clumping thresholds, and SNPs with F-statistics < 10 were excluded to minimize weak instrument bias and strengthen the relevance assumption. Potential pleiotropy and confounding effects were evaluated using sensitivity analyses. The independence and exclusion restriction assumptions were further assessed through harmonization procedures, confounder screening, MR-Egger intercept testing, Cochran’s Q heterogeneity analyses, MR-PRESSO outlier assessment, and leave-one-out sensitivity analyses to reduce the likelihood of horizontal pleiotropy and residual confounding. Additional sensitivity analyses were performed to identify pleiotropic effects and validate MR assumptions15. MR analyses were conducted to evaluate whether genetically predicted earlier AFS was associated with HIV infection risk.
3. Selection of instrumental variables
Rigorous quality-control procedures were implemented before SNP selection. SNPs significantly associated with AFS at the genome-wide significance threshold (P < 5 × 10⁻8) were extracted using the extract_instruments() function in the TwoSampleMR package with linkage disequilibrium thresholds of r2 < 0.001 and a clumping distance > 10,000 kb16. F-statistics were calculated for all selected SNPs, and weak instrumental variables with F < 10 were excluded17.
The F-statistic was calculated as follows:

where: 
In these equations, N denotes the sample size of the selected dataset, k denotes the number of SNPs used for MR analysis, β denotes the SNP effect estimate on AFS, SD denotes the standard deviation of β, and MAF denotes the minor allele frequency. SNPs meeting all predefined criteria were retained as final instrumental variables for MR analysis.
4. Removal of confounding and palindromic SNPs
All selected SNPs were reviewed for potential associations with confounding traits before harmonization. SNPs associated with HIV-related phenotypes or potential confounding traits with r2 > 0.80 were excluded23,24. Exposure and outcome datasets were harmonized using the harmonise_data() function, and palindromic SNPs with intermediate allele frequencies were removed to prevent strand ambiguity23.
Palindromic SNPs were defined as variants containing A/T or G/C alleles with intermediate allele frequencies ranging from 0.01 to 0.3024.
5. Causal effect estimation
MR analyses were performed using the mr() function with inverse-variance weighting (IVW), MR-Egger regression, weighted median, weighted mode, and simple mode to estimate the relationship between AFS and HIV infection25. Consistency across MR methods was evaluated to assess the robustness of causal estimates and potential pleiotropic bias. IVW was used as the primary analytical method because it combined SNP-specific Wald ratios through meta-analysis24.
Causal effect estimates were reported as odds ratios (ORs), beta coefficients (β), and 95% confidence intervals (CIs). Heterogeneity and sensitivity analyses were subsequently conducted. Cochran’s Q statistics were calculated using the mr_heterogeneity() function, and leave-one-out analyses were performed using the mr_leaveoneout() function to determine the influence of individual SNPs on causal estimates26,27.
MR-Egger intercept testing was performed using the mr_pleiotropy_test() function, and MR-PRESSO global testing was conducted to assess horizontal pleiotropy and identify outlier SNPs. Corrected estimates were generated following outlier removal using MR-PRESSO procedures28.
6. Statistical analysis
All statistical analyses were conducted using R software (version 4.1.0; R Foundation for Statistical Computing, Vienna, Austria) with the TwoSampleMR, LDlinkR, devtools, and MR-PRESSO packages for SNP extraction, linkage disequilibrium clumping, harmonization, Mendelian randomization analysis, and sensitivity testing. FinnGen (RRID not available), MR-PRESSO (RRID not available), and LDlinkR (RRID not available) were used to support the analytical workflow. All statistical tests were two-sided, and statistical significance was defined as P < 0.05.
Statistical power was estimated using the web-based mRnd calculator based on sample size and instrumental variable effect estimates29. Final analytical outputs included OR estimates, 95% confidence intervals, heterogeneity statistics, pleiotropy assessments, and sensitivity analyses, which were collectively interpreted to evaluate the robustness and consistency of the association between genetically predicted AFS and HIV infection risk. All analyses were performed using established functions from the TwoSampleMR, LDlinkR, devtools, and MR-PRESSO packages in R software. No custom analytical scripts were developed for the present study.
Access restricted. Please log in or start a trial to view this content.
The Mendelian randomization (MR) analyses suggested a potential association between genetically predicted age at first sexual intercourse (AFS) and HIV infection risk. The observed association remained directionally consistent following sensitivity analyses and outlier evaluation using the MR-PRESSO approach. The inverse variance weighting (IVW) analysis demonstrated that genetically predicted later AFS was associated with lower HIV infection risk (OR = 0.192, 95% CI = 0.062–0.592, P = 0.004). This estimate sugg...
Access restricted. Please log in or start a trial to view this content.
This study utilized publicly available GWAS summary statistics and applied a two-sample Mendelian randomization (MR) framework to evaluate the potential association between genetically predicted age at first sexual intercourse (AFS) and HIV infection risk. The findings suggested that genetically predicted earlier AFS was associated with increased HIV susceptibility, indicating a potential inverse relationship between AFS and HIV infection risk. Consistent findings across inverse variance weighting (IVW), weighted me...
Access restricted. Please log in or start a trial to view this content.
The authors declare no conflicts of interest.
The authors acknowledge no additional funding sources or institutional support for this study.
Access restricted. Please log in or start a trial to view this content.
| Name | Company | Catalog Number | Comments |
|---|---|---|---|
| devtools Package | R Foundation for Statistical Computing, Vienna, Austria | R package | Used for package installation, data processing, and preparation of the Mendelian randomization workflow. |
| FinnGen GWAS Dataset | FinnGen Consortium | N/A | Outcome GWAS resource used for HIV infection summary statistics. |
| GWAS Summary Data | UK Biobank Consortium, FinnGen Consortium | Public GWAS datasets | Publicly accessible GWAS summary statistics used for SNP extraction and Mendelian randomization analyses. |
| LDlinkR Package | National Cancer Institute (NCI), National Institutes of Health (NIH), USA | R package | Used for linkage disequilibrium assessment and SNP validation. |
| mRnd Calculator | Web-based tool | https://shiny.cnsgenomics.com/mRnd/ | Used for statistical power calculation in Mendelian randomization analysis. |
| MR-PRESSO Package | GitHub repository | N/A | Used for pleiotropy assessment and outlier correction in Mendelian randomization analysis. |
| R Software | R Foundation for Statistical Computing, Vienna, Austria | Version 4.1.0; RRID:SCR_001905 | Used for statistical analysis and Mendelian randomization workflow execution. |
| SNP Data | UK Biobank Consortium, FinnGen Consortium | Public SNP datasets | SNPs used as instrumental variables for Mendelian randomization analysis. |
| TwoSampleMR Package | MR-Base platform | RRID:SCR_018317 | Used for two-sample Mendelian randomization analysis, harmonization, and sensitivity testing. |
| UK Biobank GWAS Dataset | UK Biobank Consortium | N/A | Exposure GWAS resource used for age at first sexual intercourse (AFS) summary statistics. |
Request permission to reuse the text or figures of this JoVE article
Request Permission