A subscription to JoVE is required to view this content. Sign in or start your free trial.

Method Article

A Reproducible Seurat-Based Protocol for Single-Cell RNA Sequencing Analysis of Peripheral Blood Mononuclear Cell CD4+ T Cells During Malaria Reinfection

106 views

DOI:

10.3791/70858

July 31st, 2026

In This Article

Summary

Here, we present a reproducible Seurat-based protocol for the analysis of single-cell RNA sequencing data from peripheral blood mononuclear cell CD4⁺ T cells to characterize transcriptional heterogeneity and functional immune programs during malaria reinfection. This protocol enables consistent identification, comparison, and biological interpretation of dynamic CD4⁺ T-cell states and immune responses across conditions.

Abstract

Here, we present a reproducible Seurat-based protocol to analyze PBMC CD4⁺ T-cell single-cell RNA sequencing data across malaria reinfection timepoints. This protocol demonstrates a reproducible Seurat-based workflow for analyzing PBMC CD4⁺ T-cell scRNA-seq data across malaria reinfection timepoints, using representative publicly available datasets to demonstrate its application: a Plasmodium-specific TCR-transgenic CD4⁺ T-cell dataset (GSE233703) and a polyclonal CD4⁺ T-cell dataset comparing reinfection-associated timepoints (GSE233713; D27₍₃₎ versus D30). The workflow includes standardized preprocessing, integration, clustering, and downstream transcriptomic analyses within a unified computational framework. The method enables systematic computation of module scores for predefined immune and CD4⁺ T-cell programs, identification of cluster-specific marker genes, timepoint-resolved differential expression analysis, and downstream Gene Ontology and KEGG pathway enrichment. Application of this workflow identifies distinct CD4⁺ T-cell functional states and reveals dynamic transcriptional changes across malaria reinfection timepoints. The protocol generates standardized visualization outputs and tabulated results and provides practical guidance on parameter selection and troubleshooting, facilitating consistent and reproducible analysis of CD4⁺ T-cell scRNA-seq datasets in malaria and related immunological contexts. This protocol enables reproducible and biologically interpretable analysis of immune responses and can be applied to similar single-cell datasets in immunological research.

Introduction

Malaria remains a major global health burden, with repeated infections shaping host immunity in complex and incompletely understood ways1. CD4⁺ T cells play a central role in antimalarial immune responses by coordinating effector cytokine production, supporting B-cell help, and regulating inflammation2,3. During malaria reinfection, CD4⁺ T cells undergo dynamic transcriptional reprogramming, reflecting shifts among effector, regulatory, memory, proliferative, and exhausted states that influence both parasite control and immunopathology4,5. Accurately resolving this heterogeneity is essential for understanding immune protection, immune dysfunction, and the durability of naturally acquired or vaccine-induced immunity to Plasmodium infection. Here, we present a reproducible Seurat-based protocol to analyze CD4⁺ T-cell scRNA-seq data across malaria reinfection timepoints.

Single-cell RNA sequencing (scRNA-seq) enables high-resolution characterization of immune heterogeneity and has identified parasite-responsive CD4⁺ T-cell subsets, exhaustion programs, and regulatory networks in malaria6,7,8,9. However, analytical variability in quality control, normalization, clustering, and integration can limit reproducibility and complicate cross-study comparisons10,11. However, a standardized and biologically guided workflow specifically optimized for analyzing CD4⁺ T-cell dynamics during malaria reinfection is lacking.

Existing computational frameworks for scRNA-seq analysis, including Seurat and Scanpy, provide comprehensive toolsets for preprocessing, clustering, and downstream interpretation of single-cell data12,13,14. Seurat, implemented in R, offers tightly integrated workflows for normalization, data integration, and visualization, including variance-stabilizing approaches such as SCTransform that improve signal detection in heterogeneous immune datasets13. Scanpy, implemented in Python, provides scalable solutions optimized for large datasets and efficient memory usage, making it particularly suitable for high-throughput or cloud-based analyses12,14. Despite these advances, there remains a need for standardized, reproducible workflows that explicitly address biological questions in infection settings while maintaining transparency, adaptability, and consistency across datasets. The present protocol addresses this gap by combining the robustness of Seurat-based preprocessing with structured biological interpretation tailored to CD4⁺ T-cell responses during malaria reinfection. Compared to existing general-purpose workflows, this protocol emphasizes reproducibility, biologically informed parameter selection, and consistent cross-timepoint analysis tailored to infection models.

A key feature of this protocol is its emphasis on reproducibility and practical usability. Quality-control thresholds are not fixed but are derived using data-adaptive approaches based on median absolute deviation, allowing thresholds for transcript complexity, sequencing depth, and mitochondrial content to scale with dataset-specific distributions. This design makes the workflow applicable across datasets of varying sizes, typically ranging from several thousand to tens of thousands of cells, and across a wide range of sequencing depths commonly encountered in droplet-based scRNA-seq experiments. Guidance embedded within the workflow supports appropriate parameter selection for dimensionality reduction, clustering resolution, and integration, ensuring that analyses remain both biologically meaningful and technically robust. Nevertheless, the workflow depends on data quality and sequencing depth and may require adjustment for datasets with extreme sparsity or batch effects.

This protocol is designed for intermediate to advanced users with basic familiarity in R and single-cell analysis, while remaining accessible to motivated beginners through its structured, stepwise implementation and fully reproducible outputs. The workflow generates standardized tables and figures at each stage, including quality-control summaries, clustering outputs, differential expression results, and enrichment analyses, thereby facilitating transparency, validation, and reuse in collaborative or multi-study contexts. This protocol is particularly suitable for studies investigating immune heterogeneity across timepoints or conditions in infection and immunology research.

The method provides a reproducible, end-to-end Seurat-based workflow for CD4⁺ T-cell scRNA-seq analysis across malaria reinfection timepoints. It integrates adaptive quality control, variance stabilization, dimensional reduction, clustering, and multi-sample integration when appropriate13,15. To enhance biological interpretability, the workflow incorporates immune and CD4⁺ T-cell subset gene module scoring to quantify functional programs including Th1, Tfh, Tr1, Treg, central memory, effector memory, proliferation, cytotoxicity, and exhaustion16,17,18. Complementary cluster marker identification, timepoint differential expression analysis, and pathway enrichment using Gene Ontology and KEGG databases are included to support robust annotation and comparison of T-cell states19,20. Although demonstrated using publicly available Plasmodium scRNA-seq datasets, this protocol is broadly applicable to other malaria reinfection models and immunological perturbations where reproducible and interpretable single-cell analysis of CD4⁺ T cells is required. Overall, this protocol provides a reproducible and biologically interpretable framework for single-cell analysis of CD4⁺ T-cell responses, supporting robust investigation of immune dynamics in malaria and related systems.

Access restricted. Please log in or start a trial to view this content.

Protocol

Ethics statement:

All data used in this study were obtained from publicly available datasets (GSE233703 and GSE233713). The original studies complied with institutional and ethical guidelines for animal experimentation. This study involved secondary bioinformatics analysis of publicly available and de-identified single-cell RNA sequencing datasets obtained from the Gene Expression Omnibus (GEO) repository (GSE233703 and GSE233713). No new human participants, clinical samples, or identifiable patient information were involved in this study. According to the institutional and national guidelines for research involving publicly available anonymized datasets, additional ethical approval and informed consent were not required for this bioinformatics analysis. The original studies associated with these datasets were conducted in accordance with relevant institutional ethical standards and applicable guidelines for biomedical research.

1. Reproducible Seurat-based CD4⁺ T-cell scRNA-seq analysis workflow for GSE233703 and GSE233713

NOTE: This workflow is provided as Supplementary File S1. Also refer to a schematic which gives an overview of the entire workflow (Figure 1).

  1. Define the scope and intended users before beginning the protocol
  2. Define the intended users
    1. Use this protocol if the user has intermediate to advanced familiarity with R and single-cell RNA sequencing (scRNA-seq) analysis.
    2. Apply this workflow to murine spleen CD4⁺ T-cell datasets generated in 10x-style matrix format. Use the stepwise structure and expected outputs to verify each stage before proceeding.
      NOTE: Ensure appropriate data handling practices and secure storage when working with large sequencing datasets.
  3. Define scRNA-seq
    1. Treat single-cell RNA sequencing (scRNA-seq) as a transcriptomic method²¹ that quantifies gene expression in individual cells.
    2. Use scRNA-seq to identify discrete cell states, transitional populations, and heterogeneous immune programs within complex tissues.
  4. Prepare software and package requirements
  5. Install core software
    1. Install R 4.2 or later.
    2. Open the RStudio interface. Set the working directory using setwd().
    3. Execute scripts sequentially using the source() function. Record the exact versions of R, RStudio, and the operating system in the project notes.
  6. Install R packages
    1. Install the required CRAN packages: Seurat, Matrix, tidyverse, patchwork, pheatmap, RColorBrewer, cluster, glmGamPoi, and ggplot2.
    2. Install the required Bioconductor packages: clusterProfiler, org.Mm.eg.db, enrichplot, and DESeq2.
    3. Install optional packages only if needed: DoubletFinder for doublet removal and monocle3 for trajectory analysis. Load all required packages at the start of the analysis session.
  7. Record versions
    1. Save package and session information at the end of the workflow using sessionInfo() or an equivalent function.
    2. Report key analysis components explicitly in the manuscript, including R version, Seurat version, DESeq2 version, and clusterProfiler version.
  8. Run example installation commands mentioned in the Supplementary file 2
  9. Confirm hardware and storage requirements
  10. Confirm minimum resources
    1. Use a workstation with at least 16 GB RAM, 4 CPU cores, and 20 GB free disk space for routine analysis of datasets containing several thousand to tens of thousands of cells.
  11. Confirm recommended resources
    1. Use 32 GB RAM or more for integrated analyses, repeated plotting, or optional doublet detection.
    2. Increase future.globals.maxSize if large objects or integrated datasets produce memory-related errors.
    3. Apply example memory setting
  12. Run the following commands to set memory options mentioned in the Supplementary file 2

2. Create the project structure

  1. Create the root project directory
    1. Create a project directory for the analysis.
    2. Create subdirectories named data/, scripts/, and results_spleen_cd4/.
  2. Use standardized output structure
    1. Ensure that the workflow writes output to the following directories:
      results_spleen_cd4/GSE233703/fig/
      results_spleen_cd4/GSE233703/tables/
      results_spleen_cd4/GSE233703/rds/
      results_spleen_cd4/GSE233713/fig/
      results_spleen_cd4/GSE233713/tables/
      results_spleen_cd4/GSE233713/rds/
      results_spleen_cd4/post_markers/
  3. Use consistent file naming
    1. Rename GEO input files to match the paths expected by the script.
    2. Use the following exact file names:
      data/GSE233703_matrix.mtx.gz
      data/GSE233703_genes.tsv.gz
      data/GSE233703_barcodes.tsv.gz
      data/GSE233713_d27_3_matrix.mtx.gz
      data/GSE233713_d27_3_features.tsv.gz
      data/GSE233713_d27_3_barcodes.tsv.gz
      data/GSE233713_d30_matrix.mtx.gz
      data/GSE233713_d30_features.tsv.gz
      data/GSE233713_d30_barcodes.tsv.gz
    3. Name optional metadata files as follows:
      data/GSE233703_cell_metadata.csv
      data/GSE233713_cell_metadata.csv
  4. Confirm metadata requirements
    1. Confirm that every metadata file contains a barcode column. Add optional columns such as sample, timepoint, and replicate when available.
    2. Use exact barcode matches between metadata and count matrices.
      Caution: Confirm that matrix triplets and metadata files exist at the expected paths before starting import.

3. Import count matrices and validate input integrity

  1. Read 10x-style matrices
    1. Read each matrix.mtx.gz file as a sparse matrix.
    2. Read the corresponding features (or genes) file and barcode file as tab-delimited tables.
    3. Assign gene symbols to matrix rows using the second column of the features file when available. Enforce unique gene symbols using make.unique().
    4. Assign barcode identifiers to matrix columns.’
  2. Run example import function
    1. Execute the code mentioned in the Supplementary file 2, to import the matrix and assign identifiers.
  3. Validate matrix integrity
    1. Verify that the number of matrix rows equals the number of features. Verify that the number of matrix columns equals the number of barcodes.
    2. Stop the workflow if any mismatch is detected.
      NOTE: If mismatches are detected, verify file integrity and ensure correct alignment of feature and barcode files before re-running the step.
      Checkpoint: Proceed only if row counts match features and column counts match barcodes.

4. Define marker and module panels

  1. Define malaria-relevant CD4⁺ T-cell panels
    1. Define named gene panels for: Th1, Tfh, Tr1, Treg, Tcm, Tem, Exhaustion, Proliferation, Cytotoxic, Activation_early, Interferon_response, Immune_regulation
    2. Store these panels in a named R list for downstream module scoring.
  2. Run the R code mentioned in Supplementary file 2 to define gene panels.
  3. Define marker-validation genes
    1. Define a separate validation panel containing canonical marker genes such as Foxp3, Bcl6, Cxcr5, Il21, Ifng, Ctla4, Pdcd1, Lag3, Tbx21, Tcf7, and Lef1.

5. Create Seurat objects and compute quality-control metrics

  1. Initialize Seurat objects
    1. Create a Seurat object for each dataset using min.cells = 3 and min.features = 0. Do not impose arbitrary feature cutoffs at import.
    2. Add dataset, sample, and timepoint metadata. Merge optional metadata using barcode matching.
  2. Run example Seurat object initialization
    1. Execute the R code mentioned in the supplementary file 2 to create a Seurat object and assign metadata.
  3. Compute quality-control metrics
    1. Calculate the mitochondrial transcript fraction using the murine prefix ^mt-.
    2. Quantify the following metrics
      nFeature_RNA
      nCount_RNA
      percent.mt
  4. Run the example code mentioned in the supplementary file 2.
  5. Visualize pre-filter quality control
    1. Generate violin plots for nFeature_RNA, nCount_RNA, and percent.mt. Generate feature-scatter plots for nCount_RNA versus nFeature_RNA and nCount_RNA versus percent.mt.
    2. Save pre-filter figures using standardized names such as QC_pre_filter_AllCells_vln.png and QC_pre_filter_AllCells_scatter.png.
      ​Caution: Expect broad pre-filter distributions with low-quality tails and possible high-count outliers.

6. Derive adaptive QC thresholds and filter low-quality cells

  1. Derive dataset-specific thresholds
    1. Log-transform nFeature_RNA + 1 and nCount_RNA + 1. Compute the median and median absolute deviation (MAD) for both transformed variables.
    2. Define the following thresholds:
      min_features = 10^(median - 3 × MAD) - 1
      max_features = 10^(median + 3 × MAD) - 1
      min_counts = 10^(median - 3 × MAD) - 1
      max_counts = 10^(median + 3 × MAD) – 1
    3. Define the mitochondrial threshold as the 95th percentile of percent.mt plus 3 × MAD, constrained between 5 % and 20 %.
    4. Ensure that dataset_id, sample_id, and out_dir are correctly specified before running the function.
      ​qc_thr <- derive_qc_thresholds(seu, dataset_id = "GSE233703", sample_id = "AllCells", out_dir = "results_spleen_cd4/GSE233703")
  2. Apply filtering
    1. Retain cells satisfying all adaptive criteria:
      nFeature_RNA >= min_features
      nFeature_RNA <= max_features
      nCount_RNA >= min_counts
      nCount_RNA <= max_counts
      percent.mt <= max_percent_mt
    2. Run the example code mentioned in the supplementary file 2.
  3. Save filtering outputs
    1. Save QC_thresholds_*.csv and cell_counts_summary_*.csv.
      PAUSE POINT: Save intermediate outputs and resume analysis from this step if needed.
    2. Generate and save post-filter QC violin and scatter plots.
      ​Checkpoint: Expect tighter post-filter distributions, removal of low-complexity cells, and reduction of extreme outliers.

7. Normalize data and perform PCA

  1. Normalize and variance-stabilize
    1. Normalize the RNA assay before cell-cycle scoring. Run cell-cycle scoring if enabled. Use SCTransform() for variance stabilization.
    2. Regress mitochondrial content only if biologically justified and explicitly enabled.
    3. Regress cell-cycle scores only if required for the study design.
  2. Run example normalization as mentioned in supplementary file 2.
  3. Define parameter values
    1. Use n_variable_features = 3000.
    2. Use dims_max_for_pca = 50.
    3. State these values explicitly in the manuscript.
  4. Run PCA and select principal components
    1. Run PCA on the normalized assay. Confirm successful PCA execution by inspecting variance explained and principal component loadings.
    2. Compute variance explained by each principal component.
  5. Select PCs using all three criteria:
    1. Retain PCs explaining at least 1 % variance,
    2. Ensure cumulative variance reaches approximately 80 %,
    3. Constrain final selection between 10 and 40 PCs.
    4. Save PCA_variance_table_*.csv, PCA_selection_rationale_*.csv, and PCA_Elbow_*.png.
  6. Run example PCA
    1. Example as mentioned in supplementary file 2.
      Checkpoint: Expect an elbow plot with a visible decrease in marginal variance gain after the selected cutoff.

8. Construct neighborhoods, select resolution, and cluster cells

  1. Build graph structure
    1. Construct a shared nearest neighbor’s graph using the selected PCs.
    2. Run an initial low-resolution clustering if doublet detection requires cluster labels.
  2. Optionally remove doublets
    1. Run DoubletFinder only if installed and compatible.
      ​NOTE: Perform doublet detection only for datasets with high cell counts where multiplet artifacts are expected.
    2. Recompute normalization and PCA after doublet removal.
  3. Select clustering resolution
    1. Evaluate resolutions 0.2, 0.4, 0.6, 0.8, 1.0, and 1.2.8.3.2
    2. Compute mean silhouette width for each tested resolution. Select the resolution with the highest silhouette score among solutions with at least two clusters.
    3. Save resolution_sweep_*.csv, resolution_selection_rationale_*.csv, and resolution_sweep_*.png.
  4. Run example resolution selection by executing the code mentioned in the Supplementary file 2
  5. Run UMAP and final clustering.
    1. Run UMAP using the selected PCs.
    2. Rebuild the nearest-neighbor graph. Cluster cells using the selected resolution.
    3. Save UMAP plots labeled by cluster and grouped by sample or timepoint. Execute the following code mentioned in the Supplementary file 2.
      Caution: Expect stable cluster separation and interpretable UMAP structure consistent with major immune states.

9. Perform SCT-based integration for GSE233713

  1. Prepare separate objects
    1. Create separate Seurat objects for D27_3 and D30.
    2. Apply quality control and filtering independently to each sample. Normalize each sample separately with SCTransform().
  2. Justify integration
    1. Use SCT-based integration to reduce technical differences between timepoints while preserving shared biological structure.
    2. Do not assume that integration is automatically beneficial. Validate it explicitly.
  3. Integrate samples
    1. Select integration features using SelectIntegrationFeatures(). Prepare objects with PrepSCTIntegration().
    2. Find anchors with FindIntegrationAnchors(normalization.method = "SCT"). Integrate datasets with IntegrateData(normalization.method = "SCT").
  4. Run example integration mentioned in the Supplementary file 2
  5. Validate integration
    1. Generate pre-integration and post-integration UMAP plots grouped by timepoint. Interpret improved mixing of cells across timepoints as evidence of successful integration.
    2. Compute neighbor mixing before and after integration. Calculate cluster composition by timepoint and generate stacked composition plots.
    3. Save the following outputs:
      UMAP_preintegration_*
      UMAP_postintegration_*
      integration_diagnostics_*
      cluster_timepoint_composition_*
      ​Caution: Expect reduced timepoint segregation after integration, increased neighbor mixing, and multi-timepoint contribution to most clusters without complete loss of biologically meaningful structure.

10. Annotate clusters and validate marker structure

  1. Score malaria-relevant modules
    1. Run AddModuleScore() for the predefined malaria CD4⁺ T-cell panels.
    2. Save cluster-level module means and module rankings.
  2. Run example module scoring mentioned in the supplementary file 2
  3. Assign predicted cluster labels
    1. Assign the top-ranked module to each cluster as the predicted label.
    2. Save:
      cluster_module_score_means_*
      cluster_module_score_rankings_*
      cluster_predicted_labels_*
  4. Validate clusters with canonical markers
    1. Run marker-validation DotPlots and FeaturePlots using the canonical marker panel.
    2. Save:
      DotPlot_marker_validation_*
      FeaturePlot_marker_validation_*
      ​Caution: Expect concordant expression of multiple canonical genes per functional state, not isolated single-gene signals.

11. Identify markers and perform differential expressions.

  1. Find cluster markers
    1. Run FindAllMarkers() using positive markers only.
    2. Save markers_all_clusters_*.csv.
  2. Execute the code mentioned in the supplementary file to run the example marker identification
  3. Use replicate-aware differential expression when available
    1. Check whether valid replicate metadata are available. If replicates are available, aggregate counts per replicate and condition and run pseudobulk DE with DESeq2.
    2. Save DE_pseudobulk_*.
  4. Use exploratory cell-level differential expression when replicates are absent
    1. If replicate metadata are absent or insufficient, run single-cell DE as exploratory analysis.
    2. Save DE_WARNING_* to document the exploratory status. Save exploratory results as DE_exploratory_celllevel_*.
  5. Generate global differential expression outputs
    1. Create volcano plots for DE results and save Volcano_*.
    2. Save significant DE results as filtered tables.
  6. Generate functional enrichment outputs
    1. Run GO Biological Process enrichment and save GO_BP_*. Run KEGG enrichment and save KEGG_*.
    2. Run GSEA using ranked log2 fold-changes and save GSEA_GO_*. Save corresponding barplots and dotplots.
  7. Generate cluster-specific differential expression outputs
    1. Run DE within each cluster between timepoints.
    2. Save DE_cluster_* and DE_cluster_specific_combined_*.
  8. Generate immune-focused differential expression outputs
    1. Extract DE results for selected immune genes such as Ifng, Cxcl10, Ctla4, Il10, Foxp3, Bcl6, Cxcr5, Pdcd1, Lag3, Havcr2, Il21, and Tbx21.
    2. Save DE_immune_focus_*.
    3. Generate and save the following outputs:
      DotPlot_selected_genes_by_timepoint_*
      Heatmap_immune_focus_*
      Caution Expect coherence between global DE, enrichment outputs, cluster-specific DE, and immune-focused signatures.

12. Save final outputs and archive the session

  1. Save Seurat objects
    1. Save final Seurat objects in .rds format for each dataset.
  2. Save session information
    1. Write sessionInfo() to a text file in the output directory.
  3. Execute the code mentioned in the Supplementary file 2 to run example session export
  4. Verify output completeness
    1. Confirm that expected fig/, tables/, and rds/ directories contain the corresponding files.
    2. Archive scripts, session information, and outputs together.
      Caution: Do not proceed to report writing until QC summaries, PCA outputs, resolution-selection outputs, integration diagnostics, module annotation files, DE outputs, and enrichment files are all present and internally consistent.

Access restricted. Please log in or start a trial to view this content.

Results

Sequencing quality and cell-level quality control (Antigen-specific (PcAS-reactive) TCR-transgenic CD4⁺ T cells (GSE233703))

Pre-filter QC distributions (Figure 2A) showed heterogeneous transcript complexity, with most cells displaying moderate gene and UMI counts and a smaller subset showing high-count outlier profiles consistent with potential multiplets. Scatterplots (Figure 2B) showed a strong positive relationshi...

Access restricted. Please log in or start a trial to view this content.

Discussion

This study presents a standardized and reproducible Seurat-based workflow for analyzing CD4⁺ T-cell transcriptional dynamics during malaria reinfection. The protocol integrates adaptive quality control, normalization, dimensionality reduction, clustering, dataset integration, marker validation, module scoring, and differential expression analysis within a unified computational framework. Together, these analytical steps enabled reproducible identification and interpretation of biologically meaningful CD4⁺ T-c...

Access restricted. Please log in or start a trial to view this content.

Disclosures

The authors have nothing to disclose.

Materials

List of materials used in this article
NameCompanyCatalog NumberComments
Name of Material / EquipmentCompany / SourceCatalog NumberComments / Description
10x Genomics–formatted count matricesNCBI GEON/AMatrix Market files (matrix.mtx, features.tsv, barcodes.tsv)
clusterProfiler (R package)BioconductorN/ARRID:SCR_016884; Functional enrichment analysis (GO, KEGG)
enrichplot (R package)BioconductorN/ARRID:SCR_017030; Visualization of enrichment analysis results
GEO Dataset GSE233703NCBI Gene Expression OmnibusGSE233703PcAS-specific TCR-transgenic CD4+ T-cell scRNA-seq dataset
GEO Dataset GSE233713NCBI Gene Expression OmnibusGSE233713Polyclonal CD4+ T-cell scRNA-seq dataset (D273 vs D30)
GitHub (optional)GitHub Inc.N/ARRID:SCR_002630; Version control and reproducibility (optional)
glmGamPoi (R package)BioconductorN/ARRID:SCR_021001; Accelerated SCTransform model fitting
Matrix (R package)CRANN/ARRID:SCR_008389; Sparse matrix handling for scRNA-seq data
Operating systemMicrosoft / Apple / LinuxN/AWindows 10+, macOS, or Linux supported
org.Mm.eg.db (R package)BioconductorN/ARRID:SCR_002643; Mouse gene annotation database
patchwork (R package)CRANN/ARRID:SCR_018787; Multi-panel figure assembly
PDF viewerAnyN/AViewing QC plots, UMAPs, and heatmaps
Personal computer or workstationAnyN/AMinimum 16–32 GB RAM recommended for integration
pheatmap (R package)CRANN/ARRID:SCR_016418; Heatmap visualization of gene expression
R Statistical Software (version ≥ 4.2)R Foundation for Statistical ComputingN/ARRID:SCR_001905; Core computational environment
RStudio DesktopPosit SoftwareN/ARRID:SCR_000432; Integrated development environment for R
Seurat (R package, v4 or later)Satija LabN/ARRID:SCR_016341; Single-cell RNA-seq analysis
tidyverse (R package suite)CRANN/ARRID:SCR_019186; Data manipulation and visualization

References

  1. World Health Organization. WHO malaria policy advisory group (MPAG) meeting report, 18–20 April 2023. Geneva: World Health Organization; 2023.
  2. Stevenson MM, Riley EM. Innate immunity to malaria. Nat Rev Immunol. 2004;4(3):169-80.
  3. Langhorne J, Ndungu FM, Sponaas AM, Marsh K. Immunity to malaria: more questions than answers. Nat Immunol. 2008;9(7):725-32.
  4. Perez-Mazliah D, Langhorne J. CD4 T-cell subsets in malaria: TH1/TH2 revisited. Front Immunol. 2015;5:671.
  5. Illingworth J, et al. Chronic exposure to Plasmodium falciparum is associated with phenotypic evidence of B and T cell exhaustion. J Immunol. 2013;190(3):1038-47.
  6. Tang F, et al. mRNA-Seq whole-transcriptome analysis of a single cell. Nat Methods. 2009;6(5):377-82.
  7. Lönnberg T, et al. Single-cell RNA-seq and computational analysis using temporal mixture modeling resolves TH1/TFH fate bifurcation in malaria. Sci Immunol. 2017;2(9):eaal2192.
  8. Butler NS, et al. Therapeutic blockade of PD-L1 and LAG-3 rapidly clears established blood-stage Plasmodium infection. Nat Immunol. 2012;13(2):188-95.
  9. Soon MS, Haque A. Recent insights into CD4+ Th cell differentiation in malaria. J Immunol. 2018;200(6):1965-75.
  10. Slovin S, et al. Single-cell RNA sequencing analysis: a step-by-step overview. RNA Bioinformatics. 2021:343-65.
  11. Vieth B, Parekh S, Ziegenhain C, Enard W, Hellmann I. A systematic evaluation of single cell RNA-seq analysis pipelines. Nat Commun. 2019;10(1):4667.
  12. Wolf FA, Angerer P, Theis FJ. SCANPY: large-scale single-cell gene expression data analysis. Genome Biol. 2018;19(1):15.
  13. Hafemeister C, Satija R. Normalization and variance stabilization of single-cell RNA-seq data using regularized negative binomial regression. Genome Biol. 2019;20(1):296.
  14. Stuart T, et al. Comprehensive integration of single-cell data. Cell. 2019;177(7):1888-902.
  15. Satija R, Farrell JA, Gennert D, Schier AF, Regev A. Spatial reconstruction of single-cell gene expression data. Nat Biotechnol. 2015;33(5):495-502.
  16. Crotty S. T follicular helper cell differentiation, function, and roles in disease. Immunity. 2014;41(4):529-42.
  17. Wherry EJ, Kurachi M. Molecular and cellular insights into T cell exhaustion. Nat Rev Immunol. 2015;15(8):486-99.
  18. Belkaid Y, Rouse BT. Natural regulatory T cells in infectious disease. Nat Immunol. 2005;6(4):353-60.
  19. Ashburner M, et al. Gene ontology: tool for the unification of biology. Nat Genet. 2000;25(1):25-9.
  20. Kanehisa M, Goto S. KEGG: kyoto encyclopedia of genes and genomes. Nucleic Acids Res. 2000;28(1):27-30.
  21. Gulati GS, et al. Profiling cell identity and tissue architecture with single-cell and spatial transcriptomics. Nat Rev Mol Cell Biol. 2025;26(1):11-31.
  22. Schofield L, Grau GE. Immunological processes in malaria pathogenesis. Nat Rev Immunol. 2005;5(9):722-35.
  23. Plebanski M, Hill AV. The immunology of malaria infection. Curr Opin Immunol. 2000;12(4):437-41.
  24. Crotty S. T follicular helper cell biology: a decade of discovery and diseases. Immunity. 2019;50(5):1132-48.
  25. Vinuesa CG, Linterman MA, Yu D, MacLennan IC. Follicular helper T cells. Annu Rev Immunol. 2016;34:335-68.
  26. Wherry EJ. T cell exhaustion. Nat Immunol. 2011;12(6):492-9.
  27. Maizels RM, Smith KA. Regulatory T cells in infection. Adv Immunol. 2011;112:73-136.
  28. Choudhary S, Satija R. Comparison and evaluation of statistical error models for scRNA-seq. Genome Biol. 2022;23(1):27.
  29. Li M, et al. Rediscovering publicly available single-cell data with the DISCO platform. Nucleic Acids Res. 2025;53(D1):D932-8.
  30. Islam MT, Xing L. Cartography of genomic interactions enables deep analysis of single-cell expression data. Nat Commun. 2023;14(1):679.
  31. Luecken MD, Theis FJ. Current best practices in single-cell RNA-seq analysis: a tutorial. Mol Syst Biol. 2019;15(6):e8746.

Access restricted. Please log in or start a trial to view this content.

Reprints and Permissions

Tags

Seurat WorkflowPBMC AnalysisTranscriptomic AnalysisCluster Marker GenesDifferential ExpressionGene OntologyKEGG Pathway

This article has been published

Video Coming Soon