$$\rightleftharpoonup{xx}$$
$$\longleftharp{xx}$$,
$$\longrightharp{xx}$$,
As we all know, bladder cancer (BLCA) is one of the most aggressive and metastatic malignant tumors in the world, and there are still issues with the current biomarkers for bladder cancer, such as inaccuracy and so on1. To identify the prognosis of BLCA patients and predict the outcomes of BLCA patients, the search for biomarkers for bladder cancer and the establishment of prognostic models are of great importance. Although people have developed some research methods for biomarkers, most of these methods are currently limited to transcriptomics, which inevitably leads to heterogeneity in samples2. Moreover, studying transcriptomics data alone often fails to investigate the roles played by different cell subpopulations, as well as the functions of genes in different pathways at the single-cell level, which makes past research less precise and effective3. Considering that single-cell analysis might be challenging for beginners, an online platform has been introduced for single-cell analysis to help them learn the skills quickly4. Last but not least, even if a single gene is found to be a suitable biomarker through transcriptome analysis and single-cell analysis, there is no guarantee that the biomarker will be applicable across all cohorts, so it is necessary to construct prognostic models associated with the biomarker to make the conclusions more universally applicable5. The 101 machine learning algorithm refers to the construction of 101 prognosis models using a combination of 10 different machine learning algorithms, with the goal of identifying the optimal prognosis model. The inclusion of algorithms such as random forest, XGBoost, and SVM, which are more capable of handling complex data and exhibit greater stability, has resulted in this combined algorithm demonstrating remarkable stability. In order to make this prognostic model more accurate and relevant to the biomarker, a correlation analysis is first conducted between all the genes and the biomarker, then select around 20-30 genes based on requirements, and subsequently use a 101 machine learning algorithm to build the prognostic model, followed by a series of analyses to finalize the results.
Compared to the biomarkers predicted by simple transcriptomics analysis of the past6, the biomarkers of single-cell and transcriptomics analysis allow for an unbiased breakdown of tissues or samples into their basic cellular components. It enables the clear differentiation of different cell types (such as T cells, B cells, and macrophages) and the discovery of new subpopulations (such as exhausted T cells and regulatory T cells), as well as the capture of continuous dynamic processes (such as cell differentiation trajectories)4. It is akin to arranging fruits and milk separately, allowing for a clear visualization of each component and its quantity. Compared with single-cohort Cox models7, the prognostic model constructed using 101 machine learning is also more accurate and scientifically sound, as it can automatically learn about the complex, non-linear interactions between variables from the data. For instance, the impact of a specific genetic mutation might only be significant in patients of certain ages and tumor sizes. Machine learning models such as random forests and neural networks can automatically capture these high-order interactions without the need for manual specification8. The signaling key applicability constraints that the dataset used for analysis must include the vast majority of genes, with the number of genes not being too low, and the number of genes used for constructing prognostic models not being too high, generally maintained around 20-30, being optimal. The quality control standard of the single cell set should be greater than 1000, the UMI count per cell should be greater than 1000, and the gene number per cell should be greater than 5009.
Here, a step-by-step approach is provided for the identification of novel biomarkers from public transcriptomic datasets and single-cell datasets, taking the role of TRPM4 in BLCA as an example. Multiple research methods are employed and diverse datasets -- including the Cancer Genome Atlas-Bladder Cancer (TCGA-BLCA) dataset, single-cell dataset GSE145281, and BLCA datasets GSE32894 and GSE31684 -- to enhance the precision and applicability of biomarkers in oncology from multiple perspectives and to advance the study of TRPM4 in BLCA. The TCGA-BLCA dataset was obtained from the University of California, Santa Cruz Xena (UCSC Xena) website. GSE32894 and GSE31684 were downloaded from the Gene Expression Omnibus (GEO), and single-cell data GSE145281 was sourced from Tumor Immune Single-Cell Hub 2 (TISCH2). Additionally, data on BLCA molecular subtypes and treatment responses were extracted from supplementary spreadsheet files of a relevant article. Bioinformatics analysis, single-cell analysis, and machine learning are then conducted, establishing an integrated single-gene research methodology.