Blog · Research · SeNMo · 10 min read

SeNMo: learning pan-cancer prognosis from sparse multi-omics data

A self-normalizing network trained on five omics layers plus clinical variables from TCGA, then adapted to predict survival, cancer type and TLS ratio.

Overview figure of the SeNMo self-normalizing multi-omics framework
Overview of SeNMo, from the preprint. Reproduced from Waqas et al. (2024), arXiv, CC BY 4.0.
Paper authors
Asim Waqas, Aakash Tripathi, Sabeen Ahmed, Ashwin Mukund, Hamza Farooq, Joseph O. Johnson, Paul A. Stewart, Mia Naeini, Matthew B. Schabath, Ghulam Rasool
Affiliations
  • Moffitt Cancer Center and Research Institute
  • University of South Florida
  • University of Minnesota
  • Huntsman Cancer Institute, University of Utah
Published
International Journal of Molecular Sciences · Jul 30, 2025Co-developed SeNMo; co-first author
Contents
  1. Why multi-omics is hard
  2. The approach
  3. How it works
  4. One vector per patient
  5. Self-normalizing blocks
  6. Training objective
  7. Data and evaluation
  8. Results
  9. Survival across modalities
  10. External cohorts
  11. Cancer type and TLS
  12. What we learned
  13. The broader survival thread
  14. Resources

A tumor’s molecular profile is spread across several assays: RNA sequencing, DNA methylation arrays, miRNA sequencing, protein arrays and somatic mutation calls. Each is high-dimensional, each has its own scale, and in real cohorts many patients are missing one or more of them. Yet the questions people ask are simple: how long is this patient likely to live, where did the tumor start, and what does the immune environment look like?

SeNMo, our self-normalizing multi-omics network, tries to answer those questions with one model. It is a deep network trained on five omics layers plus four clinical variables from 33 cancer types in The Cancer Genome Atlas (TCGA), and the same backbone is then fine-tuned for other cohorts and tasks.1 Asim Waqas led the study as first and corresponding author; I co-developed SeNMo and share first authorship with him.

In short

  • Problem: omics data are wide, noisy and often incomplete, and most models cover one data type or one cancer.
  • What we built: SeNMo, one network trained on 80,697 molecular and clinical features from 33 TCGA cancer types.
  • Survival: a concordance index (C-index) of 0.758 on 2,754 held-out patients, where 0.5 is chance and 1.0 is a perfect ranking.
  • Other tasks: 99.8% accuracy on 33-way cancer type classification and, after fine-tuning, tertiary lymphoid structure (TLS) predictions that did not differ significantly from manual annotations.

Why multi-omics is hard

Multi-omics is a textbook “big P, small n” problem. TCGA methylation data alone has 485,576 CpG sites per sample, and gene expression has 60,483 features, while individual cancer types in our TCGA cohort range from 51 to 1,260 cases. The data also carry correlated features, different measurement scales, technical noise and missing values, within and across datasets.

Standard architectures each struggle here. Plain feedforward networks overfit and are sensitive to perturbations during training. CNNs rely on a spatial-invariance assumption and a fixed input size, and they handle sparse omics inputs inefficiently. Transformers are built around attention over sequences, which suits highly sparse molecular vectors poorly. Many existing multi-omics models also cover only one omics layer or one cancer type, or need heavy feature engineering for each modality.

The approach

We combined three ideas. First, use our MINDS framework to assemble a machine-learning-ready pan-cancer cohort from TCGA through the GDC and the UCSC Xena portal.2 Second, reduce every modality with simple, reproducible filters and concatenate the result into one fixed-length vector per patient (early fusion). Third, learn from that vector with a network whose layers keep activations close to zero mean and unit variance, so a deep model can train stably on very wide, sparse input.

Five-panel figure. A: horizontal bar chart of case counts for 33 TCGA cancer types, from breast invasive carcinoma at about 1,250 down to diffuse large B-cell lymphoma. B: per-patient survival lines up to about 30 years, colored by cancer type. C: data sources (NCI Genomic Data Commons, UCSC Xena, cBioPortal) feed the MINDS cloud system, which supplies DNA methylation, gene expression, miRNA, mutation, protein and clinical data to the SeNMo model with hidden layers 1024, 512, 256, 128, 48, 48, 48 and a prediction head for survival, TLS ratio, cancer type and stage. D: per-modality preprocessing from 566,520 raw features to 80,697 unified features. E: table of five study settings across learning regime, data and task.
(A) Number of cases for each TCGA cancer type and (B) survival in years for every patient, marked as censored or dead. (C) Overview of the SeNMo framework: data from public and private sources are collected with MINDS and curated into the multi-omics dataset, and the learned weights are reused for prognostic and diagnostic tasks. (D) Feature processing for the six data types, unified across cancers and then across modalities into an 80,697-feature multi-omics matrix. (E) Study design: a baseline trained on 33 TCGA cancer types for overall survival, then varied in data and task (red borders mark a change from the baseline). Reproduced from Waqas et al. (2025), International Journal of Molecular Sciences, CC BY 4.0.

How it works

One vector per patient

Each molecular modality went through the same sequence of filters, with thresholds set per modality. We removed features that were NaN in all samples, dropped near-constant features with a 0.998 threshold, removed duplicate features and filtered low-variance ones. For gene expression we also kept only features with log-transformed expression above 7 (127 FPKM). Missing values within a modality and cancer type were imputed with the feature mean; missing features across cancers and modalities were zero-padded.

ModalityRaw featuresAfter preprocessing
DNA methylation485,5764,006 to 4,762 per cancer type; 52,396 combined
Gene expression (RNA-seq)60,483772 to 3,671 per cancer type; 8,794 combined
miRNA expression1,880967 to 1,460 per cancer type; 1,730 combined
Protein expression (RPPA)487472
Somatic mutation18,09017,253
Feature dimensions before and after preprocessing. Ranges are per cancer type across the 33 TCGA cohorts; combined counts are the union across cancers. The final per-patient input has 80,697 features, including four clinical variables (age, gender, race and stage).

Self-normalizing blocks

Each SeNMo block is a linear layer followed by a scaled exponential linear unit (SELU) and Alpha-Dropout. SELU is defined as

SELU(x)=λ{xif x>0α (ex−1)if x≤0\mathrm{SELU}(x) = \lambda \begin{cases} x & \text{if } x > 0 \\ \alpha\,(e^{x} - 1) & \text{if } x \le 0 \end{cases}

with λ=1.05071\lambda = 1.05071 and α=1.6733\alpha = 1.6733, the values the paper reports. Klambauer and colleagues showed that with these constants, activations that start near zero mean and unit variance converge to that fixed point as they pass through many layers.3 Alpha-Dropout is a dropout variant that preserves this property. The final SeNMo encoder has seven hidden layers, ends in a 48-dimensional patient embedding and has 83.33 million parameters.

Diagram of an input feature vector feeding seven SeNMo layer blocks of decreasing size (1024, 512, 256, 128, 48, 48, 48 dimensions), branching into a one-dimensional overall survival output and a 33-dimensional primary cancer type output. A box notes that each block is a linear unit plus SELU activation plus AlphaDropout.
Architecture of the SeNMo encoder network: seven hidden layers, each a linear unit followed by SELU activation and Alpha-Dropout, with the number of neurons in every layer shown. The trained model has 83.33 million parameters, and the same model is used for regression (overall survival) and classification (primary cancer type). Reproduced from Waqas et al. (2025), International Journal of Molecular Sciences, CC BY 4.0.

Training objective

Overall survival is modeled with the negative Cox partial log-likelihood. For a batch of NN patients with predicted log-hazards θi\theta_i and event indicators δi\delta_i:

Lcox=−1N∑i=1N(θi−log⁡∑j=1NeθjRij) δi\mathcal{L}_{\text{cox}} = -\frac{1}{N} \sum_{i=1}^{N} \Big( \theta_i - \log \sum_{j=1}^{N} e^{\theta_j} R_{ij} \Big)\, \delta_i

where Rij=1R_{ij} = 1 if patient jj is still at risk at patient ii‘s event time. The full objective is a weighted sum L=λcLcox+λceLce+λrLreg\mathcal{L} = \lambda_c \mathcal{L}_{\text{cox}} + \lambda_{ce} \mathcal{L}_{\text{ce}} + \lambda_r \mathcal{L}_{\text{reg}}, adding cross-entropy for cancer type and an L1 penalty on the weights. For TLS ratio regression we used a Huber loss.

Data and evaluation

For survival, each cancer type was split 80/20 into a training-validation set and a held-out test set. The pooled training-validation cohort had 11,050 patients, trained with 10-fold cross-validation, and the pooled test set had 2,754 patients. That gave ten checkpoints. We report both their mean test C-index and an ensemble that averages their predictions.

We ran about 400 hyperparameter configurations with Weights & Biases. Training the final model took about 11 hours on one Tesla V100 (32 GB), and fine-tuning on a cohort of around 150 patients took about 15 minutes. External evaluation used two lung squamous cell carcinoma (LSCC) cohorts: the public CPTAC-LSCC proteogenomic cohort and an independent 108-patient cohort from Moffitt Cancer Center.

Results

C-index, held-out TCGA test set (ensemble)
0.758
patients in the pooled test set
2,754
33-class cancer type accuracy
99.8%
C-index, CPTAC-LSCC after fine-tuning (ensemble)
0.730

Survival across modalities

The full six-modality model scored a mean C-index of 0.757 across the ten checkpoints and 0.758 as an ensemble. To see how much each data layer contributes, we retrained on subsets of the modalities. Performance fell off gradually rather than collapsing: gene expression alone still reached 0.728.

DNA methylationcheckpoint mean 0.6440.650
miRNAcheckpoint mean 0.6860.702
Gene expressioncheckpoint mean 0.7180.728
3-modalcheckpoint mean 0.7250.726
4-modalcheckpoint mean 0.7460.751
5-modalcheckpoint mean 0.7460.749
6-modal (SeNMo)checkpoint mean 0.7570.758
0.60.8

C-index (held-out test, ensemble)

Overall survival C-index on the held-out pan-cancer TCGA test set by input modalities (ensemble of ten checkpoints; the mean of single checkpoints is in the small note). 3-modal = gene expression, methylation and miRNA; 4-modal adds protein; 5-modal adds mutation; 6-modal adds clinical variables.

Splitting the pooled test set into terciles of predicted hazard gave clearly separated Kaplan-Meier curves (log-rank p=1.16×10−46p = 1.16 \times 10^{-46} for low versus high risk). The integrated Brier score, which reflects both discrimination and calibration, was 0.125 in cross-validation and 0.158 on the held-out test set, against a non-informative reference of 0.25.

Kaplan-Meier plot of survival probability against survival time up to about 25 years. The low-risk curve (green) stays highest, near 0.5 after 15 years; the intermediate curve (blue) falls below 0.4 by about 10 years; the high-risk curve (orange) drops fastest, to about 0.1 by 10 years.
Kaplan-Meier curves for SeNMo risk stratification on the pan-cancer data, with patients split into low, intermediate and high-risk terciles at the 33rd and 66th percentiles of predicted hazard. Log-rank p-values are 1.66 × 10⁻⁵ for low vs. intermediate, 1.16 × 10⁻⁴⁶ for low vs. high and 1.92 × 10⁻²² for intermediate vs. high; shaded areas are 95% confidence intervals. Reproduced from Waqas et al. (2025), International Journal of Molecular Sciences, CC BY 4.0.

Evaluated one cancer at a time, the pan-cancer model gave statistically significant C-indices for 29 of 33 cancer types, with the best result in pheochromocytoma and paraganglioma (0.900 test, 0.929 ensemble). Four cohorts (GBM, LAML, PRAD and TGCT) were not significant at first. A ten-epoch fine-tune brought GBM, LAML and PRAD to 0.650, 0.626 and 0.542 (ensemble), while TGCT remained a failure case.

Box plots of C-index per cancer type, sorted from PCPG near 0.91 down through ACC, UVM, LGG and others to OV near 0.52, with DLBC and READ after. Green fine-tuned boxes on the right for GBM, LAML, PRAD, Moffitt-LSCC and CPTAC-LSCC sit between about 0.56 and 0.70.
Tumor-specific C-indices for SeNMo across 32 TCGA cancer types and the two external lung squamous cohorts. Pink box plots are zero-shot predictions on each held-out TCGA subset; green box plots are after a brief ten-epoch fine-tune, used where baseline performance was not significant or the cohort was unseen (CPTAC-LSCC, Moffitt-LSCC). Triangles mark ensemble estimates. Reproduced from Waqas et al. (2025), International Journal of Molecular Sciences, CC BY 4.0.

External cohorts

Zero-shot transfer did not work well. On CPTAC-LSCC the model was at chance level, and on the Moffitt cohort it was only modestly above. A short fine-tune changed that:

CohortZero-shotAfter 10-epoch fine-tune
CPTAC-LSCC0.48 / 0.500.677 / 0.730
Moffitt-LSCC0.581 / 0.5900.647 / 0.656
Overall survival C-index on the two external lung squamous cell carcinoma cohorts, reported as test / ensemble. Fine-tuning ran for ten epochs with all layers unfrozen.

Fine-tuning an MLP backbone

Our first attempts at fine-tuning ran into catastrophic forgetting: the model would not converge on either new cohort, especially when we froze some hidden layers and trained the rest at lower learning rates. What worked was unfreezing every layer and training for just 10 epochs with a very small learning rate (4×10−54 \times 10^{-5}), high weight decay and dropout of 0.35.

Cancer type and TLS

For diagnosis, we framed tissue of origin as a 33-way classification. We removed stage from the clinical inputs because its distribution differs across cancers and would leak the label. SeNMo reached 99.8% accuracy on the test set for both single-checkpoint and ensemble inference. The remaining errors were between biologically related pairs such as colon and rectal adenocarcinoma.

For the immune microenvironment, we fine-tuned the backbone on the Moffitt LSCC cohort to regress the tissue-level TLS ratio (segmented TLS area divided by total tissue area). The labels came from H&E and CD20-stained whole-slide images analyzed in Visiopharm with manual review. Using cross-validation on 76 patients and a held-out set of 20, SeNMo’s predictions showed no significant difference from the manual annotations (paired comparison, p=0.064p = 0.064). Patients split into high and low TLS groups by SeNMo’s predictions had a stronger survival separation (log-rank p=2.5×10−4p = 2.5 \times 10^{-4}) than the split by manual labels (p=0.019p = 0.019).

Multi-panel figure. A donut chart of 103 samples: 68 training, 8 validation, 20 held-out test and 7 missing TLS labels. A box plot comparing manual and SeNMo TLS ratios with p = 0.064. Violin plots of high and low TLS groups for targets and predictions. Two Kaplan-Meier plots of surviving fraction over 12 years for low and high TLS groups, labeled Target (p = 0.019) and Predictions (p = 2.5 × 10⁻⁴).
TLS ratio prediction by SeNMo on the Moffitt LSCC cohort: the split of 103 samples into training, validation and held-out test sets, with samples missing TLS labels excluded; TLS ratios from manual annotations (M) and SeNMo predictions, with no significant difference (p = 0.064); violin plots of annotated and predicted ratios split into high and low groups at the median, both significantly separated; and Kaplan-Meier curves for high versus low TLS ratio from the annotations (p = 0.019) and from SeNMo's predictions (p = 2.5 × 10⁻⁴). Reproduced from Waqas et al. (2025), International Journal of Molecular Sciences, CC BY 4.0.

What we learned

One backbone trained for survival also carried enough signal for cancer type classification and, after fine-tuning, TLS regression. More modalities generally helped, and performance degraded gracefully as layers were withheld, down to gene expression alone. That matters because incomplete data is common in routine oncology practice. The ingestion step is also modular: an RNA-seq matrix processed with limma-voom or the DESeq2 variance-stabilizing transform can go into the same encoder without changing the network.

The limits are just as clear. Four cancer types started with low C-indices and TGCT never recovered, so a shared model can still miss features specific to one tumor type. External validation covered two cohorts, both lung squamous cell carcinoma, so broader claims about generalization need more histologies. The TLS labels came from an image-analysis pipeline with its own measurement noise. Zero-shot transfer was weak, and the backbone was prone to forgetting during fine-tuning.

The paper’s future directions are contrastive pre-training with synthetic modality dropout, spatial transcriptomics and digital pathology features, new inputs such as copy-number data, and prospective validation in molecular tumor boards.

The broader survival thread

SeNMo started as an AACR 2024 abstract and an arXiv preprint before the journal version.56 Two related projects in our group ask how best to combine modalities for survival prediction.

RMSurv. Dominic Flack (Rochester Institute of Technology), whom I mentored from August 2024 to February 2025, led Robust Multimodal Fusion for Survival Prediction in Cancer Patients with Dimah Dera; Asim Waqas, Ghulam Rasool and I contributed.7 RMSurv takes the opposite approach to SeNMo’s early fusion. It trains a separate discrete-time survival model for each modality (twenty 6-month bins over 10 years), then sets late-fusion weights by searching over a synthetic dataset that keeps the real cross-modality correlations and each modality’s cross-validated C-index. The weights can also vary by time bin (TD-RMSurv). Pathology reports were one of its modalities, embedded with GatorTron through HONeYBEE.

Five-column grid. Rows for clinical data, miRNA, pathology report, DNA methylation, protein expression and gene expression each show a unimodal survival curve, its normalized version and a bar chart of time-dependent weights over 10 years; the last column shows the single combined, normalized survival curve.
Schematic of RMSurv's late fusion. Up to six unimodal models are trained separately and output their own survival predictions, which are normalized, combined with linear weights (optionally time-dependent) and normalized again. Reproduced from Flack et al. (2025), Cancer Informatics, CC BY-NC 4.0.

On the six-modality TCGA lung adenocarcinoma dataset, RMSurv improved on the best unimodal C-index by 0.0273, compared with 0.0143 for an existing late-fusion method and 0.0072 for early fusion. On the 11,060-case pan-cancer dataset, TD-RMSurv reached a cross-validated C-index of 0.7533 with six modalities. The learned weights are also interpretable: in one combined LUAD + LUSC fold, the model put 90% of the weight on clinical data in the first six-month bin and shifted weight to other modalities later. The code is public.8

PARADIGM. Asim Waqas led this project as first author, and I contributed. PARADIGM uses SeNMo as its omics encoder. It embeds each modality with a foundation model, builds a graph whose nodes are patients and trains a graph neural network for survival. In the pan-squamous preprint, SeNMo produces the 48-dimensional omics embedding, UNI encodes slides and GatorTron encodes clinical text and pathology reports. A graph convolutional network was trained on five TCGA squamous cell carcinomas with 7-fold cross-validation, and the framework was also evaluated on a 103-patient Moffitt lung SCC cohort.9 The AACR 2025 abstract extends the analysis to five adenocarcinomas and reports significant improvements over unimodal and multimodal MLP, Transformer and XGBoost models.10

Multimodal survival prediction from precomputed embeddings continues in EAGLE (Efficient Alignment of Generalized Latent Embeddings), which adds attention-based fusion and interpretable attribution analysis. SeNMo’s TCGA molecular embeddings are distributed alongside the HONeYBEE embedding resource.

Resources

  • Paper: open access in the International Journal of Molecular Sciences.1
  • Code: training, fine-tuning, testing and ensemble scripts at lab-rasool/SeNMo.11
  • Weights: the ten pan-cancer checkpoints on Hugging Face and on Zenodo in two parts.12
  • Embeddings: TCGA molecular embeddings in the molecular configuration (senmo split) of Lab-Rasool/TCGA.
  • Data curation: MINDS, used to build the TCGA cohorts.

References

  1. Waqas A, Tripathi A, Ahmed S, Mukund A, Farooq H, Johnson JO, Stewart PA, Naeini M, Schabath MB, Rasool G. Self-Normalizing Multi-Omics Neural Network for Pan-Cancer Prognostication. Int J Mol Sci. 2025;26(15):7358. doi:10.3390/ijms26157358. Open access at PMC12347193. ↩ ↩2

  2. Tripathi A, Waqas A, Venkatesan K, Yilmaz Y, Rasool G. Building Flexible, Scalable, and Machine Learning-Ready Multimodal Oncology Datasets. Sensors. 2024;24(5):1634. doi:10.3390/s24051634. Code: github.com/lab-rasool/MINDS. ↩

  3. Klambauer G, Unterthiner T, Mayr A, Hochreiter S. Self-Normalizing Neural Networks. arXiv:1706.02515, 2017. arxiv.org/abs/1706.02515. ↩

  4. Liu J, Lichtenberg T, Hoadley KA, et al. An Integrated TCGA Pan-Cancer Clinical Data Resource to Drive High-Quality Survival Outcome Analytics. Cell. 2018;173:400-416. doi:10.1016/j.cell.2018.02.052. ↩

  5. Waqas A, Tripathi A, Ahmed S, Mukund A, Stewart P, Naeini M, Farooq H, Rasool G. SeNMo: A self-normalizing deep learning model for enhanced multi-omics data analysis in oncology. Cancer Res. 2024;84(6 Suppl):908. doi:10.1158/1538-7445.AM2024-908. ↩

  6. Waqas A, Tripathi A, Ahmed S, et al. Self-Normalizing Foundation Model for Enhanced Multi-Omics Data Analysis in Oncology. arXiv:2405.08226, 2024. arxiv.org/abs/2405.08226. ↩

  7. Flack D, Tripathi A, Waqas A, Rasool G, Dera D. Robust Multimodal Fusion for Survival Prediction in Cancer Patients. Cancer Inform. 2025;24:11769351251376192. doi:10.1177/11769351251376192. ↩

  8. RMSurv source code: github.com/dom416/RMSurv. ↩

  9. Waqas A, Tripathi A, Stewart P, Naeini M, Schabath MB, Rasool G. Embedding-based Multimodal Learning on Pan-Squamous Cell Carcinomas for Improved Survival Outcomes. arXiv:2406.08521, 2024. arxiv.org/abs/2406.08521. ↩

  10. Waqas A, Tripathi A, Naeini M, Stewart PA, Schabath MB, Rasool G. PARADIGM: an embeddings-based multimodal learning framework with foundation models and graph neural networks [abstract]. Cancer Res. 2025;85(8 Suppl 1):991. doi:10.1158/1538-7445.AM2025-991. ↩

  11. SeNMo source code: github.com/lab-rasool/SeNMo. ↩

  12. SeNMo model checkpoints and sample embeddings, Part 1: doi:10.5281/zenodo.14219799; Part 2: doi:10.5281/zenodo.14286190. ↩

Cite the paper

Waqas A, Tripathi A, Ahmed S, Mukund A, Farooq H, Johnson JO, et al. Self-Normalizing Multi-Omics Neural Network for Pan-Cancer Prognostication. Int J Mol Sci. 2025;26(15):7358. doi:10.3390/ijms26157358

@article{waqas2025senmo,
  title   = {Self-Normalizing Multi-Omics Neural Network for Pan-Cancer Prognostication},
  author  = {Waqas, Asim and Tripathi, Aakash and Ahmed, Sabeen and Mukund, Ashwin and Farooq, Hamza and Johnson, Joseph O. and Stewart, Paul A. and Naeini, Mia and Schabath, Matthew B. and Rasool, Ghulam},
  journal = {International Journal of Molecular Sciences},
  year    = {2025},
  volume  = {26},
  number  = {15},
  pages   = {7358},
  doi     = {10.3390/ijms26157358}
}