Blog · Research · MINDS · 9 min read
MINDS: linking public cancer repositories into ML-ready cohorts
A metadata-first cloud framework that consolidates 41,499 open-access cancer cases into one queryable warehouse and fetches images and omics on demand.

Contents
A multimodal oncology model starts with a practical question: for these patients, where is everything? Clinical and genomic records for a lung cancer cohort may sit in the NCI Genomic Data Commons (GDC), the CT, MRI and PET scans from the same studies take a different path into an imaging archive, and proteomics lives in yet another portal. Each one has its own API, schema and query method. In the paper we put it plainly: piecing together data manually across these silos is painstakingly difficult.
MINDS, the Multimodal Integration of Oncology Data System, is our answer to that problem. It pulls the structured clinical and biospecimen records from public cancer repositories into one patient-centric database. Researchers define a cohort with SQL, and MINDS then fetches the matching slides, scans and omics files from the repositories that host them.1
In short
- Problem: cancer data for the same patients is split across repositories, each with its own API and schema.
- What we built: MINDS keeps only the small clinical and biospecimen records in one SQL database, and downloads images, slides and omics files from their source repositories only when a cohort needs them.
- Result: 41,499 open-access cases fit in a 25.85 MB extract. Typical single-table queries average under 5 seconds, and queries across 8 or more tables finish within 15 seconds.
I developed MINDS. Asim Waqas and I share first authorship, and I am the paper’s corresponding author. Kavya Venkatesan and Yasin Yilmaz are co-authors, and Ghulam Rasool supervised the work.

Why the data is scattered
Public cancer data is not missing. It is spread across systems built for different modalities. In the NCI Center for Cancer Genomics pipeline, tissue moves through collection and processing, genome characterization and genomic data analysis before the results are released through the GDC.2 Imaging from the same studies follows a separate route. The NCI Cancer Research Data Commons (CRDC) was set up to connect these resources under the FAIR principles, and its repositories include the GDC, the Proteomics Data Commons (PDC) and the Imaging Data Commons (IDC).34
The Cancer Data Aggregator (CDA) adds a unified search across those commons.5 At the time of the paper, it had two limits that mattered for building ML datasets. Its mapping was not real-time: in September 2023, a CDA query for patients whose primary diagnosis site was lung returned 4,870 cases, while the GDC data portal held 12,267. It is also confined to CRDC-managed data, so resources such as cBioPortal cannot be queried through it.
The paper lists five gaps in existing approaches:
- consolidation has mostly covered structured clinical records, not imaging, omics or pathology;
- repositories expose proprietary APIs and schemas, with no standard interface across them;
- on-premises systems limit how far storage and compute can scale;
- dynamic dataset extracts make it hard to trace which data trained which model;
- access controls are coarse-grained.
A metadata-first design
We designed MINDS around five requirements. It should minimize large-scale unstructured storage, scale both horizontally and vertically, add new data sources easily, track data from ingestion to training, and place audit checkpoints at every stage of the pipeline.
The first requirement shapes everything else. MINDS stores no whole slide images or radiology scans. It keeps only structured and semi-structured records, such as clinical and biospecimen metadata and identifiers. That makes it a compact, harmonized index of patients rather than another copy of the data commons: small enough to query interactively, with case IDs that tell the download tools which files to fetch later. Unstructured data stays in its source repository until a user builds a cohort that needs it.
Underneath, MINDS is a two-tier data lakehouse on Amazon Web Services: a data lake that takes raw records as they arrive, and a data warehouse that holds the refined relational tables used for analysis.
How it works
Acquisition and updates
For the first version we used the GDC as the primary structured source. The paper describes it as holding clinical, biospecimen and molecular data for over 86,000 cases across 78 projects, with more than 3 petabytes of data from programs such as TCGA and TARGET as of 2023. We pull the TSV and JSON files for all public cases. On the clinical side these cover clinical, exposure, family history, follow-up and pathology detail records. On the biospecimen side they cover aliquot, analyte, portion, sample and slide records. The files are staged in cloud storage (an Amazon S3 bucket) that forms the data lake.
Acquisition repeats on a schedule. Extract-transform-load (ETL) routines run every 12 hours, using the GDC REST API to fetch cases, files and metadata added since the last timestamp. New arrivals trigger an update job, so only the difference between the bucket and the lake is processed. The paper names UCSC Xena and cBioPortal as further platforms the structured pipeline can integrate.
Harmonizing to one model
Raw downloads come in different shapes: JSON clinical documents describe patients differently from TSV biospecimen exports. We use the GDC data model as the integration schema. It is a directed acyclic graph whose nodes are entities such as cases, samples, files and read groups, and whose edges are relationships such as a sample that is derived_from a case.
The ETL runs on AWS Glue, Amazon’s managed ETL service. Its crawlers read the GDC node schemas from their YAML definitions, parse the incoming JSON and map each property to a column in a Glue table. In this way the graph becomes a relational view. Child records keep the IDs of their parent case and file, so the links can be rebuilt with joins. For sources that do not fit the existing tables, Glue schema evolution can extend a definition or create a new table.
The five-step Glue crawler workflow
- Open access-controlled connections to the source databases.
- Apply any custom classifiers to catalog the data and generate metadata.
- Use the built-in classifiers for the remaining extract-transform-load work.
- Merge the catalog tables into a single database, resolving conflicts and removing duplicates.
- Upload the catalog to a data store for analytics.

Warehouse and serving
The cataloged tables are loaded into Amazon Redshift, the analytical warehouse. A separate MySQL database handles inserts and updates as new data arrives, so ad hoc analysis does not slow the update path. One shared catalog describes the tables for every downstream tool, from SQL queries to dashboards.

There are two ways to use the result. Dashboards built with Amazon QuickSight let researchers explore distributions and filter on fields such as age, gender, ethnicity, tumor grade, treatment type, year of diagnosis and survival. For pipelines, an open-source Python toolkit runs the cohort workflow. A SQL query returns a pandas DataFrame of matching clinical records, or a list of unique case IDs that the toolkit uses to call the GDC, IDC and PDC APIs. Downloads go into a top-level /raw folder with one subfolder per case, together with JSON manifests that record file IDs, types and sources. Each cohort dataset gets a unique version ID, and any change creates a new version.

Security and portability
Data is encrypted at rest with 256-bit AES keys and in transit with TLS 1.2. Access policies work at the level of individual resources and actions, and column-level controls can hide sensitive fields. Audit logs record activity, and replicated, versioned storage protects against data loss.
The design aims to avoid depending on a single provider. The paper maps each AWS component to a Google Cloud equivalent, for example BigQuery in place of Redshift and Looker in place of QuickSight. It also describes an on-premise mode: a Docker setup that bootstraps a MySQL database with the consolidated schema, and a lightweight Flask app in place of the dashboards.
What we measured
- open-access cases consolidated
- 41,499
- structured extract held in MINDS
- 25.85 MB
- average single-table query
- < 5 s
- queries spanning 8+ tables
- ≤ 15 s
The 41,499 cases come from open-access data, which can be used without access restrictions. Table 1 of the paper sets the size of the MINDS extract against the full holdings of each commons.
| Data source | Storage size | Cases |
|---|---|---|
| MINDS | 25.85 MB | 41,499 |
| PDC | 36 TB | 3,081 |
| GDC | 3.78 PB (17.68 TB public) | 86,962 |
| IDC | 40.96 TB | 63,788 |
The cases table accounts for 10.24 MB of the extract, followed by genomic variant calls and read groups. Because the extract is this small, cohort filtering can run on a single commodity machine before any raw data is transferred.
The cohort also mixes data from different technology eras. TCGA contributes 11,315 cases with molecular profiling from earlier platforms, and Foundation Medicine contributes 18,004 cases from contemporary next-generation sequencing panels. The paper argues that this mix helps models learn biological patterns that hold across platforms rather than artifacts specific to one assay.
For query speed, we timed more than 24,000 SQL invocations, from simple filters to queries joining multiple data domains. Typical single-table queries averaged under 5 seconds. Queries spanning 8 or more tables, across metadata for thousands of cases, finished within 15 seconds. Redshift scales out because cohort queries are read-only: doubling the number of cluster nodes halved the runtime of intensive 8-table cohort queries.
How to read these numbers
The 25.85 MB extract holds metadata and identifiers. The images, sequences and slides stay in their repositories, and the storage saving comes from deferring those transfers until a cohort needs them. The paper presents its first cohort-construction timings as preliminary experiments.
Limitations
The paper’s own discussion names the main limits:
- Public data only. MINDS currently uses open-access cases. Including controlled data, for example through privacy-preserving methods, is listed as future work.
- No raw imagery in the warehouse. Ingesting raw clinical images and video directly is a primary constraint. We proposed learned embeddings as a compact way to bring imaging content into the system.
- Unstructured notes. Bringing in free-text physician notes would need pretrained clinical language models.
- Polling, not events. Updates arrive through scheduled polling. Webhooks and event notifications would give more real-time incremental updates without excessive API calls.
- One cloud in practice. The implementation runs on AWS. The Google Cloud mapping is a design, and replication on Google Cloud or Azure is listed as future work.
Where it went next
In a JNCCN 2024 abstract, Asim Waqas, Ghulam Rasool and I described the next phase of MINDS around three goals: building benchmark datasets from multimodal data, handling large-scale multivariate time-series data for temporal predictive models, and acting as a store of contextual knowledge for language models through embeddings from clinical foundation models such as GatorTron.6 That abstract names TCGA, The Cancer Imaging Archive (TCIA) and institutional electronic health records as sources, and says our ongoing work aims to produce curated benchmark datasets spanning over 20 cancer types.
MINDS is now distributed on PyPI as med-minds.7 The cloud version is in closed beta, and the database can also be recreated locally on PostgreSQL. The repository lists the projects being brought into MINDS, including imaging collections, for a total of 115,974 patients.8
MINDS also supplies data to our later work. In HONeYBEE, we obtained the raw TCGA data through MINDS.9
Resources
- Paper: Sensors 24(5):1634, open access, published 2 March 2024.1
- Preprint: arXiv 2310.01438, first posted 30 September 2023.10
- Code:
lab-rasool/MINDSon GitHub, with themed-mindsPython package and documentation.8 - Follow-up abstract: JNCCN 22(2.5), BIO24-030.6
The snippet below is adapted from the README’s usage section. It assumes a local PostgreSQL database configured through a .env file:
import med_minds
# Populate or refresh the local database
med_minds.update()
# Query harmonized clinical records into a pandas DataFrame
query = "SELECT * FROM clinical WHERE project_project_id = 'TCGA-LUAD' LIMIT 10"
df = med_minds.query(query)
# Build a cohort from the query and download its files
query_cohort = med_minds.build_cohort(query=query, output_dir="./data")
query_cohort.stats()
query_cohort.download(threads=12, exclude=["Slide Image"])
References
-
Tripathi A, Waqas A, Venkatesan K, Yilmaz Y, Rasool G. Building Flexible, Scalable, and Machine Learning-Ready Multimodal Oncology Datasets. Sensors. 2024;24(5):1634. doi:10.3390/s24051634. Open-access full text: PMC10933897. ↩ ↩2
-
Grossman RL, Heath AP, Ferretti V, et al. Toward a shared vision for cancer genomic data. N Engl J Med. 2016;375:1109-1112. doi:10.1056/NEJMp1607591. GDC data portal: portal.gdc.cancer.gov. ↩
-
NCI Proteomic Data Commons: pdc.cancer.gov. ↩
-
NCI Imaging Data Commons portal: portal.imaging.datacommons.cancer.gov. ↩
-
NCI Cancer Data Aggregator: datacommons.cancer.gov. ↩
-
Tripathi A, Waqas A, Rasool G. BIO24-030: Unifying Multimodal Data, Time Series Analytics, and Contextual Medical Memory: Introducing MINDS as an Oncology-Centric Cloud-Based Platform. J Natl Compr Canc Netw. 2024;22(2.5). doi:10.6004/jnccn.2023.7305. ↩ ↩2
-
med-mindson the Python Package Index: pypi.org/project/med-minds. ↩ -
MINDS source code and README: github.com/lab-rasool/MINDS. Documentation: lab-rasool.github.io/MINDS. ↩ ↩2
-
Tripathi A, Waqas A, Schabath MB, Yilmaz Y, Rasool G. HONeYBEE: enabling scalable multimodal AI in oncology through foundation model-driven embeddings. npj Digit Med. 2025;8:622. doi:10.1038/s41746-025-02003-4. See the Data availability section. ↩
-
Tripathi A, Waqas A, Venkatesan K, Yilmaz Y, Rasool G. Building Flexible, Scalable, and Machine Learning-ready Multimodal Oncology Datasets. arXiv:2310.01438. arxiv.org/abs/2310.01438. ↩
Cite the paper
Tripathi A, Waqas A, Venkatesan K, Yilmaz Y, Rasool G. Building flexible, scalable, and machine learning-ready multimodal oncology datasets. Sensors. 2024;24(5):1634. doi:10.3390/s24051634
@article{tripathi2024minds,
title = {Building Flexible, Scalable, and Machine Learning-Ready Multimodal Oncology Datasets},
author = {Tripathi, Aakash and Waqas, Asim and Venkatesan, Kavya and Yilmaz, Yasin and Rasool, Ghulam},
journal = {Sensors},
year = {2024},
volume = {24},
number = {5},
pages = {1634},
doi = {10.3390/s24051634}
}