Blog · Research · CLEVER & pathology reports · 11 min read
Consensus reasoning with local LLMs for pathology report extraction
Eleven local LLMs extract, three reasoning models grade, one adjudicates: structured, auditable data from free-text surgical pathology reports.

Contents
A surgical pathology report is where a cancer diagnosis is written down: tumor site, laterality (which side of the body), histology (tumor type), stage, grade and behavior, plus the biomarkers that steer treatment. Cancer registries, trial screening and research cohorts need those values as structured fields. Pathologists write them as narrative, in formats that vary by institution, subspecialty and individual habit, often with a value like stage buried in the prose.
We tested a consensus-based alternative to asking a single large language model (LLM) for the answer. Eleven open models, all running on local hardware, independently extract the six registry variables and justify each value. Three reasoning models grade every extraction against the source report, and a final reasoning model settles on one value per field with a written rationale. Pathologists then scored samples of the final output.
In short
- Problem: registry fields and biomarkers sit in free-text pathology reports, and a single LLM can confidently get them wrong.
- Approach: eleven local LLMs extract, three reasoning LLMs grade, one adjudicates, and every value comes with a rationale. No data leaves the institution.
- Results: mean automated accuracy of 84.9% on 6,100 TCGA reports and 88.2% on 510 Moffitt reports; pathologists scored the consensus output at 88% to 100% across variables.1
- Limits: biomarkers were harder (70.6% automated accuracy), and expert review covered samples, not every report.
I am first author and developed the LLM workflows; Asim Waqas and I designed and implemented the framework and ran the analyses together, and Ghulam Rasool supervised as principal investigator. The work continued in CLEVER, the LLM-agent system I went on to develop at Moffitt.
Why reports resist structure
US surveillance programs such as SEER and the National Program of Cancer Registries require consistent capture of these six fields.1 Rule-based natural language processing can map free text to fields, but such systems are slow to build, sensitive to wording and generalize poorly across cancer types and documentation styles.
LLMs can extract clinical entities with little or no task-specific training, but they also hallucinate: a fabricated stage, a wrong site, an invalid combination of attributes, often stated confidently on weak evidence. In registry abstraction or trial matching, that is risky.
Privacy adds a third constraint: sending protected health information to commercial cloud APIs raises compliance concerns, so every model had to run inside the institutional firewall.
A three-stage consensus
Rather than search for one best model, we treated disagreement between models as information, mirroring how diagnostic consensus forms in practice: independent reads, reasoned comparison, reconciliation.
Stage 1: eleven extractors
Each report goes, in full, to eleven locally hosted models from 3B to 70B parameters, drawn from the LLaMA, Mistral, Qwen, Phi-4, Tulu 3 and DeepSeek-R1 families (full list below). Each returns site, laterality, histology, stage, grade and behavior as JSON, with a brief justification per field. When a value is absent or cannot be inferred, the model must say “Unknown” and explain why. Stage was excluded for brain tumors, where the paper notes it lacks clinical relevance.
For the Moffitt cohort the extractors also read organ-specific biomarkers: ER, PR, HER2 and Ki-67 for breast; IDH1/2, MGMT, ATRX, TP53, 1p/19q codeletion, TERT, EGFR and CDKN2A/B for brain; and PD-L1, EGFR, ALK, ROS1, BRAF, KRAS, MET, RET, HER2 and NTRK for lung.
The full model zoo
| Label in paper | Base model | Parameters | Stage(s) |
|---|---|---|---|
| LLaMA 3.2-small | Llama 3.2 Instruct | 3B | 1 |
| Mistral-small | Mistral Instruct v0.3 | 7B | 1 |
| Qwen-small | Qwen2.5 Instruct | 7B | 1 |
| Tulu-small | Llama 3.1 Tulu 3 | 8B | 1 |
| DeepSeek-R1-small | DeepSeek-R1-0528-Qwen3 | 8B | 1 |
| LLaMA 3.1-small | Llama 3.1 Instruct | 8B | 1 |
| Phi-medium | Phi-4 | 14B | 1 |
| DeepSeek-R1-medium | DeepSeek-R1-Distill-Qwen | 32B | 1 |
| Mixtral-medium | Mixtral Instruct v0.1 | 8x7B | 1 |
| LLaMA 3.3-large | Llama 3.3 Instruct | 70B | 1 |
| DeepSeek-R1-large | DeepSeek-R1-Distill-Llama | 70B | 1, 2, 3 |
| Qwen3-medium | Qwen3 | 32B | 2 |
| QWQ-medium | QwQ | 32B | 2 |
All models ran at FP16. Stages: 1 extraction, 2 evaluation, 3 aggregation.1
Stage 2: three graders
Every extracted value, together with the justification that produced it, goes to three reasoning models acting as evaluators: DeepSeek-R1-large (70B), Qwen3-32B and QwQ-32B. Each sees the original report and assigns a strict binary score, 1 for correct and 0 for incorrect, based on textual evidence, internal consistency and clinical plausibility. A correct value with an invalid justification does not earn a 1, and “Unknown” is correct only when the report truly lacks the information.2
Stage 3: one adjudicator
A second prompt to DeepSeek-R1-large receives the report, all eleven candidates and all evaluator scores, and selects one value per field. Its instructions amount to a filtered majority vote. Writing for model ‘s answer to variable , and for the set of extractors whose answer to was graded correct in stage 2, the target is
with ties broken by the source model’s overall accuracy, and “Unknown” plus an explanation when is empty.2 The aggregator also writes a short rationale per field naming which models contributed.
An auditable record
One JSON record per report keeps the raw candidates, evaluator judgments and final answer side by side. The sketch below abbreviates the preprint’s example record for a TCGA lung wedge resection; inner field names are placeholders.2
{
"note": "…full report text…",
"model_responses": { "phi4": { "site": "…", "justification": "…" }, "…": "…" },
"evaluation_…": "…stage-2 judgments from each evaluator…",
"final_aggregation_deepseek_r1_70B": {
"final_extraction": {
"site": "lung",
"laterality": "left",
"histology": "adenocarcinoma",
"stage": "T2NXMX",
"grade": "Grade III",
"behavior": "malignant"
},
"rationale": "…which models agreed, and why…"
}
}
A reviewer can read the rationale, see which extractors agreed and check the justification against the report.

Data and evaluation
- TCGA reports across 10 organ sites
- 6,100
- Moffitt reports: breast, brain, lung
- 510
- local extractor LLMs
- 11
- reasoning evaluators
- 3
TCGA. We randomly selected 6,100 reports from the TCGA-Reports corpus released by Jenna Kefeli and Nicholas Tatonetti,3 covering kidney, breast, lung, uterus, brain, prostate, liver, bladder, cervix and pancreas. For expert review, 138 reports from five organs (bladder, prostate, liver, cervix and kidney) were each scored by three pathologists, 414 evaluations in total.
Moffitt. 510 de-identified surgical pathology reports from patients treated between 2009 and 2018, linked to institutional cancer registry records: breast (215), brain or central nervous system (160) and lung (135). For expert review, 30 reports per organ (90 total) were randomly selected and scored by pathologists including specialists in each organ system. The expert evaluators across both data sets were Ehsan Ullah, Asma Khan, Farah Khalil, Wei-Shen Chen, Zarifa Gahramanli Ozturk, Daryoush Saeed-Vafa and Marilyn Bui.
Two tiers of scoring. Automated accuracy is the stage-2 evaluators’ binary score, averaged. Expert accuracy scores the stage-3 consensus output as 1 (correct), 0.5 (partially correct) or 0 (incorrect). We compared models, variables and organs with one- to three-way ANOVA and Tukey’s HSD tests, and measured inter-evaluator agreement with Pearson and intraclass correlation.
Results
Automated accuracy
- TCGA, six variables (± 7.3%)
- 84.9%
- Moffitt, six variables (± 7.2%)
- 88.2%
- Moffitt biomarkers (± 7.9%)
- 70.6%
- Correlation between evaluators
- ≥ 0.93
On TCGA, histology (88%), site (87%), stage (84%) and behavior (84%) were the best-extracted variables averaged across all models. Phi-4, DeepSeek-R1 70B and 32B, and LLaMA 3.3 were the strongest extractors at 86% or above, while Mixtral, Mistral-7B, the 8B DeepSeek-R1 and LLaMA 3.2 trailed. The three evaluators gave mean scores of 0.88 (DeepSeek-R1-large), 0.85 (Qwen3-32B) and 0.82 (QwQ-32B), with pairwise correlations of 0.94, 0.94 and 0.96.1

Model, variable and organ each had a significant effect on accuracy, and so did their three-way interaction (all P < .001).1 In plain terms, no single model performed uniformly across variables, and organ-specific language affected accuracy.
Moffitt reports scored higher on average, with Phi-4, Tulu 3, DeepSeek-R1 32B, LLaMA 3.3, DeepSeek-R1 70B and LLaMA 3.1 on top and evaluator correlations of 0.934, 0.981 and 0.939. Unlike TCGA, there was no significant model-by-organ or three-way interaction; the paper suggests institution-specific documentation practices may make extraction more consistent.
Pathologist review
- TCGA (138 reports)
- Moffitt (90 reports)
Expert accuracy (%), scale 1 / 0.5 / 0

Biomarkers
Biomarkers were the hardest task. As Marilyn Bui explained to The Pathologist, results are often spread across comments, ancillary studies, addenda or separate molecular reports, can involve several specimens, assays or time points, and are written in less standardized language.4

The breast panel scored lowest, with estrogen receptor (ER) and progesterone receptor (PR) status below 60%. The paper links this to complex reporting formats and multiple specimen types in breast reports.
What we learned
Many “errors” live in the source. Pathologists often traced a wrong answer to ambiguous documentation rather than model failure. Laterality was marked wrong where it does not apply (cervix, bladder) or where bilateral disease was described without naming the specimen’s side. Stage errors followed mixed staging schemes or incomplete TNM (tumor, node, metastasis) statements. Often the reviewers judged the model’s inference clinically plausible.
Organ-specific failure modes. In Moffitt breast reports, sentinel node notation such as (sn) or (i-) was dropped from predicted stage. Brain cases tripped on site and laterality when the specimen label differed from the surgical target. In lung and metastatic cases, models sometimes confused the biopsy site (a lymph node or the liver) with the primary site.
A right answer can have a weak rationale. Even when values were correct, rationales often failed to highlight the key clinical cues, limiting auditability. The paper suggests rationale filtering.
Diversity is the point. As Ghulam Rasool put it to The Pathologist, “no single model performs best across all variables, organs, and reporting styles.”4
Reading the automated numbers
Automated accuracy is assigned by LLM evaluators, not by comparison to a human gold standard. The evaluators agreed closely (r ≥ 0.93) but still differed significantly from one another (F = 2633.79, P < .001 on TCGA), and the pathologist review covers only a sample of reports. Treat the expert scores as the clinical check and the automated scores as a scalable proxy.
The paper names three limitations. Generalization to rare cancers, atypical document formats and resource-constrained settings is untested. The adjudicator’s judgment of justification quality may carry its own biases, especially with ambiguous language. And expert validation was necessarily limited to subsets because manual review is labor-intensive.1
From study to CLEVER
| When | Milestone |
|---|---|
| 2024 | Digital Pathology Association oral (first author Ghulam Rasool) on extraction with local and private LLMs |
| Mar 2025 | USCAP abstract: LLaMA 3, LLaMA 3.1 and Mixtral extract seven variables; LLaMA 3.1 above 84% average accuracy, with a review interface beside the report5 |
| Apr 2025 | Consensus framework posted to medRxiv2 |
| Oct 2025 | CLEVER talk on multi-agent, temporal-aware variable extraction (Gillies Machine Learning Workshop) |
| Dec 2025 | Consensus study published online in Laboratory Investigation1 |
| Mar 2026 | USCAP abstract on multi-agent neuro-oncology biomarker extraction6 |
| Apr 2026 | AACR abstract on social determinants of health7; The Pathologist coverage4 |
CLEVER is where this work went next. I developed it at Moffitt as a system that uses LLM agents to extract variables, patient timelines and ICD codes from medical records while retaining document evidence and reasoning for review, and I supported its institution-scale deployment for record processing, variable extraction and document quality review. The principle carries over from the consensus study: keep the evidence and reasoning a reviewer needs.
Two follow-up abstracts apply multi-agent designs to new targets. At USCAP 2026, with Osama Elzaafarany, Sepideh Mokhtari, Michael Vogelbaum, Marilyn Bui and Ghulam Rasool, we described three agents (document assessment, extraction, and per-patient synthesis with confidence scores and evidence citations) run on 42 pathology PDFs from 27 glioblastoma patients. Across nine variables (diagnosis plus eight biomarkers), expert validation against ground-truth annotations gave 83.1% average accuracy; MGMT and PTEN were lowest at 77.8% each, mainly because of OCR noise.6 At AACR 2026, Asim Waqas led a study of temporally aware extraction of social determinants of health in two retrospective cohorts: 100 cancer patients with HIV (3,460 documents) and 524 peri-operative pain management patients (15,484 documents), producing time-stamped extractions with evidence provenance.7
Where it goes next
The paper’s own agenda: uncertainty estimation, active learning, more variables such as treatment response and recurrence status, and a pediatric evaluation with synoptic templates curated alongside pediatric pathologists. Prospective validation in live clinical environments, including laboratory information systems and registry pipelines, is the critical next step for measuring effects on accuracy, workflow and downstream decisions.1 Local deployment remains a thread: in June 2026 I co-presented a SIIM Learning Lab with Ghulam Rasool and Asim Waqas on running local LLMs behind institutional firewalls.
Resources
- Paper: Laboratory Investigation 2026;106(2):104272, doi:10.1016/j.labinv.2025.104272. Prompts for all three stages and an example JSON record are in the supplementary material.
- Preprint: medRxiv 2025.04.22.25326217.
- Data: TCGA reports via the TCGA-Reports repository. The Moffitt data set is not public; access may be requested from the corresponding author, subject to IRB approval and a data use agreement.
- Code: per the paper, the extraction framework and analysis code are available on reasonable request to the corresponding author.
- Coverage: AI Tackles Pathology Report Complexity, The Pathologist, April 8, 2026.
References
-
Tripathi A, Waqas A, Venkatesan K, Ullah E, Khan A, Khalil F, Chen WS, Ozturk ZG, Saeed-Vafa D, Bui MM, Schabath MB, Rasool G. Using Consensus-Based Reasoning and Large Language Models to Extract Structured Data From Surgical Pathology Reports. Laboratory Investigation. 2026;106(2):104272. Published online December 16, 2025. https://doi.org/10.1016/j.labinv.2025.104272 ↩ ↩2 ↩3 ↩4 ↩5 ↩6 ↩7 ↩8
-
Tripathi A, Waqas A, Ullah E, et al. Employing Consensus-Based Reasoning with Locally Deployed LLMs for Enabling Structured Data Extraction from Surgical Pathology Reports. medRxiv. 2025. Version 1 posted April 25, 2025; version 3 posted November 7, 2025. The stage 1 to 3 prompt descriptions and the example record cited here are from version 3. https://doi.org/10.1101/2025.04.22.25326217 ↩ ↩2 ↩3 ↩4
-
Kefeli J, Tatonetti N. TCGA-Reports: A machine-readable pathology report resource for benchmarking text-based AI models. Patterns. 2024;5(3):100933. https://doi.org/10.1016/j.patter.2024.100933. Data: github.com/tatonetti-lab/tcga-path-reports ↩
-
Allerton J. AI Tackles Pathology Report Complexity. The Pathologist. April 8, 2026. https://www.thepathologist.com/issues/2026/articles/april/ai-tackles-pathology-report-complexity/ ↩ ↩2 ↩3
-
Tripathi A, Waqas A, Venkatesan K, Ullah E, Bui MM, Rasool G. AI-Driven Extraction of Key Clinical Data from Pathology Reports to Enhance Cancer Registries. Abstract 1391, USCAP 114th Annual Meeting. Laboratory Investigation. 2025;105(3 Suppl):103629. https://doi.org/10.1016/j.labinv.2024.103629 ↩
-
Tripathi A, Elzaafarany O, Mokhtari S, Vogelbaum M, Bui M, Rasool G. Multi-Agent System for Automated Extraction of Neuro-Oncology Biomarkers from Pathology Reports. Abstract 521, USCAP 115th Annual Meeting. Laboratory Investigation. 2026;106(3 Suppl):104806. https://doi.org/10.1016/j.labinv.2025.104806 ↩ ↩2
-
Waqas A, Tripathi AG, Bowles K, Miner B, Islam JY, Coghill AE, Jones A, Schabath MB, Rasool G. Multi-agent AI orchestration for temporal-aware extraction of social determinants of health from unstructured clinical records in cancer populations. Abstract 26, AACR Annual Meeting 2026. Cancer Research. 2026;86(7 Suppl):26. https://doi.org/10.1158/1538-7445.AM2026-26 ↩ ↩2
Cite the paper
Tripathi A, Waqas A, Venkatesan K, Ullah E, Khan A, Khalil F, et al. Using Consensus-Based Reasoning and Large Language Models to Extract Structured Data From Surgical Pathology Reports. Lab Invest. 2026;106(2):104272. doi:10.1016/j.labinv.2025.104272
@article{tripathi2026consensus,
title = {Using Consensus-Based Reasoning and Large Language Models to Extract Structured Data From Surgical Pathology Reports},
author = {Tripathi, Aakash and Waqas, Asim and Venkatesan, Kavya and Ullah, Ehsan and Khan, Asma and Khalil, Farah and Chen, Wei-Shen and Ozturk, Zarifa Gahramanli and Saeed-Vafa, Daryoush and Bui, Marilyn M. and Schabath, Matthew B. and Rasool, Ghulam},
journal = {Laboratory Investigation},
year = {2026},
volume = {106},
number = {2},
pages = {104272},
doi = {10.1016/j.labinv.2025.104272}
}