 |
The Ultimate Crossover: Unifying Pixels, Genes, and Language in Spatial Biology
Pathology has traditionally been split: researchers either examine tissue morphology on H&E slides or sequence molecular profiles. To truly understand the tumor microenvironment, AI models need to simultaneously capture both the structure and the molecular signature. But integrating gigapixel images with sparse, high-dimensional gene matrices is difficult.
Three new papers are solving this by reimagining how AI processes multimodal spatial data.
• 𝙏𝙧𝙚𝙖𝙩𝙞𝙣𝙜 𝙂𝙚𝙣𝙚𝙨 𝘼𝙨 𝙇𝙖𝙣𝙜𝙪𝙖𝙜𝙚: Both 𝙒𝙚𝙞𝙦𝙞𝙣𝙜 𝘾𝙝𝙚𝙣 𝙚𝙩 𝙖𝙡. (OmiCLIP/Loki) and 𝙏𝙞𝙖𝙣𝙮𝙪 𝙇𝙞𝙪 𝙚𝙩 𝙖𝙡. (spEMO) take the approach of treating transcriptomics as text. OmiCLIP strings the top-expressed genes of a tissue patch into a sentence and uses contrastive learning to align this genomic text directly with the corresponding H&E image. Meanwhile, spEMO leverages pre-existing Large Language Models (LLMs) to embed biological text descriptions of genes and proteins, fusing them with pathology foundation models. This fused embedding allows spEMO to autonomously generate
highly accurate clinical medical reports that surpass human consistency.
• 𝙐𝙣𝙞𝙫𝙚𝙧𝙨𝙖𝙡 𝙎𝙥𝙖𝙩𝙞𝙖𝙡 𝙂𝙧𝙖𝙥𝙝𝙨: 𝙌𝙪𝙚𝙣𝙩𝙞𝙣 𝘽𝙡𝙖𝙢𝙥𝙚𝙮 𝙚𝙩 𝙖𝙡. focus on the physical cellular neighborhood with Novae. Rather than fusing images and text, Novae is a graph attention network trained on nearly 30 million cells across 18 tissues. While OmiCLIP and spEMO rely heavily on H&E image integration, Novae focuses entirely on creating a universal spatial transcriptomics embedding across different technologies. It learns to map cells within their spatial contexts regardless of the specific gene panel or technology used, natively correcting batch effects without needing external clustering tools like Harmony.
• 𝙕𝙚𝙧𝙤-𝙎𝙝𝙤𝙩 𝘾𝙖𝙥𝙖𝙗𝙞𝙡𝙞𝙩𝙞𝙚𝙨: A major similarity across all three architectures is their push toward zero-shot or highly adaptable inference without massive retraining. Whether it is the Loki platform retrieving a molecular profile directly from an unseen H&E image, spEMO projecting spatial spots into a joint image-text space to identify spatial domains, or Novae inferring hierarchical cross-slide spatial domains on the fly, these models are moving the field away from narrow, task-specific pipelines.
𝙏𝙝𝙚 𝙏𝙖𝙠𝙚𝙖𝙬𝙖𝙮: The future of spatial biology is not just about collecting more modalities; it is about building AI architectures that inherently understand how tissue morphology, genomics, and clinical language
intersect.
A visual–omics foundation model to bridge histopathology with spatial transcriptomics
spEMO: Leveraging Multi-Modal Foundation Models for Analyzing Spatial Multi-Omic and Histopathology Data
Novae: a graph-based foundation model for spatial transcriptomics data
|
|
|
|
|
|
 |
Multi-Modal Mamba Modeling for Survival Prediction (M4Survive): Adapting Joint Foundation Model Representations
Accurately predicting patient survival in oncology requires a comprehensive understanding of tumor biology. Yet, most clinical AI models focus on just one piece of the puzzle, evaluating either a radiology scan or a pathology slide in isolation.
Traditional single-modality approaches often fail to leverage the complementary insights provided by combining macroscopic radiological scans with microscopic pathological assessments. While fusing these diverse imaging modalities is essential to capture the complex interplay of a tumor, doing so is computationally expensive and difficult due to the differing structures of the data.
A new preprint by 𝙃𝙤 𝙃𝙞𝙣 𝙇𝙚𝙚 𝙚𝙩 𝙖𝙡. introduces 𝙈𝟰𝙎𝙪𝙧𝙫𝙞𝙫𝙚 (𝙈𝙪𝙡𝙩𝙞-𝙈𝙤𝙙𝙖𝙡 𝙈𝙖𝙢𝙗𝙖 𝙈𝙤𝙙𝙚𝙡𝙞𝙣𝙜 𝙛𝙤𝙧 𝙎𝙪𝙧𝙫𝙞𝙫𝙖𝙡 𝙋𝙧𝙚𝙙𝙞𝙘𝙩𝙞𝙤𝙣), a framework that bridges this gap by learning joint foundation model representations using highly efficient adapter networks.
Here are the key innovations:
• 𝘿𝙮𝙣𝙖𝙢𝙞𝙘 𝙁𝙤𝙪𝙣𝙙𝙖𝙩𝙞𝙤𝙣
𝙈𝙤𝙙𝙚𝙡 𝙁𝙪𝙨𝙞𝙤𝙣: Rather than training a massive multi-modal model from scratch, 𝙈𝟰𝙎𝙪𝙧𝙫𝙞𝙫𝙚 dynamically fuses heterogeneous embeddings from leading foundation models in the field, including 𝙈𝙚𝙙𝙄𝙢𝙖𝙜𝙚𝙄𝙣𝙨𝙞𝙜𝙝𝙩, 𝘽𝙞𝙤𝙢𝙚𝙙𝘾𝙇𝙄𝙋, 𝙋𝙧𝙤𝙫-𝙂𝙞𝙜𝙖𝙋𝙖𝙩𝙝, 𝙖𝙣𝙙 𝙐𝙉𝙄𝟮-𝙝. This approach creates a correlated latent space specifically optimized for estimating survival risk.
• 𝙈𝙖𝙢𝙗𝙖-𝘽𝙖𝙨𝙚𝙙 𝘼𝙙𝙖𝙥𝙩𝙚𝙧𝙨: To handle the integration of these distinct and heavy embeddings, the framework utilizes Mamba-based adapter networks. This enables effective multi-modal learning while rigorously preserving computational efficiency.
• 𝙎𝙪𝙥𝙚𝙧𝙞𝙤𝙧 𝘼𝙘𝙘𝙪𝙧𝙖𝙘𝙮: In experimental benchmark evaluations, this dynamic framework successfully outperforms both unimodal baselines and traditional static multi-modal models in survival prediction accuracy.
𝙏𝙝𝙚 𝙏𝙖𝙠𝙚𝙖𝙬𝙖𝙮: 𝙈𝟰𝙎𝙪𝙧𝙫𝙞𝙫𝙚 demonstrates that the future of predictive analytics
and precision oncology relies not just on building individual foundation models, but on efficiently fusing them to create a holistic view of patient health.
|
|
|
|
|
|
 |
The Era of Virtual Spatial Omics: Predicting Molecular Maps from H&E
Spatial omics provides unprecedented detail into the tumor microenvironment, but its high cost and technical complexity keep it confined to small research cohorts. What if we could generate these spatial maps directly from a standard H&E slide?
Recent breakthroughs in deep learning have made this a reality. By leveraging foundation models, researchers are now translating routine histology images into high-resolution spatial transcriptomics and proteomics, opening the door for massive-scale biomarker discovery. Three new papers showcase the power of this virtual approach.
Here is how they compare:
• 𝗧𝗿𝗮𝗻𝘀𝗰𝗿𝗶𝗽𝘁𝗼𝗺𝗶𝗰𝘀 𝘃𝘀. 𝗣𝗿𝗼𝘁𝗲𝗼𝗺𝗶𝗰𝘀: While 𝗣𝗮𝘁𝗵𝟮𝗦𝗽𝗮𝗰𝗲 and 𝗗𝗲𝗲𝗽𝗦𝗽𝗼𝘁 focus on predicting spatial gene expression (RNA), the 𝗛𝗘𝗫 framework shifts the focus to
𝘱𝘳𝘰𝘵𝘦𝘰𝘮𝘪𝘤𝘴. 𝗛𝗘𝗫 predicts 40 targeted protein biomarkers (like CODEX), arguing that proteins are often more closely related to cellular functions and clinical outcomes than RNA transcripts.
• 𝗔𝗿𝗰𝗵𝗶𝘁𝗲𝗰𝘁𝘂𝗿𝗮𝗹 𝗜𝗻𝗻𝗼𝘃𝗮𝘁𝗶𝗼𝗻𝘀 𝗳𝗼𝗿 𝗥𝗲𝘀𝗼𝗹𝘂𝘁𝗶𝗼𝗻: Spatial spots often contain multiple cells, muddying the signal. 𝗗𝗲𝗲𝗽𝗦𝗽𝗼𝘁 tackles this by using a deep-set neural network, treating each transcriptomic spot as a bag of sub-spots to capture local morphology alongside global tissue context. Its successor, 𝗗𝗲𝗲𝗽𝗦𝗽𝗼𝘁𝟮𝗖𝗲𝗹𝗹, pushes this even further to virtual single-cell resolution. Alternatively, 𝗣𝗮𝘁𝗵𝟮𝗦𝗽𝗮𝗰𝗲 utilizes spatial smoothing and
targeted cell-type deconvolutions to extract localized cell abundance directly from the inferred gene expression.
• 𝗠𝗮𝘀𝘀𝗶𝘃𝗲 𝗦𝗰𝗮𝗹𝗲 𝗮𝗻𝗱 𝗖𝗹𝗶𝗻𝗶𝗰𝗮𝗹 𝗜𝗻𝘁𝗲𝗴𝗿𝗮𝘁𝗶𝗼𝗻: Because virtual omics are highly cost-effective, they enable unprecedented scale. 𝗗𝗲𝗲𝗽𝗦𝗽𝗼𝘁 generated a massive resource of 56 million virtual spots across 3,780 TCGA patients. 𝗣𝗮𝘁𝗵𝟮𝗦𝗽𝗮𝗰𝗲 applied its predictions to large breast cancer cohorts to identify SpatioTypes that predict chemotherapy and trastuzumab response. 𝗛𝗘𝗫 took it a step further with its 𝗠𝗜𝗖𝗔 integration framework, fusing H&E images with virtual proteomics to significantly outperform
traditional clinical risk factors in predicting immunotherapy response.
𝘛𝘩𝘦 𝘛𝘢𝘬𝘦𝘢𝘸𝘢𝘺: Virtual spatial omics will not completely replace physical sequencing, but it acts as a powerful, scalable bridge. By transforming archival H&E slides into multi-layered molecular maps, we are unlocking population-scale data essential for true precision medicine.
DeepSpot: Leveraging Spatial Context for Enhanced Spatial Transcriptomics Prediction from H&E Images
DeepSpot2Cell: Predicting Virtual Single-Cell Spatial Transcriptomics from H&E images using Spot-Level Supervision
AI-Driven Spatial Transcriptomics Unlocks Large-Scale Breast Cancer Biomarker Discovery from Histopathology
AI-enabled virtual spatial proteomics from histopathology for interpretable biomarker discovery in lung cancer
|
|
|
|
|
|
 |
Rethinking Tissue Architecture: AI Beyond the Isolated Patch
Computational pathology models often divide tissue slides into isolated 2D tiles for processing. Yet, biology doesn't operate in isolated boxes; tissue is a continuous, interconnected environment and a complex 3D volume.
To accurately predict spatial transcriptomics and molecular signatures, models need to understand cellular neighborhoods, subtle spatial frequencies, and 3D depth. Four new papers introduce advanced architectures designed to capture this::
• 𝘾𝙖𝙥𝙩𝙪𝙧𝙞𝙣𝙜 𝘾𝙚𝙡𝙡𝙪𝙡𝙖𝙧 𝙉𝙚𝙞𝙜𝙝𝙗𝙤𝙧𝙝𝙤𝙤𝙙𝙨: Both 𝙂𝙤𝙣𝙜 𝙚𝙩 𝙖𝙡. and𝙈𝙖𝙧𝙠𝙚𝙮 𝙚𝙩 𝙖𝙡. focus on how localized regions interact, but at different scales. 𝙂𝙤𝙣𝙜 𝙚𝙩 𝙖𝙡. introduce 𝘼𝙄𝘿𝙊.𝙏𝙞𝙨𝙨𝙪𝙚, an architecture that explicitly feeds multiple neighboring cells into an asymmetrical encoder-decoder to effectively learn cross-cell dependencies at the single-cell level. Conversely, 𝙈𝙖𝙧𝙠𝙚𝙮 𝙚𝙩
𝙖𝙡. focus on macro-level spatial interpretability across the whole slide with 𝙖𝙈𝙄𝙇. By replacing standard black-box pooling with an additive aggregation function, 𝙖𝙈𝙄𝙇 ensures every individual tissue patch contributes quantifiably to a slide-level gene signature, generating highly granular spatial heatmaps without needing patch-level annotations.
• 𝙉𝙚𝙬 𝘽𝙖𝙘𝙠𝙗𝙤𝙣𝙚𝙨 𝘼𝙣𝙙 𝘿𝙞𝙢𝙚𝙣𝙨𝙞𝙤𝙣𝙨: While the first two papers modify context and aggregation, 𝘾𝙝𝙤 𝙚𝙩 𝙖𝙡. and 𝙕𝙝𝙪 𝙚𝙩 𝙖𝙡. alter the core dimensions of how features are processed. 𝘾𝙝𝙤 𝙚𝙩 𝙖𝙡. present 𝙈𝙑𝙃𝙮𝙗𝙧𝙞𝙙, arguing that standard vision transformers struggle with biomarker prediction because they fail to capture subtle, low-frequency morphological patterns. By integrating state space models tuned with negative real eigenvalues, their hybrid backbone explicitly biases the network to preserve these critical low-frequency biological
signals.
• 𝘽𝙧𝙚𝙖𝙠𝙞𝙣𝙜 𝙏𝙝𝙚 2𝘿 𝘽𝙖𝙧𝙧𝙞𝙚𝙧: Meanwhile, 𝙕𝙝𝙪 𝙚𝙩 𝙖𝙡. shatter the 2D limitation entirely with 𝘼𝙎𝙄𝙂𝙉. Recognizing that full 3D spatial transcriptomics is prohibitively expensive, their graph network extends spatial relationships into the z-axis. It imputes a 3D spatial transcriptomic volume by combining a stack of H&E sections with just a single 2D spatial transcriptomic slide.
𝙏𝙝𝙚 𝙏𝙖𝙠𝙚𝙖𝙬𝙖𝙮: The next generation of pathology AI will not just rely on training with more data. It requires architectures that inherently reflect the physical reality of human tissue—whether through cell-neighborhood inputs, frequency-biased state
space models, or true 3D spatial graphs.
AIDO.Tissue: Spatial Cell-Guided Pretraining for Scalable Spatial Transcriptomics Foundation Model
Spatial Mapping of Gene Signatures in Hematoxylin and Eosin-Stained Images: A Proof of Concept for Interpretable Predictions Using Additive Multiple Instance Learning
MVHybrid: Improving Spatial Transcriptomics Prediction with Hybrid State Space-Vision Transformer Backbone in Pathology Vision Foundation Models
ASIGN: An Anatomy-aware Spatial Imputation Graphic Network for 3D Spatial Transcriptomics
|
|
|
|
|
|
 |
HONeYBEE: enabling scalable multimodal AI in oncology through foundation model-driven embeddings
Oncology data is inherently multimodal—combining radiology scans, pathology slides, genomics, and clinical notes. Yet, most AI models are trapped in single-modality silos, missing the complete biological picture and underutilizing complementary information.
Integrating these heterogeneous data types into a unified patient representation is complex due to fragmented tools and rigid code dependencies. 𝘼𝙖𝙠𝙖𝙨𝙝 𝙏𝙧𝙞𝙥𝙖𝙩𝙝𝙞 𝙚𝙩 𝙖𝙡. published a comprehensive solution, 𝙃𝙊𝙉𝙚𝙔𝘽𝙀𝙀 (Harmonized ONcologY Biomedical Embedding Encoder). This
open-source framework generates and integrates patient-level embeddings using domain-specific foundation models.
Here are the key innovations from their evaluation of over 11,400 patients across 33 cancer types:
• 𝙐𝙣𝙞𝙛𝙞𝙚𝙙 𝙈𝙪𝙡𝙩𝙞𝙢𝙤𝙙𝙖𝙡 𝙋𝙞𝙥𝙚𝙡𝙞𝙣𝙚: 𝙃𝙊𝙉𝙚𝙔𝘽𝙀𝙀 processes five distinct data types—clinical text, pathology reports, radiologic images, whole slide images (WSIs), and molecular profiles—through specialized preprocessing pipelines. Crucially, its modular design accommodates patients with missing data modalities without requiring complete-case cohorts.
• 𝙏𝙝𝙚 𝙋𝙤𝙬𝙚𝙧 𝙊𝙛 𝘾𝙡𝙞𝙣𝙞𝙘𝙖𝙡 𝙏𝙚𝙭𝙩: In an interesting reality check, clinical embeddings derived from structured and unstructured data actually showed the strongest single-modality performance, achieving 98.5% classification accuracy and the highest overall survival prediction concordance indices. The authors note this reflects the expert-curated nature of clinical documentation in datasets like TCGA, which effectively summarizes information dispersed across other raw modalities.
• 𝙁𝙪𝙨𝙞𝙤𝙣 𝙁𝙤𝙧 𝙎𝙪𝙧𝙫𝙞𝙫𝙖𝙡: While clinical data dominated, multimodal fusion strategies (such as concatenation and Kronecker product) provided critical complementary benefits.
For specific cancers, fusing information from molecular, pathology, and imaging modalities significantly improved overall survival predictions beyond what clinical features could capture alone.
• 𝙇𝙇𝙈𝙨 𝙋𝙪𝙩 𝙏𝙤 𝙏𝙝𝙚 𝙏𝙚𝙨𝙩: The team compared four large language models to evaluate text embeddings. They found that general-purpose models (like Qwen3) actually outperformed specialized medical models (like GatorTron) on standard clinical text. However, task-specific fine-tuning proved essential across all models to achieve high performance on messy, heterogeneous data like pathology reports.
𝙏𝙝𝙚 𝙏𝙖𝙠𝙚𝙖𝙬𝙖𝙮: The future of precision oncology relies not just on building individual foundation models, but on creating scalable, open-source infrastructure that can standardize and unify these distinct representations into a cohesive clinical picture.
|
|
|
|
|
|
|