 |
|
The Ultimate Crossover: Unifying Pixels, Genes, and Language in Spatial Biology
Pathology has traditionally been split: researchers either examine tissue morphology on H&E slides or sequence molecular profiles. To truly understand the tumor microenvironment, AI models need to simultaneously capture both the structure and the molecular signature. But integrating gigapixel images with sparse, high-dimensional gene matrices is difficult.
Three new papers are solving this by reimagining how AI processes multimodal spatial data.
โข ๐๐ง๐๐๐ฉ๐๐ฃ๐ ๐๐๐ฃ๐๐จ ๐ผ๐จ ๐๐๐ฃ๐๐ช๐๐๐: Both ๐๐๐๐ฆ๐๐ฃ๐ ๐พ๐๐๐ฃ ๐๐ฉ ๐๐ก. (OmiCLIP/Loki) and ๐๐๐๐ฃ๐ฎ๐ช ๐๐๐ช ๐๐ฉ ๐๐ก. (spEMO) take the approach of treating transcriptomics as text. OmiCLIP strings the top-expressed genes of a tissue patch into a sentence and uses contrastive learning to align this genomic text directly with the corresponding H&E image. Meanwhile, spEMO leverages pre-existing Large Language Models (LLMs) to embed biological text descriptions of genes and proteins, fusing them with pathology foundation models. This fused embedding allows spEMO to autonomously generate highly accurate clinical medical reports that surpass human consistency.
โข ๐๐ฃ๐๐ซ๐๐ง๐จ๐๐ก ๐๐ฅ๐๐ฉ๐๐๐ก ๐๐ง๐๐ฅ๐๐จ: ๐๐ช๐๐ฃ๐ฉ๐๐ฃ ๐ฝ๐ก๐๐ข๐ฅ๐๐ฎ ๐๐ฉ ๐๐ก. focus on the physical cellular neighborhood with Novae. Rather than fusing images and text, Novae is a graph attention network trained on nearly 30 million cells across 18 tissues. While OmiCLIP and spEMO rely heavily on H&E image integration, Novae focuses entirely on creating a universal spatial transcriptomics embedding across different technologies. It learns to map cells within their spatial contexts regardless of the specific gene panel or technology used, natively correcting batch effects without needing external clustering tools like Harmony.
โข ๐๐๐ง๐ค-๐๐๐ค๐ฉ ๐พ๐๐ฅ๐๐๐๐ก๐๐ฉ๐๐๐จ: A major similarity across all three architectures is their push toward zero-shot or highly adaptable inference without massive retraining. Whether it is the Loki platform retrieving a molecular profile directly from an unseen H&E image, spEMO projecting spatial spots into a joint image-text space to identify spatial domains, or Novae inferring hierarchical cross-slide spatial domains on the fly, these models are moving the field away from narrow, task-specific pipelines.
๐๐๐ ๐๐๐ ๐๐๐ฌ๐๐ฎ: The future of spatial biology is not just about collecting more modalities; it is about building AI architectures that inherently understand how tissue morphology, genomics, and clinical language intersect.
A visualโomics foundation model to bridge histopathology with spatial transcriptomics
spEMO: Leveraging Multi-Modal Foundation Models for Analyzing Spatial Multi-Omic and Histopathology Data
Novae: a graph-based foundation model for spatial transcriptomics data
|
|
|
|
|
|
 |
|
Multi-Modal Mamba Modeling for Survival Prediction (M4Survive): Adapting Joint Foundation Model Representations
Accurately predicting patient survival in oncology requires a comprehensive understanding of tumor biology. Yet, most clinical AI models focus on just one piece of the puzzle, evaluating either a radiology scan or a pathology slide in isolation.
Traditional single-modality approaches often fail to leverage the complementary insights provided by combining macroscopic radiological scans with microscopic pathological assessments. While fusing these diverse imaging modalities is essential to capture the complex interplay of a tumor, doing so is computationally expensive and difficult due to the differing structures of the data.
A new preprint by ๐๐ค ๐๐๐ฃ ๐๐๐ ๐๐ฉ ๐๐ก. introduces ๐๐ฐ๐๐ช๐ง๐ซ๐๐ซ๐ (๐๐ช๐ก๐ฉ๐-๐๐ค๐๐๐ก ๐๐๐ข๐๐ ๐๐ค๐๐๐ก๐๐ฃ๐ ๐๐ค๐ง ๐๐ช๐ง๐ซ๐๐ซ๐๐ก ๐๐ง๐๐๐๐๐ฉ๐๐ค๐ฃ), a framework that bridges this gap by learning joint foundation model representations using highly efficient adapter networks.
Here are the key innovations:
โข ๐ฟ๐ฎ๐ฃ๐๐ข๐๐ ๐๐ค๐ช๐ฃ๐๐๐ฉ๐๐ค๐ฃ ๐๐ค๐๐๐ก ๐๐ช๐จ๐๐ค๐ฃ: Rather than training a massive multi-modal model from scratch, ๐๐ฐ๐๐ช๐ง๐ซ๐๐ซ๐ dynamically fuses heterogeneous embeddings from leading foundation models in the field, including ๐๐๐๐๐ข๐๐๐๐๐ฃ๐จ๐๐๐๐ฉ, ๐ฝ๐๐ค๐ข๐๐๐พ๐๐๐, ๐๐ง๐ค๐ซ-๐๐๐๐๐๐๐ฉ๐, ๐๐ฃ๐ ๐๐๐๐ฎ-๐. This approach creates a correlated latent space specifically optimized for estimating survival risk.
โข ๐๐๐ข๐๐-๐ฝ๐๐จ๐๐ ๐ผ๐๐๐ฅ๐ฉ๐๐ง๐จ: To handle the integration of these distinct and heavy embeddings, the framework utilizes Mamba-based adapter networks. This enables effective multi-modal learning while rigorously preserving computational efficiency.
โข ๐๐ช๐ฅ๐๐ง๐๐ค๐ง ๐ผ๐๐๐ช๐ง๐๐๐ฎ: In experimental benchmark evaluations, this dynamic framework successfully outperforms both unimodal baselines and traditional static multi-modal models in survival prediction accuracy.
๐๐๐ ๐๐๐ ๐๐๐ฌ๐๐ฎ: ๐๐ฐ๐๐ช๐ง๐ซ๐๐ซ๐ demonstrates that the future of predictive analytics and precision oncology relies not just on building individual foundation models, but on efficiently fusing them to create a holistic view of patient health.
|
|
|
|
|
|
 |
|
The Era of Virtual Spatial Omics: Predicting Molecular Maps from H&E
Spatial omics provides unprecedented detail into the tumor microenvironment, but its high cost and technical complexity keep it confined to small research cohorts. What if we could generate these spatial maps directly from a standard H&E slide?
Recent breakthroughs in deep learning have made this a reality. By leveraging foundation models, researchers are now translating routine histology images into high-resolution spatial transcriptomics and proteomics, opening the door for massive-scale biomarker discovery. Three new papers showcase the power of this virtual approach.
Here is how they compare:
โข ๐ง๐ฟ๐ฎ๐ป๐๐ฐ๐ฟ๐ถ๐ฝ๐๐ผ๐บ๐ถ๐ฐ๐ ๐๐. ๐ฃ๐ฟ๐ผ๐๐ฒ๐ผ๐บ๐ถ๐ฐ๐: While ๐ฃ๐ฎ๐๐ต๐ฎ๐ฆ๐ฝ๐ฎ๐ฐ๐ฒ and ๐๐ฒ๐ฒ๐ฝ๐ฆ๐ฝ๐ผ๐ focus on predicting spatial gene expression (RNA), the ๐๐๐ซ framework shifts the focus to ๐ฑ๐ณ๐ฐ๐ต๐ฆ๐ฐ๐ฎ๐ช๐ค๐ด. ๐๐๐ซ predicts 40 targeted protein biomarkers (like CODEX), arguing that proteins are often more closely related to cellular functions and clinical outcomes than RNA transcripts.
โข ๐๐ฟ๐ฐ๐ต๐ถ๐๐ฒ๐ฐ๐๐๐ฟ๐ฎ๐น ๐๐ป๐ป๐ผ๐๐ฎ๐๐ถ๐ผ๐ป๐ ๐ณ๐ผ๐ฟ ๐ฅ๐ฒ๐๐ผ๐น๐๐๐ถ๐ผ๐ป: Spatial spots often contain multiple cells, muddying the signal. ๐๐ฒ๐ฒ๐ฝ๐ฆ๐ฝ๐ผ๐ tackles this by using a deep-set neural network, treating each transcriptomic spot as a bag of sub-spots to capture local morphology alongside global tissue context. Its successor, ๐๐ฒ๐ฒ๐ฝ๐ฆ๐ฝ๐ผ๐๐ฎ๐๐ฒ๐น๐น, pushes this even further to virtual single-cell resolution. Alternatively, ๐ฃ๐ฎ๐๐ต๐ฎ๐ฆ๐ฝ๐ฎ๐ฐ๐ฒ utilizes spatial smoothing and targeted cell-type deconvolutions to extract localized cell abundance directly from the inferred gene expression.
โข ๐ ๐ฎ๐๐๐ถ๐๐ฒ ๐ฆ๐ฐ๐ฎ๐น๐ฒ ๐ฎ๐ป๐ฑ ๐๐น๐ถ๐ป๐ถ๐ฐ๐ฎ๐น ๐๐ป๐๐ฒ๐ด๐ฟ๐ฎ๐๐ถ๐ผ๐ป: Because virtual omics are highly cost-effective, they enable unprecedented scale. ๐๐ฒ๐ฒ๐ฝ๐ฆ๐ฝ๐ผ๐ generated a massive resource of 56 million virtual spots across 3,780 TCGA patients. ๐ฃ๐ฎ๐๐ต๐ฎ๐ฆ๐ฝ๐ฎ๐ฐ๐ฒ applied its predictions to large breast cancer cohorts to identify SpatioTypes that predict chemotherapy and trastuzumab response. ๐๐๐ซ took it a step further with its ๐ ๐๐๐ integration framework, fusing H&E images with virtual proteomics to significantly outperform traditional clinical risk factors in predicting immunotherapy response.
๐๐ฉ๐ฆ ๐๐ข๐ฌ๐ฆ๐ข๐ธ๐ข๐บ: Virtual spatial omics will not completely replace physical sequencing, but it acts as a powerful, scalable bridge. By transforming archival H&E slides into multi-layered molecular maps, we are unlocking population-scale data essential for true precision medicine.
DeepSpot: Leveraging Spatial Context for Enhanced Spatial Transcriptomics Prediction from H&E Images
DeepSpot2Cell: Predicting Virtual Single-Cell Spatial Transcriptomics from H&E images using Spot-Level Supervision
AI-Driven Spatial Transcriptomics Unlocks Large-Scale Breast Cancer Biomarker Discovery from Histopathology
AI-enabled virtual spatial proteomics from histopathology for interpretable biomarker discovery in lung cancer |
|
|
|
|
|
 |
|
Rethinking Tissue Architecture: AI Beyond the Isolated Patch
Computational pathology models often divide tissue slides into isolated 2D tiles for processing. Yet, biology doesn't operate in isolated boxes; tissue is a continuous, interconnected environment and a complex 3D volume.
To accurately predict spatial transcriptomics and molecular signatures, models need to understand cellular neighborhoods, subtle spatial frequencies, and 3D depth. Four new papers introduce advanced architectures designed to capture this::
โข ๐พ๐๐ฅ๐ฉ๐ช๐ง๐๐ฃ๐ ๐พ๐๐ก๐ก๐ช๐ก๐๐ง ๐๐๐๐๐๐๐ค๐ง๐๐ค๐ค๐๐จ: Both ๐๐ค๐ฃ๐ ๐๐ฉ ๐๐ก. and๐๐๐ง๐ ๐๐ฎ ๐๐ฉ ๐๐ก. focus on how localized regions interact, but at different scales. ๐๐ค๐ฃ๐ ๐๐ฉ ๐๐ก. introduce ๐ผ๐๐ฟ๐.๐๐๐จ๐จ๐ช๐, an architecture that explicitly feeds multiple neighboring cells into an asymmetrical encoder-decoder to effectively learn cross-cell dependencies at the single-cell level. Conversely, ๐๐๐ง๐ ๐๐ฎ ๐๐ฉ ๐๐ก. focus on macro-level spatial interpretability across the whole slide with ๐๐๐๐. By replacing standard black-box pooling with an additive aggregation function, ๐๐๐๐ ensures every individual tissue patch contributes quantifiably to a slide-level gene signature, generating highly granular spatial heatmaps without needing patch-level annotations.
โข ๐๐๐ฌ ๐ฝ๐๐๐ ๐๐ค๐ฃ๐๐จ ๐ผ๐ฃ๐ ๐ฟ๐๐ข๐๐ฃ๐จ๐๐ค๐ฃ๐จ: While the first two papers modify context and aggregation, ๐พ๐๐ค ๐๐ฉ ๐๐ก. and ๐๐๐ช ๐๐ฉ ๐๐ก. alter the core dimensions of how features are processed. ๐พ๐๐ค ๐๐ฉ ๐๐ก. present ๐๐๐๐ฎ๐๐ง๐๐, arguing that standard vision transformers struggle with biomarker prediction because they fail to capture subtle, low-frequency morphological patterns. By integrating state space models tuned with negative real eigenvalues, their hybrid backbone explicitly biases the network to preserve these critical low-frequency biological signals.
โข ๐ฝ๐ง๐๐๐ ๐๐ฃ๐ ๐๐๐ 2๐ฟ ๐ฝ๐๐ง๐ง๐๐๐ง: Meanwhile, ๐๐๐ช ๐๐ฉ ๐๐ก. shatter the 2D limitation entirely with ๐ผ๐๐๐๐. Recognizing that full 3D spatial transcriptomics is prohibitively expensive, their graph network extends spatial relationships into the z-axis. It imputes a 3D spatial transcriptomic volume by combining a stack of H&E sections with just a single 2D spatial transcriptomic slide.
๐๐๐ ๐๐๐ ๐๐๐ฌ๐๐ฎ: The next generation of pathology AI will not just rely on training with more data. It requires architectures that inherently reflect the physical reality of human tissueโwhether through cell-neighborhood inputs, frequency-biased state space models, or true 3D spatial graphs.
AIDO.Tissue: Spatial Cell-Guided Pretraining for Scalable Spatial Transcriptomics Foundation Model
Spatial Mapping of Gene Signatures in Hematoxylin and Eosin-Stained Images: A Proof of Concept for Interpretable Predictions Using Additive Multiple Instance Learning
MVHybrid: Improving Spatial Transcriptomics Prediction with Hybrid State Space-Vision Transformer Backbone in Pathology Vision Foundation Models
ASIGN: An Anatomy-aware Spatial Imputation Graphic Network for 3D Spatial Transcriptomics
|
|
|
|
|
|
 |
|
HONeYBEE: enabling scalable multimodal AI in oncology through foundation model-driven embeddings
Oncology data is inherently multimodalโcombining radiology scans, pathology slides, genomics, and clinical notes. Yet, most AI models are trapped in single-modality silos, missing the complete biological picture and underutilizing complementary information.
Integrating these heterogeneous data types into a unified patient representation is complex due to fragmented tools and rigid code dependencies. ๐ผ๐๐ ๐๐จ๐ ๐๐ง๐๐ฅ๐๐ฉ๐๐ ๐๐ฉ ๐๐ก. published a comprehensive solution, ๐๐๐๐๐๐ฝ๐๐ (Harmonized ONcologY Biomedical Embedding Encoder). This open-source framework generates and integrates patient-level embeddings using domain-specific foundation models.
Here are the key innovations from their evaluation of over 11,400 patients across 33 cancer types:
โข ๐๐ฃ๐๐๐๐๐ ๐๐ช๐ก๐ฉ๐๐ข๐ค๐๐๐ก ๐๐๐ฅ๐๐ก๐๐ฃ๐: ๐๐๐๐๐๐ฝ๐๐ processes five distinct data typesโclinical text, pathology reports, radiologic images, whole slide images (WSIs), and molecular profilesโthrough specialized preprocessing pipelines. Crucially, its modular design accommodates patients with missing data modalities without requiring complete-case cohorts.
โข ๐๐๐ ๐๐ค๐ฌ๐๐ง ๐๐ ๐พ๐ก๐๐ฃ๐๐๐๐ก ๐๐๐ญ๐ฉ: In an interesting reality check, clinical embeddings derived from structured and unstructured data actually showed the strongest single-modality performance, achieving 98.5% classification accuracy and the highest overall survival prediction concordance indices. The authors note this reflects the expert-curated nature of clinical documentation in datasets like TCGA, which effectively summarizes information dispersed across other raw modalities.
โข ๐๐ช๐จ๐๐ค๐ฃ ๐๐ค๐ง ๐๐ช๐ง๐ซ๐๐ซ๐๐ก: While clinical data dominated, multimodal fusion strategies (such as concatenation and Kronecker product) provided critical complementary benefits. For specific cancers, fusing information from molecular, pathology, and imaging modalities significantly improved overall survival predictions beyond what clinical features could capture alone.
โข ๐๐๐๐จ ๐๐ช๐ฉ ๐๐ค ๐๐๐ ๐๐๐จ๐ฉ: The team compared four large language models to evaluate text embeddings. They found that general-purpose models (like Qwen3) actually outperformed specialized medical models (like GatorTron) on standard clinical text. However, task-specific fine-tuning proved essential across all models to achieve high performance on messy, heterogeneous data like pathology reports.
๐๐๐ ๐๐๐ ๐๐๐ฌ๐๐ฎ: The future of precision oncology relies not just on building individual foundation models, but on creating scalable, open-source infrastructure that can standardize and unify these distinct representations into a cohesive clinical picture.
|
|
|
|
|
|
|
Enjoy this newsletter? Here are more things you might find helpful:
Pixel Clarity Call - A free 30-minute conversation to cut through the noise and see where your vision AI project really stands. Weโll pinpoint vulnerabilities, clarify your biggest challenges, and decide if an assessment or diagnostic could save you time, money, and credibility.
Book now |
|
|
|
Did someone forward this email to you, and you want to sign up for more? Subscribe to future emails
This email was sent to _t.e.s.t_@example.com. Want to change to a different address? Update subscription
Want to get off this list? Unsubscribe
My postal address: Pixel Scientia Labs, LLC, PO Box 98412, Raleigh, NC 27624, United States |
|
|
|
|