Share


The Ultimate Crossover: Unifying Pixels, Genes, and Language in Spatial Biology


Pathology has traditionally been split: researchers either examine tissue morphology on H&E slides or sequence molecular profiles. To truly understand the tumor microenvironment, AI models need to simultaneously capture both the structure and the molecular signature. But integrating gigapixel images with sparse, high-dimensional gene matrices is difficult.

Three new papers are solving this by reimagining how AI processes multimodal spatial data.

โ€ข ๐™๐™ง๐™š๐™–๐™ฉ๐™ž๐™ฃ๐™œ ๐™‚๐™š๐™ฃ๐™š๐™จ ๐˜ผ๐™จ ๐™‡๐™–๐™ฃ๐™œ๐™ช๐™–๐™œ๐™š: Both ๐™’๐™š๐™ž๐™ฆ๐™ž๐™ฃ๐™œ ๐˜พ๐™๐™š๐™ฃ ๐™š๐™ฉ ๐™–๐™ก. (OmiCLIP/Loki) and ๐™๐™ž๐™–๐™ฃ๐™ฎ๐™ช ๐™‡๐™ž๐™ช ๐™š๐™ฉ ๐™–๐™ก. (spEMO) take the approach of treating transcriptomics as text. OmiCLIP strings the top-expressed genes of a tissue patch into a sentence and uses contrastive learning to align this genomic text directly with the corresponding H&E image. Meanwhile, spEMO leverages pre-existing Large Language Models (LLMs) to embed biological text descriptions of genes and proteins, fusing them with pathology foundation models. This fused embedding allows spEMO to autonomously generate highly accurate clinical medical reports that surpass human consistency.

โ€ข ๐™๐™ฃ๐™ž๐™ซ๐™š๐™ง๐™จ๐™–๐™ก ๐™Ž๐™ฅ๐™–๐™ฉ๐™ž๐™–๐™ก ๐™‚๐™ง๐™–๐™ฅ๐™๐™จ: ๐™Œ๐™ช๐™š๐™ฃ๐™ฉ๐™ž๐™ฃ ๐˜ฝ๐™ก๐™–๐™ข๐™ฅ๐™š๐™ฎ ๐™š๐™ฉ ๐™–๐™ก. focus on the physical cellular neighborhood with Novae. Rather than fusing images and text, Novae is a graph attention network trained on nearly 30 million cells across 18 tissues. While OmiCLIP and spEMO rely heavily on H&E image integration, Novae focuses entirely on creating a universal spatial transcriptomics embedding across different technologies. It learns to map cells within their spatial contexts regardless of the specific gene panel or technology used, natively correcting batch effects without needing external clustering tools like Harmony.

โ€ข ๐™•๐™š๐™ง๐™ค-๐™Ž๐™๐™ค๐™ฉ ๐˜พ๐™–๐™ฅ๐™–๐™—๐™ž๐™ก๐™ž๐™ฉ๐™ž๐™š๐™จ: A major similarity across all three architectures is their push toward zero-shot or highly adaptable inference without massive retraining. Whether it is the Loki platform retrieving a molecular profile directly from an unseen H&E image, spEMO projecting spatial spots into a joint image-text space to identify spatial domains, or Novae inferring hierarchical cross-slide spatial domains on the fly, these models are moving the field away from narrow, task-specific pipelines.

๐™๐™๐™š ๐™๐™–๐™ ๐™š๐™–๐™ฌ๐™–๐™ฎ: The future of spatial biology is not just about collecting more modalities; it is about building AI architectures that inherently understand how tissue morphology, genomics, and clinical language intersect.


A visualโ€“omics foundation model to bridge histopathology with spatial transcriptomics


spEMO: Leveraging Multi-Modal Foundation Models for Analyzing Spatial Multi-Omic and Histopathology Data

Novae: a graph-based foundation model for spatial transcriptomics data


Multi-Modal Mamba Modeling for Survival Prediction (M4Survive): Adapting Joint Foundation Model Representations


Accurately predicting patient survival in oncology requires a comprehensive understanding of tumor biology. Yet, most clinical AI models focus on just one piece of the puzzle, evaluating either a radiology scan or a pathology slide in isolation.

Traditional single-modality approaches often fail to leverage the complementary insights provided by combining macroscopic radiological scans with microscopic pathological assessments. While fusing these diverse imaging modalities is essential to capture the complex interplay of a tumor, doing so is computationally expensive and difficult due to the differing structures of the data.

A new preprint by ๐™ƒ๐™ค ๐™ƒ๐™ž๐™ฃ ๐™‡๐™š๐™š ๐™š๐™ฉ ๐™–๐™ก. introduces ๐™ˆ๐Ÿฐ๐™Ž๐™ช๐™ง๐™ซ๐™ž๐™ซ๐™š (๐™ˆ๐™ช๐™ก๐™ฉ๐™ž-๐™ˆ๐™ค๐™™๐™–๐™ก ๐™ˆ๐™–๐™ข๐™—๐™– ๐™ˆ๐™ค๐™™๐™š๐™ก๐™ž๐™ฃ๐™œ ๐™›๐™ค๐™ง ๐™Ž๐™ช๐™ง๐™ซ๐™ž๐™ซ๐™–๐™ก ๐™‹๐™ง๐™š๐™™๐™ž๐™˜๐™ฉ๐™ž๐™ค๐™ฃ), a framework that bridges this gap by learning joint foundation model representations using highly efficient adapter networks.

Here are the key innovations:

โ€ข ๐˜ฟ๐™ฎ๐™ฃ๐™–๐™ข๐™ž๐™˜ ๐™๐™ค๐™ช๐™ฃ๐™™๐™–๐™ฉ๐™ž๐™ค๐™ฃ ๐™ˆ๐™ค๐™™๐™š๐™ก ๐™๐™ช๐™จ๐™ž๐™ค๐™ฃ: Rather than training a massive multi-modal model from scratch, ๐™ˆ๐Ÿฐ๐™Ž๐™ช๐™ง๐™ซ๐™ž๐™ซ๐™š dynamically fuses heterogeneous embeddings from leading foundation models in the field, including ๐™ˆ๐™š๐™™๐™„๐™ข๐™–๐™œ๐™š๐™„๐™ฃ๐™จ๐™ž๐™œ๐™๐™ฉ, ๐˜ฝ๐™ž๐™ค๐™ข๐™š๐™™๐˜พ๐™‡๐™„๐™‹, ๐™‹๐™ง๐™ค๐™ซ-๐™‚๐™ž๐™œ๐™–๐™‹๐™–๐™ฉ๐™, ๐™–๐™ฃ๐™™ ๐™๐™‰๐™„๐Ÿฎ-๐™. This approach creates a correlated latent space specifically optimized for estimating survival risk.

โ€ข ๐™ˆ๐™–๐™ข๐™—๐™–-๐˜ฝ๐™–๐™จ๐™š๐™™ ๐˜ผ๐™™๐™–๐™ฅ๐™ฉ๐™š๐™ง๐™จ: To handle the integration of these distinct and heavy embeddings, the framework utilizes Mamba-based adapter networks. This enables effective multi-modal learning while rigorously preserving computational efficiency.

โ€ข ๐™Ž๐™ช๐™ฅ๐™š๐™ง๐™ž๐™ค๐™ง ๐˜ผ๐™˜๐™˜๐™ช๐™ง๐™–๐™˜๐™ฎ: In experimental benchmark evaluations, this dynamic framework successfully outperforms both unimodal baselines and traditional static multi-modal models in survival prediction accuracy.

๐™๐™๐™š ๐™๐™–๐™ ๐™š๐™–๐™ฌ๐™–๐™ฎ: ๐™ˆ๐Ÿฐ๐™Ž๐™ช๐™ง๐™ซ๐™ž๐™ซ๐™š demonstrates that the future of predictive analytics and precision oncology relies not just on building individual foundation models, but on efficiently fusing them to create a holistic view of patient health.


The Era of Virtual Spatial Omics: Predicting Molecular Maps from H&E


Spatial omics provides unprecedented detail into the tumor microenvironment, but its high cost and technical complexity keep it confined to small research cohorts. What if we could generate these spatial maps directly from a standard H&E slide?

Recent breakthroughs in deep learning have made this a reality. By leveraging foundation models, researchers are now translating routine histology images into high-resolution spatial transcriptomics and proteomics, opening the door for massive-scale biomarker discovery. Three new papers showcase the power of this virtual approach.

Here is how they compare:

โ€ข ๐—ง๐—ฟ๐—ฎ๐—ป๐˜€๐—ฐ๐—ฟ๐—ถ๐—ฝ๐˜๐—ผ๐—บ๐—ถ๐—ฐ๐˜€ ๐˜ƒ๐˜€. ๐—ฃ๐—ฟ๐—ผ๐˜๐—ฒ๐—ผ๐—บ๐—ถ๐—ฐ๐˜€: While ๐—ฃ๐—ฎ๐˜๐—ต๐Ÿฎ๐—ฆ๐—ฝ๐—ฎ๐—ฐ๐—ฒ and ๐——๐—ฒ๐—ฒ๐—ฝ๐—ฆ๐—ฝ๐—ผ๐˜ focus on predicting spatial gene expression (RNA), the ๐—›๐—˜๐—ซ framework shifts the focus to ๐˜ฑ๐˜ณ๐˜ฐ๐˜ต๐˜ฆ๐˜ฐ๐˜ฎ๐˜ช๐˜ค๐˜ด. ๐—›๐—˜๐—ซ predicts 40 targeted protein biomarkers (like CODEX), arguing that proteins are often more closely related to cellular functions and clinical outcomes than RNA transcripts.

โ€ข ๐—”๐—ฟ๐—ฐ๐—ต๐—ถ๐˜๐—ฒ๐—ฐ๐˜๐˜‚๐—ฟ๐—ฎ๐—น ๐—œ๐—ป๐—ป๐—ผ๐˜ƒ๐—ฎ๐˜๐—ถ๐—ผ๐—ป๐˜€ ๐—ณ๐—ผ๐—ฟ ๐—ฅ๐—ฒ๐˜€๐—ผ๐—น๐˜‚๐˜๐—ถ๐—ผ๐—ป: Spatial spots often contain multiple cells, muddying the signal. ๐——๐—ฒ๐—ฒ๐—ฝ๐—ฆ๐—ฝ๐—ผ๐˜ tackles this by using a deep-set neural network, treating each transcriptomic spot as a bag of sub-spots to capture local morphology alongside global tissue context. Its successor, ๐——๐—ฒ๐—ฒ๐—ฝ๐—ฆ๐—ฝ๐—ผ๐˜๐Ÿฎ๐—–๐—ฒ๐—น๐—น, pushes this even further to virtual single-cell resolution. Alternatively, ๐—ฃ๐—ฎ๐˜๐—ต๐Ÿฎ๐—ฆ๐—ฝ๐—ฎ๐—ฐ๐—ฒ utilizes spatial smoothing and targeted cell-type deconvolutions to extract localized cell abundance directly from the inferred gene expression.

โ€ข ๐— ๐—ฎ๐˜€๐˜€๐—ถ๐˜ƒ๐—ฒ ๐—ฆ๐—ฐ๐—ฎ๐—น๐—ฒ ๐—ฎ๐—ป๐—ฑ ๐—–๐—น๐—ถ๐—ป๐—ถ๐—ฐ๐—ฎ๐—น ๐—œ๐—ป๐˜๐—ฒ๐—ด๐—ฟ๐—ฎ๐˜๐—ถ๐—ผ๐—ป: Because virtual omics are highly cost-effective, they enable unprecedented scale. ๐——๐—ฒ๐—ฒ๐—ฝ๐—ฆ๐—ฝ๐—ผ๐˜ generated a massive resource of 56 million virtual spots across 3,780 TCGA patients. ๐—ฃ๐—ฎ๐˜๐—ต๐Ÿฎ๐—ฆ๐—ฝ๐—ฎ๐—ฐ๐—ฒ applied its predictions to large breast cancer cohorts to identify SpatioTypes that predict chemotherapy and trastuzumab response. ๐—›๐—˜๐—ซ took it a step further with its ๐— ๐—œ๐—–๐—” integration framework, fusing H&E images with virtual proteomics to significantly outperform traditional clinical risk factors in predicting immunotherapy response.

๐˜›๐˜ฉ๐˜ฆ ๐˜›๐˜ข๐˜ฌ๐˜ฆ๐˜ข๐˜ธ๐˜ข๐˜บ: Virtual spatial omics will not completely replace physical sequencing, but it acts as a powerful, scalable bridge. By transforming archival H&E slides into multi-layered molecular maps, we are unlocking population-scale data essential for true precision medicine.


DeepSpot: Leveraging Spatial Context for Enhanced Spatial Transcriptomics Prediction from H&E Images


DeepSpot2Cell: Predicting Virtual Single-Cell Spatial Transcriptomics from H&E images using Spot-Level Supervision


AI-Driven Spatial Transcriptomics Unlocks Large-Scale Breast Cancer Biomarker Discovery from Histopathology


AI-enabled virtual spatial proteomics from histopathology for interpretable biomarker discovery in lung cancer

Rethinking Tissue Architecture: AI Beyond the Isolated Patch


Computational pathology models often divide tissue slides into isolated 2D tiles for processing. Yet, biology doesn't operate in isolated boxes; tissue is a continuous, interconnected environment and a complex 3D volume.

To accurately predict spatial transcriptomics and molecular signatures, models need to understand cellular neighborhoods, subtle spatial frequencies, and 3D depth. Four new papers introduce advanced architectures designed to capture this::

โ€ข ๐˜พ๐™–๐™ฅ๐™ฉ๐™ช๐™ง๐™ž๐™ฃ๐™œ ๐˜พ๐™š๐™ก๐™ก๐™ช๐™ก๐™–๐™ง ๐™‰๐™š๐™ž๐™œ๐™๐™—๐™ค๐™ง๐™๐™ค๐™ค๐™™๐™จ: Both ๐™‚๐™ค๐™ฃ๐™œ ๐™š๐™ฉ ๐™–๐™ก. and๐™ˆ๐™–๐™ง๐™ ๐™š๐™ฎ ๐™š๐™ฉ ๐™–๐™ก. focus on how localized regions interact, but at different scales. ๐™‚๐™ค๐™ฃ๐™œ ๐™š๐™ฉ ๐™–๐™ก. introduce ๐˜ผ๐™„๐˜ฟ๐™Š.๐™๐™ž๐™จ๐™จ๐™ช๐™š, an architecture that explicitly feeds multiple neighboring cells into an asymmetrical encoder-decoder to effectively learn cross-cell dependencies at the single-cell level. Conversely, ๐™ˆ๐™–๐™ง๐™ ๐™š๐™ฎ ๐™š๐™ฉ ๐™–๐™ก. focus on macro-level spatial interpretability across the whole slide with ๐™–๐™ˆ๐™„๐™‡. By replacing standard black-box pooling with an additive aggregation function, ๐™–๐™ˆ๐™„๐™‡ ensures every individual tissue patch contributes quantifiably to a slide-level gene signature, generating highly granular spatial heatmaps without needing patch-level annotations.

โ€ข ๐™‰๐™š๐™ฌ ๐˜ฝ๐™–๐™˜๐™ ๐™—๐™ค๐™ฃ๐™š๐™จ ๐˜ผ๐™ฃ๐™™ ๐˜ฟ๐™ž๐™ข๐™š๐™ฃ๐™จ๐™ž๐™ค๐™ฃ๐™จ: While the first two papers modify context and aggregation, ๐˜พ๐™๐™ค ๐™š๐™ฉ ๐™–๐™ก. and ๐™•๐™๐™ช ๐™š๐™ฉ ๐™–๐™ก. alter the core dimensions of how features are processed. ๐˜พ๐™๐™ค ๐™š๐™ฉ ๐™–๐™ก. present ๐™ˆ๐™‘๐™ƒ๐™ฎ๐™—๐™ง๐™ž๐™™, arguing that standard vision transformers struggle with biomarker prediction because they fail to capture subtle, low-frequency morphological patterns. By integrating state space models tuned with negative real eigenvalues, their hybrid backbone explicitly biases the network to preserve these critical low-frequency biological signals.

โ€ข ๐˜ฝ๐™ง๐™š๐™–๐™ ๐™ž๐™ฃ๐™œ ๐™๐™๐™š 2๐˜ฟ ๐˜ฝ๐™–๐™ง๐™ง๐™ž๐™š๐™ง: Meanwhile, ๐™•๐™๐™ช ๐™š๐™ฉ ๐™–๐™ก. shatter the 2D limitation entirely with ๐˜ผ๐™Ž๐™„๐™‚๐™‰. Recognizing that full 3D spatial transcriptomics is prohibitively expensive, their graph network extends spatial relationships into the z-axis. It imputes a 3D spatial transcriptomic volume by combining a stack of H&E sections with just a single 2D spatial transcriptomic slide.

๐™๐™๐™š ๐™๐™–๐™ ๐™š๐™–๐™ฌ๐™–๐™ฎ: The next generation of pathology AI will not just rely on training with more data. It requires architectures that inherently reflect the physical reality of human tissueโ€”whether through cell-neighborhood inputs, frequency-biased state space models, or true 3D spatial graphs.


AIDO.Tissue: Spatial Cell-Guided Pretraining for Scalable Spatial Transcriptomics Foundation Model


Spatial Mapping of Gene Signatures in Hematoxylin and Eosin-Stained Images: A Proof of Concept for Interpretable Predictions Using Additive Multiple Instance Learning

MVHybrid: Improving Spatial Transcriptomics Prediction with Hybrid State Space-Vision Transformer Backbone in Pathology Vision Foundation Models


ASIGN: An Anatomy-aware Spatial Imputation Graphic Network for 3D Spatial Transcriptomics


HONeYBEE: enabling scalable multimodal AI in oncology through foundation model-driven embeddings


Oncology data is inherently multimodalโ€”combining radiology scans, pathology slides, genomics, and clinical notes. Yet, most AI models are trapped in single-modality silos, missing the complete biological picture and underutilizing complementary information.

Integrating these heterogeneous data types into a unified patient representation is complex due to fragmented tools and rigid code dependencies. ๐˜ผ๐™–๐™ ๐™–๐™จ๐™ ๐™๐™ง๐™ž๐™ฅ๐™–๐™ฉ๐™๐™ž ๐™š๐™ฉ ๐™–๐™ก. published a comprehensive solution, ๐™ƒ๐™Š๐™‰๐™š๐™”๐˜ฝ๐™€๐™€ (Harmonized ONcologY Biomedical Embedding Encoder). This open-source framework generates and integrates patient-level embeddings using domain-specific foundation models.

Here are the key innovations from their evaluation of over 11,400 patients across 33 cancer types:

โ€ข ๐™๐™ฃ๐™ž๐™›๐™ž๐™š๐™™ ๐™ˆ๐™ช๐™ก๐™ฉ๐™ž๐™ข๐™ค๐™™๐™–๐™ก ๐™‹๐™ž๐™ฅ๐™š๐™ก๐™ž๐™ฃ๐™š: ๐™ƒ๐™Š๐™‰๐™š๐™”๐˜ฝ๐™€๐™€ processes five distinct data typesโ€”clinical text, pathology reports, radiologic images, whole slide images (WSIs), and molecular profilesโ€”through specialized preprocessing pipelines. Crucially, its modular design accommodates patients with missing data modalities without requiring complete-case cohorts.

โ€ข ๐™๐™๐™š ๐™‹๐™ค๐™ฌ๐™š๐™ง ๐™Š๐™› ๐˜พ๐™ก๐™ž๐™ฃ๐™ž๐™˜๐™–๐™ก ๐™๐™š๐™ญ๐™ฉ: In an interesting reality check, clinical embeddings derived from structured and unstructured data actually showed the strongest single-modality performance, achieving 98.5% classification accuracy and the highest overall survival prediction concordance indices. The authors note this reflects the expert-curated nature of clinical documentation in datasets like TCGA, which effectively summarizes information dispersed across other raw modalities.

โ€ข ๐™๐™ช๐™จ๐™ž๐™ค๐™ฃ ๐™๐™ค๐™ง ๐™Ž๐™ช๐™ง๐™ซ๐™ž๐™ซ๐™–๐™ก: While clinical data dominated, multimodal fusion strategies (such as concatenation and Kronecker product) provided critical complementary benefits. For specific cancers, fusing information from molecular, pathology, and imaging modalities significantly improved overall survival predictions beyond what clinical features could capture alone.

โ€ข ๐™‡๐™‡๐™ˆ๐™จ ๐™‹๐™ช๐™ฉ ๐™๐™ค ๐™๐™๐™š ๐™๐™š๐™จ๐™ฉ: The team compared four large language models to evaluate text embeddings. They found that general-purpose models (like Qwen3) actually outperformed specialized medical models (like GatorTron) on standard clinical text. However, task-specific fine-tuning proved essential across all models to achieve high performance on messy, heterogeneous data like pathology reports.

๐™๐™๐™š ๐™๐™–๐™ ๐™š๐™–๐™ฌ๐™–๐™ฎ: The future of precision oncology relies not just on building individual foundation models, but on creating scalable, open-source infrastructure that can standardize and unify these distinct representations into a cohesive clinical picture.

Enjoy this newsletter? Here are more things you might find helpful:


Pixel Clarity Call - A free 30-minute conversation to cut through the noise and see where your vision AI project really stands. Weโ€™ll pinpoint vulnerabilities, clarify your biggest challenges, and decide if an assessment or diagnostic could save you time, money, and credibility.

Book now

Did someone forward this email to you, and you want to sign up for more? Subscribe to future emails
This email was sent to _t.e.s.t_@example.com. Want to change to a different address? Update subscription
Want to get off this list? Unsubscribe
My postal address: Pixel Scientia Labs, LLC, PO Box 98412, Raleigh, NC 27624, United States


Email Marketing by ActiveCampaign