Share

Multimodal Alignment Improves Generalizability of Genomic Biomarker Prediction in Computational Pathology


Predicting genomic biomarkers directly from digitized histopathology whole-slide images offers a cost-effective and scalable alternative to traditional molecular assays. However, the field faces a structural bottleneck: every time a new genomic biomarker is discovered or quantified, researchers must prospectively collect large, labeled datasets to train new predictive models.

A new paper by ๐™€๐™ ๐™–๐™ฉ๐™š๐™ง๐™ž๐™ฃ๐™– ๐™๐™š๐™™๐™š๐™ ๐™ค๐™ฅ ๐™š๐™ฉ ๐™–๐™ก. addresses this challenge with ๐™ˆ๐˜ผ๐™๐˜ฝ๐™‡๐™€, a multimodal contrastive pretraining strategy.

Here are the key innovations from the research:

โ€ข ๐˜ฝ๐™ž๐™ค๐™ก๐™ค๐™œ๐™ž๐™˜๐™–๐™ก๐™ก๐™ฎ ๐™„๐™ฃ๐™›๐™ค๐™ง๐™ข๐™š๐™™ ๐˜ผ๐™ก๐™ž๐™œ๐™ฃ๐™ข๐™š๐™ฃ๐™ฉ: Instead of relying solely on visual data, ๐™ˆ๐˜ผ๐™๐˜ฝ๐™‡๐™€ integrates structured biomarker knowledge directly into the representation learning of histopathology images.

โ€ข ๐™‡๐™‡๐™ˆ๐™จ ๐˜ผ๐™ฃ๐™™ ๐™‹๐™‡๐™ˆ๐™จ ๐˜ผ๐™จ ๐˜ผ๐™ฃ๐™˜๐™๐™ค๐™ง๐™จ: The framework functions by aligning representations derived from histopathology images with representations of genomic biomarkers that are generated by a large language model (LLM) and a protein language model (PLM).

โ€ข ๐™Š๐™ช๐™ฉ-๐™Š๐™›-๐˜ฟ๐™ž๐™จ๐™ฉ๐™ง๐™ž๐™—๐™ช๐™ฉ๐™ž๐™ค๐™ฃ ๐™‚๐™š๐™ฃ๐™š๐™ง๐™–๐™ก๐™ž๐™ฏ๐™–๐™ฉ๐™ž๐™ค๐™ฃ: By grounding the visual data in biological semantics, this alignment enables data-efficient generalization to novel, out-of-distribution biomarkers without requiring massive new training datasets.

โ€ข ๐™๐™š๐™–๐™ก-๐™’๐™ค๐™ง๐™ก๐™™ ๐™‘๐™–๐™ก๐™ž๐™™๐™–๐™ฉ๐™ž๐™ค๐™ฃ: The team validated their approach using the MSK-IMPACT cohort, grounding their experiments in real-world data across over 40,000 patients and multiple biomarker panel versions.

๐™๐™๐™š ๐™๐™–๐™ ๐™š๐™–๐™ฌ๐™–๐™ฎ: As precision oncology expands, relying on building new models from scratch for every new biomarker is inefficient. By leveraging language and protein models for multimodal alignment, we can create adaptable pathology AI capable of generalizing to future discoveries.


Unbottlenecking the MIL Pipeline: Four Approaches to Whole Slide AI


Extracting rich features from a pathology slide is only the first step. The real challenge is aggregating tens of thousands of isolated tile embeddings into a single, accurate clinical result.

Multiple Instance Learning (MIL) is the standard framework for this task, but traditional MIL pipelines suffer from rigid linear transformations, loss of spatial context, and overly simplistic attention-pooling mechanisms. Four new papers demonstrate that instead of simply building larger patch-level foundation models, we need to completely redesign how those patches are contextualized, transformed, and aggregated.

Here is how their approaches to solving the MIL bottleneck compare:

โ€ข ๐™๐™๐™š ๐™ˆ๐™ž๐™ญ๐™ฉ๐™ช๐™ง๐™š-๐™Š๐™›-๐™€๐™ญ๐™ฅ๐™š๐™ง๐™ฉ๐™จ ๐˜ผ๐™ฅ๐™ฅ๐™ง๐™ค๐™–๐™˜๐™: Two papers adapt the MoE paradigm to solve different bottlenecks in the MIL pipeline. ๐™ˆ๐˜ผ๐™ˆ๐™ˆ๐™Š๐™๐™ƒ focuses on the linear layer that transforms general features into task-specific ones ๐™—๐™š๐™›๐™ค๐™ง๐™š aggregation. Using a parameter-efficient mixture of mini-experts, it applies tailored low-rank transformations to each patch's phenotype. Conversely, ๐™ˆ๐™ค๐˜ผ (Mixture of Aggregators) applies MoE directly to the aggregation stage. Instead of a single pooling mechanism, a router dynamically selects the top-2 most relevant aggregators for each specific slide to better capture morphological heterogeneity.

โ€ข ๐˜พ๐™ค๐™ฃ๐™ฉ๐™š๐™ญ๐™ฉ๐™ช๐™–๐™ก๐™ž๐™ฏ๐™–๐™ฉ๐™ž๐™ค๐™ฃ ๐˜ฝ๐™š๐™›๐™ค๐™ง๐™š ๐˜ผ๐™œ๐™œ๐™ง๐™š๐™œ๐™–๐™ฉ๐™ž๐™ค๐™ฃ: Standard MIL treats tiles as isolated bags of features. ๐™๐™„๐˜พ๐™Š๐™‰ is a transformer-based tile contextualizer that uses a masked modeling objective to infuse local tiles with global slide context before they ever reach the aggregator. An aggregator trained on TICON embeddings using just 11K WSIs successfully outperformed slide-level models pretrained on up to 350K WSIs.

โ€ข ๐™‚๐™š๐™ฃ๐™š๐™ง๐™–๐™ก๐™ž๐™ฏ๐™–๐™—๐™ก๐™š ๐™Ž๐™–๐™ข๐™ฅ๐™ก๐™ž๐™ฃ๐™œ ๐˜ผ๐™ฃ๐™™ ๐™€๐™ฃ๐™จ๐™š๐™ข๐™—๐™ก๐™ž๐™ฃ๐™œ: While the other papers focus on specialized architectural modules, nnMIL focuses on robust training dynamics. It introduces random sampling at both the patch and feature levels to enable large-batch optimization. This is paired with a lightweight aggregator that performs sliding-window inference for ensemble slide-level predictions.

๐™๐™๐™š ๐™๐™–๐™ ๐™š๐™–๐™ฌ๐™–๐™ฎ: The next leap in computational pathology relies on the connective tissue of the MIL pipeline. Whether through contextual transformers, dynamic expert routing, or smarter sampling, overcoming the aggregation bottleneck is the key to unlocking the full potential of pathology foundation models.


MoA: Mixture of Aggregators Improves Slide-Level Diagnosis in Computational Pathology

TICON: A Slide-Level Tile Contextualizer for Histopathology Representation Learning

Mixture of Mini Experts: Overcoming the Linear Layer Bottleneck in Multiple Instance Learning

nnMIL: A generalizable multiple instance learning framework for computational pathology

Efficient Universal Perception Encoder


Running advanced AI models on smart edge devices presents a core dilemma. Users expect versatile, multi-task experiences, but these devices are fundamentally constrained by limited compute power.

To bridge this gap, we need vision encoders that are incredibly small yet capable of outputting powerful, versatile representations. Historically, researchers have tried agglomerative methodsโ€”taking multiple large, domain-expert foundation models and distilling them directly down into a single, small encoder. However, this direct scale-down struggles to efficiently capture the full breadth of knowledge required for diverse downstream tasks.

A new paper by ๐˜พ๐™๐™š๐™ฃ๐™˜๐™๐™š๐™ฃ ๐™•๐™๐™ช ๐™š๐™ฉ ๐™–๐™ก. introduces a more effective solution: the Efficient Universal Perception Encoder (EUPE).

Here are the key innovations from their research:

โ€ข ๐™๐™๐™š ๐™Ž๐™˜๐™–๐™ก๐™š-๐™๐™ฅ ๐˜ผ๐™ฃ๐™™ ๐™Ž๐™˜๐™–๐™ก๐™š-๐˜ฟ๐™ค๐™ฌ๐™ฃ ๐™Ž๐™ฉ๐™ง๐™–๐™ฉ๐™š๐™œ๐™ฎ: Instead of distilling directly from multiple teachers into a small model, the authors add a crucial intermediate step. They demonstrate the importance of first scaling up by distilling multiple domain-expert models into one massive proxy teacher. Only then do they scale down, distilling from this single proxy into the efficient edge encoder.

โ€ข ๐™๐™ฃ๐™˜๐™ค๐™ข๐™ฅ๐™ง๐™ค๐™ข๐™ž๐™จ๐™ž๐™ฃ๐™œ ๐™‹๐™š๐™ง๐™›๐™ค๐™ง๐™ข๐™–๐™ฃ๐™˜๐™š: This unique distillation process yields significantly more versatile representations. EUPE achieves on-par or better performance across diverse task domains compared to individual domain experts of the exact same size, and it successfully outperforms previous agglomerative encoders.

๐™๐™๐™š ๐™‹๐™ค๐™ฉ๐™š๐™ฃ๐™ฉ๐™ž๐™–๐™ก ๐™๐™ค๐™ง ๐™‹๐™–๐™ฉ๐™๐™ค๐™ก๐™ค๐™œ๐™ฎ: While the authors focus on general smart edge devices, applying this universal, compute-efficient framework to digital pathology presents a massive opportunity. Processing gigapixel tissue slides currently relies on massive foundation models that require energy-intensive, cloud-based GPUs. The EUPE strategy could enable the field to combine and distill multiple top-tier modelsโ€”such as Virchow v2, UNI2, and H-Optimusโ€”into a massive proxy, and then scale it down into a much smaller, yet incredibly powerful model.

๐™๐™๐™š ๐™๐™–๐™ ๐™š๐™–๐™ฌ๐™–๐™ฎ: We do not necessarily have to choose between versatility and efficiency on the edge. By rethinking the distillation pipeline and using an intermediate proxy teacher, we can deploy highly capable perception models on compute-constrained devicesโ€”potentially unlocking powerful AI everywhere from our pockets to remote pathology clinics.


Cracks in the Foundation: How Data-Hungry and Sensitive to Domain Shift are Vision Foundation Models for Computational Pathology?


Vision Foundation Models (VFMs) are widely touted as the ultimate solution to data scarcity and poor generalization in computational pathology. But when stress-tested against real-world clinical variables, are they actually as robust and data-efficient as promised?

A preprint by ๐˜ผ๐™ฃ๐™Ÿ๐™– ๐™’๐™ž๐™ฉ๐™ฉ๐™š ๐™š๐™ฉ ๐™–๐™ก. rigorously evaluated six VFMs on a protocol-variant prostate cancer dataset comprising over 37,000 spot images. The authors introduced six controlled domain shiftsโ€”including variations in staining duration, section thickness, scanner type, and sampling locationโ€”to test the models on clinically relevant tasks like ISUP grading and 5-year relapse prediction.

The results reveal some critical limitations in the current generation of models:

โ€ข ๐™๐™๐™š ๐™„๐™ข๐™ฅ๐™ค๐™ง๐™ฉ๐™–๐™ฃ๐™˜๐™š ๐™Š๐™› ๐˜ฟ๐™š๐™˜๐™ค๐™™๐™š๐™ง ๐˜ฟ๐™š๐™จ๐™ž๐™œ๐™ฃ: The authors found that "downstream performance depends strongly on the chosen decoder architecture". While "simple probing approaches such as KNN, which are commonly used in the evaluation of foundation models, were insufficient for clinically relevant tasks", decoder-based approaches proved to be essential.

โ€ข ๐™๐™๐™š ๐™ˆ๐™ฎ๐™ฉ๐™ ๐™Š๐™› ๐˜ฟ๐™–๐™ฉ๐™– ๐™€๐™›๐™›๐™ž๐™˜๐™ž๐™š๐™ฃ๐™˜๐™ฎ: One of the primary appeals of foundation models is their presumed ability to perform well with very few labeled examples. However, this study demonstrated that "the presumed data efficiency of VFMs did not hold: stable decoder performance typically required more than 1000 training samples."

โ€ข ๐™‘๐™ช๐™ก๐™ฃ๐™š๐™ง๐™–๐™—๐™ž๐™ก๐™ž๐™ฉ๐™ฎ ๐™๐™ค ๐˜ฟ๐™ค๐™ข๐™–๐™ž๐™ฃ ๐™Ž๐™๐™ž๐™›๐™ฉ๐™จ: Even with massive pre-training, "none of the models demonstrated sufficient robust generalization under protocol-level domain shifts." The models exhibited performance reductions of 4 to 13% in key tasks when faced with standard laboratory variations.

๐™๐™๐™š ๐™๐™–๐™ ๐™š๐™–๐™ฌ๐™–๐™ฎ: While larger foundation models exhibit better peak accuracy and somewhat greater robustness, they do not fully address the critical issues of data efficiency and domain shift. To build truly reliable clinical tools, the field still needs substantial labeled datasets, robust decoder architectures, and improved domain adaptation methods.

Did someone forward this email to you, and you want to sign up for more? Subscribe to future emails
This email was sent to _t.e.s.t_@example.com. Want to change to a different address? Update subscription
Want to get off this list? Unsubscribe
My postal address: Pixel Scientia Labs, LLC, PO Box 98412, Raleigh, NC 27624, United States


Email Marketing by ActiveCampaign