AI Fixes Cancer Trial Failures Through Precise Patient Selection

Original Title: 🔬 Training Transformers to solve 95% failure rate of Cancer Trials — Ron Alfa & Daniel Bear, Noetik

The 95% Failure Rate in Cancer Trials Isn't About Bad Drugs, It's About Misguided Patient Selection--and AI is Poised to Fix It.

This conversation with Ron Alfa and Daniel Bear of Noetik reveals a profound, yet often overlooked, bottleneck in drug development: the inability to precisely match treatments to the right patients. The staggering 95% failure rate in cancer clinical trials, they argue, is not a reflection of flawed pharmacology but a consequence of treating cancer as a monolithic entity. By leveraging advanced AI, particularly transformer-based models trained on vast, multimodal human tumor data, Noetik is building "virtual cells" and "world models" that can dissect cancer's true heterogeneity. This insight offers a critical advantage to pharmaceutical companies and researchers: a path to dramatically improve success rates, repurpose existing drugs, and ultimately save lives, by shifting focus from discovering new drugs to understanding precisely who existing drugs will help. Anyone invested in accelerating therapeutic breakthroughs, from drug developers to AI engineers seeking impactful applications, will find a compelling case for this data-driven, patient-centric approach.

The Hidden Cost of "One Size Fits All" in Oncology

The prevailing narrative around cancer treatment often centers on the discovery of novel drugs. Yet, Ron Alfa and Daniel Bear of Noetik present a contrarian, and arguably more impactful, thesis: the primary failure point in clinical trials isn't the drug itself, but the flawed methodology of patient selection. For decades, pharmaceutical companies have relied on preclinical models--cell lines and animal studies--that, while useful, fail to capture the intricate biological diversity of human tumors. This disconnect leads to trials where a drug might show promise in a dish but falters in patients because the underlying biology driving response or resistance is poorly understood.

"Most of those drugs fail, we'd argue, is because we're bad at selecting which patients those drugs are going to work in."

-- Ron Alfa

This isn't a minor oversight; it's a systemic issue that results in billions of dollars and years of research being poured into treatments that ultimately fail. The consequence? Promising molecules are shelved, and patients are left with fewer effective options. Noetik's approach directly confronts this by building AI models trained on rich, multimodal data derived from actual human tumors. This data includes HE staining (standard pathology images), immunofluorescence (protein markers), and spatial transcriptomics (gene expression mapped spatially). By processing this data at scale, their models aim to identify true biological subtypes of cancer--subtypes that may not be apparent through traditional pathological classifications. This deeper understanding allows for a more precise matching of drugs to patient populations, transforming the trial process from a broad, often unsuccessful, net-casting exercise into a highly targeted, data-informed strategy. The implication is that many "failed" drugs might actually be effective, simply administered to the wrong patients.

When the Lab Doesn't Speak to the Clinic: The Translation Gap

The chasm between laboratory findings and clinical efficacy is a well-documented challenge in drug development. Daniel Bear elaborates on how traditional preclinical models, such as immortalized cell lines, often possess abnormal genomes and gene expression patterns that bear little resemblance to actual human tumors. This fundamental disconnect means that even drugs that perform well in these artificial systems frequently fail when tested in human patients. The downstream effect of this translational gap is immense: drugs that could potentially help specific patient subsets are discarded because the preclinical data provided a misleading picture.

"The problem is these cell lines as an abstraction do not relate in any way to to human patients."

-- Daniel Bear

Noetik's strategy bypasses this limitation by prioritizing data generated from real human tumors. Their commitment to collecting thousands of human tumor samples and meticulously processing them into multimodal datasets addresses the core issue of relevance. This data forms the foundation for their "virtual cell" and "world models," which simulate patient biology and treatment responses. The advantage here is twofold: first, it enables the identification of patient subgroups that are most likely to respond to a given treatment, thereby increasing the probability of clinical trial success. Second, it allows for the intelligent repurposing of existing drugs that may have failed due to poor patient stratification in previous trials. For pharmaceutical companies, this represents a significant opportunity to de-risk development pipelines and unlock value from previously shelved assets. The long-term payoff is a more efficient and effective drug development ecosystem, driven by a deeper, data-informed understanding of cancer biology.

The Data Moat: Why Scale and Specificity Matter in Bio-AI

The success of AI in biology, particularly in areas like cancer research, is heavily contingent on the quality and scale of the data used for training. Ron Alfa highlights Noetik's "data moat"--their extensive, proprietary dataset comprising millions of spatially resolved cells from human tumors, paired with multimodal information. This isn't just about quantity; it's about the deliberate design and collection of data that directly addresses the complexity of cancer. While public repositories exist, they often lack the depth, breadth, and multimodal integration that Noetik provides.

"We've generated now, you know, more than a hundred million cells spatially resolved and spatial transcriptomics that's all paired with H&E and protein as well, at least an order of magnitude larger than any of the other datasets that we've seen out there."

-- Ron Alfa

This scale and specificity are crucial. Dropping to even 40% or 10% of their data, they've observed, significantly degrades model performance, especially in generalizing to different cancer types. The consequence of insufficient or poorly curated data is the inability for models to learn the subtle, complex, non-linear patterns that differentiate patient responses. This is where Noetik's custom-built transformer architectures, like TARIO-2, come into play. These models are specifically designed to handle multimodal data and leverage autoregressive training objectives, similar to large language models, to predict sequences and understand spatial context. The advantage of this approach is the potential for models to become increasingly predictive and generalizable as more data is incorporated. For pharmaceutical partners, this means access to sophisticated models that have been rigorously trained on the most relevant data, offering a significant head start in identifying patient cohorts and predicting treatment efficacy, thereby accelerating the journey from discovery to patient benefit.

From H&E to Insight: Unlocking Value from Existing Diagnostics

A particularly compelling aspect of Noetik's work is their ability to derive rich biological insights from Hematoxylin and Eosin (H&E) stains--the standard pathology slides that nearly every cancer patient already receives. Daniel Bear explains that their trained models, at inference time, only require an H&E image to make predictions about patient response and underlying biology. This is a game-changer because it circumvents the need for expensive, time-consuming, or less common assays like spatial transcriptomics for every patient.

"The reason that that is so powerful and flexible is again because H&E is kind of like the lingua franca of pathology and especially oncology."

-- Daniel Bear

The consequence of this strategy is democratized access to advanced predictive capabilities. Pharmaceutical companies can leverage their existing trial data, which often includes H&E slides, to retroactively analyze patient responses and refine future trial designs. This transforms a routine diagnostic into a powerful predictive tool. Noetik's models can identify clusters of responders within a patient population based solely on H&E images, and even predict gene expression patterns within those clusters. This level of interpretability is critical for building trust and understanding the biological basis of drug efficacy. The advantage for the industry is clear: a more efficient, cost-effective, and accessible way to stratify patients, de-risk clinical trials, and ultimately bring more effective treatments to market faster. It represents a significant leap from simply classifying tumors to predicting treatment outcomes based on readily available data.

Key Action Items

  • Immediate Action (Next Quarter):

    • Audit existing clinical trial data: Review all historical oncology trials for readily available H&E pathology slides of responders and non-responders.
    • Engage with Noetik for pilot analysis: Initiate a pilot project with Noetik to analyze a subset of H&E slides from a past trial to assess their models' predictive power for patient response.
    • Identify key biological questions: For internal drug development programs, clearly define the critical biological questions that, if answered, would significantly de-risk a clinical trial.
  • Short-Term Investment (Next 6-12 Months):

    • Integrate H&E data into preclinical assessments: Begin incorporating analysis of H&E slides from preclinical models to assess potential translational relevance to human biology.
    • Explore custom model fine-tuning: Investigate the feasibility and value of fine-tuning Noetik's foundation models on proprietary internal datasets to enhance predictive accuracy for specific therapeutic areas.
    • Develop internal biomarker discovery frameworks: Begin building internal capabilities or partnerships to leverage AI-driven patient stratification, moving beyond single-gene biomarkers.
  • Longer-Term Investment (12-18 Months+):

    • Establish a "virtual trial" simulation capability: Invest in capabilities, potentially through partnerships, to simulate clinical trial outcomes using AI models based on patient biology data, informing trial design before patient enrollment.
    • Invest in multimodal data generation infrastructure: For companies with significant R&D, consider strategic investments in generating high-quality, multimodal biological data from human samples, recognizing it as a critical long-term asset.
    • Foster cross-disciplinary AI-Biology teams: Build and empower teams that blend deep AI expertise with biological understanding, enabling a more holistic approach to drug development challenges. This requires a commitment to hiring individuals comfortable navigating complex, interdisciplinary landscapes.

---
Handpicked links, AI-assisted summaries. Human judgment, machine efficiency.
This content is a personally curated review and synopsis derived from the original podcast episode.