Prioritizing Sub-Angstrom Physical Accuracy Over Pattern-Matching Benchmarks

Original Title: 🔬 The Coolest Diffusion Research Isn't in LLMs — Evan Feinberg & Sergey Edunov, Genesis Molecular AI

The 1-Angstrom Threshold: Why Drug Discovery is a Physics Problem, Not Just a Pattern-Matching One

In this conversation, Evan Feinberg and Sergey Edunov of Genesis Molecular AI explain why current AI-driven drug discovery is failing, arguing that the industry is trapped by low-quality benchmarks. They show that the push for fast, high-throughput results from automated labs often creates molecules that fail in real-world testing because they lack sub-angstrom resolution. By moving from pattern matching to physical validity, Genesis has found that the true competitive advantage in drug discovery lies in the difficult, high-accuracy work that most competitors avoid. This analysis provides a guide for researchers looking to move past the current limits of generative AI by building physical constraints into agentic workflows.

The Hidden Cost of Good Enough Benchmarks

The AI community has largely based its progress on benchmarks like 2-Angstrom RMSD, a metric Feinberg calls slop. In protein-ligand interactions, this threshold is misleading: it allows models to produce structures that look correct while failing to model the precise atomic interactions, such as hydrogen bonds, that actually dictate binding.

If your model is sitting at 1.8, 1.9 Angstrom RMSD, that is slop, most likely.

-- Evan Feinberg

When models are optimized for these loose benchmarks, they essentially hallucinate details to satisfy the metric. Because these models serve as inputs for downstream physics-based simulations, these errors multiply. A model that is mostly correct at the start becomes wrong by the time a medicinal chemist tries to synthesize the molecule. The systemic failure here is the gap between the AI researcher and the actual user. By optimizing for a leaderboard rather than physical reality, the community has built tools that look sophisticated in papers but offer no utility in the lab.

How Induced Fit Creates a Competitive Moat

A major bottleneck in drug discovery is induced fit, the process where a protein changes shape to accommodate a ligand. Traditional methods treat the protein as a static structure because modeling the dynamics is computationally expensive. Most teams accept this limitation, choosing static models that fail when the target protein is dynamic.

Where PEARL was exceptionally good is figuring out how to move this loop. We are basically correct for every single pose.

-- Sergey Edunov

Feinberg and Edunov argue that by focusing on sub-angstrom resolution through diffusion-based models, they can capture these induced-fit dynamics without the high cost of traditional molecular dynamics simulations. This creates a lasting advantage: while competitors stay in the static paradigm, Genesis is working on a much harder, more accurate task. This shows how immediate discomfort, such as investing in higher-resolution models that are harder to train, creates a barrier to entry that others cannot easily cross.

The Agentic Feedback Loop: Closing the Gap

The transition from chatbots to agents in the broader AI space has been rapid, but Edunov and Feinberg warn that agents are only as useful as the models they use. In drug discovery, an agent that makes small errors will amplify them across hundreds of iterations, resulting in useless outputs.

The system responds to this by requiring a tighter link between computation and wet-lab validation. By partnering with organizations like Incyte, Genesis has created a feedback loop where models are continuously fine-tuned on real-world experimental data. This is the agentic discovery frontier: moving from a model that predicts to a system that reasons, iterates, and validates across a 24/7 cycle. The competitive advantage here is not just the model; it is the speed of the design-make-test-analyze cycle.

Key Action Items

  • Audit Your Benchmarks: Stop optimizing for metrics like 2A RMSD if your work requires sub-angstrom precision. (Immediate)
  • Prioritize Physical Priors: When training models, include physical constraints instead of relying purely on data-driven pattern matching. This prevents overfitting and improves generalizability. (Over the next quarter)
  • Shift to Agentic Workflows: Identify repetitive tasks in your R&D pipeline and build agents to handle them, but ensure your underlying models are battle-tested first. (12-18 months)
  • Focus on Vertical Integration: If you are building AI for physical sciences, consider the benefit of in-house data generation. Testing your own models on real-world experiments is an unbeatable source of truth. (Long-term investment)
  • Seek Out Hard Targets: Avoid the easy targets of well-characterized proteins. The real value and the real opportunity for AI lie in targets that have been labeled undruggable due to their complexity. (Ongoing)

---
Handpicked links, AI-assisted summaries. Human judgment, machine efficiency.
This content is a personally curated review and synopsis derived from the original podcast episode.