SHARE:

Biology-native data infrastructure for the AI era: how the next generation of biotechs is being built

Biological data infrastructure is one of the hottest topics in applied science today, and for good reason. With computing costs dropping and models maturing every quarter, the competitive edge in global drug development no longer hinges solely on having the best algorithm — it hinges on innovating across the entire data stack that feeds those algorithms.

Developing a drug has always been synonymous with patience, money, and a generous dose of luck. Over five years of work to reach a clinical candidate, nearly 90% failure rate in clinical phases, and an R&D cost that doubles every nine years. Not exactly the most encouraging equation, right? The limiting factor in drug development was never a lack of hypotheses, but rather the scarcity of resources to evaluate them effectively and efficiently.

But something is changing, and pretty fast. Artificial intelligence has entered this conversation full force, and the first real results are already showing up. Between 2012 and 2022, roughly 200 companies using AI for drug discovery collectively raised about 18 billion dollars in investments. Now we are seeing the fruits of those efforts reaching the clinic.

In June 2025, Insilico Medicine published positive Phase IIa results in Nature Medicine for rentosertib, a first-in-class TNIK inhibitor for idiopathic pulmonary fibrosis. It was the first drug to generate clinical proof of concept where both the target was discovered and the molecule was designed entirely by generative AI. Only 78 molecules were tested — compared to the thousands required by the traditional process — all within 18 months and at less than 10% of the average cost per approved drug.

This is not science fiction. It is a clear signal that the game is changing. And with heavyweights like Google DeepMind, Anthropic, GSK, and Eli Lilly pouring billions into this race, the question is no longer whether AI will transform drug development, but who will lead that transformation — and with what kind of infrastructure.

In early 2026, GSK struck a deal with NOETIK committing 50 million dollars upfront for access to oncology foundation models, while Eli Lilly signed an eight-figure annual access partnership with Chai Discovery for biologics design. Isomorphic Labs, the Google DeepMind spinout behind AlphaFold, established deep partnerships with Lilly, Novartis, and J&J with a potential value exceeding 3 billion dollars. And in April 2026, Anthropic acquired Coefficient Bio, an eight-month-old startup founded by computational biologists from Evozyne, Genentech, and Prescient Design, for 400 million dollars in stock — signaling that frontier AI labs are making direct bets on drug discovery.

The answer, as you will see below, has much less to do with the models themselves and much more to do with something that still gets too little attention: the biology-native data infrastructure that powers all of it. 🧬

The Cambrian explosion of biological AI models

Although computational chemistry tools emerged in the 1980s, the modern era of AI for biotech effectively began with the rise of deep learning in the 2010s, when it became clear that neural networks could learn meaningful representations of molecular structure from data. The watershed moment came when DeepMind’s AlphaFold2 and the Baker lab’s RoseTTAFold solved the problem of predicting the 3D structure of a protein from its amino acid sequence alone.

From there, the number of biological AI models grew exponentially. By 2024, more than 350 biological AI models had been published, including AlphaFold3, ESM3, Boltz-1, BindCraft, Evo, scGPT, and H-Optimus-0. These models demonstrate AI’s ability to perform tasks ranging from generative protein design to genomic modeling, perturbation modeling, and pathology image analysis. Between 2015 and 2025, the number of new AI models for biology released per year jumped from fewer than ten to more than 380.

More recently, models like JAM-2, BoltzGen, Latent-X2, Chai-2, and IsoDDE have continued bringing us closer to designing biologics with pharmaceutical potential directly from a computer. Isomorphic Labs’ IsoDDE, in fact, more than doubled AlphaFold 3’s accuracy on the hardest generalization benchmarks, making it one of the most closely watched models in the AI-driven drug design space.

Receive the best innovation content in your email.

All the news, tips, trends, and resources you're looking for, delivered to your inbox.

By subscribing to the newsletter, you agree to receive communications from Método Viral. We are committed to always protecting and respecting your privacy.

Why biological data is the real competitive asset

When most people think about AI applied to healthcare, their mind goes straight to the models — the sophisticated algorithms, deep neural networks, and transformer architectures that seem to solve everything. But there is an earlier layer, quiet and absolutely critical, that determines whether those models will actually work or simply generate expensive noise. That layer is biological data infrastructure.

Much of the data that makes today’s biological AI models possible was slowly assembled over decades of publicly funded science. The more than 200,000 protein structures in the Protein Data Bank (PDB) were experimentally determined using techniques like X-ray crystallography and nuclear magnetic resonance spectroscopy. The Human Genome Project mapped human genes and DNA through sequencing efforts at research institutions around the world. And ChEMBL’s bioactivity database on millions of small molecules was accumulated through years of manual extraction from patents and scientific literature. The impact of these databases is remarkable: structural data from the PDB contributed to the development of 100% of the protein-targeted small-molecule anticancer drugs approved by the FDA between 2019 and 2023.

The biological AI models developed over the past decades reflect the readily available data, with nearly 63% of models trained on protein sequences and structures from the Uniprot and PDB databases. The most common tasks these models are used for include contextual understanding of protein or nucleotide sequences, protein folding prediction, and protein design.

The gaps that current data does not cover

Despite its scale, the PDB has a strong bias toward proteins that are stable, soluble, and amenable to crystallization. Membrane proteins, intrinsically disordered proteins, and transient protein complexes — some of the most promising therapeutic targets in oncology and neurodegeneration — often defy these criteria and remain dramatically underrepresented. Furthermore, the structures captured by the PDB are static snapshots, freezing proteins in a single conformation rather than the dynamic ensemble of shapes they adopt in a living cell. And it is often these alternative conformations that are most therapeutically relevant, such as allosteric binding sites that only become accessible upon ligand binding.

While bringing a new drug to market begins with protein structure and design tasks, initial discovery represents only a fraction of the total time and cost of the process. More than two-thirds of development time and resources are allocated to stages after initial discovery, including ADME optimization work (pharmacokinetic properties of absorption, distribution, metabolism, and excretion) done in preclinical studies, as well as safety and efficacy studies conducted in clinical trials.

To advance a drug from a hit to a development candidate, much more is needed than confirming a molecule binds to its target. The process also requires understanding developability, immunogenicity, off-target effects, thermostability, solubility, and aggregation propensity — properties for which large, high-quality public datasets simply do not exist.

There is no PDB equivalent for understanding cellular phenotypes in response to perturbations, or even proteomic data across different disease states. Linking cellular data to clinical data presents an even larger gap, as patient omic profiles tied to treatment outcomes and clinical trial responses sit isolated in hospital systems and biopharma databases. This makes it nearly impossible to train models capable of predicting which patients will respond to a therapy before they are enrolled in a trial. The most commercially important predictions are precisely those where the data infrastructure is weakest.

Data generated before the AI era has significant limitations

Much of the biological data available today was generated before the explosion of biological AI models, meaning it often lacks the characteristics that make it useful for machine learning. Annotations are frequently incomplete or unstandardized, and important contexts — like cellular environment or laboratory equipment used — are rarely captured or encoded in the datasets. In many cases, biological datasets simply do not have the scale for models to draw statistically significant conclusions or make unbiased predictions.

To truly unlock AI for drug development, companies need to invest on two fronts: first, in generating multimodal and unprecedented biological measurements that expand our understanding of diseases, and second, in building datasets with the scale, consistency, and contextualization needed to train models that generalize across diverse biological settings.

Companies like Peptone are combining atomic-level biophysics with supercomputing to generate proprietary structural data on intrinsically disordered proteins. Inductive Bio is assembling one of the largest and most diverse ADMET datasets in the industry to train its Beacon models, which recently placed first among more than 370 submissions in the OpenADMET-ExpansionRx prediction challenge. Converge Bio is generating large-scale datasets to train and validate its own models for antibody design and sequence optimization. And NOETIK is building one of the most comprehensive oncology datasets available, pairing tumor multi-omics with longitudinal treatment outcomes.

Agentic AI in R&D workflows

While the cost of bringing a drug to market has kept rising, the cost of computing has decreased exponentially since the 1950s, consistent with Moore’s Law. Tasks along the development pipeline that are computationally expensive today will be dramatically cheaper within a few years. Companies that build their tech stack to be rapidly adaptable to AI’s evolving capabilities will have a growing structural advantage over those that treat AI as a fixed investment.

Although building proprietary molecular modeling tools in-house may have been a differentiator a decade ago, the abundance of ready-to-use in silico tools has changed that narrative. Structural predictors, ADMET models, and molecular dynamics simulators have matured enormously and are now broadly accessible through open-source repositories and closed architectures.

Companies should build their infrastructure from day one to be able to test, deploy, and leverage the newest tools, rather than being anchored to a single tech stack. Today, this modular infrastructure might look like a system that autonomously orchestrates the best tools for specific tasks, whether a literature review or the execution of a bioinformatics pipeline.

Cheaper computing has made long-context inference economically viable, enabling AI agents to synthesize more than a thousand papers and 40,000 lines of code in a single run. Combined with techniques that boost AI accuracy and efficiency, like chain-of-thought reasoning and multi-agent frameworks, it has become increasingly realistic that AI can significantly compress both the cost and timeline of the R&D cycle. 🤖

Agentic AI scientists will be able to mine preprint servers, patent registries, and public biological databases to identify non-obvious connections, generate novel hypotheses, perform in silico data analyses, design lab experiments, and write reports — all while maintaining the team’s research context and a historical record of experiments that empowers scientists to make smarter, faster decisions.

Soon, it will be standard to adopt an AI operating system spanning the entire drug development process, leveraging AI’s ability to retain extensive context to unify analyses and results in a single research environment, rather than leaving them scattered across separate point solutions.

A growing wave of companies is building in this direction. Anthropic now offers connectors to integrate Claude with platforms like Benchling, PubMed, ChEMBL, and ClinicalTrials.gov. K-Dense and Edison Scientific are developing autonomous AI scientist platforms that can plan, execute, and iterate on complex research workflows end to end. Phylo is taking a complementary approach with its Integrated Biology Environment, a unified space where scientists can collaborate with AI agents without switching between fragmented interfaces.

Companies like Potato and Convoke are building the operating systems for biopharma, with Potato serving as infrastructure for autonomous experiment design and execution, and Convoke acting as the system of record and action to accelerate regulatory and document workflows for bringing drugs to market.

Lab automation and the closed-loop discovery cycle

Even companies using the most advanced AI models hit the limitations of experimental data generation. Despite dramatic advances in structural prediction and molecular modeling, many in silico results — like binding affinity predictions — still need to be validated in the lab before any development decision can be made with confidence. On top of that, in vivo efficacy is essentially unpredictable from first principles, with late-stage failures disproportionately driven by pharmacokinetic and toxicity properties that in silico models did not flag.

Tools we use daily

Given that experimental results are the ultimate source of biological truth, it is crucial that these models continuously incorporate lab feedback to stay grounded in accuracy.

Unfortunately, the experimental cycles between a model’s output and the data needed to update its predictions often take weeks to months. Lab experiments are slow, failure-prone, and dependent on skilled human labor, making them one of the biggest bottlenecks for shortening development timelines. The iterative design-test-build-analyze cycle that characterizes lead compound optimization can take up to three years on its own and accounts for nearly a quarter of the total development timeline. These timelines are further extended by the fact that experimental validation is often outsourced to Contract Research Organizations (CROs), where coordination overhead, queuing, and data quality inconsistencies can add weeks or months to each iteration cycle.

From point automation to the autonomous lab

Although Hamilton liquid-handling robots and Chemspeed automated synthesis platforms have existed in labs for decades, they are optimized for high-throughput execution of specific point tasks, not for automating and integrating entire experimental workflows. Most current lab automation still requires significant human intervention to transfer materials between instruments, troubleshoot issues, and interpret results before the next experimental step can begin.

Automating robotic lab equipment has historically required dedicated automation engineers to configure instruments and continuously write new scripts for different workflows. Natural language interfaces for robot control could effectively democratize automation capabilities, enabling scientists with no background in robotics or software engineering to execute, monitor, and iterate on experiments remotely and autonomously. Advances in robotics and physical AI could further orchestrate the material and data transfers that are still done by humans today. Computer vision-native systems, for example, can already autonomously read and interpret cell microscopy images and feed structured data directly into model pipelines without a scientist needing to manually extract and input the results. 🔬

The progress toward autonomous labs gives companies significant leverage in both speed and operating cost. A model trained on five design-test-analyze cycles in the time a competitor completes one will compound its biological understanding much faster, and that compounding translates directly into better models, better molecules, and a structural advantage that is very hard to recover for anyone relying on traditional CRO timelines.

Companies across the lab automation landscape are pursuing this from different angles. Medra is building an instrument-agnostic robotics platform where general-purpose robots interact with existing lab equipment through physical controls and software interfaces. Automata takes a lab orchestration approach with its LINQ platform, providing modular hardware and software that connect different instruments into end-to-end automated and coordinated workflows. Dash Bio is using robotics to become a faster, more automated CRO. And Lila Sciences represents one of the most vertically integrated approaches, building a fully automated lab for end-to-end drug discovery and development.

The future of life sciences will run on AI

The landscape taking shape over the next few years points to a deep convergence of biological science, data engineering, and artificial intelligence. The companies building large biology-native datasets, AI-centric development stacks, and lab automation platforms that enable rapid closed-loop experimentation will define the next generation of life sciences companies.

This market organizes itself into three interdependent layers:

  • Data layer: companies generating data at the scale, modality, and fidelity that AI requires to produce meaningful discoveries across the entire drug development process.
  • Software and agentic workflow layer: AI platforms that orchestrate and unify end-to-end R&D workflows, from hypothesis generation to regulatory submission.
  • Physical infrastructure layer: lab automation platforms that compress timelines at every stage, enabling rapid feedback cycles between models and experiments.

Together, these three layers represent much of the value chain taking shape in AI-driven drug development. The speed of new drug discovery will depend less on how many scientists an organization has and more on how well-structured and rich its biological data infrastructure is. The first clinical results already show that this is not a distant promise: it is a reality under construction, and the foundational blocks are being laid right now, at this very moment, by teams that understood that quality data, intelligent workflows, and integrated automation are the true competitive differentiator in this race. 🚀

Picture of Rafael

Rafael

Operations

I transform internal processes into delivery machines — ensuring that every Viral Method client receives premium service and real results.

Fill out the form and our team will contact you within 24 hours.

Related publications

AI SDR Agent on WhatsApp: How SMBs Can Cut Costs and Scale Sales

Respond 21x faster your leads and scale your sales operation with a fraction of the cost of expanding your sales

Robot Detects Unusual Browser Activity Using JavaScript and Cookies

Learn why sites require JavaScript and cookies for unusual activity and how to fix blocks with quick, simple steps

Productivity with Agentic Artificial Intelligence in execution and workflows.

Agentic AI: how to operationalize AI agents to improve workflows, metrics, and governance, turning pilots into real productivity gains.

Receive the best innovation content in your email.

All the news, tips, trends, and resources you're looking for, delivered to your inbox.

By subscribing to the newsletter, you agree to receive communications from Método Viral. We are committed to always protecting and respecting your privacy.

Rafael

Online

Atendimento

Website Pricing Calculator

Find out how much the ideal website for your business costs

Website Pages

How many pages do you need?

Drag to select from 1 to 20 pages

In just 2 minutes, automatically find out how much a custom website for your business costs

More than 0+ companies have already calculated their quote

Fale com um consultor

Preencha o formulário e nossa equipe entrará em contato.