Anthropic’s AI Protein Design Run: 354 Binders, 15 Targets

Anthropic’s AI Protein Design Run: 354 Binders, 15 Targets

Anthropic published results on August 18, 2026 from a protein binder design campaign it says was run almost entirely by its own models. Adaptyv Bio, the automated contract research organisation near Lausanne that carried out the binding assays, published a separate case study the following day.

How the campaign was run

According to Anthropic’s account, the models were given a design prompt of roughly 30,000 tokens, internet access, a library of protein-design papers, GPU compute, and no cap on token use or the number of sub-agents they could spawn. Human involvement was limited to approving network requests, maintaining infrastructure and placing DNA orders. Anthropic attributes the remaining steps to the models: epitope selection, backbone and sequence generation, optimisation cycles, and screening for expression and solubility.

Two models were tested: Opus 4.8, and Mythos Preview, which Anthropic describes as an unreleased frontier model. Thirty designs were requested per target under two configurations, all targets addressed together in a single 48-hour run with up to 12,500 H100 hours, or one target per parallel 24-hour session with up to 2,500 H100 hours each.

Anthropic states that the models did not generate structures using any protein model of Anthropic’s own. They operated publicly available structure design, sequence design and co-folding tools. The company’s claim is about orchestration: target reading, epitope choice, pipeline assembly, hyperparameter selection, filtering and iteration.

Reported results

Of 1,320 designs, 354 were classified as binders. Anthropic reports overall hit rates of 26.7% for Mythos Preview and 22.6% for Opus 4.8 in multi-target mode, rising to 35.1% for Mythos Preview running one target at a time. The company cites 10–15% as the current field norm, a figure it describes as an aggregate drawn from public data. Per-target results ranged from roughly 90% to zero. Binders with K_D below 10 nM, Anthropic’s threshold for high affinity, were obtained against at least six targets.

Sixteen targets were selected and fifteen are reported. Mature GDF-8 was excluded after it aggregated and bound non-specifically in the assay. Binders were confirmed against 14 of the 15 remaining targets.

The two arms were not run against an identical panel. 15-PGDH and latent GDF-8 were not included in multi-target mode, and Opus 4.8 was additionally run in single-target mode against three targets, so the 26.7% and 35.1% figures do not describe the same set of targets.

The 354 total is a reconciled figure rather than any single laboratory’s count. Adaptyv returned results for 1,296 designs and classified 336 as binders; Twist Bioscience tested 1,260. Of the 1,235 designs measured at both sites, the two laboratories agreed on 89%, with a Cohen’s kappa of 0.71. Reconciliation removed one design from the binder count and added nineteen. Two plates carry caveats in the published materials: the SpCas9 plate at Twist lacked a positive control in the capture format, and the RBX1 plate used the CUL1–RBX1 complex, in which CUL1 occludes part of the RBX1 surface. RBX1 is among the results Anthropic highlights.

Adaptyv’s case study reports 95% expression across the designs. Restricting the comparison to de novo minibinders, Adaptyv says the campaign’s hit rates would have won five of the six protein design competitions it has run. On TREM2 it reports an 80% hit rate against 38.3% among human entrants. Best measured affinity for 15-PGDH improved from 1.7 µM to 33.4 nM, and for RBX1 from 25.7 nM to 3.9 nM. On Nipah glycoprotein, the best human competition result was not beaten. Adaptyv’s post closes by directing readers to its own API.

Anthropic says it intends to carry out further characterisation to confirm both the hit rates and the affinities, and describes the current figures as first-pass.

Failures and one reversal

Maltose-binding protein produced no confirmed binder across 90 designs. Anthropic characterises the target as large, flexible and hydrophilic. BBF-14, a β-barrel with no natural counterpart that is used as a design benchmark for that reason, yielded three binders at sub-micromolar to micromolar affinity.

On TNFα, where the druggable groove lies between two subunits of a trimer, Opus 4.8 produced binders and Mythos Preview did not. The Opus binders were cross-reactive across human, cynomolgus and mouse TNFα. Anthropic says it does not know why the less capable model succeeded where the more capable one failed.

What was and was not measured

Seven of the targets sit in or adjacent to oncology: PD-L1, EGFR, VEGF-A, TrkA, IL-7Rα, RBX1 and 15-PGDH.

The measurement performed was affinity to purified antigen by surface plasmon resonance, at five concentrations in duplicate. No data were reported on cell-surface engagement, receptor blockade, epitope competition against approved biologics, formatting, immunogenicity or in vivo behaviour. Adaptyv’s case study sets out the further stages it considers necessary, developability, cell-based function, organoid work and translational evidence, none of which were part of this campaign.

Reception

The release drew wide coverage across technology and trade press within 72 hours, most of it restating Anthropic’s figures. Substantive assessment has clustered around arguments, some supportive and some critical.

The strongest favourable reading has come from Adaptyv Bio, which ran the assays. Its case study concludes that the models showed expert-level skill at protein design, matching or exceeding human experts on many tasks. Adaptyv is a paid participant in the campaign rather than an independent evaluator, and its post closes by directing readers to its own API.

A second line of support, advanced in outlets including XenoSpectrum, holds that the significant result is architectural rather than molecular: no new protein foundation model was built, and the novelty lies in a single agent handling target research, tool selection, candidate narrowing and ranking under a protocol set by human experts. On this reading the campaign is a claim about orchestration, and the hit rates are evidence for it.

The most direct criticism came from Martin Shkreli, the former pharmaceutical executive, who wrote on X that:

this is not impressive work!!!
affinities are quite low for peptidics. you can take business end of any mab & cut everything, too.
these aren’t useful probe molecules because notice none of them are intracellular! if i needed an extracellular probe, that’s what a mab is for!!!

Several structural caveats sit behind all positions. The work was designed and reported by Anthropic, which has released prompts, sequences, in silico design files and raw data on Hugging Face and announced a Proteinbase deposition; the assay work was performed externally and, according to Anthropic, blinded to model identity. The 10–15% baseline is an aggregate Anthropic drew from publicly documented campaigns in the Proteinbase database, and the head-to-head comparisons are against competition entrants rather than a resourced industrial group working the same targets with comparable compute. The campaign used thousands of H100 hours per target. Most targets were established benchmarks; two were novel, one of which was the excluded GDF-8.

Next steps

The campaign was open-loop: the models designed, the designs were tested, and the results were reported without the models ever seeing them. No multi-round autonomous campaign, in which an agent iterates on its own assay results, has been published. Adaptyv exposes both an API and an MCP server that would permit one.

Anthropic says protein design remains blocked for general access in Fable 5 on dual-use grounds, and that a scientist access programme is planned but not launched. Independent scrutiny is therefore currently limited to re-analysis of the released data rather than replication.

Read more biotech insights on OncoDaily Biotech.

Written by: Semiramida Nina Markosyan, Editor, OncoDaily Canada