09Sources

References

Every paper, every Protein Data Bank entry, every map, every licence and every link this site rests on. If a number appears anywhere on evoai.bio and it is not computed by the engine in front of you, it comes from something on this page.

01Papers

The three papers

The work this site reads. None of it is this site's work, and none of it is restated as this site's result.

02Model

The model, its code and its data

Published by the Arc Institute. This site does not run any of it.

Evo 2 code

The open source implementation and the released weights.

Open the repository

OpenGenome2

The training dataset named in the Nature paper.

Open the dataset

Stanford Report: AI designs a novel E. coli killer

The story on the phage work, and one of the two links the owner fixed for the footer of every page.

Read the story

Stanford Engineering: welcome Evo

The 2024 announcement of the first model, for context on where the work started.

Read the announcement

03Entries

Every Protein Data Bank entry

18 entries, in the order the library lists them, each with its deposited title, its method and resolution, its own primary citation and its baked lattice size.

04Maps

The electron microscopy maps

The two maps behind the comparison this site is built on, and the identifier of every other map in the set.

EMD-77391PDB 36CR

Structure of the Evo-Phi36 bacteriophage

The map associated with the written phage, at 2.9 A.

Open the EMDB entry

EMD-77390PDB 36CQ

Structure of the PhiX174 bacteriophage

The map associated with the natural template, at 2.76 A.

Open the EMDB entry

The other four maps in the set. Four more entries were solved by electron microscopy and carry a map of their own. The identifiers are here because they are part of the record, and no other page on this site prints them.

EMD-27397, PDB 8DES EMD-6035, PDB 3J7W EMD-6324, PDB 3JA7 EMD-28656, PDB 8EXA

04bSequence

The pinned genome

The only data on this site that does not come from the Protein Data Bank. The genome scan reads it and nothing else.

NCBI
Nucleotide

NC_001422.15,386 bases

Escherichia phage phiX174, complete genome

The reference sequence the GENOME simulation scans for GC content, GC skew, tetranucleotide counts and Shine Dalgarno hits. Fetched from the NCBI Nucleotide database on 2026-09-17 and baked into this repository, so the scan reads the same bytes on every device. Its content hash is printed with every genome scan result.

Open the NCBI record

05Licence

Licence and attribution

Structure data is used under the wwPDB dedication. Credit belongs to the depositors and to the primary publication of each entry.

RCSB
CC0

Licence.

Attribution.

Nothing on this site alters a deposited record, and nothing on this site is submitted back to the Protein Data Bank. A label bought with research credit is a label attached by a wallet on this site. It never replaces a deposited title and it never implies anything was renamed anywhere else.

RCSB PDB usage policy RCSB PDB

06Engine

The engine that produced the numbers

The identifier below covers every byte of the engine source. Change one constant and it changes, and the old identifier stays registered so old runs stay replayable.

engine id

A result is a canonical text block whose keccak256 is the hash that goes on chain. The block starts with a fixed header line, then one key and value per metric, sorted, newline separated. There is no JSON in the hashed path and no floating point anywhere in it.

Verify a run in your own browser Open the run console

07Limits

What this site does not claim

The short version of every honesty note, in one place, so a sceptical reader does not have to collect them.

The radius of gyration and the radii this site prints are lattice estimates at the baked cell pitch. They are not atomic values, and the pitch is always printed with them.

A lattice difference is a shape overlap in normalised lattice units. It is not an atomic superposition and it is not an RMSD.

A hydropathy profile is the published Kyte and Doolittle scale applied to a published sequence. It is not a claim about membrane spanning, topology or function.

A genome scan is a pattern count over a published sequence. It is not gene calling, and a hit is not a promise that a gene starts there.

A symmetry overlap measures deposited coordinates at the baked cell pitch. It is not a claim that any capsid is or is not icosahedral.

The chain records what was claimed and when. It does not verify that the computation ran. That is what replication is for, and why the replay button exists.

This site does not run the Evo models, and nothing it prints is model output.