Mission Control
What are we engineering today?
Your constructs
Open “My designs” →Loading…
Pick up where you left off
Lineage
Your loop — the flywheel turning
How it works →The Field Atlas — the commons, growing
Open full atlas →Loading the commons…
§ Structure
Predicted structure
Turing loads the wild-type model as soon as you name a protein.
§ Begin
Design a smart mutation library
Provide a wild-type sequence — get a library of multi-mutant variants ranked by predicted fitness, codon-optimized for your expression host, and ready to order from your synthesis vendor. From there: build the construct in the Plasmid Editor, edit it with CRISPR if the project calls for it, and log what you measure to design a sharper round two.
§ 1 · Sequence
Provide a wild-type
FASTA, SnapGene, GenBank, raw DNA, or raw protein. Auto-detected.
§ Check it against your own results
Paste the wild-type and a table of variants you have already measured. TuringDNA scores them blind and reports how well its ranking matches yours — then designs the next round from your numbers. Nothing is sent anywhere but this engine.
One variant per line: the mutation, then the number.
Tabs, commas or spaces all work, and a header row is fine.
Doubles as A24G,L88F.
Multiple CDS features detected — pick the gene to evolve
The longest CDS in a plasmid is usually the antibiotic resistance gene. Pick your gene of interest.
§ 2 · Search
Tune the parameters
Sensible defaults — change if you know what you're after.
§ 3 · Generate
Run the engine
Your individual sequences are never shared, exposed, or reproduced; only anonymous, aggregated signals improve our models. Privacy.
§ Awaiting input
The library appears here.
Provide a sequence above and click Generate. Results render as a sortable variant table with mutation map, PCR primers, and one-click synthesis ordering.
§ 4 · Library
Variants
How to read this
Fitness — the summed ESM-2 log-likelihood gain over wild-type (ΣΔLL) across the variant’s mutations. Higher means the model finds that combination more evolutionarily plausible. It’s a prior, not a verdict — screen experimentally.
Mutations — substitutions vs. wild-type (WT·position·new, e.g. W58L). GC % — GC of the codon-optimized DNA (40–60% is PCR-friendly). Tm — primer melting temp. bp — amplicon length.
Filter box: C49 = variants mutating residue 49 · W58L = that exact substitution · gc>50, tm>58, fitness>2, bp<800 = numeric ranges.
Stay shallow — cap at 3–4 mutations per variant unless you have a structural reason.
§ What Turing learned
Per-substitution score, ΔLL — ESM-2’s prior for that one mutation, before → after your measured results were folded in. Not the same scale as the table’s summed ΣΔLL.
§ How well round 1 predicted your results
Mutation landscape
Where mutations land across the protein. x = residue position (1–N) · bar height = variants sharing that position · colour = the mean ΔLL there, signed, in ESM-2 log-likelihood units. Full scale and domain in the key below.
§ Interaction Radar
Will these mutations play well together?
Your library ranks each substitution on its own. This checks the combinations — ESM-2 re-scores every multi-mutant in context and flags where mutations antagonize, before you spend a cent synthesizing a dud.
Fitness = ΣΔLL — the ESM-2 log-likelihood change vs. wild type, summed over this variant’s substitutions. Unnormalised: no fixed zero and no maximum, and because the table shows the top-ranked slice of the search, negatives rarely survive into it. Ranks variants within this library — not across proteins or runs.
| Rank ↕ | Mutations ↕ | Fitness ΣΔLL ↕ | GC % ↕ | Tm (°C) ↕ | bp ↕ |
|---|
§ Learn from results
Tested these on the bench? Teach the engine.
Enter a measured value for the variants you assayed (any number — activity, expression, stability; higher = better). Leave the rest blank. TuringDNA fits a model on your results and proposes a smarter round 2 — conditioned on your data, not just the prior.
Wild-type, measured in the same run
The anchor. Instruments, buffers, operators and days differ in ways nothing here can see — expressing every variant relative to the wild-type from the same run cancels most of it. Leave this blank and your results are still valid for your own comparison, but they cannot be placed beside another lab’s.
Conditions what might explain a difference between two labs
DNA Design
Score changes to DNA itself — promoters, splice sites, UTRs, any non-coding stretch — against a genome foundation model. Directed Evolution asks whether a protein change is tolerated; this asks whether a sequence change is, in the context it actually sits in.
The stretch of DNA your change sits in — a promoter, a splice junction, a UTR. Context is the point: the same substitution scores differently depending on what surrounds it, which is exactly what a lookup table cannot tell you.
One per line, as <WT base><position><new base>
— e.g. T10G. Positions are 1-based against the reference above.
A variant whose WT base doesn't match the reference is refused, not
scored, and listed separately so a typo can't quietly shrink your table.
The model continues from what you paste. Prometheus tier only — generation is the one real capability difference between the DNA tiers (the checkpoint is identical).
How to read this
Δ log-likelihood — how much more, or less, likely the model finds your sequence after the change. Negative means the model is more surprised by the variant than by the reference: the change breaks a pattern the model learned from real genomes. Near zero means it's unremarkable. Positive means the variant looks more typical than what's there now.
This is zero-shot — no training on your system, no measured data. It ranks candidates; it does not predict expression level, and it is not a substitute for a reporter assay. Treat it as a prior for what to test first.
Reference log-likelihood is the model's score for your unchanged sequence, shown so the deltas have something to sit against.
Therapeutic Compiler
Give it a pathogenic variant; it lowers that to base-editing guides or a prime-editing pegRNA through named passes, and stops at the first one that can't be satisfied. What it refuses to compile — and why — is the point.
Somatic design and assessment only. Germline and embryo applications are refused in code, not just in copy. Output is a design record for humans to evaluate — it is not IND-ready, not a clinical decision, and not a clearance of any strategy for use. Immunogenicity, pharmacokinetics, dosing, manufacturing and delivery efficacy are not modelled and not reported.
Correction runs patient → wild-type. Reversed, you would be designing an editor that installs the disease, so the two fields are named for what they are rather than ref/alt. Leave an allele empty for an insertion or deletion.
Wild-type sequence around the variant. Without it the lesion can still be classified from alleles alone — but nothing is checked against real sequence, and an off-by-one coordinate still type-checks at the allele level.
§ Begin · CRISPR
Design a CRISPR edit.
Paste a gene or region. Get ranked sgRNAs for a SpCas9 / Cas12a knockout — or switch to base editing to install a precise C→T / A→G change (including premature-stop knockouts) with predicted amino-acid outcomes. Free for any account. If a project calls for an edit beyond what a variant library covers — a knockout, a precise base change — design it here, on the same construct.
§ 1 · Target
Paste your DNA
The gene or region you want to edit. ACGT only; FASTA headers OK; up to 1 Mbp.
§ 2 · Guides
Ranked sgRNAs
—
How these numbers were computed
Loading methods…
| Rank | Strand | Position | Spacer | PAM | Composite | On-target | Self-off | KO score | Top indel | FS % | Dominance | Base editor | Edit | Editability | AA change | Outcome | Genome off | Exon context | Sense oligo | Antisense oligo | Order | GC% | Flags |
|---|
Tested these guides at the bench? Enter the measured editing efficiency (% edits / indels). Your guides stay private — only anonymous, aggregated patterns improve the on-target model for everyone.
Primer Analysis
Paste your candidate primers, BLAST them all in one go, and let the fitness model rank them — best amplification of your template, fewest off-target matches. No more BLASTing one primer at a time. Pairs well with Directed Evolution's output — check or refine the PCR primers for the variants you're ordering.
One per line or FASTA. Optional names and forward/reverse tags
(e.g. GeneA_F / GeneA_R) — pairs are auto-detected and scored
as amplicons.
The DNA region you want to amplify. Stays on our server — used to check each primer's binding site, uniqueness and the predicted amplicon size.
How to read this
Fitness — the combined score used to rank your candidates (higher = better pick): Tm, GC, 3′-end stability and specificity.
On template — whether the primer binds your template uniquely. Off-target products — extra amplicons in-silico PCR predicts beyond the intended one (0 = specific). Tm — nearest-neighbor melting temp; aim for a matched pair.
These primers are yours — TuringDNA scores and ranks them; it doesn’t invent them.
The specificity scan runs full in-silico PCR against your chosen genome — the same check NCBI Primer-BLAST performs — so the result is complete here. Predictions are a strong guide; as always, validate critical results at the bench.
Paste the DNA you want to amplify (and, optionally, the region the product must span). TuringDNA designs Tm-matched forward/reverse pairs entirely on-server — no NCBI, nothing leaves.
Paste all the primers you want to run in one tube. We flag 3′ cross-dimers between any two of them, and amplicon sizes too close to resolve on a gel.
In-silico PCR against a genome/FASTA you upload — the same check as the NCBI scan, but it runs on our server, so your primers never leave, and it works for any organism (incl. non-model). Up to ~12 Mb.
Ran a PCR with these primers? Tell us if it amplified. Only anonymous, aggregated patterns (Tm/GC/3′-end buckets) improve the model — your primers stay private.
Plasmid Editor
Import a GenBank or FASTA record (or paste raw DNA), see the annotated circular & linear map, run restriction analysis across 1,000+ enzymes, simulate a digest with a virtual gel, and save constructs to your library. Often the next stop after Directed Evolution — bring a designed insert here to map it, check the digest, and get it synthesis-ready.
§ One workbench
Your construct is the centre of gravity. Import it above to map and cut it here — then, without leaving your sequence, hand any region to the other tools.
Features
Restriction analysis
Tap enzymes to show their cut sites on the map; pick a few and digest to size the fragments.
Restriction sites from Biopython/REBASE (1,000+ enzymes), circular-aware. Auto-annotation flags ORFs + common motifs; GenBank features are kept as-is.
Align two sequences
Compare two DNA or protein sequences — percent identity, mismatches and gaps. Global aligns end-to-end (Needleman–Wunsch); Local finds the best shared region (Smith–Waterman).
Clone & assemble
Paste your fragments — FASTA, or one per line as name: SEQUENCE. Overlapping ends assemble scarlessly; bare junctions get homology arms designed into the primers.
Did your assembly work at the bench? Logging it teaches the model which methods/junctions succeed. Only anonymous, aggregated patterns are kept.
§ Memory
What Turing remembers
Facts Turing carries into new conversations, so you don't re-explain them. It can suggest one; it only keeps what you approve. Sequences are never stored here — only facts about the work.
§ Library
My plasmids
§ Documentation
How TuringDNA works
Turing, and the four tools it runs — what each does, what to trust, how to cite it.
Overview
TuringDNA is driven by Turing — tell it what you want (fetch a gene, evolve it, design guides, check primers) and it calls the right tool, in the right order, in one conversation. Every tool is also available directly, for hands-on work without the chat.
Your individual sequences stay private — never shared, exposed, or reproduced, and they only leave this machine on the explicit opt-in network steps listed below. Only anonymous, aggregated signals improve our models (Privacy Policy, Terms).
Turing
A conversational agent that runs the other four tools for you — describe what you want and it calls the right one, chaining them in a single reply when a task needs more than one (fetch a gene, then evolve it, for example).
- Tools it can call. fetch_sequence (gene symbol/accession → real DNA, via Ensembl/NCBI), design_variant_library (directed evolution), design_crispr_guides, design_primers — the same engines the four direct tools use below, not a separate model.
- Sign-in. Chatting is free; running a tool (fetching a real sequence, generating a design) requires a free account, same as the direct tools.
- Memory. Remembers what it's fetched or designed earlier in the same conversation — refer back to “it” or “that gene” instead of re-pasting.
- Honesty. If something's out of scope — a capability that isn't wired up, or a request needing bench validation it can't do — it says so instead of fabricating a result.
Turing calls the same scoring/design code the direct tools call — the same caveats apply. See each tool's section below for what to trust and how to cite it.
Plasmid Editor
Import and work with circular or linear constructs end to end.
- Import. GenBank · FASTA · EMBL · SnapGene-style features · raw DNA. Annotations are read from the record where present.
- Maps & editing. Annotated circular and linear maps, plus a sequence view with feature colouring, find, translate, reverse-complement and in-place editing.
- Restriction. Scans the REBASE enzyme set (1,000+ enzymes), separates cutters from non-cutters, and simulates a digest on a virtual agarose gel.
- Cloning. Simulates Gibson assembly, Golden Gate, and restriction-ligation — junctions, fragment order, and the assembled product.
- Export. GenBank or FASTA; save constructs to your per-user library for 45 days.
Cloning and digests are in-silico predictions. Confirm overhangs, junctions and orientation before you commit reagents.
CRISPR
Design SpCas9 guide RNAs against a pasted region (or a gene resolved by symbol/accession) — for knockout or base editing.
- Guides. Every NGG-PAM 20-mer with an on-target activity score (Doench-style sequence features), strand and position.
- Off-target. CFD scoring (Doench 2016) against other sites in the input. Base-edit mode adds CBE/ABE windows, the predicted edit, the amino-acid consequence, bystander flags and CRISPR-STOP knockouts.
- Knockout. Predicted indel spectrum, frameshift %, out-of-frame dominance, and a loss-of-function likelihood.
- Cloning oligos. Ready-to-order sense/antisense pairs for the standard vectors (BbsI/BsmBI; Cas12a geometry), copied straight to a vendor order.
Off-target is screened against the genome when you choose an organism: your top-ranked guides are searched with a seed index and CFD-scored. Coverage is the complete genome for E. coli and coding sequence for human (GRCh38) and mouse (GRCm39) — intronic and intergenic off-targets sit outside that index, so a whole-genome tool is still the right call if your application depends on them. Leave the organism unset and only the sequence you pasted is checked.
Primer Analysis
Bring your candidate primers; TuringDNA evaluates and ranks them — it does not invent primers from scratch.
- Scoring. Nearest-neighbor Tm (SantaLucia), GC content, 3′-end stability, and hairpin / self-dimer heuristics for every primer.
- In-silico PCR. Binds each pair against your template and reports the intended amplicon plus any off-target products and their sizes.
- Specificity (opt-in). Checks candidates against NCBI nt and separates intended from off-target hits.
- Ranking. Combines the above into a recommended pair — no BLASTing one primer at a time.
DNA Design
The DNA-level counterpart to Directed Evolution. Everything else here works on protein (ESM-2) or on deterministic sequence bookkeeping; this is the only tool that reads nucleotides with a learned model — Evo 2 (Arc Institute, 7B, Apache 2.0), trained autoregressively on genomes.
- Scoring. Δ log-likelihood for each point substitution against a reference you supply: how much more, or less, likely the model finds the sequence after your change. Negative = the change breaks a pattern the model learned from real genomes.
- Context is the point. The same substitution scores differently depending on what surrounds it. That is what a curated parts table cannot do, and it is why this works on non-coding sequence — promoters, splice sites, UTRs, terminators — that the rest of the toolkit only annotates.
- Generation. Continues a sequence you paste. Prometheus tier only; the checkpoint is identical on both tiers, generation access is the difference.
- Refusals are shown. A variant whose wild-type base disagrees with your reference is listed as not scored, with the reason — never folded into a quietly shorter table.
Zero-shot: no training on your system and no measured data. It ranks candidates and gives you a prior for what to test first — it does not predict expression level and does not replace a reporter assay. Every call runs on a rented GPU; unlike protein's Achilles tier there is no free local path for a 7B model, so DNA scoring needs an account.
Directed Evolution
Design a smart mutation library for an existing protein using ESM-2 zero-shot scoring. Provide a wild-type sequence (or send a CDS from the Plasmid Editor); receive a library of multi-mutant variants ranked by predicted evolutionary fitness, codon-optimized for your host, ready to order.
This is not de novo protein design. It's directed-evolution library generation: improving an existing protein along an existing axis (stability, expression, mild activity tuning).
- Parse & translate. FASTA · SnapGene · GenBank · EMBL · raw DNA · raw protein. Plasmid files surface a CDS picker; raw DNA with multiple stops triggers 6-frame ORF discovery.
- Zero-shot scoring. ESM-2 computes ΔLL = log P(mutant | x_WT) − log P(WT | x_WT) for every position × 19 substitutions (Meier et al. 2021 wild-type marginal scheme).
- Combinatorial search. Simulated annealing over the top-percentile pool of single-site mutations. Multiple restarts; cumulative ΣΔLL as fitness; stop-codon and duplicate-position penalties.
- Codon optimization. Reverse-translate with a host codon-usage table. Synonymously scrub BsaI / BsmBI / NotI sites for Golden Gate compatibility.
| Mutations / variant | Approx. functional retention |
|---|---|
| 1–2 | 70–85% |
| 3–4 | 50–70% |
| 5–6 | 30–55% |
| 7–8 | 15–40% |
| 9+ | < 25% |
Stay shallow. Cap Max mutations / variant at 3–4 unless you have a structural reason. Screen, don't trust.
What ESM-2 sees: evolutionary plausibility from ~65 M UniRef50 sequences — conservation, coevolution, sequence context. What it can't see: structure explicitly, rare-but-essential active-site roles, epistasis between selected mutations; membrane proteins, IDPs and multi-domain assemblies are weakest.
Filter syntax (variant table): C49 = variants mutating residue 49 · W58L = exactly that substitution · gc>50, tm>58, fitness>2, bp<800 = numeric ranges.
What runs locally vs over the network
- Local: plasmid parsing, annotation, restriction & cloning simulation; CRISPR design, CFD off-target and oligo generation; primer scoring and in-silico PCR; ESM-2 inference, simulated annealing and codon optimization; all CSV / Excel / GenBank / FASTA export. Sequences never leave this machine for these steps.
- Opt-in network: NCBI BLAST (sequence identification and primer specificity — sends that sequence to NCBI); AlphaFold-DB / ESMFold structure embeds; Europe PMC literature lookup (IP Radar); synthesis-vendor redirects. Each is explicit and off by default.
Citations
Meier et al. (2021). Language models enable zero-shot prediction of the effects of mutations on protein function. NeurIPS 34.
Lin et al. (2023). Evolutionary-scale prediction of atomic-level protein structure. Science 379:1123–1130.
Doench et al. (2016). Optimized sgRNA design to maximize activity and minimize off-target effects of CRISPR-Cas9. Nat. Biotechnol. 34:184–191.
Hsu et al. (2013). DNA targeting specificity of RNA-guided Cas9 nucleases. Nat. Biotechnol. 31:827–832.
Komor et al. (2016) & Gaudelli et al. (2017). Programmable base editing of C·G and A·T pairs. Nature 533:420 / 551:464.
Allawi & SantaLucia (1997). Thermodynamics and NMR of internal G·T mismatches in DNA. Biochemistry 36:10581–10594.
Licenses of bundled components
- ESM-2 weights — MIT (Meta/FAIR)
- Transformers, Accelerate — Apache 2.0 (Hugging Face)
- PyTorch — BSD-style (Meta)
- Biopython (parsing, restriction, Tm) — BSD-derived
- Mol* viewer — MIT (PDBe / RCSB)
- Inter, JetBrains Mono — SIL Open Font License
- Lucide icons — ISC