TuringDNA / Research engine
Research workspace
Continue your research
Most recently workedYour constructs
View saved work →Loading…
Scientific tools Design, build, and verify
Lineage
Research activity
How it works →Model benchmark & contributionsInspect the evidence
Field atlasPublished and contributed measurements
Substitution evidence
Open full atlas →Loading the commons…
Structure
Predicted structure
Loads as soon as you name a protein.
Begin
Design a variant library
A wild-type sequence in; a ranked, codon-optimized variant library out.
1 · Sequence
Provide a wild-type
FASTA, SnapGene, GenBank, raw DNA, or raw protein. Auto-detected.
Check it against your own results
Paste wild-type plus variants you have already measured. We score blind and report the match.
One per line: mutation, then value. Tabs, commas or spaces. Doubles as A24G,L88F.
Multiple CDS features detected — pick the gene to evolve
The longest CDS is usually the resistance gene — pick your gene of interest.
2 · Search
Configure the library
Set library size, mutation limits, and expression host.
3 · Generate
Run the engine
Your sequences are never shared or reproduced; only aggregate signals train our models. Privacy.
Awaiting input
The library appears here.
Provide a sequence above and click Generate.
4 · Library
Variants
How to read this
Fitness — the summed ESM-2 log-likelihood gain over wild-type (ΣΔLL) across the variant’s mutations. Higher means the model finds that combination more evolutionarily plausible. It’s a prior, not a verdict — screen experimentally.
Mutations — substitutions vs. wild-type (WT·position·new, e.g. W58L). GC % — GC of the codon-optimized DNA (40–60% is PCR-friendly). Tm — primer melting temp. bp — amplicon length.
Filter box: C49 = variants mutating residue 49 · W58L = that exact substitution · gc>50, tm>58, fitness>2, bp<800 = numeric ranges.
Cap at 3–4 mutations unless you have a structural reason.
What Turing learned
Per-substitution score, ΔLL — ESM-2’s prior for that one mutation, before → after your measured results were folded in. Not the same scale as the table’s summed ΣΔLL.
How well round 1 predicted your results
Mutation landscape
Where mutations land. x = residue position (1–N) · height = variants there · colour = mean ΔLL, signed (ESM-2 log-likelihood).
Interaction Radar
Will these mutations play well together?
Your library ranks substitutions singly. This re-scores every multi-mutant in context and flags where they antagonize.
Fitness = ΣΔLL — the ESM-2 log-likelihood change vs. wild type, summed over this variant’s substitutions. Unnormalised: no fixed zero and no maximum, and because the table shows the top-ranked slice of the search, negatives rarely survive into it. Ranks variants within this library — not across proteins or runs.
| Rank ↕ | Mutations ↕ | Fitness ΣΔLL ↕ | GC % ↕ | Tm (°C) ↕ | bp ↕ |
|---|
Learn from results
Tested these on the bench? Teach the engine.
A measured value per assayed variant — any scale, higher = better. Leave the rest blank.
Wild-type, measured in the same run
Wild-type from the same run. Without it, results hold for your own comparison but not across labs.
Conditions what might explain a difference between two labs
DNA Design
Score non-coding DNA — promoters, splice sites, UTRs — against a genome model.
The DNA around your change — promoter, splice junction, UTR.
One per line, e.g. T10G. 1-based against the reference. Mismatched WT bases are refused, not scored.
The model continues from what you paste. Prometheus tier only.
How to read this
Δ log-likelihood — how much more, or less, likely the model finds your sequence after the change. Negative means the model is more surprised by the variant than by the reference: the change breaks a pattern the model learned from real genomes. Near zero means it's unremarkable. Positive means the variant looks more typical than what's there now.
This is zero-shot — no training on your system, no measured data. It ranks candidates; it does not predict expression level, and it is not a substitute for a reporter assay. Treat it as a prior for what to test first.
Reference log-likelihood is the model's score for your unchanged sequence, shown so the deltas have something to sit against.
Therapeutic Compiler
Turn a pathogenic variant into base-editing guides or a prime-editing pegRNA.
Somatic design and assessment only. Germline and embryo applications are refused in code, not just in copy. Output is a design record for humans to evaluate — it is not IND-ready, not a clinical decision, and not a clearance of any strategy for use. Immunogenicity, pharmacokinetics, dosing, manufacturing and delivery efficacy are not modelled and not reported.
Correction runs patient → wild-type. Leave an allele empty for an indel.
Wild-type sequence around the variant. Without it, nothing is checked against real sequence.
Begin · CRISPR
Design a CRISPR edit.
Paste a gene. Get ranked sgRNAs for knockout, or switch to base editing.
1 · Target
Paste your DNA
ACGT only; FASTA headers OK; up to 1 Mbp.
2 · Guides
Ranked sgRNAs
—
How these numbers were computed
Loading methods…
| Rank | Strand | Position | Spacer | PAM | Composite | On-target | Self-off | KO score | Top indel | FS % | Dominance | Base editor | Edit | Editability | AA change | Outcome | Genome off | Exon context | Sense oligo | Antisense oligo | Order | GC% | Flags |
|---|
Tested at the bench? Enter measured efficiency (% edits/indels). Guides stay private; only aggregate patterns are kept.
Primer Analysis
BLAST your candidate primers in one go and rank them by predicted amplification.
One per line or FASTA. Name pairs GeneA_F / GeneA_R to auto-detect.
The region to amplify. Stays on our server.
How to read this
Fitness — the combined score used to rank your candidates (higher = better pick): Tm, GC, 3′-end stability and specificity.
On template — whether the primer binds your template uniquely. Off-target products — extra amplicons in-silico PCR predicts beyond the intended one (0 = specific). Tm — nearest-neighbor melting temp; aim for a matched pair.
These are your primers, scored and ranked. To generate new pairs, use the Design tab.
The specificity scan runs full in-silico PCR against your chosen genome — the same check NCBI Primer-BLAST performs — so the result is complete here. Predictions are a strong guide; as always, validate critical results at the bench.
Paste the DNA to amplify. Tm-matched pairs designed on-server — nothing leaves.
All primers for one tube. We flag 3′ cross-dimers and amplicon sizes too close to resolve.
In-silico PCR against a genome you upload. Runs on our server, any organism, up to ~12 Mb.
Ran a PCR with these? Tell us if it amplified. Aggregate patterns only; primers stay private.
Plasmid Editor
Import GenBank, FASTA or raw DNA. Annotated map, 1,000+ enzymes, assembly simulation.
One workbench
Your construct is the centre of gravity. Import it above to map and cut it here — then, without leaving your sequence, hand any region to the other tools.
Features
Restriction analysis
Tap enzymes to show cut sites, then digest to size the fragments.
Restriction sites from Biopython/REBASE (1,000+ enzymes), circular-aware. Auto-annotation flags ORFs + common motifs; GenBank features are kept as-is.
Align two sequences
Compare two sequences. Global aligns end-to-end; Local finds the best shared region.
Clone & assemble
Fragments as FASTA, or name: SEQUENCE per line. Overlapping ends assemble scarlessly.
Knock a cassette into a chromosomal locus by yeast homologous recombination — knock-in (a payload at a site) or knock-out (the flanks straddle the deleted ORF). Paste the genomic DNA flanking the site; the engine can't see the genome, so it designs from what you give it and flags that the arms must be BLAST-checked.
Did it work at the bench? Logging teaches the model which junctions succeed. Aggregate only.
Library
My plasmids
Documentation
How TuringDNA works
What each tool does, what to trust, how to cite it.
Overview
TuringDNA is driven by Turing — tell it what you want (fetch a gene, evolve it, design guides, check primers) and it calls the right tool, in the right order, in one conversation. Every tool is also available directly, for hands-on work without the chat.
Your individual sequences stay private — never shared, exposed, or reproduced, and they only leave this machine on the explicit opt-in network steps listed below. Only anonymous, aggregated signals improve our models (Privacy Policy, Terms).
Turing
A conversational agent that runs the other four tools for you — describe what you want and it calls the right one, chaining them in a single reply when a task needs more than one (fetch a gene, then evolve it, for example).
- Tools it can call. fetch_sequence (gene symbol/accession → real DNA, via Ensembl/NCBI), design_variant_library (directed evolution), design_crispr_guides, design_primers — the same engines the four direct tools use below, not a separate model.
- Sign-in. Chatting is free; running a tool (fetching a real sequence, generating a design) requires a free account, same as the direct tools.
- Memory. Remembers what it's fetched or designed earlier in the same conversation — refer back to “it” or “that gene” instead of re-pasting.
- Honesty. If something's out of scope — a capability that isn't wired up, or a request needing bench validation it can't do — it says so instead of fabricating a result.
Turing calls the same scoring/design code the direct tools call — the same caveats apply. See each tool's section below for what to trust and how to cite it.
Plasmid Editor
Import and work with circular or linear constructs end to end.
- Import. GenBank · FASTA · EMBL · SnapGene-style features · raw DNA. Annotations are read from the record where present.
- Maps & editing. Annotated circular and linear maps, plus a sequence view with feature colouring, find, translate, reverse-complement and in-place editing.
- Restriction. Scans the REBASE enzyme set (1,000+ enzymes), separates cutters from non-cutters, and simulates a digest on a virtual agarose gel.
- Cloning. Simulates Gibson assembly, Golden Gate, and restriction-ligation — junctions, fragment order, and the assembled product.
- Export. GenBank or FASTA; save constructs to your per-user library for 45 days.
Cloning and digests are in-silico predictions. Confirm overhangs, junctions and orientation before you commit reagents.
CRISPR
Design SpCas9 guide RNAs against a pasted region (or a gene resolved by symbol/accession) — for knockout or base editing.
- Guides. Every NGG-PAM 20-mer with an on-target activity score (Doench-style sequence features), strand and position.
- Off-target. CFD scoring (Doench 2016) against other sites in the input. Base-edit mode adds CBE/ABE windows, the predicted edit, the amino-acid consequence, bystander flags and CRISPR-STOP knockouts.
- Knockout. Predicted indel spectrum, frameshift %, out-of-frame dominance, and a loss-of-function likelihood.
- Cloning oligos. Ready-to-order sense/antisense pairs for the standard vectors (BbsI/BsmBI; Cas12a geometry), copied straight to a vendor order.
Off-target is screened against the genome when you choose an organism: your top-ranked guides are searched with a seed index and CFD-scored. Coverage is the complete genome for E. coli and coding sequence for human (GRCh38) and mouse (GRCm39) — intronic and intergenic off-targets sit outside that index, so a whole-genome tool is still the right call if your application depends on them. Leave the organism unset and only the sequence you pasted is checked.
Primer Analysis
Bring your candidate primers; TuringDNA evaluates and ranks them — it does not invent primers from scratch.
- Scoring. Nearest-neighbor Tm (SantaLucia), GC content, 3′-end stability, and hairpin / self-dimer heuristics for every primer.
- In-silico PCR. Binds each pair against your template and reports the intended amplicon plus any off-target products and their sizes.
- Specificity (opt-in). Checks candidates against NCBI nt and separates intended from off-target hits.
- Ranking. Combines the above into a recommended pair — no BLASTing one primer at a time.
DNA Design
The DNA-level counterpart to Directed Evolution. Everything else here works on protein (ESM-2) or on deterministic sequence bookkeeping; this is the only tool that reads nucleotides with a learned model — Evo 2 (Arc Institute, 7B, Apache 2.0), trained autoregressively on genomes.
- Scoring. Δ log-likelihood for each point substitution against a reference you supply: how much more, or less, likely the model finds the sequence after your change. Negative = the change breaks a pattern the model learned from real genomes.
- Context is the point. The same substitution scores differently depending on what surrounds it. That is what a curated parts table cannot do, and it is why this works on non-coding sequence — promoters, splice sites, UTRs, terminators — that the rest of the toolkit only annotates.
- Generation. Continues a sequence you paste. Prometheus tier only; the checkpoint is identical on both tiers, generation access is the difference.
- Refusals are shown. A variant whose wild-type base disagrees with your reference is listed as not scored, with the reason — never folded into a quietly shorter table.
Zero-shot: no training on your system and no measured data. It ranks candidates and gives you a prior for what to test first — it does not predict expression level and does not replace a reporter assay. Every call runs on a rented GPU; unlike protein's Achilles tier there is no free local path for a 7B model, so DNA scoring needs an account.
Directed Evolution
Design a variant library for an existing protein using ESM-2 zero-shot scoring. Provide a wild-type sequence (or send a CDS from the Plasmid Editor); receive a library of multi-mutant variants ranked by predicted evolutionary fitness, codon-optimized for your host, ready to order.
This is not de novo protein design. It's directed-evolution library generation: improving an existing protein along an existing axis (stability, expression, mild activity tuning).
- Parse & translate. FASTA · SnapGene · GenBank · EMBL · raw DNA · raw protein. Plasmid files surface a CDS picker; raw DNA with multiple stops triggers 6-frame ORF discovery.
- Zero-shot scoring. ESM-2 computes ΔLL = log P(mutant | x_WT) − log P(WT | x_WT) for every position × 19 substitutions (Meier et al. 2021 wild-type marginal scheme).
- Combinatorial search. Simulated annealing over the top-percentile pool of single-site mutations. Multiple restarts; cumulative ΣΔLL as fitness; stop-codon and duplicate-position penalties.
- Codon optimization. Reverse-translate with a host codon-usage table. Synonymously scrub BsaI / BsmBI / NotI sites for Golden Gate compatibility.
| Mutations / variant | Approx. functional retention |
|---|---|
| 1–2 | 70–85% |
| 3–4 | 50–70% |
| 5–6 | 30–55% |
| 7–8 | 15–40% |
| 9+ | < 25% |
Stay shallow. Cap Max mutations / variant at 3–4 unless you have a structural reason. Screen, don't trust.
What ESM-2 sees: evolutionary plausibility from ~65 M UniRef50 sequences — conservation, coevolution, sequence context. What it can't see: structure explicitly, rare-but-essential active-site roles, epistasis between selected mutations; membrane proteins, IDPs and multi-domain assemblies are weakest.
Filter syntax (variant table): C49 = variants mutating residue 49 · W58L = exactly that substitution · gc>50, tm>58, fitness>2, bp<800 = numeric ranges.
What runs locally vs over the network
- Local: plasmid parsing, annotation, restriction & cloning simulation; CRISPR design, CFD off-target and oligo generation; primer scoring and in-silico PCR; ESM-2 inference, simulated annealing and codon optimization; all CSV / Excel / GenBank / FASTA export. Sequences never leave this machine for these steps.
- Opt-in network: NCBI BLAST (sequence identification and primer specificity — sends that sequence to NCBI); AlphaFold-DB / ESMFold structure embeds; Europe PMC literature lookup (IP Radar); synthesis-vendor redirects. Each is explicit and off by default.
Citations
Meier et al. (2021). Language models enable zero-shot prediction of the effects of mutations on protein function. NeurIPS 34.
Lin et al. (2023). Evolutionary-scale prediction of atomic-level protein structure. Science 379:1123–1130.
Doench et al. (2016). Optimized sgRNA design to maximize activity and minimize off-target effects of CRISPR-Cas9. Nat. Biotechnol. 34:184–191.
Hsu et al. (2013). DNA targeting specificity of RNA-guided Cas9 nucleases. Nat. Biotechnol. 31:827–832.
Komor et al. (2016) & Gaudelli et al. (2017). Programmable base editing of C·G and A·T pairs. Nature 533:420 / 551:464.
Allawi & SantaLucia (1997). Thermodynamics and NMR of internal G·T mismatches in DNA. Biochemistry 36:10581–10594.
Licenses of bundled components
- ESM-2 weights — MIT (Meta/FAIR)
- Transformers, Accelerate — Apache 2.0 (Hugging Face)
- PyTorch — BSD-style (Meta)
- Biopython (parsing, restriction, Tm) — BSD-derived
- Mol* viewer — MIT (PDBe / RCSB)
- Inter, JetBrains Mono — SIL Open Font License
- Lucide icons — ISC