DNA to Protein Translator (Transcription & Translation)
DNA → mRNA → protein in any frame and any NCBI genetic code, with an ORF finder.
Spaces, numbers and line breaks are removed. Your sequence never leaves this device.
Protein
mRNA (5′→3′)
Codon by codon
- AUG (Met, the usual start codon)
- Stop
- Ambiguous or context-dependent
Open reading frames
| # | Strand | Frame | Start | End | nt | aa | Start codon | Protein |
|---|
About the DNA to Protein Translator (Transcription & Translation)
Paste a DNA or mRNA sequence to see its mRNA (or either DNA strand) and its protein. You can translate one reading frame, all three forward frames, all six frames (both strands), or the protein a ribosome would make from the first AUG to the next stop codon. Choose any of NCBI’s 27 genetic codes: the standard code, the vertebrate and yeast mitochondrial codes, the bacterial and plastid code, the ciliate codes, and the rest. Amino acids can be shown in one-letter or three-letter code.
The ORF finder lists open reading frames on both strands. You can choose the start codons it accepts and a minimum length, and download the ORFs as FASTA or CSV. A codon-by-codon view marks starts, stops and ambiguous codons. The tool reads plain text, FASTA (with a menu for multi-record files) and GenBank records, and IUPAC ambiguity codes such as N or R are translated wherever the amino acid is certain. Everything runs in your browser — your sequence is never uploaded.
How to use it
- Paste the sequence (or press Load example for the human insulin mRNA). Spaces, numbers and FASTA headers are handled; for a multi-record FASTA file, pick the record to translate.
- Say what the sequence is: the coding strand or mRNA (the usual case), or the template strand written 3′→5′ or 5′→3′.
- Choose the reading frame — +1, +2, +3, from the first AUG, three forward frames or all six — and the genetic code (table 1 for nuclear genes of most organisms, table 2 for vertebrate mitochondria, table 11 for bacteria and plastids).
- Read the protein and the mRNA — or switch that box to the coding or template DNA strand — copy them, or download the protein as FASTA. Below them, check the codon-by-codon view and the ORF finder; set its start codons and minimum length, then download the ORFs.
Examples
NM_000207.3 (the example), from the first AUG
First AUG at position 60 → MALWMRLLPLLALLALWGPDPAAAFVNQHLCG… (110 amino acids)
The same protein as NCBI’s NP_000198.1; the ORF finder reports it at 60–392.
TACTTTATC, written 3′→5′
mRNA AUGAAAUAG → Met-Lys-Stop (MK*)
Human COX1 (NC_012920.1, 5904–7445)
With table 2, the full 513-residue protein ending at an AGA stop. With table 1, it is broken by false stops at UGA, which mitochondria read as Trp.
E. coli lacI, which begins GUG, with table 11
Frame +1 starts with V (Val); the ORF from that GUG starts with M, as in NCBI’s protein
NCBI: an initiator codon is translated as methionine whatever it is.
GCN, RAY, SAR, NNN
A (all four codons are Ala), B (Asp or Asn), Z (Glu or Gln), X (unknown)
Common uses
- Biology homework on transcription and translation, including template-strand questions.
- Checking that a cloned or synthetic gene is in frame and has no premature stop codons.
- Translating mitochondrial, bacterial, plastid or ciliate genes with the right genetic code.
- Finding candidate protein-coding regions (ORFs) in a new sequence, on both strands.
From DNA to protein
The coding strand of a gene has the same sequence as its mRNA, with T in place of U. The template strand is its complement, the strand RNA polymerase reads. To get the mRNA from a template written 3′→5′, take the complement base by base. From a template written 5′→3′, take the reverse complement.
The ribosome reads the mRNA 5′→3′ in codons of three bases, starting at an initiator codon (usually AUG) and stopping at a stop codon. A sequence has three reading frames on each strand:
- +1, +2 and +3 start at its 1st, 2nd and 3rd base;
- −1, −2 and −3 do the same on the reverse complement.
Positions in the tool are always counted on the coding strand.
The genetic codes
The tables are NCBI’s “The Genetic Codes” (Elzanowski & Ostell), taken from NCBI’s gc.prt file, version 4.6. They cover tables 1–6, 9–16 and 21–33, which differ in a few codons. For example, UGA is Trp in vertebrate mitochondria (table 2), where AGA and AGG are stops and AUA is Met. As NCBI specifies, an initiator codon — AUG, or an alternative such as GUG, UUG or CUG — is translated as Met when it starts a protein (the first-AUG translation and the ORFs). Inside a frame it keeps its usual meaning: GUG is Val.
In tables 27, 28 and 31 some codons (UAA, UAG or UGA) can be read either as an amino acid or as a stop, depending on context. The frame translations show the amino acid and the codon view marks the codon; the translation from the first AUG and the ORF finder end there, as at a stop.
Open reading frames
An ORF here runs from a start codon to the next stop codon in the same frame, stop included. Its length is counted in nucleotides, stop included, and it must reach the minimum length you choose. You can accept:
- ATG only;
- ATG and the table’s alternative initiators (for the standard code, also CTG and TTG);
- any sense codon, which gives stop-to-stop ORFs.
For each stop the longest ORF is reported — the one from the first start after the previous stop. You can also ask for ORFs that run off the end without a stop, and hide ORFs lying entirely inside a longer one on any strand or frame. Minus-strand ORFs are given in coding-strand coordinates, so their start is the higher number. The table shows the 200 longest; the FASTA and CSV downloads contain up to 2,000.
Ambiguity codes and gaps
IUPAC-IUB ambiguity codes are expanded to every codon they could stand for:
- if all those codons give one amino acid, it is shown — GCN is Ala;
- if they give only stops, a stop is shown;
- if they give Asp or Asn, the result is B (Asx);
- if they give Glu or Gln, the result is Z (Glx);
- anything else is X (the IUPAC-IUB amino-acid symbols).
Gap characters (- and .) are removed before translating, and a note says how many.
Sources
- NCBI, The Genetic Codes, by A. Elzanowski and J. Ostell, and its data file gc.prt (version 4.6).
- NC-IUB, Nomenclature for Incompletely Specified Bases in Nucleic Acid Sequences (Cornish-Bowden, Nucleic Acids Res. 13, 3021–3030).
- IUPAC-IUB JCBN, Nomenclature and Symbolism for Amino Acids and Peptides, 3AA-1.
- The translations were checked against NCBI’s own proteins for human insulin (NM_000207.3 → NP_000198.1), human mitochondrial COX1 and ND5 (NC_012920.1, table 2) and E. coli lacI (NC_000913.3, table 11).
Limitations
- Introns are not removed. Translate a spliced sequence (mRNA or CDS), not genomic DNA with introns.
- Special recoding is not modelled: selenocysteine at UGA, pyrrolysine at UAG, ribosomal frameshifts, and stop codons completed by polyadenylation in mitochondrial mRNAs.
- Input is limited to 2,000,000 characters. The codon view shows up to 6,000 codons and the ORF list up to 2,000 ORFs.
- An ORF is only a candidate: whether it is really translated depends on promoters, ribosome binding sites, splicing and other signals not checked here.
Privacy
Everything happens in your browser. What you enter or open here is not uploaded or stored by MySmartCoPilot.
Frequently asked questions
How do I translate DNA into protein?
Write the mRNA (the coding strand with U for T), split it into codons from the start, and look each codon up in the genetic code until a stop codon. For AUG GCC CUG UGG → Met Ala Leu Trp (MALW), the start of human insulin. Paste the sequence here and choose the frame, or From the first AUG.
Is my sequence the coding strand or the template strand?
mRNA and coding (CDS) sequences in databases such as GenBank are written as the coding (sense) strand, 5′→3′, with T for U — choose that unless you were told otherwise. Textbook problems often give the template strand 3′→5′ (for example 3′-TAC…-5′). In that case choose Template strand, written 3′→5′: its complement is the mRNA.
Which reading frame is the right one?
For a coding sequence that begins with its start codon, frame +1. Otherwise the right frame is usually the one with a long stretch free of stop codons. Choose All 6 frames or use the ORF finder, which lists the longest open reading frames on both strands.
Why does my mitochondrial gene have stop codons in the middle?
Mitochondria use different genetic codes. In vertebrate mitochondria (table 2) UGA codes for tryptophan rather than stop, AUA codes for methionine, and AGA/AGG are stops. Choose the right table, for example 2 for human or other vertebrate mtDNA, 5 for invertebrates or 3 for yeast.
Why does an ORF start with M when the first codon is GUG or UUG?
Alternative initiator codons are read by the initiator tRNA, so the first amino acid is methionine. NCBI’s convention is to translate the initiator codon as Met whatever it is. Elsewhere in the frame the codon has its normal meaning (GUG = Val).
What do X, B, Z and * mean in the protein?
* is a stop codon. X (Xaa) means the codon contains ambiguity codes that could give different amino acids. B (Asx) means it is Asp or Asn, and Z (Glx) means Glu or Gln. These are the IUPAC-IUB symbols.