Reverse Complement Generator (DNA/RNA)
Reverse-complement any DNA or RNA — FASTA, GenBank or plain — with IUPAC codes.
Spaces, line breaks and GenBank numbering are removed. Nothing leaves your device.
Reverse complement
The result appears here as you type.
| # | Header | Length | GC |
|---|
About the Reverse Complement Generator (DNA/RNA)
The two strands of a DNA double helix run in opposite directions and pair base by base: A with T, and C with G. The reverse complement of a sequence is the other strand, written in the usual 5′→3′ direction. You need it to design a reverse primer, to read a gene on the minus strand, to check whether a site is palindromic, or to search a sequence in both orientations.
This generator gives the reverse complement, the complement alone, or the reverse alone. It handles one sequence or many, and DNA or RNA. Every IUPAC ambiguity code is complemented correctly (R↔Y, K↔M, B↔V, D↔H; S, W and N stay the same), and lower-case (soft-masked) letters stay lower-case. FASTA headers are kept. GenBank and EMBL records, or pasted GenBank lines with their numbers and spaces, are cleaned up automatically. Everything runs in your browser — your sequences are never uploaded.
How to use it
- Paste your sequence: plain DNA or RNA, one or more FASTA records, a GenBank or EMBL record, or lines copied from a GenBank file (the numbers and spaces are removed).
- Choose Reverse complement, Complement or Reverse.
- Optionally choose the output letters (keep, DNA with T, or RNA with U) and the line length, tick One sequence per line for a list of primers, or add “reverse complement” to the FASTA headers.
- Copy the result or download it as a FASTA or text file. Problems are reported with the line and column, so you can find a stray character quickly.
Examples
5′-ATGCGT-3′
5′-ACGCAT-3′
Complement TACGCA, read backwards.
GAATTC (EcoRI site)
GAATTC
Many restriction sites are their own reverse complement.
AUGGC
GCCAU
Sequences with U and no T stay RNA; choose DNA (T) to convert.
RYKMBVDHNSW
WSNDHBVKMRY
>seq1 aaccGGTT
>seq1 AACCggtt
Lower-case letters stay lower-case, so masked regions are kept.
fwd ACGTTG rev, GGCAAT T7 TAATACGACTCACTATAGGG
fwd CAACGT rev, ATTGCC T7 CCCTATAGTGAGTCGTATTA
With One sequence per line, a name before the sequence is kept as typed.
Common uses
- Writing the reverse primer from the binding site you want on the top strand.
- Reading a gene or feature annotated on the minus (complementary) strand of a genome.
- Checking whether a restriction site, a hairpin stem or a probe is palindromic.
- Getting the template (antisense) strand of an mRNA for homework or lab notes.
- Converting many FASTA records at once while keeping their headers and masking.
The complement rules
Watson–Crick pairing gives A↔T (A↔U in RNA) and C↔G. For uncertain positions the calculator uses the IUPAC-IUB nomenclature for incompletely specified bases (Cornish-Bowden, Nucleic Acids Res. 13), Table 2:
- R (A or G) ↔ Y (C or T);
- K (G or T) ↔ M (A or C);
- B (not A) ↔ V (not T);
- D (not C) ↔ H (not G);
- S (G or C), W (A or T) and N (any base) complement to themselves.
Gaps written as - or . stay in place in the complement and move with the sequence when it is reversed.
Reverse, complement or both?
Sequences are conventionally written 5′→3′. The complement pairs each base but leaves the order alone, so it reads 3′→5′. The reverse complement also flips the order, so the other strand reads 5′→3′, as it would in a database. That is usually what you want, for example for a primer. Reverse alone just reads the sequence backwards, without pairing; it is not a biological strand, but it is sometimes asked for in exercises.
Formats it reads
- Plain text: spaces, line breaks, tabs and digits are removed, so lines pasted from a GenBank file such as “1 gatcctccat atacaacggt” work.
- FASTA: one or many records. Each header is kept, and the output keeps each record’s line length unless you choose another. Lines starting with ; are treated as comments.
- GenBank and EMBL flat files: the name and definition become a FASTA header, and the sequence after ORIGIN (or SQ) is used. The result is written as FASTA with 60 bases per line.
- One sequence per line: each line is a separate sequence, optionally after a name and a space, tab, comma or semicolon. A name is recognised by a letter that is not a nucleotide code (fwd, M13F), a digit next to a letter (R1, T7) or a tab, comma or semicolon after it; once one line has a name, the first part of every line is read as a name. A line such as “R ACGT” on its own is read as the sequence RACGT, so write “R1 ACGT” or “R, ACGT”.
Input can be up to 2,000,000 characters. Anything that is not a nucleotide code is reported with its line and column. A protein sequence pasted by mistake is recognised by letters such as E, F, I, L, P and Q.
Sources
- NC-IUB, Nomenclature for Incompletely Specified Bases in Nucleic Acid Sequences (A. Cornish-Bowden, Nucleic Acids Res. 13, 3021–3030), Tables 1 and 2.
Limitations
- Only standard bases and IUPAC codes are accepted. Modified bases (inosine, methylated bases and so on) must be replaced, for example by N, before converting.
- Very long inputs (above 2,000,000 characters, about 2 Mb) are refused; split them into smaller files.
- Features and annotations in GenBank or EMBL records are not carried over — only the name, definition and sequence.
Privacy
Everything happens in your browser. What you enter or open here is not uploaded or stored by MySmartCoPilot.
Frequently asked questions
What is the reverse complement of a DNA sequence?
It is the sequence of the opposite strand, read 5′→3′. Replace each base by its partner (A↔T, C↔G), then reverse the order. For ATGCGT the complement is TACGCA, and reversing gives the reverse complement ACGCAT.
Why reverse as well as complement?
The strands are antiparallel: where one runs 5′→3′ left to right, its partner runs 5′→3′ right to left. Reversing the complement writes the partner strand in the standard 5′→3′ direction, the way databases and primer orders expect.
How are N, R, Y and the other ambiguity codes complemented?
Each code stands for a set of bases, and its complement is the code for the complementary set. R (A or G) becomes Y (C or T), K becomes M, B becomes V and D becomes H, and the reverse. S (G or C), W (A or T) and N (any base) are their own complements.
Does it keep lower-case letters and FASTA headers?
Yes. Lower-case input, as in soft-masked sequences, stays lower-case in the output, and each FASTA header is kept as it was. You can add “reverse complement” to the headers if you like.
How do I get the RNA complement?
Paste RNA (with U), and the complement of A is written as U automatically. To turn a DNA sequence into an RNA result, choose RNA (U) for the output letters. The DNA (T) choice converts the other way.
Can I paste a GenBank record?
Yes — either the whole record, from LOCUS to //, or just the numbered sequence lines from below ORIGIN. The numbers and spaces are removed. A full record’s name and definition are used as the FASTA header.