Regex Generator from Examples
Type the strings you want to match; get a tested regex, explained part by part.
One example per line · up to 2,000 examples
Options
Regular expression
Generating…
Opens the Regex Tester with your examples as the test text and copies the JavaScript pattern — paste it into the pattern box there.
Use it in code
Verbose form, with comments
The x flag makes the engine ignore the spaces, line breaks and # comments, so the pattern means the same as the one above.
What each part means
Check it
Your examples are checked here as soon as there is a pattern.
Checked with your browser’s JavaScript engine. In Python, .NET and Rust \d, \w and \s also match digits, letters and spaces of other scripts, so a string such as “٣” can match there.
About the Regex Generator from Examples
Paste the strings you want to match, one per line, and get a regular expression that matches exactly those strings — shared beginnings and endings factored out, single characters merged into classes such as [a-c]. Turn on the options to make it more general: digits become \d, letters \w, repeated parts {n} or {min,max}, and case can be ignored.
Every pattern is checked against your examples in your browser as it is generated, and you can add strings that must not match to see whether a generalised pattern goes too far. Pick JavaScript, Python, PHP, Java, C#/.NET, Go or Rust to get the pattern in that flavour’s syntax with a line of code, a commented verbose version where the flavour has one, and a plain-English breakdown of every part.
How to use it
- Type or paste your examples into Examples, one per line — or load one of the example sets. Repeated and blank lines are ignored.
- Read the pattern under Regular expression. It matches every example and, as long as no generalising option is on, nothing else.
- To make it match similar strings too, turn on Digits → \d, Words → \w or Whitespace → \s, and Detect repeats to write
aaaasa{3}. Ignore case matches upper and lower case alike. - Choose the flavour of the language you will use it in, then copy the pattern or the ready-made code. Put strings that must not match under Should not match to check that the pattern is not too broad.
- Use Copy & open in Regex Tester to refine it further: the pattern goes to your clipboard and your examples become the test text there.
Examples
cat cats catalog catalogue
^cat(?:alog(?:ue)?|s)?$
The shared beginning cat is written once; (?:…)? makes the rest optional.
2026-10-04 1999-01-31 2000-02-29
^\d{4}-\d{2}-\d{2}$This matches any date in the same shape, including impossible ones such as 2026-99-99 — a regex generated from examples checks the shape, not the calendar.
report.pdf invoice.pdf notes.txt summary.txt
^(?:(?:invoice|report)\.pdf|(?:summary|notes)\.txt)$
Shared endings are factored out too, and the dot is escaped so that it only matches a dot.
b ba baa baaa
^ba{0,3}$With Detect repeats on, counts that follow each other merge into one range.
Common uses
- Writing a validation pattern for codes, IDs or file names from a list of real values.
- Turning a list of allowed words (log levels, country codes, product names) into one pattern for a filter, a router or a firewall rule.
- Getting a first draft of a pattern to refine in the Regex Tester instead of starting from a blank line.
- Learning how a regex is put together, with every part explained in plain English.
How the pattern is built
The method is the one grex documents. Your examples become a deterministic finite automaton — a tree of their characters, one branch per difference. The automaton is minimised by merging the states that can be followed by exactly the same endings, and the result is solved into a single expression: each state becomes an alternation of its characters followed by what comes after them, with shared beginnings and endings factored out (abc and bc become a?bc).
This page has its own implementation, so details can differ from grex: b, ba, baa, baaa with repeats gives ^ba{0,3}$ here and ^b(?:a{1,3})?$ in grex, which match the same strings. Because the automaton is built from plain characters and an ending is only factored out when that cannot make the pattern ambiguous, the alternatives never start with the same character — so backtracking engines do not have to try long branches twice.
What the options do
- Digits → \d, Words → \w, Whitespace → \s replace a character with its class only when every flavour agrees:
0–9;A–Z,a–z,0–9and_; space, tab, line feed, carriage return and form feed. A digit such as٣or a letter such aséstays as it is, because JavaScript, Go and Java (by default) do not count it while Python, .NET and Rust do. - Non-digits → \D, Non-word → \W, Non-whitespace → \S take only characters that no flavour counts in the class. The more specific option wins: with both Digits and Words on,
5becomes\dandabecomes\w. - Detect repeats writes a part that occurs several times in a row once, with a count:
aaa→a{3},abab→(?:ab){2}. Times in a row sets how often it must occur (2 or more) and Shortest part how long the repeated part must be. Neighbouring counts merge:a,aa,aaa→a{1,3}. - Ignore case lower-cases the examples and adds the case-insensitive flag (
i, or(?i)at the start). - Capturing groups writes
( )instead of(?: )so you can read the parts back. - ^ start and $ end anchor the pattern to the whole text; turn them off to find the examples inside longer text.
- Escape non-ASCII writes characters such as
éand emoji as escapes (\u00E9,\u{1F600}…) in the chosen flavour’s syntax. Invisible characters such as a no-break space are always escaped.
Flavours
- JavaScript: a
/…/flagsliteral; theuflag is added when an example contains characters such as emoji, which JavaScript otherwise sees as two halves. - Python, Java, .NET, Go, Rust: a pattern string, with
(?i)at the start for ignore case; Java gets(?iu)when non-ASCII letters must ignore case too. - PHP (PCRE):
/…/delimiters with modifiers;uis added for non-ASCII text, without which PCRE works on bytes. - .NET matches UTF-16 code units, so emoji are kept out of character classes and grouped before a quantifier.
- Verbose form: Python, PHP, Java, .NET and Rust can read a pattern spread over lines with
#comments (thexflag). JavaScript and Go have no such mode; the breakdown under the pattern explains it instead. - In Python, PHP, Java and .NET
$also matches just before a final line break, soabc\npasses a check forabc. When validating input, write\Z(Python) or\z(the others) instead of$.
Limitations
- The pattern never generalises beyond what the examples show: it has no
+or*, and the counts in{min,max}come from your examples. Edit it by hand (or in the Regex Tester) for open-ended rules such as “one or more digits”. - A generalised pattern checks the shape of the text, not its meaning:
\d{4}-\d{2}-\d{2}also accepts 2026-99-99. - Up to 2,000 different examples, 1,000 characters each and 100,000 characters in total. Blank lines are ignored, so the empty string cannot be one of the examples.
- The check runs your browser’s JavaScript engine. Other flavours get the same pattern in their own syntax; the page does not run Python, PHP, Java, .NET, Go or Rust.
- Very large example sets give very long patterns. Above 20,000 characters the verbose form and the breakdown are left out.
Privacy
Everything happens in your browser. What you enter or open here is not uploaded or stored by MySmartCoPilot.
Frequently asked questions
Will the regex match strings that are not in my list?
Not while the generalising options are off: the pattern matches exactly your examples (with Ignore case, also their upper- and lower-case variants). Digits → \d, Words → \w and the other class options make it match similar strings on purpose. Put strings that must stay out under Should not match to see which ones a pattern lets through.
Why is there no + or * in the pattern?
The generator only knows the examples you gave, so it writes exact counts and ranges such as \d{2,4} taken from them. If any number of digits should be allowed, replace the count with + yourself — or add longer and shorter examples so the range covers what you need.
Which flavour should I choose?
The one of the language or tool that will run the pattern: JavaScript for browsers and Node.js, Python for the re module, PHP for preg_match, Java for java.util.regex, C# for System.Text.RegularExpressions, Go for the regexp package and Rust for the regex crate. Most other tools (editors, grep, databases) understand the PCRE or JavaScript form.
Why does Digits → \d leave some digits as they are?
Because \d means different things in different languages. In JavaScript, Go and Java it is only 0–9; in Python, .NET and Rust it also matches digits of other scripts such as ٣ (Arabic-Indic three). The generator only writes \d where every flavour agrees, so the pattern behaves the same wherever you use it.
How is this different from grex?
It follows the method grex documents (a minimised automaton solved into an expression) and offers similar options, but it is a separate implementation that runs in this page, adds a check against your examples and counter-examples, writes the pattern for seven flavours and explains it. Results can differ in details such as a{0,3} versus (?:a{1,3})?.
Is my data uploaded?
No. The examples are processed by JavaScript in your browser, in a background worker of this page, and work offline once the page has loaded.