Fuzzy Match (Fuzzy Lookup) Tool
A fuzzy VLOOKUP: the closest match, its score and the runner-up for every row.
Try before you buy.
- Free preview: the match counts, and the first rows of the result with their best matches and scores (up to 10; a single row is shown with its results hidden).
- Locked until you unlock it: download and copy.
- Unlock: Premium pass, ₹799 for 30 days, a one-time payment that never renews.
Ways to unlock shows how to get the full result.
Printing this result is locked in the free preview.
1. Your list
The rows to look up — for example your customer or product list.
Or paste the table (from Excel, Google Sheets or a CSV)
2. Look up in
Where the matches come from — for example a CRM or supplier export.
Or paste the table (from Excel, Google Sheets or a CSV)
3. How to match
Rows below it keep their closest value and score, but nothing is brought across.
4. Result
Locked in the free preview. Opens the ways to unlock this result.
Locked in the free preview. Batch runs unlock with a pass.
Locked in the free preview. Query results unlock with a pass.
About the Fuzzy Match (Fuzzy Lookup) Tool
When the same customer, supplier or product is written differently in two files — “Sharma Traders Pvt. Ltd.” and “SHARMA TRADERS PRIVATE LIMITED”, “Jon Smith” and “John Smith” — an exact VLOOKUP finds nothing. This tool finds the closest value of the other list for every row, gives its similarity score (0 to 100) and the runner-up, and brings the lookup table’s columns across for rows that pass your threshold, like a fuzzy VLOOKUP. It can also find near-duplicates within one list and group them.
Choose the similarity — token sort, token set, Jaro–Winkler, Levenshtein, ratio or Soundex — and what to ignore before comparing: case, accents, punctuation, and noise words such as Pvt, Ltd or Inc. Long lists are matched in a background thread with blocking, so 10,000 × 10,000 rows take seconds, all in your browser.
How to use it
- Open or paste Your list (the rows to look up) and the list to Look up in — CSV, Excel or cells copied from a spreadsheet. For near-duplicates, choose Find near-duplicates in one list and give one list.
- Choose the column to match in each list, the similarity (token sort suits names in any word order) and the threshold: the score a best match needs to count.
- Tick what to ignore — case, accents, punctuation, noise words — and edit the noise words for your data.
- Tick the columns of the lookup list to bring across, and read the counts: exact matches, matches, close calls to check and rows below the threshold.
- Read the result table. Download it as CSV, TSV or Excel, or copy it — with a Premium pass, or after unlocking this result; without one the page shows a free preview.
Examples
Sharma Traders Pvt. Ltd. ↔ SHARMA TRADERS PRIVATE LIMITED
100 (exact after ignoring case, punctuation and the noise words Pvt, Ltd, Private, Limited)
Jon Smith ↔ John Smith (token sort ratio)
94.7: 2 × 9 shared characters ÷ (9 + 10)
MARTHA ↔ MARHTA (Jaro–Winkler)
96.1: Jaro 0.944 plus the bonus for the shared first three letters
Common uses
- Matching customer or supplier names between an accounting export and a CRM or bank statement.
- Bringing prices, codes or account numbers from one list into another when the names are not spelt the same.
- Cleaning a contact or product list: grouping near-duplicates before merging them.
- Checking survey answers or free-text entries against a list of known values.
- Matching product titles between two marketplaces or suppliers.
The similarity measures
Every measure gives 0 (nothing alike) to 100 (the same after normalising):
- Token sort ratio — sorts the words of each value alphabetically, then compares the whole text: word order does not matter (“Traders Sharma” = “Sharma Traders”).
- Token set ratio — compares the shared words with each side’s other words, so a value whose words all appear in the other scores 100 (“Acme Tools” and “Acme Tools India”). Generous: use a higher threshold.
- Ratio — 2 × the longest common subsequence ÷ the total length of both values (the normalised Indel similarity).
- Levenshtein — 100 × (1 − edits ÷ length of the longer value), where an edit inserts, deletes or replaces one character (Levenshtein’s distance).
- Jaro–Winkler — counts characters that match within a small window and swapped pairs, and adds a bonus for up to four shared first letters (Winkler’s string comparator, built for matching names in census records).
- Soundex — the share of words that sound alike in English by the American Soundex code (Smith and Smyth are both S530), by the US National Archives’ rules. Numbers and non-Latin letters are not coded.
The ratio and the token sort and token set ratios are defined as in the open-source RapidFuzz library.
Normalising and noise words
Before comparing, both values are cleaned the same way: case ignored, accents removed (é = e), apostrophes dropped (Zoë’s = Zoes), other punctuation turned into spaces, “&” read as “and”, extra spaces removed, and noise words dropped. The default noise words are common company suffixes — Pvt, Private, Ltd, Limited, LLC, LLP, Inc, Incorporated, Corp, Corporation, Co, Company, PLC, GmbH, AG, Pty, SA, SRL, BV — and “the”; edit the list for your data (for example add “traders” or “and sons”). A value made only of noise words keeps them.
Thresholds and statuses
Each row’s best match gets a status: Exact (the same after normalising), Match (at or above the threshold), Check: close runner-up (a match whose second-best candidate is within 3 points, so either could be right), Below the threshold (the closest value is shown with its score, but no columns are brought across) or No candidate. Moving the threshold re-labels the rows at once without matching again. Around 85–90 suits token sort and Jaro–Winkler on names; token set needs more, as it is generous with subsets.
Large lists
Comparing every value with every other is slow for big lists (10,000 × 10,000 is 100 million comparisons), so the tool uses blocking: each value of the lookup list is indexed by its three-letter pieces, and each row is compared only with the 60 values sharing the most of its rarest pieces. Lists of up to 1,000 × 1,000 are always compared in full; tick Compare every pair to do the same for up to 25 million pairs. Blocking can occasionally miss a match that shares few pieces (for example when two words are run together), which a full comparison would find.
Limitations
- Compares text, not meaning: “IBM” and “International Business Machines” are not alike to any of these measures.
- Soundex is designed for English surnames and ignores digits and non-Latin scripts; use another measure for codes or other languages.
- Best matches are chosen for each row on its own: two rows can get the same best match.
- Lists of up to about 100,000 rows each, as far as your device’s memory allows; the page shows the first 100 rows of the result and the downloads have every row.
Privacy
Everything happens in your browser. What you enter or open here is not uploaded or stored by MySmartCoPilot.
Frequently asked questions
What do I get without a pass?
Without a pass, Fuzzy Match (Fuzzy Lookup) Tool shows the match counts, and the first rows of the result with their best matches and scores (up to 10; a single row is shown with its results hidden). Until you unlock it, the result can’t be downloaded or copied. A Premium or Ultimate pass, a one-time payment that never renews, unlocks the full result. The pricing page lists the passes and their prices.
How do I do a fuzzy VLOOKUP?
Give your list as Your list and the table with the values to bring back as Look up in, pick the column to match in each, and tick the columns to bring across. Rows whose best match reaches the threshold get those columns; the others show their closest value and its score.
Which similarity should I choose?
For names of people or companies, token sort ratio or Jaro–Winkler. For codes and IDs with typos, Levenshtein. When one value may be a shorter form of the other (“Acme” and “Acme Tools India”), token set ratio. For English surnames that sound alike, Soundex.
What threshold should I use?
Start at 85 and look at the rows marked “Check” and “Below the threshold”. If wrong matches pass, raise it; if right ones fail, lower it. The counts update as you move it.
Can it find duplicates in one list?
Yes: choose Find near-duplicates in one list. Values at or above the threshold are grouped (if A is like B and B is like C, all three share a group), and the result lists every row with its group, the group’s size and its first value.
Is my data uploaded?
No. Both lists are read and matched in your browser, in a background thread, and stay on your device.