Log Pattern Analyzer (Group Millions of Log Lines into Templates)
Turn a huge log into a short list of message templates, errors first.
Free preview.
- Free preview: every count (entries, templates, errors, spikes), the first templates of each list (up to 10) and the timeline under a MySmartCoPilot mark.
- Locked until you unlock it: download and copy.
- Unlock: Premium pass, ₹799 for 30 days, a one-time payment that never renews.
Ways to unlock shows how to get the full result.
Printing this result is locked in the free preview.
Log files
Log
After (the second log)
Grouping options
Lower joins more lines into fewer templates.
Timeline
Spikes
What changed between the two logs
Templates
Locked in the free preview. Opens the ways to unlock this result.
Locked in the free preview. Batch runs unlock with a pass.
Locked in the free preview. Query results unlock with a pass.
About the Log Pattern Analyzer (Group Millions of Log Lines into Templates)
A production log has millions of lines but only a few hundred kinds of message. Load the log and the analyzer reads timestamps and levels, then groups the lines into templates — Connection to <IP> refused, GET /api/orders/<NUM> 200 <NUM>ms — with the Drain algorithm, turning IDs, numbers, IP addresses and times into placeholders. Templates are listed errors first, with their count, share, first and last time seen and real examples.
A timeline shows errors and warnings over time and marks the spikes, with the templates behind each one. Compare two logs (before and after a deploy, or a good day and a bad one) to see the patterns that are new, the ones that vanished and the ones whose rate changed threefold or more.
Plain text, JSON lines (also journalctl -o json and Docker’s json-file logs), syslog (RFC 5424 and RFC 3164), journalctl output and Kubernetes container logs are recognised. Files are read as a stream in a background worker, so multi-gigabyte logs and .gz files work; nothing is uploaded.
Without a pass the result is a free preview: every count, the first templates of each list (up to 10) and the timeline under a mark. The full lists, the CSV and copying unlock with a Premium pass.
How to use it
- Choose the log file (several files of one log, such as rotated
app.log.1,app.log.2, are read in order) or paste lines. Compressed.gzfiles are read directly. - Wait for the progress bar: the result shows the number of entries, templates and errors, and the time span.
- Read the templates, errors first. Open Examples to see real lines and filter by text or level; the CSV of the whole list unlocks with a Premium pass.
- Check the timeline for spikes and the templates behind them. To compare, choose Compare two logs and add the second log.
- Templates too coarse or too fine? Raise or lower the similarity under Grouping options, or turn a placeholder off.
Examples
2024-05-14 09:00:04,120 ERROR [db-1] Pool - Connection to 10.0.0.5:5432 refused 2024-05-14 09:03:51,007 ERROR [db-2] Pool - Connection to 10.0.0.6:5432 refused
error · 2× · [db-<NUM>] Pool - Connection to <IP> refused
The timestamp and level are read off first; the thread number and the address become placeholders.
Before: 30 minutes of a service log · After: the 30 minutes after the deploy
New in the second log: [http-nio-<NUM>-exec-<NUM>] pay.Gateway - Payment provider timeout after <NUM> ms (order <NUM>) Gone from the second log: [main] cache.Warmup - Cache warmed in <NUM> ms
Common uses
- Finding the handful of errors that matter in a log of millions of lines.
- Checking what a release changed: new error messages, messages that stopped, rates that jumped.
- Finding the minute an incident started and the messages that came with it.
- Choosing which messages to alert on, or which noisy messages to silence.
How the grouping works
Variable parts are replaced first: numbers by <NUM>, IPv4 and IPv6 addresses by <IP>, UUIDs by <UUID>, hex ids by <HEX>, times by <TIME>, e-mail addresses and URLs by <EMAIL> and <URL>.
Then Drain (He, Zhu, Zheng and Lyu, “Drain: An Online Log Parsing Approach with Fixed Depth Tree”, IEEE ICWS) looks the message up in a tree: first by its number of words, then by its first word (a word with a digit takes a wildcard branch); each level of depth above 4 adds the next word. In that leaf it joins the most similar template when the share of equal words reaches the similarity (0.4 by default); the words that differ become <*>. Otherwise the message starts a new template. The grouping follows Drain3, with its defaults: similarity 0.4, depth 4 and 100 children per node.
Because the first word chooses the branch, two lines whose first word differs (alice logged in, bob logged in) stay apart; mask such parts, or set the depth to 3, where only the number of words and the similarity count.
Spikes and comparisons
- The timeline divides the time span into at most 72 bars of a round width (a second up to a week). A bar is a spike when it is at least three times the median bar, at least 5 entries, and above the median by six times the median absolute deviation; the same is done for errors alone.
- In a comparison, rates are per 10,000 entries of each log, so logs of different lengths can be compared. “Much more or less frequent” lists templates whose rate changed by a factor of three or more (with at least 5 lines in one of the logs).
- Times are taken as written in the log; times without a zone are read as UTC, and syslog times without a year are shown without one.
Limitations
- Files are read as UTF-8; logs in UTF-16 or another encoding should be converted first. Only .gz compression is read; unpack .zip or .bz2 archives first.
- Grouping is a heuristic: some templates may join two kinds of message or split one. Adjust the similarity and the placeholders, and check the examples.
- Lines longer than 20,000 characters are cut for grouping. Up to 3 examples are kept per template.
- Speed depends on the device and on the lines: a laptop reads about two hundred thousand lines a second or more, so a gigabyte of logs takes roughly half a minute to a minute, and a phone takes longer.
Privacy
Everything happens in your browser. What you enter or open here is not uploaded or stored by MySmartCoPilot.
Frequently asked questions
What do I get without a pass?
Without a pass, Log Pattern Analyzer (Group Millions of Log Lines into Templates) shows every count (entries, templates, errors, spikes), the first templates of each list (up to 10) and the timeline under a MySmartCoPilot mark. Until you unlock it, the result can’t be downloaded or copied. A Premium or Ultimate pass, a one-time payment that never renews, unlocks the full result. The pricing page lists the passes and their prices.
Is my log uploaded?
No. Files are read by a background worker in your browser, line by line, and nothing leaves your device. That also means a log with personal data or secrets stays where it is.
How large can the log be?
Files are streamed, so the size of the file does not matter for memory: only the templates and their counts are kept. The limit is time — roughly half a minute to a minute per gigabyte on a laptop, longer on a phone.
Which formats are recognised?
JSON lines with common fields (time/timestamp/@timestamp, level/severity, msg/message), journalctl (-o json and the default and short-iso outputs), Docker json-file and Kubernetes CRI container lines, RFC 5424 and RFC 3164 syslog, and plain text with a timestamp and a level near the start. Stack traces are joined to the entry above them.
Why are two different messages in one template?
They have the same number of words, the same first word, and enough equal words for the similarity threshold. Raise Similarity to join a template (for example to 0.6), or set a greater tree depth, to keep them apart.
Can I turn a template into a parsing rule?
Yes: Write a grok pattern for it opens the Grok Debugger with the template’s example lines, ready for a Logstash or Elasticsearch pattern.