Server Access Log Analyzer
See who crawls your site, what breaks and what Googlebot does — without uploading the log.
Free preview.
- Free preview: the summary counts (requests, people, bots, Googlebot, 4xx and 5xx), the first rows of each table and list (up to 10) and a marked crawl chart.
- Locked until you unlock it: download, copy and opening in another app.
- Unlock: Premium pass, ₹799 for 30 days, a one-time payment that never renews.
Ways to unlock shows how to get the full result.
Printing this result is locked in the free preview.
Open access logs
Paste log lines instead
Options
Access logs contain your visitors’ IP addresses: they are read by this page on your device and never uploaded.
Reading
Overview
Findings
Nothing stands out: no server errors, missing pages or redirects worth a note for Googlebot.
Crawl activity
This log has no usable times, so requests cannot be shown over time.
Bots
Named by what their user agent claims. Most bots announce themselves; a bot that copies a browser’s user agent is counted under Browsers.
| Bot and kind | Requests | Share | IPs | 2xx | 3xx | 4xx | 5xx | Last seen |
|---|
Check that crawlers are genuine
Anyone can send a Googlebot user agent. To check an address the way Google documents it, look up its reverse DNS name — it must end in googlebot.com, google.com or googleusercontent.com (Bingbot: search.msn.com) — and then look up that name: it must give back the same IP. DNS Lookup does both for one address at a time.
Requests
Lines that could not be read
The first few, as they appear in the file (cut after 300 characters). They are not counted anywhere else.
Locked in the free preview. Opens the ways to unlock this result.
Locked in the free preview. Batch runs unlock with a pass.
Locked in the free preview. Query results unlock with a pass.
About the Server Access Log Analyzer
Your web server writes a line for every request it answers: who asked (IP address and user agent), for which URL, when, and with which status code. This analyzer reads those access logs — Apache and Nginx (Common or Combined Log Format), IIS and Amazon CloudFront (W3C extended), AWS load balancers, and JSON logs from Cloudflare, Nginx, Caddy or Traefik, plain or gzip-compressed — and shows status codes, the most requested URLs, missing pages (404), redirects, server errors, people versus bots, and how often Googlebot, Bingbot and AI crawlers came, day by day.
Logs contain your visitors’ IP addresses, so they are read on your device: files are streamed through a background thread in this page and never uploaded.
How to use it
- Get the access log of your site: in cPanel under Metrics → Raw Access; over SSH usually
/var/log/nginx/access.log,/var/log/apache2/access.log(Debian, Ubuntu) or/var/log/httpd/access_log(Red Hat family); for IIS underinetpub\logs\LogFileson the system drive; for CloudFront, load balancers and Cloudflare Logpush in the bucket you send logs to. Rotated.gzfiles can be added as they are. - Drop one or more files here, or paste a few lines. The format is detected from the first lines of each file; choose it under Options if you need to.
- Read the findings first: server errors, missing pages and redirects that Googlebot ran into, robots.txt problems and AI crawler traffic, each with the table that lists the URLs.
- Switch the tables between All requests, Browsers, All bots, Googlebot, Bingbot and AI crawlers to see what each of them got. Without a pass the page shows a free preview: the summary counts in full and the first rows of each table. With a pass, download any table as CSV, or copy a summary.
- Before you block a crawler or trust one, check that its IP addresses really belong to it: with a pass, copy the IP and user-agent pairs from Check that crawlers are genuine.
Examples
Access log summary: sample-access.log Format: Apache / Nginx combined Period: 28 Sept 2026, 00:00 – 4 Oct 2026, 23:53 (UTC) Requests: 7,930 · data sent 171 MB Status: 2xx 5,896 (74%), 3xx 1,870 (24%), 4xx 100 (1.3%), 5xx 64 (0.8%) Browsers (people): 6,294 (79%) · bots: 1,636 (21%) · other or no user agent: 0 Googlebot: 732 requests · 2xx 622, 3xx 54, 4xx 42, 5xx 14 Bingbot: 232 requests · 2xx 200, 3xx 32, 4xx 0, 5xx 0 AI crawlers: 245 requests · 2xx 245, 3xx 0, 4xx 0, 5xx 0 Most requested missing URLs: 55 × /blog/old-post-2019/ 13 × /wp-login.php 7 × /.env 7 × /admin/config.php 7 × /shop/?sort=price Findings: [Warning] Googlebot got 14 server errors (1.9% of its requests) [Warning] Googlebot asked for 4 URLs that do not exist (42 requests, 404 or 410) [Note] 30% of Googlebot’s requests were for URLs with a query string [Note] AI crawlers and assistants made 245 requests (3.1%) Counts are by user agent: bots can pretend to be Googlebot. Verify crawler IP addresses before blocking or trusting them.
The sample is a made-up week of a small site (press “Try a sample log”). The "Googlebot" requests from 203.0.113.200 in it only claim to be Googlebot: that address is not one of Google’s.
66.249.66.1 - - [04/Oct/2026:06:25:24 +0000] "GET /blog/ HTTP/1.1" 200 15342 "-" "Mozilla/5.0 (Linux; Android 6.0.1; Nexus 5X Build/MMB29P) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/141.0.7390.122 Mobile Safari/537.36 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)"
Counted as: Googlebot Smartphone (search engine) · GET /blog/ · 200 OK · 15.0 KB · 4 Oct 2026, 06:25 (UTC)
The user agent says Googlebot and the address is in Google’s published crawler ranges; the analyzer counts what the user agent claims, and the IP check tells you whether it is true.
{"ClientIP":"203.0.113.7","ClientRequestHost":"www.example.com","ClientRequestMethod":"GET","ClientRequestURI":"/shop/?sort=price","EdgeResponseStatus":404,"EdgeResponseBytes":1520,"EdgeStartTimestamp":"2026-10-04T08:15:02Z","ClientRequestUserAgent":"Mozilla/5.0 (compatible; bingbot/2.0; +http://www.bing.com/bingbot.htm)","VerifiedBotCategory":"Search Engine Crawler"}Counted as: Bingbot Desktop (search engine, verified by Cloudflare) · GET /shop/?sort=price · 404 Not Found
Common uses
- Finding the pages Googlebot cannot reach — server errors, missing pages and redirect chains — and fixing them before they drop out of the index.
- Seeing how much of your traffic and bandwidth goes to bots, SEO tools and AI crawlers, and which ones to allow or block in robots.txt.
- Checking that a migration worked: old URLs should now answer 301 and nobody should land on 404s.
- Spotting scanners that probe for /wp-login.php, /.env or /xmlrpc.php, and the addresses they come from.
Formats it reads
- Apache and Nginx: the Common and Combined Log Formats —
%h %l %u %t "%r" %>s %b "%{Referer}i" "%{User-agent}i"in Apache, the predefinedcombinedformat in Nginx — also with a virtual host in front (vhost_combined) and extra fields after the user agent, such as Nginx’s"$http_x_forwarded_for". Escaped characters (\xC3\xA9,\") are decoded. See Apache log files and Nginx log_format. - W3C extended: IIS and Amazon CloudFront standard logs; the columns are read from the
#Fields:line. Their times are UTC. - AWS Application Load Balancer access logs.
- JSON lines (one object per line, or a JSON array): Cloudflare Logpush HTTP requests with any timestamp format, Nginx
escape=json, Caddy, Traefik, Elastic and Google Cloud load balancer logs. Fields are found by their usual names (status,uriorrequest,remote_addr,user_agent,time…).
Files may be plain text or gzip (.gz). ZIP, bzip2 and xz archives need to be extracted first.
How bots and people are told apart
Each request is sorted by its user agent. Google’s crawlers are named by type from the tokens Google documents — Googlebot Smartphone and Googlebot Desktop (both send Googlebot/2.1, the smartphone one with a mobile browser string), Googlebot Image and Video, StoreBot, Google-InspectionTool, GoogleOther and AdsBot (Google’s common crawlers) — and so are the fetches people ask Google for, such as FeedFetcher-Google and Google-Read-Aloud (user-triggered fetchers). Bingbot is named the same way. Every other bot is matched against the open crawler-user-agents list and labelled by kind: search engine, AI crawler or assistant, SEO tool, link preview, monitoring, security scanner, script or HTTP library. A user agent that looks like a web browser counts as a browser, most likely a person.
A user agent is only a claim. Anyone can send "Googlebot", and bots that copy a browser’s user agent are counted as browsers. Google documents how to verify its crawlers: a reverse DNS lookup of the IP must end in googlebot.com, google.com or googleusercontent.com, and a forward lookup of that name must return the same IP; Google also publishes its crawlers’ IP ranges (verifying Google’s crawlers). For Bingbot the reverse name ends in search.msn.com, and Bing publishes bingbot.json. Cloudflare logs that include VerifiedBotCategory already say which bots Cloudflare verified, and the analyzer uses it.
What the findings are based on
- Server errors (5xx) and
429 Too Many Requestsmake Google’s crawlers slow down, and URLs that keep failing are eventually dropped from the index; indexed pages that return 404 or 410 are removed; 301 and 308 redirects are a strong signal for the target URL, 302 and 307 a weak one (Google: HTTP status codes). - When robots.txt answers with a server error, Google pauses crawling for the first 12 hours and then uses its last good copy; a 404 robots.txt means “crawl everything” (Google’s robots.txt specification).
- Duplicate URLs, such as sort and filter parameters, and long redirect chains waste crawling — something Google says mostly matters for large sites (crawl budget).
Limitations
- Counts are by user agent: the analyzer cannot verify IP addresses itself. Bots that pretend to be browsers are counted as people.
- Each table keeps up to 200,000 different URLs, 100,000 IP addresses and 30,000 user agents per view; requests beyond that are counted in the totals but not listed (the table says so).
- Custom log formats are read only when their fields keep the order of the Combined Log Format; extra fields at the start or end are fine.
- Times are shown in the zone written in the log (UTC for IIS, CloudFront and load balancers). Logs with a CDN in front only see the requests the CDN passed on.
- Very large logs (gigabytes) work, but take a while on phones; reading a week or a month at a time is faster.
Privacy
Log files are read in this browser, in a background thread of this page, and never uploaded: no request, IP address or user agent from your logs leaves your device. Nothing is stored; closing the report or the page removes it from memory.
Frequently asked questions
What do I get without a pass?
Without a pass, Server Access Log Analyzer shows the summary counts (requests, people, bots, Googlebot, 4xx and 5xx), the first rows of each table and list (up to 10) and a marked crawl chart. Until you unlock it, the result can’t be downloaded, copied or opened in another app. A Premium pass, a one-time payment that never renews, unlocks the full result. The pricing page lists the passes and their prices.
Are my log files uploaded?
No. The files are read from your device by this page, the counting runs in a background thread of your browser, and the report exists only in this tab. That matters because access logs contain your visitors’ IP addresses, which are personal data in many countries.
Can it tell a real Googlebot from a fake one?
Not from the log alone: anyone can send a Googlebot user agent. With a pass, use Check that crawlers are genuine to copy the IP addresses, then verify them as Google documents (reverse DNS, then forward DNS) — one at a time with DNS Lookup, or against Google’s published IP ranges. Cloudflare logs with VerifiedBotCategory are the exception: there Cloudflare has already checked.
Why do almost all requests come from one or two IP addresses?
Your server sits behind a proxy, load balancer or CDN, and the log records that machine’s address instead of the visitor’s. If the log has an X-Forwarded-For field, turn on Client IP from X-Forwarded-For under Options; otherwise configure the server to log the original address.
Does Google treat 404 and 410 differently?
Google’s documentation says all 4xx codes except 429 are treated the same: the content is considered gone, and indexed URLs that return them are removed from the index. Use 404 or 410 for pages you deleted on purpose, and a 301 redirect for pages that moved.
Which log should I analyze for SEO?
The log of the server that answers crawlers. If a CDN serves cached pages, your origin server never sees those requests: analyze the CDN’s logs (CloudFront, Cloudflare Logpush) for the full picture, or both.
How much log data do I need?
Crawling changes from day to day, so a few weeks show patterns better than a single day. Add several rotated files at once — access.log, access.log.1, access.log.2.gz … — and they are analyzed together.