Web Scraping and Browser Automation, Responsibly
The planned outline of 10 modules. Lessons are written, run and reviewed before they are published.
Before this track
Planned outline
- Module 1
Permission first: robots.txt, terms, law and limits
Coming soon - Module 2
HTTP, sessions, pagination and APIs
Coming soon - Module 3
Parsing HTML in Python and JavaScript
Coming soon - Module 4
Dynamic pages and browser automation
Coming soon - Module 5
Politeness, caching and robust crawls
Coming soon - Module 6
Scrapy 2.19 spiders, pipelines and operations
Coming soon - Module 7
Cleaning, validating and storing data
Coming soon - Module 8
Scheduling, monitoring, testing and safety
Coming soon - Module 9
Office automation: Excel, PDF and email
Coming soon - Module 10
Advanced collection: many sites, articles, AI
Coming soon
Related tools
Robots.txt URL Tester See which crawlers may fetch a URL — and the exact rule that decides it. Robots.txt Generator & Tester Create, check and test robots.txt rules exactly the way Googlebot reads them. HTTP Header Checker See what a server sends back, and what each header means for search. User Agent Parser See which browser, OS, device or bot sent a request: one user agent or a whole log. HTML / Web Table to CSV & Excel Copy a table from any web page, Word file or Wikipedia article — get clean CSV or Excel. Cron Expression Generator & Explainer Write, check and translate cron schedules — and see exactly when they run. HAR File Analyzer & Sanitizer See what a page loaded and why it was slow — then clean the file before sending it. PDF Table to Excel / CSV Box the table once — get every page of it as one clean spreadsheet.
Official documentation
- www.rfc-editor.org/rfc/rfc9309.html (rfc-editor.org)
- www.rfc-editor.org/rfc/rfc9110.html (rfc-editor.org)
- www.crummy.com/software/BeautifulSoup/bs4/doc (crummy.com)
- docs.scrapy.org/en/latest/intro/tutorial.html (docs.scrapy.org)
- playwright.dev/python/docs/intro (playwright.dev)
- pptr.dev (pptr.dev)
- www.selenium.dev/documentation (selenium.dev)
- cheerio.js.org/docs/intro (cheerio.js.org)
- openpyxl.readthedocs.io/en/stable (openpyxl.readthedocs.io)
- pypdf.readthedocs.io/en/stable (pypdf.readthedocs.io)
- www.meity.gov.in/static/uploads/2024/06/2bf1f0e9f04e6fb4f8fef35e82c42aa5.pdf (meity.gov.in)
- copyright.gov.in/Documents/Copyright_Act_1957.pdf (copyright.gov.in)
More in Data and databases
- Data Analysis with Python (pandas, polars, DuckDB) (coming soon)
- Probability & Statistics (coming soon)
- Excel & Google Sheets (coming soon)
- PostgreSQL in Depth (coming soon)
- All Data tracks
- How we make lessons