Topic

Tools for scraping and extracting web data · Open source

8

match your filters

8

answered today

0

arrived this week

0

stopped answering

Watching this group for 24 days, checked once a day · how we check

Scans your site the way AI crawlers like GPTBot and ClaudeBot actually see it, scores 32 checks, then an AI assistant walks you through fixes, including on WordPress.

Open source
What to know
How it works
Scans a website the way AI crawlers like GPTBot and ClaudeBot actually see it (they don't run JS), running 32 checks and producing a 0-100 score in about 5 seconds; an AI assistant then walks non-developers through fixes, including on WordPress.
What's different
Open source with a CLI, and built specifically around what AI crawlers see rather than standard SEO checks.
Best for
Site owners, including non-developers on WordPress, worried their site is invisible to AI crawlers despite ranking well on Google.
Details →

Fortress

LiveWatched 12d

A stealth browser engine that lets scraping and automation scripts avoid bot detection, for developers running Playwright or Puppeteer agents.

Open source
What to know
How it works
Drop-in Chromium that reads as ordinary Chrome to detection systems
What's different
keeps existing Playwright/Puppeteer code unchanged, open BSD-3 engine
Best for
developers building scrapers or web-acting AI agents
Details →

Set of open-source skills that let Claude Code or other AI agents scrape Google Maps business listings, extract emails and draft cold outreach.

FreeOpen source
What to know
How it works
8 installable skills covering scraping, email extraction, competitor and review analysis
Pricing
3 skills free/local; scraping skills use a paid API with free starting credits
Best for
Developers building automated local-business lead-generation workflows
Details →

Client Koi

LiveWatched 13d

A free, open-source Chrome extension that scrapes Google Maps to extract verified B2B leads with emails and social profiles by deep-crawling business websites, including automated CAPTCHA bypass.

FreeOpen source
What to know
How it works
Deep-crawls B2B business websites found via Google Maps to find verified emails and social profiles, with automated 2Captcha bypass.
Pricing
Free.
Details →

Cockroach Crawler

LiveWatched 8d

An open-source crawling toolkit that gives AI agents restricted, governed web access instead of unrestricted network access.

Open source
What to know
How it works
Static and rendered crawling, adaptive traversal, screenshots, PDFs, structured extraction via explicit origins, budgets, and runtime controls.
What's different
Focused on governed/restricted access for AI agents rather than open crawling.
Details →

Clearcote

LiveWatched 18d

An open-source, de-Googled Chromium build with fingerprint controls and SDKs for browser automation.

Open source
What to know
How it works
Provides engine-level fingerprint controls, reproducible builds, and drop-in SDKs for Playwright and Puppeteer.
Best for
Developers doing browser automation who need fingerprint control.
Details →

Reddit OpenScraper

LiveWatched 11d

An open-source scraper for pulling Reddit posts, nested comments and user histories with built-in rate limiting.

Open source
What to know
How it works
Extracts subreddit posts, comments and user histories with auto pagination
What's different
open source, built-in rate limiting to avoid blocks
Best for
Developers who need Reddit data at scale
Details →

ShieldFont

LiveWatched 2d
Open source
Details →