Skip to content

Curate a new Web scraping collection#5283

Open
aminembarki wants to merge 1 commit into
github:mainfrom
aminembarki:add-web-scraping-collection
Open

Curate a new Web scraping collection#5283
aminembarki wants to merge 1 commit into
github:mainfrom
aminembarki:add-web-scraping-collection

Conversation

@aminembarki

Copy link
Copy Markdown

Please confirm this pull request meets the following requirements:

Which change are you proposing?

  • Suggesting edits to an existing topic or collection
  • Curating a new topic or collection
  • Something that does not neatly fit into the binary options above

Curating a new topic or collection

  • I've formatted my changes as a new folder directory, named for the topic or collection as it appears in the URL on GitHub (e.g. https://github.com/topics/[NAME] or https://github.com/collections/[NAME])
  • My folder contains a *.png image (if applicable) and index.md
  • All required fields in my index.md conform to the Style Guide and API docs: https://github.com/github/explore/tree/main/docs

Web scraping is one of the most active areas of open source — the web-scraping topic alone has tens of thousands of repositories, and projects like Scrapy, Firecrawl, Crawlee, and Playwright are among the most-starred on GitHub — yet Explore has no collection for it (the closest, opensource-testing and digital-preservation, cover browser automation only for QA and archiving). This collection curates the ecosystem end to end: crawling frameworks across seven languages (Python, JavaScript, Go, Rust, Java, Ruby, and PHP), AI-era extraction tools, headless browser automation, HTML parsers, and self-hosted platforms. Every entry was checked to be actively maintained and non-archived as of July 2026; historically important but archived projects (PhantomJS, pyspider, Portia) were deliberately excluded.

For transparency: I maintain one of the smaller entries (crawlee-cloud/crawlee-cloud) and have ordered it last. Although it is young on GitHub, it is not a toy project — it runs in production today, orchestrating roughly 190 scraper runs daily. That said, I'm happy to drop it from the list if the maintainers prefer.

@aminembarki
aminembarki requested a review from a team as a code owner July 26, 2026 20:21
@github-actions

Copy link
Copy Markdown
Contributor

Maintainer triage

Collection web-scraping

Item Stars Last push Owner type Notes
scrapy/scrapy 63,421 2026-07-25 Organization
apify/crawlee 25,008 2026-07-24 Organization
apify/crawlee-python 9,367 2026-07-22 Organization
gocolly/colly 25,392 2026-06-18 Organization
D4Vinci/Scrapling 71,330 2026-07-26 User
spider-rs/spider 2,625 2026-07-23 Organization
projectdiscovery/katana 17,226 2026-07-20 Organization
firecrawl/firecrawl 156,393 2026-07-26 Organization
unclecode/crawl4ai 75,110 2026-07-25 User
ScrapeGraphAI/Scrapegraph-ai 28,640 2026-07-20 Organization
microsoft/playwright 93,506 2026-07-24 Organization
puppeteer/puppeteer 95,359 2026-07-26 Organization
SeleniumHQ/selenium 34,325 2026-07-26 Organization
seleniumbase/SeleniumBase 12,894 2026-07-26 Organization
lightpanda-io/browser 32,524 2026-07-26 Organization
browserless/browserless 13,524 2026-07-24 Organization
cheeriojs/cheerio 30,430 2026-07-25 Organization
jhy/jsoup 11,379 2026-07-26 User
sparklemotion/nokogiri 6,277 2026-07-24 Organization
PuerkitoBio/goquery 14,974 2026-07-17 Organization
symfony/panther 3,065 2026-06-04 Organization
dgtlmoon/changedetection.io 32,495 2026-07-24 User
getmaxun/maxun 16,855 2026-07-26 Organization
crawlee-cloud/crawlee-cloud 32 2026-07-26 Organization

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant