Also known as CommonCrawl, commoncrawl.org, Common Crawl Foundation
nonprofit organization eponym of a large web periodic and open crawl
Common Crawl - Open Repository of Web Crawl Data
We build and maintain an open repository of web crawl data that can be accessed and analyzed by anyone.
commoncrawl.org →Link to the official site · 4,242 chars · not written by Vinony
via Wikidata · CC0
via Wikidata sitelinks · CC0
Discovered by embedding cosine similarity (sentence-transformers MiniLM, 384-dim).