What is Common Crawl?
Common Crawl is a freely accessible archive of web pages that the nonprofit Common Crawl Foundation has been building since 2008 with its crawler CCBot, adding a new crawl roughly once a month. Each crawl contains raw data, metadata, and extracted text from around two billion pages and is available for download at no cost. This makes Common Crawl one of the most important sources from which large language models draw their knowledge of the web.

Term profile at a glance
| Attribute | Details |
|---|---|
| Part of speech | Proper noun, used without an article |
| Pronunciation | ˈkɒmən ˈkrɔːl |
| Operator | Common Crawl Foundation, a US nonprofit organization founded in 2007 |
| Crawler | CCBot, user agent CCBot/2.0 (https://commoncrawl.org/faq/) |
| Current crawl | CC-MAIN-2026-39: 2.17 billion pages from 40.4 million hosts, captured from September 4 to 17, 2026 |
| File formats | WARC (raw data), WAT (metadata and links), WET (plain text) |
| Related terms | CCBot, AI crawler, training data, robots.txt, web graph |
How does Common Crawl work?
Common Crawl works in 4 steps: The crawler selects candidates, fetches the pages, archives them in 3 file formats, and publishes the finished crawl as an open data set. A crawl runs for about two weeks; the September 2026 crawl ran from September 4 to 17.
- Selecting candidates: Common Crawl computes the harmonic centrality of every domain from its own link graph. The value measures how close a domain sits to all other domains through links. Well-linked domains receive more crawl budget and end up in the archive with more pages.
- Fetching pages: The crawler CCBot is based on Apache Nutch. It reads a domain’s robots.txt first, follows up to four redirects, and uses the sitemap listed there. CCBot does not execute JavaScript and does not set cookies.
- Archiving: Every fetched page is stored as a WARC file with the complete HTTP response, as a WAT file with metadata and links, and as a WET file with the plain text. Common Crawl also archives the robots.txt files and the error responses of each crawl.
- Publishing: The finished crawl is hosted in the public data sets of Amazon Web Services and at data.commoncrawl.org. Access is free, and a URL index makes individual pages searchable.
Common Crawl captures a sample of the web. The foundation itself describes its data set as a randomly selected subset of each website. For website owners, this means: Which pages end up in the archive changes from crawl to crawl.
Content that is loaded later via JavaScript is missing from the archive. This is the same hurdle that JavaScript SEO describes for search engines; with CCBot, there is no rendering service that resolves it later.
Which language models use Common Crawl?
At least 30 of 47 large language models released between 2019 and October 2023 were trained on a filtered version of Common Crawl, which is 64 percent. This is shown by the study “Training Data for the Price of a Sandwich” published by the Mozilla Foundation in February 2024. For OpenAI’s GPT-3, more than 80 percent of the training tokens came from Common Crawl.
A large language model usually processes Common Crawl through a prepared version. The raw data contains menus, cookie notices, duplicates, and spam; AI developers therefore rely on filtered data sets. The 3 best known are:
- C4 (Colossal Clean Crawled Corpus): Google built C4 for the language model T5. The data set is one of the most widely used filtered versions.
- Pile-CC: Pile-CC is the Common Crawl portion of the data set “The Pile” by EleutherAI.
- FineWeb: Hugging Face filtered FineWeb from numerous Common Crawl snapshots and released it openly.
Meta combined two filtered versions for the first Llama generation. The filters mainly remove pornography, boilerplate text, and duplicates. They keep the selection of pages that Common Crawl made beforehand unchanged: Content from English-language and well-linked domains is overrepresented.
Why does old content live on in AI answers through Common Crawl?
A language model knows a website in the state that the crawl captured at the time of training. If you update a page today, that changes nothing about the knowledge of a model trained on a crawl from 2024. The old version remains in the archive, in the filtered data sets built from it, and in every model trained on them.
Three consequences matter for website owners:
- Outdated information remains available: Old prices, former services, or an outdated company name remain available in earlier crawls and can show up in model answers.
- New content needs a training cycle: A new page only reaches model knowledge once a later crawl captures it and a provider retrains with it.
- Freshness comes through live search: Many AI systems compare their training knowledge with current search results via retrieval augmented generation. There, the current state of your page counts.
For a brand, this means: Consistent information over the years pays off twice. If you keep your name, services, and core statements stable, you shape training knowledge and live answers with the same information. This turns the company into a clearly recognizable entity for models.
How do you check whether your website is included in Common Crawl?
Whether your website is in Common Crawl is shown by the public URL index at index.commoncrawl.org. A single query per crawl returns all captured addresses of your domain with status code and file type:
https://index.commoncrawl.org/CC-MAIN-2026-39-index?url=your-domain.com/*&output=json
The identifier CC-MAIN-2026-39 stands for the September 2026 crawl; the list of all crawls is available at index.commoncrawl.org/collinfo.json. The index is heavily rate limited. Send your queries one after another and pause between them, otherwise the server blocks your IP address for 24 hours.
We checked this for taismo.de on September 24, 2026. The August crawl CC-MAIN-2026-34 contains 213 retrievable pages of the domain, the September crawl CC-MAIN-2026-39 only 69. The two crawls share just 7 URLs, out of around 450 addresses in the sitemap. The measurement confirms the sampling principle: Each crawl shows a different slice of the same website.
How do you control CCBot with robots.txt?
CCBot follows the robots.txt and responds to the user agent CCBot. Two lines block the crawler for the entire website:
User-agent: CCBot
Disallow: /
If CCBot should keep crawling, but more slowly, a crawl delay in seconds is enough. CCBot explicitly honors this directive, while Googlebot ignores it:
User-agent: CCBot
Crawl-delay: 2
Four points are part of controlling CCBot:
- The block works going forward: It prevents future fetches. Earlier crawls containing your pages remain published, as do the data sets built from them.
- Fake CCBots exist: Common Crawl itself warns about crawlers posing as CCBot. Genuine requests come from fixed IP ranges listed at
index.commoncrawl.org/ccbot.jsonand resolve via reverse DNS tocrawl.commoncrawl.org. - The sitemap helps when you allow it: CCBot reads every sitemap listed in the robots.txt. A well-maintained sitemap increases the chance that the important pages are captured.
- Google stays unaffected: A CCBot block only affects Common Crawl. Googlebot and your rankings in Google Search do not change as a result.
CCBot is one of many bots that collect content for AI systems. How the others differ, and which of them serve training data and which serve live answers, is explained in the article on AI crawlers.
Should you block or allow CCBot?
For companies that want to be mentioned in AI answers, allowing CCBot is the better choice in most cases. A block removes the website from the future training data of many open and commercial models. Language models then learn about the brand, its services, and its terminology only from third-party sources.
| Goal | Recommendation | Reason |
|---|---|---|
| Become visible in AI answers | Allow CCBot | Your own content flows into future training data and shapes what models know about the brand |
| Protect paid or exclusive content | Block CCBot | The archive is freely accessible, and anyone can download and reuse the texts |
| Limit server load | Set a crawl delay | CCBot stays active and fetches pages at longer intervals |
Large publishers such as the New York Times already block CCBot for their content. For a mid-sized company with a service that needs explaining, the same block has a different effect: It makes the company less visible to models without protecting a business model. That is why taismo.de allows CCBot.
Once allowed, you can influence your presence in the archive. Because Common Crawl prioritizes by harmonic centrality, backlinks from well-connected websites increase the chance that more pages of a domain are captured. The llms.txt adds a clear entry map for language models. How these levers work together is described on our page about GEO and AI visibility.
Frequently asked questions about Common Crawl
Is Common Crawl the same as the Internet Archive?
Common Crawl and the Internet Archive are two separate nonprofit projects. The Internet Archive’s Wayback Machine shows people earlier versions of individual pages in the browser, while Common Crawl provides bulk data for research and data analysis.
Does it cost anything to use Common Crawl?
Common Crawl data is free. The foundation provides it through the public data sets of Amazon Web Services and through data.commoncrawl.org; costs arise at most for your own computing and storage capacity.
Does a CCBot block hurt my Google ranking?
A CCBot block has no effect on your Google ranking. Google crawls with its own Googlebot and does not use Common Crawl as a source for search.
Welt der SEO lernen?