Skip to main content

What is training data?

Training data is the set of data from which an AI model learns language, patterns, and factual knowledge during training. For large language models, training data consists mostly of publicly accessible web text, books, code, and licensed sources. Whatever is in the training data is stored permanently in the model weights after training and ends at a cutoff date; more recent content only reaches a model at runtime through retrieval.

Training data explained: what an AI model learns in training and what it retrieves at runtime, taismo SEO wiki

Term profile at a glance

Attribute Details
Part of speech Noun, uncountable (training data)
Syllables train·ing da·ta
German equivalent Trainingsdaten
Related terms Validation data, test data, pretraining, fine-tuning, knowledge cutoff, retrieval
Field Machine learning, AI search, GEO

What types of training data are there?

Machine learning splits data into 3 datasets: training data, validation data, and test data. A common split is 80 percent training, 10 percent validation, and 10 percent test.

  • Training data: The model adjusts its parameters with it. The model sees these examples again and again during training.
  • Validation data: It is used for fine-tuning during development, for example when choosing the learning rate or the model size.
  • Test data: It stays untouched until the end and shows how well the model handles examples it has never seen.

For a large language model, a second classification applies. Language models go through 3 training phases, each with its own training data:

  1. Pretraining: The model reads raw text amounting to trillions of tokens and learns language and world knowledge from it. Meta states more than 15 trillion tokens for Llama 3.
  2. Fine-tuning: Curated question and answer pairs teach the model to follow instructions and to respond in a dialog.
  3. Preference data: People rate several answers to the same prompt. From these ratings, the model learns which answer counts as helpful (reinforcement learning from human feedback, RLHF for short).

For the visibility of a website, the first phase matters almost exclusively. Web text ends up in pretraining, and that is where it is decided which brands, terms, and relationships a model knows without any search.

Where does the training data of large language models come from?

The training data of large language models comes mostly from the public web, supplemented by books, Wikipedia, source code, and licensed content. The most important single source is Common Crawl, a nonprofit archive that has regularly fetched billions of web pages since 2008 and makes them freely available.

OpenAI last disclosed what the mix looks like for GPT-3 (Brown et al., 2020). Newer models usually name their sources only in broad categories.

Dataset Size Share of the training mix
Common Crawl (filtered) 410 billion tokens 60 %
WebText2 19 billion tokens 22 %
Books1 12 billion tokens 8 %
Books2 55 billion tokens 8 %
Wikipedia 3 billion tokens 3 %

The table shows two things. Web text makes up the largest share, and small, high-quality sources such as WebText2 and Wikipedia carry more weight in training than their size would suggest. A model therefore knows well-documented, frequently cited facts far more reliably than statements that appear on a single page only.

Every dataset ends on a fixed day, the knowledge cutoff. Everything published after that is missing from the model. The knowledge cutoff is usually several months before a model is released, because filtering, training, and testing take time.

What is the difference between training data and retrieval?

Training data shapes a model’s knowledge permanently before launch, while retrieval supplies current sources only at the moment a question is asked. The two paths differ in timing, freshness, source attribution, and how quickly a change on your website reaches AI answers.

Attribute Training data Retrieval
Timing before the model launches with every single request
Freshness ends at the knowledge cutoff as current as the search index
Source attribution none, the knowledge is merged into the model links to the retrieved pages
Effect of a change only with the next model generation after the next fetch, often within days
Responsible crawlers training crawlers such as GPTBot or ClaudeBot search crawlers such as OAI-SearchBot or Googlebot
Two paths by which content reaches AI answers: training data and retrievalTop, the path via training data: A training crawler collects web pages, these form a dataset up to the knowledge cutoff, and training stores the knowledge permanently in the model weights. Bottom, the path via retrieval: A question triggers a search, a search crawler supplies current pages, and the model answers with a source link. Both paths come together in the AI answer.Path 1: Training data, before the model launchesWeb pagese.g. via GPTBotDatasetends at the cutoffTrainingknowledge in the weightsPath 2: Retrieval, with every single questionQuestionthe user’s promptSearche.g. via OAI-SearchBotCurrent pageswith source linkAI answerboth paths combinedChanges to your website take effect via retrieval within days,via training data only with the next model generation.Fig. 1 · taismo
Fig. 1: Two paths into the AI answer. Training data shapes model knowledge up to the knowledge cutoff, retrieval fetches current pages with every question.

The method behind the second path is called retrieval augmented generation. The system searches for suitable documents, often through a semantic search with embeddings, and the model writes the answer on that basis. Google additionally breaks a question into many sub-searches through query fan-out before AI Overviews are generated.

In practice, both paths work together. With the knowledge from its training data, the model recognizes which terms and providers belong to a question, and it checks them against current sources through retrieval. A brand the model knows from training is more likely to be recognized and classified correctly when sources are selected.

How does your content get into training data and AI answers?

Your content gets into training data when training crawlers are allowed to fetch it and it recurs often enough on the web in the same form. It reaches AI answers through retrieval when search crawlers can access it and the page answers a question directly and in a citable way.

The major providers separate their crawlers by purpose. Your robots.txt can allow or block each of them individually:

Provider Crawler for training data Crawler for search and retrieval
OpenAI GPTBot OAI-SearchBot, ChatGPT-User
Anthropic ClaudeBot Claude-SearchBot, Claude-User
Google Google-Extended (token for Gemini) Googlebot

If you block GPTBot, you disappear from future OpenAI training data and still remain findable in ChatGPT search through OAI-SearchBot. Google-Extended only affects use for Gemini and, according to Google, has no influence on Google Search. The article on AI crawlers gives an overview of all bots.

Open crawlers are the prerequisite. Whether a model actually knows your brand afterward depends on 4 factors:

  1. Repetition: Facts that read the same on many independent pages stick during training. This includes brand mentions without a link.
  2. Consistency: If your website names three different founding years, the model learns none of them reliably.
  3. Unambiguity: A brand that is clearly described as an entity can be told apart from namesakes.
  4. Citability: A model is more likely to adopt short, self-contained definitions and figures than long derivations. A grounding page bundles such facts in one place.

For retrieval, an llms.txt that lists the most important pages of your website for AI systems also helps. This is exactly where GEO work comes in: Your brand should be anchored in model knowledge and be retrieved as a source for current questions.

Which rules apply to AI training data in the EU?

In the EU, two sets of rules govern the handling of AI training data: Copyright law with its exception for text and data mining, and the AI Act with transparency obligations for model providers.

  • Text and data mining (Section 44b of the German Copyright Act): Lawfully accessible works may be reproduced for automated analysis. Rights holders can reserve this use; for online content, the reservation must be machine-readable, for example via robots.txt.
  • AI Act (Article 53): Since August 2, 2025, providers of general-purpose AI models must have a policy to comply with copyright law and publish a sufficiently detailed summary of their training data.

For website owners, this means: You make the decision about using your content as training data technically in robots.txt. How a block affects your visibility in AI search depends on whether you exclude only training crawlers or search crawlers as well.

Frequently asked questions about training data

Can I have my website removed from models that have already been trained?
Individual websites cannot be extracted from a fully trained model after the fact. A block in robots.txt only affects future crawls and therefore later model generations.

How current is ChatGPT’s training data?
Every model behind ChatGPT has its own knowledge cutoff, which OpenAI states in the model documentation. ChatGPT answers questions about more recent events through its built-in web search, in other words through retrieval.

Does what I type into ChatGPT become training data?
For personal ChatGPT accounts, OpenAI may use conversations for training unless you turn this off in the data controls settings. According to OpenAI, business, enterprise, and API inputs are not used for training by default.

0%
Get insider knowledge first!
taismo Logo

© taismo GmbH

Address


Weißenfelder Str. 6
85551 Kirchheim near Munich, Germany