What is natural language processing?
Natural language processing (NLP) is a field of computer science and artificial intelligence that teaches computers to process, understand, and generate human language in written and spoken form. Natural language processing combines insights from linguistics with statistics and machine learning. Search engines, translation services, voice assistants, and large language models such as ChatGPT are all built on natural language processing.

Term profile
| Attribute | Details |
|---|---|
| Part of speech | Noun, technical term |
| Abbreviation | NLP |
| Pronunciation | ˈnætʃərəl ˈlæŋɡwɪdʒ ˈprɑːsɛsɪŋ |
| Subfields | Natural language understanding (NLU), natural language generation (NLG) |
| Related terms | Computational linguistics, large language model, embedding, named entity recognition, semantic search |
| Other meaning | In psychology, NLP stands for neuro-linguistic programming, a communication model with no connection to computer science |
How does natural language processing work?
Natural language processing breaks language down into units a computer can calculate with and then puts them back together into a meaning. Classic NLP systems work through a chain of subtasks to do this. Modern language models learn many of these steps together, but the chain still explains what happens in terms of content.
- Tokenization: The system splits a sentence into tokens, meaning words, parts of words, and punctuation marks. “Where can I buy bread in Munich?” becomes eight units.
- Parts of speech and sentence structure: Part-of-speech tagging determines the word class of each token. Dependency parsing identifies which word depends on which, for example that “bread” is the object of “buy”.
- Entity recognition: Named entity recognition marks the names of places, people, companies, and products. “Munich” is recognized as a place and linked to the known thing called Munich.
- Meaning in context: The system resolves ambiguous words through their surroundings. “Bank” means something different next to “account” than next to “river”. Technically this usually relies on an embedding that represents words and sentences as vectors of numbers.
- Performing the task: On this basis the system carries out the actual task, such as recognizing an intent, translating a text, rating a sentiment, or writing an answer.
In the example sentence, this chain leads to a clear reading: The user is looking for a place in Munich where they can buy bread, in other words a bakery nearby. In search, this translation of words into an intention is called search intent.
Which tasks does natural language processing handle?
Natural language processing covers 6 major task areas that are combined in almost every application.
- Text classification: A system assigns texts to categories, for example spam or not spam, topic A or topic B.
- Sentiment analysis: Sentiment analysis evaluates whether a statement is meant positively, negatively, or neutrally. Companies use it to analyze reviews and support requests.
- Information extraction: The system pulls facts out of running text, such as names, dates, prices, and the relationships between them.
- Machine translation: Services such as DeepL and Google Translate transfer texts between languages.
- Question answering: The system finds the passage that matches a question or writes the answer itself.
- Text generation: Language models write summaries, emails, code, and entire articles.
Many of these tasks run unnoticed in everyday life. Examples are autocorrect on your smartphone, the spam filter in your inbox, video subtitles, and the suggestions in the search box.
How has natural language processing developed?
Natural language processing has developed in 4 phases, from fixed rules through statistics and neural networks to today’s transformer models.
| Phase | Period | Characteristics and milestones |
|---|---|---|
| Rule-based | 1950s to 1980s | Linguists write grammar rules by hand. In 1954 the Georgetown-IBM experiment translates around 60 Russian sentences into English, and in 1966 Joseph Weizenbaum’s program ELIZA simulates a conversation. |
| Statistical | from the late 1980s | Systems learn probabilities from large collections of text instead of rules. Statistical machine translation shapes the 2000s. |
| Neural | 2010s | Neural networks learn word meanings as vectors. In 2013 Google’s Word2Vec shows that meaning can be represented in numbers. |
| Transformer | since 2017 | Google researchers introduce the transformer architecture in 2017. Google BERT (2018) and OpenAI’s GPT models build on it, and ChatGPT launches in November 2022. |
The leap of the last phase lies in context. For every word, a transformer weighs which other words in the sentence matter for its meaning. This gave rise to today’s large language models, which no longer run natural language processing as individual steps but as one single large model.
What is the difference between NLP, NLU, and NLG?
NLP is the umbrella term, and NLU and NLG are its two directions. Natural language understanding (NLU) describes the path from language to meaning: A system understands what a sentence means. Natural language generation (NLG) describes the opposite path from meaning to language: A system writes a text itself.
| Term | Direction | Example |
|---|---|---|
| Natural language processing (NLP) | Umbrella term for any machine processing of language | Search engine, translator, chatbot |
| Natural language understanding (NLU) | Language to meaning | A search engine recognizes the intent behind a query |
| Natural language generation (NLG) | Meaning to language | An AI system writes an answer from several sources |
Computational linguistics is the scientific discipline behind it. It studies language with the methods of computer science, while natural language processing tends to describe the practical application.
How does Google use natural language processing?
Google has been using natural language processing for more than ten years to match search queries and web pages by meaning instead of by individual words. The development happened in stages, each one a separate system:
- Hummingbird (2013): The update put the meaning of the whole query at the center instead of weighing every word on its own.
- RankBrain (2015): RankBrain was the first deep learning system in Google Search and helped above all with queries Google had never seen before.
- Google BERT (2019): BERT reads a word in the context of the words before and after it and thereby understands small words such as “for” or “without” that change the meaning of a query.
- MUM (2021): The Multitask Unified Model processes text and images and transfers knowledge between languages.
- Gemini (since 2023): Models from the Gemini family now write the answers in AI Overviews and in AI Mode in Google Search.
Together, these systems form the basis of semantic search. The Natural Language API in Google Cloud offers a glimpse of how Google reads texts by machine. For every text it shows the recognized entities, a weighting called salience between 0 and 1, and the sentiment of the text. The API is a separate product and not a window into ranking, but it is a good way to check whether a text makes its main topic clear to a machine.
What does natural language processing mean for SEO and AI visibility?
Natural language processing rewards texts that cover a topic clearly, unambiguously, and in natural language. Repeated keywords give a language model no additional information and tend to disrupt the flow of reading. The way NLP works leads to 5 rules for content:
- Name entities unambiguously: Companies, products, places, and people should always have the same name. The more clearly an entity is described, the more reliably a system links it to its knowledge, for example to the Knowledge Graph.
- Context instead of repetition: A text about bread shows its depth through terms such as sourdough, flour type, and baking time, not by repeating “buy bread” ten times.
- Answer questions directly: If the answer is in the first sentence of a section, a system recognizes it as the answer and can extract it.
- Keep sections self-contained: AI systems break pages into chunks of text and evaluate each chunk on its own. A section that makes sense without the rest of the page is more likely to be used. The guide on getting cited by AI shows what this looks like in detail.
- Mark up meaning as well: Structured data tells a machine directly what a text describes and saves it the guesswork.
For AI visibility this counts twice. Large language models are themselves applied natural language processing: They learn from training data which brand stands for what, and they read web pages by the same rules when they fetch them live. Our page on answer engine optimization (AEO, also known as GEO) and AI visibility shows how content can be prepared specifically for these systems.
A common misconception is so-called NLP keywords. Some tools list terms that frequently appear together in well-ranking texts and sell them under this name. There is no official list of such terms from Google. The lists help you find subtopics, but they are not a requirement that a text has to work through.
Frequently asked questions about natural language processing
Is natural language processing the same as artificial intelligence?
No. Natural language processing is a field of artificial intelligence that deals specifically with language. Other fields of AI deal with images, robotics, or planning, for example.
Is ChatGPT an example of natural language processing?
Yes. ChatGPT is based on a large language model and combines both directions of natural language processing: It understands the input (NLU) and writes an answer (NLG).
What does natural language processing have to do with neuro-linguistic programming?
Nothing apart from the abbreviation. Neuro-linguistic programming is a communication and coaching model from 1970s psychology. Natural language processing is a field of computer science.
Welt der SEO lernen?