Ask an ordinary database for documents containing the word “automobile” and it can find “automobile.” Ask a modern AI search system for “cheap cars suitable for a family,” and it may also find documents discussing affordable SUVs, hatchbacks and used vehicles without those exact words appearing anywhere.
Somewhere between those two searches, language has been turned into geometry.
The trick is called an embedding, and it quietly powers search engines, recommendation systems, RAG applications and a sizeable portion of the AI infrastructure currently acquiring expensive logos.
Meaning becomes a list of numbers
An embedding is a numerical representation of something: a word, sentence, document, image, product or even a user.
Instead of storing “Golden retriever” simply as text, an embedding model converts it into a vector — essentially a long list of numbers such as:
[0.18, -0.73, 0.24 ...]
Real embeddings may contain hundreds or thousands of dimensions. Humans cannot look at those numbers and meaningfully announce, “Ah yes, definitely a dog.” Computers can.
During training, the model learns to arrange related things relatively close together in this mathematical space. Google describes embeddings as dense representations that capture meaningful relationships while reducing unwieldy high-dimensional data into something more useful for machine learning.
Imagine an absurdly complicated map. Delhi and Mumbai might sit near other Indian cities. “Puppy” might be near “dog.” A paragraph about mortgage rates could sit near another discussing home-loan interest even when their wording differs.
That ability to compare meaning instead of exact text is the important bit.
How does AI decide what is “near”?
Once two pieces of information have been converted into vectors, software can mathematically compare them.
Common measures include cosine similarity, dot products and Euclidean distance. The details differ, but the basic question is wonderfully straightforward: how close are these two points?
Google's recommendation-system documentation describes essentially the same process: compute an embedding for a query, then find nearby item embeddings. At enormous scale, checking every possible vector becomes expensive, so systems use techniques such as approximate nearest-neighbour search.
That is why YouTube can recommend another video, an ecommerce site can surface related products and an AI application can retrieve a paragraph relevant to your question.
Different products. Same mathematical party trick.
This is also what makes RAG work
Suppose a company has 100,000 internal documents and wants an AI assistant to answer questions from them.
Stuffing all 100,000 into every prompt would be expensive, slow and spectacularly unnecessary.
Instead, a RAG — retrieval-augmented generation — system can split the documents into chunks and create embeddings for them. When somebody asks a question, the question itself becomes an embedding. The system searches for the nearest document chunks and gives only those relevant pieces to the language model.
Vector search has consequently become important infrastructure for modern AI systems. Google explicitly lists semantic search, recommendations, RAG and AI agents among its applications.
But an embedding is not understanding in a bottle
There is an important limitation.
Two things being mathematically similar does not make them factually equivalent. An embedding model may retrieve something topically relevant but wrong, outdated or subtly different.
Embeddings can also inherit biases from their training data, perform differently across languages and struggle when terminology from a specialist domain means something very different from ordinary usage.
And more dimensions do not automatically mean better embeddings, in roughly the same way that adding more drawers does not guarantee somebody has organised the kitchen.
Models, datasets, chunk sizes, similarity thresholds and retrieval strategies all matter.
That makes embeddings less glamorous than the AI model eventually answering the question. But they are frequently the reason the model was able to find the right information in the first place.
Generative AI gets to write the sentence.
Embeddings often quietly decide what it gets to read before writing it.