What is an Embedding?
An embedding is a way of translating text, images, or audio into an array of numbers (a vector). This numerical representation is specifically designed so that concepts with similar meanings end up mathematically close to each other, allowing computers to understand semantic relationships.
How does it work?
An embedding model processes a word (or image) and outputs a list of hundreds or thousands of numbers representing coordinates in a high-dimensional space. "Dog" and "Puppy" will have very similar coordinates. "Dog" and "Skyscraper" will be far apart.
What is it commonly confused with?
It is a misconception that embeddings represent perfect semantic proximity. Embeddings capture statistical co-occurrences from training data, meaning they can inadvertently learn biases (like associating "doctor" mathematically closer to "man"). They do not actually "understand" meaning the way a human does.
Why does it matter?
Embeddings are the backbone of modern AI memory and search. Instead of searching a database for exact keyword matches, systems can use embeddings to search for meaning. This is the critical technology that makes Retrieval-Augmented Generation (RAG) possible, allowing an AI to find a relevant document even if the user didn't type the exact right keywords.