What is Self-Supervised Learning?
Self-supervised learning is a technique that blends the autonomy of unsupervised learning with the structured feedback of supervised learning. The algorithm is given raw, unlabelled data, but it automatically generates its own labels by hiding parts of the data and attempting to predict what is missing.
How does it work?
In language models, the most common self-supervised method is "next-token prediction." The algorithm takes a sentence from the internet, hides the final word, and tries to guess it. It then uncovers the word to see if it was right, adjusting its parameters accordingly. It acts as its own supervisor.
What is a simple example?
Imagine trying to learn a language by reading a book with words randomly blacked out. You guess what the blacked-out word is based on context, and then lift the black tape to check your answer. Over thousands of pages, your understanding of the language's grammar and vocabulary improves dramatically.
Why does it matter?
Self-supervised learning is the breakthrough that enabled modern Foundation Models and LLMs. Because the model labels the data itself, researchers can train models on petabytes of raw internet data without needing an army of humans to annotate it. This unlocked a massive leap in AI capabilities.