What is a Small Language Model?
A Small Language Model (SLM) is a text-generating AI model designed to be highly efficient and require significantly less computational power and memory than its larger counterparts. The definition of "small" is relative and shifts as hardware improves, but generally refers to models that can run locally on consumer hardware like laptops or smartphones.
How does it work?
SLMs use the same underlying architectures (like the Transformer) as Large Language Models, but they are built with far fewer parameters. To make up for having less capacity to store information, developers often train SLMs on highly curated, high-quality "textbook" data, rather than the raw, unfiltered internet.
What is a simple example?
If a massive cloud-based LLM is a supercomputer, an SLM is a smartphone. The supercomputer can simulate global weather patterns, but you cannot carry it in your pocket. An SLM can't answer obscure trivia questions as reliably, but it can quickly summarize an email directly on your phone without needing an internet connection.
Why does it matter?
Massive language models are expensive to host, slow to respond, and require sending sensitive user data to cloud servers. Small language models solve these issues by allowing AI features to run locally, cheaply, and privately on the edge.