What is Batch Size?
Batch size is a hyperparameter that dictates the number of individual training examples a machine learning model processes in one chunk before pausing to update its internal mathematical weights.
How does it work?
Instead of looking at the entire dataset at once (which requires too much memory) or updating after every single example (which is chaotic), data is broken into batches. If a dataset has 1,000 images and the batch size is 100, the model processes 100 images, calculates its average error, updates its parameters, and then moves to the next batch of 100.
What are the trade-offs?
Engineers must balance several factors when choosing a batch size:
- Memory use: Larger batches require significantly more GPU memory.
- Training speed: Larger batches process faster because they take advantage of parallel computing.
- Update stability: Larger batches provide a smoother, more accurate estimate of the error direction.
- Generalization: Surprisingly, smaller batches introduce a bit of helpful "noise" that can prevent the model from getting stuck and actually improve its ability to generalize to new data.
Why does it matter?
Batch size directly impacts both the hardware cost of training and the ultimate quality of the resulting AI model. Getting it right ensures efficient use of expensive GPUs.