What is a Hyperparameter?
A hyperparameter is a high-level configuration setting established by the developer before the training process begins. It controls how the machine learning algorithm learns and shapes the overall architecture of the model.
How does it work?
Think of hyperparameters as the control dials on a complex machine. Examples include:
- Learning rate: How aggressively the model updates its knowledge.
- Batch size: How many examples it looks at before adjusting itself.
- Number of epochs: How many times it reviews the entire dataset.
- Model architecture: How many layers of neurons the neural network has.
What is it commonly confused with?
Clearly distinguish hyperparameters from standard parameters (like weights and biases).
- Parameters are learned automatically by the AI during training.
- Hyperparameters are set manually by the human engineer before training starts.
Why does it matter?
Choosing the right hyperparameters is crucial for success. If set incorrectly, the model might take weeks to train instead of hours, memorize the data instead of learning patterns (overfitting), or completely fail to learn anything useful at all.