- Cyber Success
- September 28, 2026
- IT Courses
Overfitting vs Underfitting in ML: What’s the Difference?
Almost every machine learning model that disappoints in production fails in one of two ways. It either learned too little, or it learned too much of the wrong thing. These two failures, underfitting and overfitting, are the reason a model can look excellent on your laptop and fall apart on real data. Understanding them, and being able to diagnose which one you’re facing, is one of the most valuable practical skills a beginner data scientist can build. It’s also one of the most frequently asked interview topics.
The Core Idea: Generalization
The goal of any supervised model isn’t to perform well on the data it was trained on. It’s to perform well on new data it has never seen. That ability is called generalization. Overfitting and underfitting are the two ways a model fails to generalize, sitting at opposite ends of a spectrum.
What Is Underfitting?
Underfitting happens when a model is too simple to capture the real pattern in the data. It performs poorly on the training data and poorly on new data. The model isn’t wrong about unseen examples specifically. It’s consistently wrong about everything.
A classic example is predicting house prices using only the number of rooms. Price depends on size, location, age, and more, so a model that looks at one weak feature can’t learn the real relationship. Another way to picture it is fitting a straight line through data that clearly follows a curve.
In statistical terms, underfitting is a high-bias problem. The model makes rigid, oversimplified assumptions and holds to them no matter what the data says.
What Is Overfitting?
Overfitting happens when a model learns the training data too well, including its noise and quirks, and then fails on new data. Training performance looks superb while performance on unseen data drops sharply.
The polynomial example makes this concrete: a highly complex curve that passes through every single training point of a house-price dataset achieves near-zero training error, but it has memorized the noise specific to that sample and predicts new houses badly.
In statistical terms, overfitting is a high-variance problem. The model is so flexible that it reacts to random fluctuations in the training set as though they were real signal, so its predictions change dramatically with different training samples.
Side-by-Side Comparison
Aspect | Underfitting | Overfitting |
Model complexity | Too simple | Too complex |
Training performance | Poor | Excellent |
Performance on new data | Poor | Poor (much worse than on training data) |
Bias/variance | High bias, low variance | Low bias, high variance |
Analogy | A student who only learned addition and can’t handle the exam | A student who memorized last year’s answers and can’t handle new questions |
Typical cause | Weak features, over-simple model, too much regularization | Too many parameters, too little data, noisy data, training too long |
The Bias-Variance Tradeoff
The tradeoff ties both failures together. Bias is the error from oversimplified assumptions. Variance is the error from sensitivity to small changes in the training data. Reducing one tends to increase the other, and you can’t eliminate both, only balance them.
Picture total error against model complexity as a U-shaped curve. Bias falls as the model gains capacity to capture real patterns, while variance rises as the model becomes sensitive to its particular training sample. The lowest point of the U is the sweet spot. To the left of it the model underfits, and to the right it overfits. Every diagnostic and fix below is about finding and staying near that minimum.
How to Tell Which One You Have
The simplest diagnostic is to compare error on the training set with error on a held-out validation set.
- Both errors high and close together:
- Training error low, validation error much higher:
- Both low and close together: a good fit.
A learning curve makes this visible. In overfitting, training error keeps falling while validation error stops improving and starts rising. In underfitting, both curves plateau at a poor level.
A useful reality check: a suspiciously perfect result deserves suspicion before celebration. On a churn dataset where 99% of customers don’t churn, a model that predicts “no churn” for everyone scores 99% accuracy while learning nothing. Always check the base rate, the gap between training and validation performance, and whether data leakage is sneaking future information into your features.
How to Fix Underfitting
Underfitting means the model needs more capacity or better information.
- Use a more expressive model, such as moving from linear regression to a decision tree or ensemble.
- Add or engineer more informative features.
- Reduce regularization if it’s too strong.
- Train longer, if training was cut short.
How to Fix Overfitting
Overfitting means the model needs constraints or more evidence.
- Get more training data: the most reliable fix, since more examples make noise harder to memorize.
- Regularization: L1 and L2 penalties discourage large weights and keep the model simpler.
- Early stopping: halt training when validation performance stops improving.
- Cross-validation: use it to get a more trustworthy estimate of how the model generalizes.
- Simplify the model: fewer features, shallower trees, or pruning.
- Dropout: for neural networks, randomly dropping units to discourage memorization.
- Ensembling: methods like random forests average many models to reduce variance.
The key discipline is to diagnose before you fix. Adding complexity to an overfit model, or constraining an underfit one, makes things worse.
A Short Code Illustration
from sklearn.tree import DecisionTreeRegressor
from sklearn.model_selection import train_test_split
from sklearn.metrics import mean_squared_error
X_train, X_val, y_train, y_val = train_test_split(X, y, test_size=0.2, random_state=42)
for depth in [1, 5, None]:
model = DecisionTreeRegressor(max_depth=depth, random_state=42)
model.fit(X_train, y_train)
train_err = mean_squared_error(y_train, model.predict(X_train))
val_err = mean_squared_error(y_val, model.predict(X_val))
print(f”depth={depth}: train={train_err:.2f}, validation={val_err:.2f}”)
Running this on your own data typically shows the pattern: a very shallow tree has high error on both sets (underfitting), while an unrestricted tree has near-zero training error but a noticeably higher validation error (overfitting). The depth in between usually gives the best validation score.
Common Interview Questions
- How do you detect overfitting? Compare training and validation performance. A large gap with strong training results points to overfitting.
- What is the bias-variance tradeoff? Simpler models have high bias and low variance; complex models have low bias and high variance. You aim to minimize total error.
- Name three ways to reduce overfitting. More data, regularization, and early stopping or cross-validation are the standard answers.
- Can a model with 99% accuracy still be bad? Yes, if the classes are imbalanced or there’s data leakage, so check the base rate and the validation gap.
Final Word
Overfitting and underfitting are two sides of one balancing act. A model that’s too simple misses the pattern; a model that’s too flexible memorizes the noise. The practical skill is to check the gap between training and validation performance, name the problem correctly, and apply the fix that matches it. That habit is what separates models that work in a notebook from models that work in the real world.
Cyber Success’s Data Science course in Pune builds this understanding through hands-on modelling projects, where you diagnose and fix fitting problems on realistic datasets, with placement support to help you move into your first data science role. Explore our Data Science course to build a practical machine learning foundation.
Frequently Asked Questions
What is the main difference between overfitting and underfitting?
Underfitting is when a model is too simple and performs poorly on both training and new data. Overfitting is when a model is too complex, performing very well on training data but poorly on new data.
Is high training accuracy always a good sign?
No. High training accuracy with much lower validation accuracy indicates overfitting. Always evaluate on data the model hasn’t seen.
How does the bias-variance tradeoff relate to these two problems?
Underfitting corresponds to high bias and overfitting to high variance. Reducing one usually raises the other, so the goal is to find the complexity that minimizes total error.
Does more data fix underfitting?
Usually not on its own. An underfit model is too simple to learn the pattern even with more data, so it typically needs more capacity or better features. More data is a much stronger remedy for overfitting.
What is the easiest way to spot overfitting as a beginner?
Split your data into training and validation sets and compare scores. If training performance is far better than validation performance, you’re likely overfitting.
