Post

Understanding GridSearchCV

Understanding GridSearchCV

When building a machine learning model, one common question is:

Which algorithm should I choose, and what settings should I use?

For example:

  • Should I use Linear Regression?
  • Should I use Lasso Regression with alpha=1 or alpha=2?
  • Should a Decision Tree use best splitting or random splitting?

Instead of manually trying every combination, GridSearchCV automates the entire process.

Full Code

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
from sklearn.model_selection import GridSearchCV, ShuffleSplit
from sklearn.linear_model import LinearRegression, Lasso
from sklearn.tree import DecisionTreeRegressor


def find_best_model_using_gridsearchcv(X, y):

    algos = {

        'linear_regression': {
            'model': LinearRegression(),
            'params': {}
        },

        'lasso': {
            'model': Lasso(),
            'params': {
                'alpha': [1, 2],
                'selection': ['random', 'cyclic']
            }
        },

        'decision_tree': {
            'model': DecisionTreeRegressor(),
            'params': {
                'criterion': ['squared_error', 'friedman_mse'],
                'splitter': ['best', 'random']
            }
        }

    }

    scores = []

    cv = ShuffleSplit(
        n_splits=5,
        test_size=0.2,
        random_state=0
    )

    for algo_name, config in algos.items():

        gs = GridSearchCV(
            estimator=config['model'],
            param_grid=config['params'],
            cv=cv,
            return_train_score=False
        )

        gs.fit(X, y)

        scores.append({
            'model': algo_name,
            'best_score': gs.best_score_,
            'best_params': gs.best_params_
        })

    return pd.DataFrame(scores, columns=['model', 'best_score', 'best_params'])

# run
find_best_model_using_gridsearchcv(X, y)

Step 1: Define Candidate Models

We create a dictionary containing all models and their hyperparameters.

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
algos = {

    'linear_regression': {
        'model': LinearRegression(),
        'params': {}
    },

    'lasso': {
        'model': Lasso(),
        'params': {
            'alpha': [1, 2],
            'selection': ['random', 'cyclic']
        }
    },

    'decision_tree': {
        'model': DecisionTreeRegressor(),
        'params': {
            'criterion': ['squared_error', 'friedman_mse'],
            'splitter': ['best', 'random']
        }
    }

}

Think of this as creating a menu of experiments.

ModelHyperparameters to Test
Linear RegressionNone
Lassoalpha = 1,2
Lassoselection = random, cyclic
Decision Treecriterion = squared_error, friedman_mse
Decision Treesplitter = best, random

Grid Search simply means:

Try every possible combination of hyperparameters.

For Lasso:

1
2
'alpha': [1,2]
'selection':['random','cyclic']

GridSearchCV generates:

Experimentalphaselection
11random
21cyclic
32random
42cyclic

So Lasso gives us 4 models.

Similarly for Decision Tree:

1
2
3
4
5
6
7
8
9
10
11
criterion:
[
'squared_error',
'friedman_mse'
]

splitter:
[
'best',
'random'
]

Possible combinations:

Experimentcriterionsplitter
1squared_errorbest
2squared_errorrandom
3friedman_msebest
4friedman_mserandom

Again 4 models.

Linear Regression has no parameters.

So:

  • Linear Regression → 1 model
  • Lasso → 4 models
  • Decision Tree → 4 models

Total experiments:

1
1 + 4 + 4 = 9 models

Step 3: What is Cross Validation?

1
2
3
4
5
cv = ShuffleSplit(
    n_splits=5,
    test_size=0.2,
    random_state=0
)

This means:

  • Create 5 different train-test splits
  • 80% training data
  • 20% testing data
  • Randomly shuffle before splitting

Suppose we have 1000 samples.

Split 1:

1
2
Train : 800
Test  : 200

Split 2:

1
2
Different random 800
Different random 200

Split 3:

1
Another random split

And so on.

Total:

1
5 train-test experiments

Step 4: How GridSearchCV Actually Works

Suppose we are evaluating Lasso.

GridSearchCV starts with:

1
2
alpha=1
selection='random'

Then:

Split 1

Train on 80%

Test on 20%

Score = 0.78

Split 2

Train on another 80%

Test on another 20%

Score = 0.80

Split 3

Score = 0.77

Split 4

Score = 0.79

Split 5

Score = 0.81

Average:

1
2
3
(0.78+0.80+0.77+0.79+0.81)/5

= 0.79

GridSearchCV stores:

1
2
3
4
alpha=1
selection=random

mean score = 0.79

Then it moves to:

1
2
alpha=1
selection='cyclic'

Again performs 5 splits.

Then:

1
2
alpha=2
selection='random'

Again 5 splits.

Then:

1
2
alpha=2
selection='cyclic'

Again 5 splits.

Finally it compares all averages.

alphaselectionMean CV Score
1random0.79
1cyclic0.81
2random0.76
2cyclic0.83

Best:

1
2
3
alpha=2
selection=cyclic
score=0.83

Step 5: Same Process for Decision Tree

GridSearchCV tries:

1
2
3
4
5
6
7
squared_error + best

squared_error + random

friedman_mse + best

friedman_mse + random

Each combination is trained:

1
5 times

because:

1
n_splits = 5

Then average performance is calculated.

Step 6: Understanding the Loop

Next, iterate through each algorithm.

1
for algo_name, config in algos.items():

First iteration:

1
Linear Regression

GridSearchCV runs.

Finds best score.

Stores result.

Second iteration:

1
Lasso

GridSearchCV evaluates:

1
2
3
4
5
6
4 parameter combinations
×
5 CV splits

=
20 model trainings

Third iteration:

1
Decision Tree

Again:

1
2
3
4
5
6
4 combinations
×
5 splits

=
20 trainings

Step 7: Fitting GridSearchCV

1
gs.fit(X, y)

This single line performs:

1
2
3
4
5
6
7
8
9
10
11
12
13
Generate parameter combinations
↓
Perform Cross Validation
↓
Train models
↓
Evaluate scores
↓
Average scores
↓
Choose best combination
↓
Store results

All automatically.

Step 8: Best Parameters

After training:

1
gs.best_score_

might return:

1
0.8477

And:

1
gs.best_params_

might return:

1
{'alpha':2,'selection':'cyclic'}

These values correspond to the model configuration that achieved the highest average cross-validation score.

Step 9: Final Result

The function returns:

1
pd.DataFrame(scores)

Output:

modelbest_scorebest_params
linear_regression0.8478{}
lasso0.7268{‘alpha’:2,’selection’:’random’}
decision_tree0.7190{‘criterion’:’squared_error’,’splitter’:’best’}

Visualizing the Entire Workflow

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
Define Models
      │
      ▼

Generate Hyperparameter Combinations
      │
      ▼

For each combination
      │
      ▼

Perform 5 ShuffleSplit validations
      │
      ▼

Train Model
      │
      ▼

Evaluate Score
      │
      ▼

Average Scores
      │
      ▼

Pick Best Parameters
      │
      ▼

Compare Algorithms
      │
      ▼

Return Best Model Information

GridSearchCV is an automated system that tests every hyperparameter combination, evaluates each one using cross-validation, calculates average performance, and selects the configuration that generalizes best on unseen data.

This post is licensed under CC BY 4.0 by the author.