|
|
The decision tree is called that because it creates models for classification or prediction in the shape of a tree. It splits the data into smaller parts and connects each part with a decision. This makes a tree with decision points and final answers. A decision point can have two or more paths and leads to the final answers. A final answer shows the result of the classification or decision. It uses the if-then-else rule to make predictions. As you go deeper into the tree, the rules get more complicated, which makes the model more accurate. A decision tree has:
Root Node: The process starts at the top of the tree with the entire dataset.
Internal Nodes (Decision Nodes): At each internal node, the algorithm tests a specific attribute or feature to split the data into smaller, more homogeneous subsets.
Leaf Nodes (Terminal Nodes): The process ends at the leaf nodes, which represent the final decision, class label, or predicted continuous value.

Complete House Price Prediction Example
10 houses • 3 independent variables • MSE-based splitting
| House | Size (sq ft) | Bedrooms | Age (years) | Price ($000s) |
|---|---|---|---|---|
| 1 | 1,000 | 2 | 30 | 220 |
| 2 | 1,200 | 2 | 20 | 250 |
| 3 | 1,400 | 3 | 15 | 290 |
| 4 | 1,500 | 3 | 10 | 320 |
| 5 | 1,600 | 3 | 8 | 350 |
| 6 | 1,800 | 3 | 5 | 300 |
| 7 | 2,000 | 4 | 10 | 390 |
| 8 | 2,200 | 4 | 5 | 430 |
| 9 | 2,400 | 4 | 3 | 470 |
| 10 | 2,600 | 5 | 2 | 510 |
The tree wants each resulting group to contain prices that are as close as possible to that group's mean.
For a candidate split:
Initially all 10 houses are together.
So the root predicts $353,000.
| House | Y | Mean | Error | Error² |
|---|---|---|---|---|
| 1 | 220 | 353 | -133 | 17,689 |
| 2 | 250 | 353 | -103 | 10,609 |
| 3 | 290 | 353 | -63 | 3,969 |
| 4 | 320 | 353 | -33 | 1,089 |
| 5 | 350 | 353 | -3 | 9 |
| 6 | 300 | 353 | -53 | 2,809 |
| 7 | 390 | 353 | 37 | 1,369 |
| 8 | 430 | 353 | 77 | 5,929 |
| 9 | 470 | 353 | 117 | 13,689 |
| 10 | 510 | 353 | 157 | 24,649 |
| Candidate split | Weighted MSE | |
|---|---|---|
| Size ≤ 1,100 | 6,956.76 | |
| Size ≤ 1,300 | 5,044.44 | |
| Size ≤ 1,450 | 4,085.71 | |
| Size ≤ 1,550 | 3,530.00 | |
| Size ≤ 1,750 | 3,692.00 | |
| Size ≤ 1,900 | 4,420.00 | |
| Size ≤ 2,100 | 4,260.00 | |
| Size ≤ 2,300 | 3,652.22 | |
| Size ≤ 2,500 | 3,278.89 ★ LOWEST |
The first important lesson is that the tree does not simply choose the first split that improves MSE. It searches the candidate thresholds and selects the lowest one.
| Feature | Best threshold | Lowest weighted MSE |
|---|---|---|
| Size | ≤ 2,500 sq ft | 3,278.89 |
| Bedrooms | ≤ 3 | 4,050.00 |
| Age | ≤ 4 years | 5,120.00 |
Thus this tree predicts approximately $335,560 for Houses 1–9 and $510,000 for House 10.
If this tree is being used inside gradient boosting, the next trees learn from the errors left by the previous model.
Decision Tree Regression • Housing Price Worked Example • MSE-Based Splitting
import pandas as pd from sklearn.tree import DecisionTreeRegressor from sklearn.model_selection import train_test_split from sklearn.metrics import mean_absolute_error, r2_score # ----------------------------------- # 1. Load data from Excel # ----------------------------------- data = pd.read_excel("product_data.xlsx") print("Dataset Preview:") print(data.head()) # ----------------------------------- # 2. Define features and target # ----------------------------------- X = data[['Production_Cost', 'Advertising_Spend', 'Demand_Level']] y = data['Product_Price'] # ----------------------------------- # 3. Split into training and testing # ----------------------------------- X_train, X_test, y_train, y_test = train_test_split( X, y, test_size=0.25, random_state=42 ) # ----------------------------------- # 4. Train the Decision Tree model # ----------------------------------- model = DecisionTreeRegressor( max_depth=4, random_state=42 ) ## What is random_state? #train_test_split randomly shuffles the dataset before splitting. #Without random_state: #Each run → different split #Model performance changes slightly #With random_state=42: #Same rows go to train/test every time #Results are reproducible model.fit(X_train, y_train) # ----------------------------------- # 6. Evaluate the model # ----------------------------------- y_pred = model.predict(X_test) print("MAE:", mean_absolute_error(y_test, y_pred)) print("R² score:", r2_score(y_test, y_pred)) # ----------------------------------- # 7. Predict price for a new product # ----------------------------------- new_product = pd.DataFrame({ 'Production_Cost': [68], 'Advertising_Spend': [13], 'Demand_Level': [37] }) predicted_price = model.predict(new_product) print("\nPredicted Product Price:", predicted_price[0])
USE CASE 2: Using Decision Tree with scikit-learn to predict the Student Grade. The 'Hours_Studied, 'Attendance_%', 'Previous_Score' are the independent variables.
import numpy as np from sklearn.model_selection import train_test_split from sklearn.tree import DecisionTreeRegressor from sklearn.metrics import mean_absolute_error, r2_score # ----------------------------------- # 1. Load data from Excel # ----------------------------------- #sample data can be exported to #excel from the URL # https://pythonPlaza.com/linear_school_grade_data.html data = pd.read_excel("student_data.xlsx") print("Dataset Preview:") print(data.head()) # ----------------------------------- # 2. Define features and target # ----------------------------------- X = data[['Hours_Studied', 'Attendance_%', 'Previous_Score']] y = data['Final_Grade'] # ----------------------------------- # 3. Split into training and testing # ----------------------------------- X_train, X_test, y_train, y_test = train_test_split( X, y, test_size=0.25, random_state=42 ) # ----------------------------------- # 4. Train the Decision Tree model # ----------------------------------- model = DecisionTreeRegressor( max_depth=4, random_state=42 ) ## What is random_state? #train_test_split randomly shuffles the dataset before splitting. #Without random_state: #Each run → different split #Model performance changes slightly #With random_state=42: #Same rows go to train/test every time #Results are reproducible model.fit(X_train, y_train) # ----------------------------------- # 6. Evaluate the model # ----------------------------------- y_pred = model.predict(X_test) print("MAE:", mean_absolute_error(y_test, y_pred)) print("R² score:", r2_score(y_test, y_pred)) Example: Predict a new student’s grade # New student: [hours_studied, attendance %, previous_score] new_student = np.array([[6, 85, 78]]) predicted_grade = model.predict(new_student) print("Predicted final grade:", predicted_grade[0])
USE CASE 3: Using Decision Tree with scikit-learn to predict the Profit Optimization. The Price (P), Advertising (A), Units Sold (Q) are the independent variables, and Profit is the dependent variable.
import numpy as np from sklearn.model_selection import train_test_split from sklearn.tree import DecisionTreeRegressor from sklearn.metrics import mean_absolute_error, r2_score # ----------------------------------- # 1. Load data from Excel # ----------------------------------- #sample data can be exported to #excel from the URL Get the Profit Optimization data in Excel data = pd.read_excel("profit_optimization.xlsx") print("Dataset Preview:") print(data.head()) # ----------------------------------- # 2. Define features and target Price (P) # ----------------------------------- X = data[['Price', 'Advertising', 'Units_Sold']] y = data['Profit'] # ----------------------------------- # 3. Split into training and testing # ----------------------------------- X_train, X_test, y_train, y_test = train_test_split( X, y, test_size=0.25, random_state=42 ) # ----------------------------------- # 4. Train the Decision Tree model # ----------------------------------- model = DecisionTreeRegressor( max_depth=4, random_state=42 ) ## What is random_state? #train_test_split randomly shuffles the dataset before splitting. #Without random_state: #Each run → different split #Model performance changes slightly #With random_state=42: #Same rows go to train/test every time #Results are reproducible model.fit(X_train, y_train) # ----------------------------------- # 6. Evaluate the model # ----------------------------------- y_pred = model.predict(X_test) print("MAE:", mean_absolute_error(y_test, y_pred)) print("R² score:", r2_score(y_test, y_pred)) #Predict profit for a new business strategy # Example: Price = 15, Advertising = 165, Units Sold = 460 new_strategy = np.array([[15, 165, 460]]) predicted_profit = model.predict(new_strategy) print("Predicted profit:", predicted_profit[0])
USE CASE 4: Using Decision Tree with scikit-learn to predict the Patient Response. The Dosage (mg), Age (yrs), Weight (lbs) are the independent variables, and Patient Response is the dependent variable.
import numpy as np from sklearn.model_selection import train_test_split from sklearn.tree import DecisionTreeRegressor from sklearn.metrics import mean_absolute_error, r2_score # ----------------------------------- # 1. Load data from Excel # ----------------------------------- #sample data can be exported to #excel from the URL Get the Patient Response Data in Excel data = pd.read_excel("patient_dosage_response.xlsx") print("Dataset Preview:") print(data.head()) # ----------------------------------- # 2. Define features and target Price (P) # ----------------------------------- X = data[['Dosage', 'Age', 'Weight']] y = data['Patient_Response'] # ----------------------------------- # 3. Split into training and testing # ----------------------------------- X_train, X_test, y_train, y_test = train_test_split( X, y, test_size=0.25, random_state=42 ) # ----------------------------------- # 4. Train the Decision Tree model # ----------------------------------- model = DecisionTreeRegressor( max_depth=4, random_state=42 ) ## What is random_state? #train_test_split randomly shuffles the dataset before splitting. #Without random_state: #Each run → different split #Model performance changes slightly #With random_state=42: #Same rows go to train/test every time #Results are reproducible model.fit(X_train, y_train) # ----------------------------------- # 6. Evaluate the model # ----------------------------------- y_pred = model.predict(X_test) print("MAE:", mean_absolute_error(y_test, y_pred)) print("R² score:", r2_score(y_test, y_pred)) Predict response for a new patient # New patient: Dosage=72mg, Age=36yrs, Weight=172lbs new_patient = np.array([[72, 36, 172]]) predicted_response = model.predict(new_patient) print("Predicted patient response:", predicted_response[0])