In machine learning one of the core-steps we’re usually required to, is to split the data into train set and test set.

The train set is usually relevant for supervised model, and the test set is usually relevant for unsupervised models, where we test the result given by the supervised model.

For the sake of the discussion, we’ll assume the proportion of the train set is 80% and the proportion of the test set is 20%. However, this ratio between the train set and the test set can vary.

We’ll re-create the same dataset we’re using.

As usual, we’ll use the set.seed() command with the value 1000, for all the randomized results, just so the same result could be re-evaluated, as needed.

This is the code for the generation of the data:

set.seed(1000)
Gender <- sample(size = 1000, replace = T, x = c("Male", "Female"))
City <- sample(size = 1000, replace = T, 
               x = c("NY", "LA", "Boston", "Chicago", "Denver", "London", "Paris", "Rome"))
Fav_Pet <- sample(size = 1000, replace = T, 
                  x = c("Cat", "Dog", "Fish", "Squirl"))

Salary_yearly <- sample(x = seq(55000, 150000), size = 1000, replace = T)
Weight_kg <- ifelse(Gender == "Male", sample(x = seq(80,120), size = 1000, replace = T), 
                    sample(x = seq(45,70), size = 1000, replace = T))  
Height_m <- ifelse(Gender == "Male", sample(x = seq(175,210), size = 1000, replace = T) / 100, 
                    sample(x = seq(155,180), size = 1000, replace = T) / 100)  

data <- data.frame(Gender, City, Fav_Pet, Height_m, Weight_kg, Salary_yearly)

This is how the data looks:

Method 1:

Randomized numbers:

These are the relevant phases for this method:

  1. Add a column with randomized numbers from 0 to 1 to the dataset.
    The numbers should have decimals digits, and shouldn’t be integers.
  2. Organize the data in ascending or descending order.
  3. Adding another column to the dataset, with a running number from 1 to n, where n is length of the dataset.
  4. Using the subset function to split the dataset, to train set for the running number lower than 80% and test set for all the rest.

This is the relevant code for that:

# Method 1:
set.seed(1000)
data$rand <- runif(n = nrow(data), min = 0, max = 1)
data <- arrange(data, by_group = data$rand)
data$running <- seq(1, nrow(data))
train_set <- subset(data, data$running<=0.8*nrow(data))
test_set <- subset(data, data$running>0.8*nrow(data))

And this is the train set result:

Train set, n=800

And this is the test-set result:

Test set, n=200

Method 2:

Using “caTools” package:

These are the relevant phases for this method:

  1. Install the package caTools() as needed, and load it.
  2. Split the data with the sample.split() command with the output TRUE or FALSE based upon the requested ratio.
  3. Use to command subset(), for creation of both the train set and the test set.

This is the relevant code:

# Method 2:
install.packages("caTools")
library(caTools)
set.seed(1000)
data$split <- sample.split(data$Gender, SplitRatio = 0.8)
train_set <- subset(data, data$split == TRUE) 
test_set <- subset(data, data$split == FALSE)

This is the train set result:

And this is the test set:

Method 3:

Using “dplyr” functions sample_frac() and anti_join():

  1. Use the command sample_frac() to create the initial sample.
  2. Use the command anti_join() to keep only the lines that aren’t included in the initial sample drawn from the data.

This is the relevant code:

# Method 3:
set.seed(1000)
training_set <- sample_frac(tbl = data, size = 0.8, replace = F)
test_set <- anti_join(data, training_set)

And this is the train set:

And this is the test set:

Of course, completely other, or quite similar solutions can be applied for this problem.