Skip to main content
100%

tut17

✓ Published0🌍 Public
BBenHeubl
Last edited Feb 13, 2020
Created on Feb 13, 2020

This example builds a classification-ready dataset from income survey data, showing how raw income brackets are recoded into binary “Under 30k” versus “Over 30k” categories. It then splits the cleaned data into training and test sets using base R’s `sample` function. The code uses `dplyr`’s `mutate`, `if_else`, and `mutate_if` for transformations, along with `forcats::fct_explicit_na` to label missing factors as “Unknown.” The data is sourced from a remote CSV via `read.csv` and processed entirely in R, with no visualization rendering or plotting involved.

AI-generated description

Similar vizzes