Skip to main content
100%

Gist 1dcc8a9bd27317b89fdde2c22c049100

✓ Published0🌍 Public
RRobColeman
Last edited Aug 24, 2016
Created on Aug 24, 2016

This example presents a Scala-based feature engineering pipeline for machine learning models, specifically for predicting eCPM (effective cost per mille) in advertising. The code defines a system of feature computers that transform prediction requests into sparse feature vectors using Apache Spark's MLlib library (`SparkVectors.sparse`) and the `MurmurHash3` hashing function. It shows how categorical features like app publisher and model identifiers are hashed into indexed blocks within a numeric vector, with each block's position and size tracked through `FeatureComputerParams`. The implementation uses pattern matching in `FeatureComputerRouter` to select appropriate computations and aggregates values into a map that can be converted to Spark's sparse vector format for model training.

AI-generated description

Similar vizzes