Skip to main content
100%

Helpers for TDunnings Java TDigest library

✓ Published0🌍 Public
RRobColeman
Last edited Jan 28, 2016
Created on Jan 28, 2016

This example demonstrates how to use the TreeDigest implementation from the Java TDigest library to compute approximate cumulative distribution functions (CDFs) and probability density functions (PDFs) from large datasets. It shows three Scala workflows: serializing and deserializing digests, merging per-element digests across an Apache Spark RDD using `reduce`, and using Kryo serialization to pass TreeDigest objects directly between Spark tasks. The code samples from exponential and normal distributions via Apache Commons Math, then compares CDF values computed locally versus through Spark, highlighting the library’s integration with distributed data processing.

AI-generated description

Similar vizzes