FLINK-2030: Implement an online histogram with Merging and equalization features
Github user sachingoel0101 commented on a diff in the pull request:

--- Diff: docs/libs/ml/statistics.md ---
+* This will be replaced by the TOC
+## Description
+ The statistics utility provides features such as building histograms over data.
+## Methods
+ The Statistics utility provides two major functions: createHistogram and
+ createDiscreteHistogram.
+
+### Creating a histogram
+ There are two types of histograms:
+   1. **Continuous Histograms**: These histograms are formed on a data set X: DataSet[Double]
+   when the values in X are from a continuous range. These histograms support
+   quantile and sum  operations. Here quantile(q) refers to a value $x_q$ such
that $|x: x + \leq x_q| = q * |X|$. Further, sum(s) refers to the number of elements $x \leq s$,
which can
+    be construed as a cumulative probability value at $s$[Of course, *scaled* probability].
The idea is, histograms are meant to represent a probability distribution, and there is
a very clear correspondence. A probability mass function corresponds to a discrete valued
histogram and a probability density function corresponds to a continuous valued histogram.

Implement an online histogram with Merging and equalization features
>          Components: Machine Learning Library
>            Reporter: Sachin Goel
>            Assignee: Sachin Goel
>            Priority: Minor
>              Labels: ML
> For the implementation of the decision tree in https://issues.apache.org/jira/browse/FLINK-1727,
we need to implement an histogram with online updates, merging and equalization features.
A reference implementation is provided in [1]
> [1].http://www.jmlr.org/papers/volume11/ben-haim10a/ben-haim10a.pdf

