Statistics
Given a handful of numbers, where is the middle, and how spread out are they around it? This program summarises a slice of measurements: how many there are, the smallest and largest, the mean, two kinds of variance and the median. Then it does the same for awkward inputs — repeated values, a single value and no values at all — because a summary that only works on nice data is not finished.
It is the checkpoint for Part 18: Algorithms, and most of the work is done by that part's functions: MinIndex, MaxIndex, Sum and Sort.
How it is put together
One function, Summarize, does all the work, and Main calls it four times with different data. Inside, the order of the steps matters:
flowchart LR
s(["Summarize"]) --> ext["count, smallest, largest<br/>(fine for zero values)"]
ext --> empty{"count == 0?"}
empty -- "yes" --> stop(["stop: nothing to divide"])
empty -- "no" --> mean["mean<br/>first pass"]
mean --> var["variance<br/>second pass"]
var --> sort["Sort, then median<br/>(changes the slice)"]| Statistic | How it is found | Lessons it uses |
|---|---|---|
| Smallest, largest | MinIndex, MaxIndex → uint? | Min and max, Presence |
| Mean | Sum(values, 0.0) / n | Fold, Convert |
| Variance, deviation | A loop of squared distances, Sqrt | For, Math |
| Median | Sort, then the middle element(s) | Sort, Writable slice |
| Four data sets | Views of arrays, including an empty one | Slice, Array |
Extremes that cope with nothing
MinIndex and MaxIndex do not return a value — they return the position of one, as an optional. An empty slice has no smallest element, and none says so without any special case:
match MinIndex(values) {
at? => PrintLine(" smallest {}", values[at]),
none => PrintLine(" smallest none")
}
Everything after this point divides by the count, so the function stops early when there is nothing to divide:
if count == 0 {
PrintLine(" mean, variance and median need at least one value");
PrintLine();
return;
}
Mean and variance: two passes
The variance is the average squared distance from the mean, so the mean has to be known before any distance can be measured — hence two passes over the data. The distances are squared so that values above and below the mean both count, instead of cancelling each other out:
let n = count as float64;
let mean = Sum(values, 0.0) / n;
var squares = 0.0;
for value in values {
let distance = value - mean;
squares += distance * distance;
}
Sum is the fold of + from the Fold lesson, already written. Its second argument is the starting value, 0.0, which also fixes the result's type as float64.
Population or sample?
There are two conventions for "average squared distance", and a summary should say which it uses:
| Name | Divide by | Use it when |
|---|---|---|
| Population variance | n | the data is everything there is |
| Sample variance | n − 1 | the data is a sample of something larger |
Dividing by n − 1 corrects for the sample's mean having been computed from the sample itself. With one value there is nothing to correct with — n − 1 is zero — so the program reports the sample variance as undefined rather than printing a number. The standard deviation is the square root of either variance, which brings the answer back into the units of the data.
The median sorts the data
The median is the middle value once the data is in order, or the average of the two middle values when the count is even:
Sort(values);
let middle = count / 2;
if count % 2 == 1 {
PrintLine(" median {} (the middle value)", values[middle]);
} else {
let low = values[middle - 1];
let high = values[middle];
PrintLine(" median {} (between {} and {})", (low + high) / 2.0, low, high);
}
Sort works in place, so Summarize takes a writable slice, var float64[..], and the caller's array is left sorted afterwards. That is why the median comes last: the values line at the top of each summary still shows the data in its original order.
Four kinds of input
Main keeps each data set in a var array and passes a view of it. The last call passes readings[..0], an empty view of an array that already exists — the simplest way to make "no values" without a new type:
var single: float64[1] = [42.0];
Summarize("a single value", single[..]);
// An empty view of an array, so there is nothing to summarise.
Summarize("no values", readings[..0]);
The program
The whole lesson is one package in the Examples repository. Its comments explain every step.
// Summarising a set of numbers: where the middle is, and how spread out the numbers are around it.
//
// The mean is the total divided by the count. The variance is the average squared distance from
// the mean, and the standard deviation is its square root, which brings the answer back to the
// units of the data. There are two conventions for that average, and a summary should say which
// one it uses:
//
// population variance divide by n the data is everything there is
// sample variance divide by n - 1 the data is a sample of something larger
//
// Dividing by n - 1 corrects for the sample's mean being computed from the sample itself. With one
// value there is nothing to correct with, so the sample variance is undefined rather than zero.
//
// The median is the middle value once the data is sorted, or the average of the two middle values
// when the count is even. `Sort` works in place, so `Summarize` takes a writable slice and leaves
// the caller's data in order.
import Algorithms::{ MaxIndex, MinIndex, Sort, Sum };
import Io::{ Print, PrintLine };
import Math::Sqrt;
func Summarize(label: char8[..], values: var float64[..]) {
PrintLine("{}", label);
Print(" values ");
for value in values {
Print(" {}", value);
}
PrintLine();
let count = values.length;
PrintLine(" count {}", count);
// The extremes answer with an optional index, so an empty slice is handled right here.
match MinIndex(values) {
at? => PrintLine(" smallest {}", values[at]),
none => PrintLine(" smallest none")
}
match MaxIndex(values) {
at? => PrintLine(" largest {}", values[at]),
none => PrintLine(" largest none")
}
// Everything else divides by the count, and there is no mean of nothing.
if count == 0 {
PrintLine(" mean, variance and median need at least one value");
PrintLine();
return;
}
// Two passes: the mean has to be known before the distances from it can be measured. The
// distances are squared so that values above and below both count, instead of cancelling.
// `Sum` is the fold of `+` from the Fold lesson, already written.
let n = count as float64;
let mean = Sum(values, 0.0) / n;
var squares = 0.0;
for value in values {
let distance = value - mean;
squares += distance * distance;
}
PrintLine(" mean {:.4}", mean);
let population = squares / n;
PrintLine(" population variance {:.4}, deviation {:.4}", population, Sqrt(population));
if count > 1 {
let sample = squares / (n - 1.0);
PrintLine(" sample variance {:.4}, deviation {:.4}", sample, Sqrt(sample));
} else {
PrintLine(" sample variance undefined for one value");
}
// The median needs the values in order, so it comes last.
Sort(values);
let middle = count / 2;
if count % 2 == 1 {
PrintLine(" median {} (the middle value)", values[middle]);
} else {
let low = values[middle - 1];
let high = values[middle];
PrintLine(" median {} (between {} and {})", (low + high) / 2.0, low, high);
}
PrintLine();
}
func Main() -> int {
// An odd count: one value sits exactly in the middle.
var readings: float64[9] = [12.5, 9.0, 15.25, 11.0, 8.75, 14.0, 10.5, 13.25, 11.75];
Summarize("nine readings", readings[..]);
// An even count with repeated values. Duplicates are ordinary data: each one counts.
var scores: float64[6] = [4.0, 7.0, 4.0, 1.0, 7.0, 7.0];
Summarize("six scores with repeats", scores[..]);
// One value: no spread at all, and no sample variance.
var single: float64[1] = [42.0];
Summarize("a single value", single[..]);
// An empty view of an array, so there is nothing to summarise.
Summarize("no values", readings[..0]);
return 0;
}
Besides Io, its Rux.toml lists Algorithms and Math under [Dependencies].
Run it
cd Examples/Projects/Statistics
rux run
nine readings
values 12.5 9.0 15.25 11.0 8.75 14.0 10.5 13.25 11.75
count 9
smallest 8.75
largest 15.25
mean 11.7778
population variance 4.3117, deviation 2.0765
sample variance 4.8507, deviation 2.2024
median 11.75 (the middle value)
six scores with repeats
values 4.0 7.0 4.0 1.0 7.0 7.0
count 6
smallest 1.0
largest 7.0
mean 5.0000
population variance 5.0000, deviation 2.2361
sample variance 6.0000, deviation 2.4495
median 5.5 (between 4.0 and 7.0)
a single value
values 42.0
count 1
smallest 42.0
largest 42.0
mean 42.0000
population variance 0.0000, deviation 0.0000
sample variance undefined for one value
median 42.0 (the middle value)
no values
values
count 0
smallest none
largest none
mean, variance and median need at least one value
Common mistakes
Summarize sorts, so its parameter is var float64[..]. An array declared with let gives only a read-only view, and the call fails: error: argument 2 to 'Summarize' has type 'float64[..]', but parameter 'values' requires 'var float64[..]'. Dropping the var from the parameter just moves the problem to Sort, which needs a var T[..] too.Without
Sort(values); the program still runs, but the "median" is whatever sits in the middle of the unsorted data: 8.75 for the nine readings instead of 11.75.Remove the
count == 0 check and the empty summary prints a mean of NaN — zero divided by zero — and then stops with Panic: index out of range: middle - 1 on a count of 0 is not −1 but a huge unsigned number.Try it yourself
- Add the range (largest minus smallest) to the summary. Where must it go so that it still works for an empty slice?
- Add the mode, the most frequent value. Because the slice is sorted by then, equal values sit side by side, so one pass that counts runs is enough.
- Print the values again after the median, to see that the caller's array really has been sorted.
- Summarise a copy instead: make
Summarizetake a read-onlyfloat64[..]and sort a copy of the data in a local array, so the caller's data keeps its order. What limits the size of that local array?
Learn more
- Min and max, Fold and Sort
- Writable slice — why sorting needs
var T[..] - Sqrt in the API reference
- Next projects: Guess, Age, Password and Launch, the checkpoints for Utilities