UUID
A UUID — universally unique identifier — is 16 bytes used as a name that nobody else will pick. Databases key rows by them, upload services name files with them, and logs tag every request with one, so the entries of a single request can be found among millions.
What makes them useful is that nobody hands them out. Two programs that never talk to each other can each make UUIDs, and in practice they will never make the same one. This lesson makes one, reads and writes the text form, and compares two:
import Uuid::{ Parse, Random, Uuid, UuidParseError };
The text form
A UUID is written as 32 hexadecimal digits in groups of 8, 4, 4, 4 and 12, separated by hyphens:
123e4567-e89b-42d3-a456-426614174000
^ ^
| variant: a = the standard layout
version 4: random
The usual kind, version 4, is 122 random bits plus 6 fixed bits that record the version and the layout. 122 random bits is about 5 × 10³⁶ possibilities — enough that a collision between honestly made UUIDs is, in practice, never going to happen.
Reading one
Parse reads the text form and returns Uuid ! UuidParseError. Like ParseDate in Date, the error is a variant whose every case carries the byte where the text went wrong:
func Reason(error: UuidParseError) -> char8[..] {
return match error {
.WrongLength(_) => "wrong length",
.Misplaced(_) => "hyphen in the wrong place",
.NotHexadecimal(_) => "not a hexadecimal digit",
.BadPrefix(_) => "not a urn:uuid: prefix"
};
}
func Show(text: char8[..]) {
match Parse(text) {
.Success(id) => PrintLine("{:<45} -> {}", text, id),
.Failure(error) => PrintLine("{:<45} -> {} at byte {}", text, Reason(error), error.Offset())
}
}
Parse is strict. It accepts the hyphenated form, in either case, with or without the urn:uuid: prefix — and refuses everything else:
| Text | Result |
|---|---|
123e4567-e89b-42d3-a456-426614174000 | accepted — the canonical form |
123E4567-E89B-42D3-A456-426614174000 | accepted — upper case means the same |
urn:uuid:123e4567-e89b-42d3-a456-426614174000 | accepted — the standard URN prefix |
123e4567e89b42d3a456426614174000 | wrong length — the hyphens are not optional |
123e4567-e89b-42d3-a456-42661417400g | not a hexadecimal digit, at byte 35 |
{123e4567-e89b-42d3-a456-426614174000} | wrong length — braces are not part of the form |
A strict reader is the safe default: text that is almost a UUID is more likely a mistake than a UUID, and refusing it with a position makes the mistake easy to find.
Writing and comparing
Printing with {} always writes the one canonical form, in lower case. So all three accepted spellings above print identically, and comparison is on the 16 bytes, not on the text:
let lower = Parse("123e4567-e89b-42d3-a456-426614174000") catch { else => Uuid::Nil() };
let upper = Parse("123E4567-E89B-42D3-A456-426614174000") catch { else => Uuid::Nil() };
PrintLine("same identifier: {}", lower.Equals(upper));
PrintLine("the nil UUID: {}", Uuid::Nil());
The catch with a fallback is safe here because both texts are known to be valid. The fallback, Uuid::Nil(), is the special UUID with every bit zero — useful as a "no identifier" placeholder, and never produced by Random.
Making one
Random() makes a version 4 UUID from the operating system's entropy, so like every request for entropy in Entropy it can fail, and returns Uuid ! EntropyError:
match Random() {
.Success(id) => PrintLine("a new one: {} version {}", id, id.Version()),
.Failure(_) => PrintLine("no entropy for a new UUID")
}
Version() reads the version digit back: 4.
The program
The whole lesson is one package in the Examples repository. Its comments explain every step.
// A UUID (universally unique identifier) is 16 bytes used as a name that nobody else will pick:
// a database row, an uploaded file, a request in a log. Its text form is 32 hexadecimal digits
// in groups of 8-4-4-4-12, such as 123e4567-e89b-42d3-a456-426614174000.
//
// The usual kind, version 4, is 122 random bits plus 6 bits marking the version and layout. With
// that many random bits, two programs that never talk to each other will still, in practice,
// never produce the same one. `Random()` makes one from the operating system's entropy, so like
// any request for entropy it returns `Uuid ! EntropyError`.
//
// `Parse` reads the text form and returns `Uuid ! UuidParseError`. It is strict: upper case is
// accepted, since it means the same, but braces, missing hyphens or stray characters are
// refused, with the byte where the text went wrong. Printing with `{}` always writes the one
// canonical form, in lower case.
import Io::PrintLine;
import Uuid::{ Parse, Random, Uuid, UuidParseError };
func Reason(error: UuidParseError) -> char8[..] {
return match error {
.WrongLength(_) => "wrong length",
.Misplaced(_) => "hyphen in the wrong place",
.NotHexadecimal(_) => "not a hexadecimal digit",
.BadPrefix(_) => "not a urn:uuid: prefix"
};
}
func Show(text: char8[..]) {
match Parse(text) {
.Success(id) => PrintLine("{:<45} -> {}", text, id),
.Failure(error) => PrintLine("{:<45} -> {} at byte {}", text, Reason(error), error.Offset())
}
}
func Main() -> int {
Show("123e4567-e89b-42d3-a456-426614174000");
Show("123E4567-E89B-42D3-A456-426614174000");
Show("urn:uuid:123e4567-e89b-42d3-a456-426614174000");
Show("123e4567e89b42d3a456426614174000");
Show("123e4567-e89b-42d3-a456-42661417400g");
Show("{123e4567-e89b-42d3-a456-426614174000}");
// Two spellings, one identifier: comparison is on the 16 bytes, not on the text.
let lower = Parse("123e4567-e89b-42d3-a456-426614174000") catch { else => Uuid::Nil() };
let upper = Parse("123E4567-E89B-42D3-A456-426614174000") catch { else => Uuid::Nil() };
PrintLine("same identifier: {}", lower.Equals(upper));
PrintLine("the nil UUID: {}", Uuid::Nil());
// A fresh random one, different on every run.
match Random() {
.Success(id) => PrintLine("a new one: {} version {}", id, id.Version()),
.Failure(_) => PrintLine("no entropy for a new UUID")
}
return 0;
}
Besides Io, its Rux.toml lists Uuid under [Dependencies].
Run it
cd Examples/Utilities/Uuid
rux run
123e4567-e89b-42d3-a456-426614174000 -> 123e4567-e89b-42d3-a456-426614174000
123E4567-E89B-42D3-A456-426614174000 -> 123e4567-e89b-42d3-a456-426614174000
urn:uuid:123e4567-e89b-42d3-a456-426614174000 -> 123e4567-e89b-42d3-a456-426614174000
123e4567e89b42d3a456426614174000 -> wrong length at byte 32
123e4567-e89b-42d3-a456-42661417400g -> not a hexadecimal digit at byte 35
{123e4567-e89b-42d3-a456-426614174000} -> wrong length at byte 36
same identifier: true
the nil UUID: 00000000-0000-0000-0000-000000000000
a new one: 12c16353-3376-4e49-a1ec-403586becf8d version 4
Every line is the same on each run except the last, which is a sample: a new random UUID each time.
Every line is the same on each run except the last, which is a sample: a new random UUID each time.
Common mistakes
let id: Uuid = Random(); fails with error: cannot assign 'Uuid ! EntropyError' to 'Uuid'. Making a UUID can fail when the system has no entropy, so handle that first.123e4567-… and 123E4567-… are different strings and the same identifier. Parse both and compare the values — with Equals or == — rather than the text.catch { else => Uuid::Nil() } is fine for a literal the program knows is valid. On user input it turns every typo into the nil UUID, which then looks like a real value. Report the UuidParseError instead.The
match in Reason must name every case. Drop .BadPrefix and the compiler says error: match on 'UuidParseError' is not exhaustive; missing UuidParseError::BadPrefix.Try it yourself
- Show a UUID with its last two digits missing, and one with a hyphen one place too far right. Which reasons and bytes do you get?
- Print
Uuid::Max(), the UUID with every bit set. - Import
TimeOrderedand make a version 7 UUID, which starts with the current time. Make two and compare their first groups. - Make five random UUIDs in a loop and print them. Do any two share even their first group?
Learn more
- Entropy — where the random bits come from, and why it can fail
- Variant match — matching the cases of
UuidParseError - Catch fallback —
catch { else => … }and when it is safe - Hash — fingerprints derived from data, the opposite of an identifier made from nothing
20.8 Hash
Fingerprint bytes with Fnv1a64Of, XxHash64Of and Crc32Of, fast hashes that catch accidents but are not cryptographic.
Overview
Read and write JSON and TOML: parse a document into a tree, build one and write it out, stream a large one event by event, and report a bad document by where it went wrong.