Unicode
Encoding found two ways to count the same text: code units and scalars. There is a third, and it is the one a reader means. Ask a person how many characters are in café and they will say four — however the computer happens to store it. This lesson counts the same text all three ways, and finds text where all three answers differ.
Three levels of text
| Level | What it is | Counted here with |
|---|---|---|
| bytes | how UTF-8 stores the text | utf8.length |
| scalars | the characters Unicode numbers, one char32 each | utf32.length |
| graphemes | what a person would point at and call one character | CountGraphemes |
The program passes each text twice — as UTF-8 for the byte count, and as UTF-32, where one unit is one scalar:
func Count(utf8: char8[..], utf32: char32[..]) {
PrintLine("{} bytes {}, scalars {}, graphemes {}",
utf8, utf8.length, utf32.length, CountGraphemes(utf32));
}
CountGraphemes comes from the Unicode package, which knows where graphemes begin and end. Like everything in that package, it works on char32[..].
When the counts differ
Count("cafe", c32"cafe");
Count("caf\u{E9}", c32"caf\u{E9}");
Count("cafe\u{301}", c32"cafe\u{301}");
Count("\u{1F1FA}\u{1F1E6}", c32"\u{1F1FA}\u{1F1E6}");
| Text | Spelled as | Bytes | Scalars | Graphemes |
|---|---|---|---|---|
cafe | four ASCII letters | 4 | 4 | 4 |
café | é as one scalar, U+00E9 | 5 | 4 | 4 |
café | e + U+0301, a combining accent | 6 | 5 | 4 |
| 🇺🇦 | two regional indicators | 8 | 2 | 1 |
Plain ASCII — all three counts agree, which is why the difference is so easy to miss.
é as one scalar — two bytes, but one scalar and one grapheme.
e plus a combining accent — U+0301 is a scalar of its own that attaches to the letter before it. On screen it looks exactly like the line above, yet every count differs: six bytes, five scalars, four graphemes.
A flag — the Ukrainian flag is two "regional indicator" letters, U and A, four bytes each. A screen draws the pair as one picture, so it is one grapheme.
flowchart LR
t(["café, written with<br/>a combining accent"]) --> g["4 graphemes<br/>c · a · f · é"]
g --> s["5 scalars<br/>c · a · f · e · U+0301"]
s --> b["6 bytes<br/>63 61 66 65 CC 81"]Which count is right?
It depends on the question:
| Question | Count |
|---|---|
| How much room does it take in a file or a buffer? | bytes |
| What does an encoder or a case mapping work through? | scalars |
| Where does a cursor step? Does it fit "maximum 20 characters"? | graphemes |
Bytes measure storage, scalars are what an encoding works with, and graphemes are what a person sees. A text field that limits a name to 20 "characters" by counting bytes would reject a perfectly short name written in Ukrainian, and a cursor that stepped by scalars would stop between a letter and its accent.
The program
The whole lesson is one package in the Examples repository. Its comments explain every step.
// The Encoding lesson found two ways to count the same text: code units and scalars. There is a
// third, and it is the one a reader means. Text has three levels:
//
// bytes how UTF-8 stores it
// scalars the characters Unicode numbers, one `char32` each
// graphemes what a person would point at and call one character
//
// A grapheme can be made of several scalars. An accent may be its own scalar that attaches to
// the letter before it, and a flag is two "regional indicator" letters that a screen draws as
// one picture. The Unicode package knows where graphemes begin and end.
import Io::PrintLine;
import Unicode::CountGraphemes;
// The same text twice: as UTF-8 for the byte count, and as UTF-32, where one unit is one scalar.
func Count(utf8: char8[..], utf32: char32[..]) {
PrintLine("{} bytes {}, scalars {}, graphemes {}",
utf8, utf8.length, utf32.length, CountGraphemes(utf32));
}
func Main() -> int {
// Plain ASCII: all three counts agree, which is why the difference is easy to miss.
Count("cafe", c32"cafe");
// `é` as one scalar: two bytes, but one scalar and one grapheme.
Count("caf\u{E9}", c32"caf\u{E9}");
// `e` followed by U+0301, a combining acute accent. It looks the same as the line above,
// yet every count differs: six bytes, five scalars, four graphemes.
Count("cafe\u{301}", c32"cafe\u{301}");
// The Ukrainian flag: two regional indicators, four bytes each, drawn as one grapheme.
Count("\u{1F1FA}\u{1F1E6}", c32"\u{1F1FA}\u{1F1E6}");
// Which count is right depends on the question. Bytes measure storage, scalars are what an
// encoding works with, and graphemes are what a cursor steps over or a "maximum 20
// characters" rule should count.
return 0;
}
Besides Io, its Rux.toml lists Unicode under [Dependencies].
Run it
cd Examples/Text/Unicode
rux run
cafe bytes 4, scalars 4, graphemes 4
café bytes 5, scalars 4, graphemes 4
café bytes 6, scalars 5, graphemes 4
🇺🇦 bytes 8, scalars 2, graphemes 1
Common mistakes
CountGraphemes takes char32[..]. CountGraphemes(utf8) fails with error: argument 1 to 'CountGraphemes' has type 'char8[..]', but parameter 'text' requires 'char32[..]'.The two
café lines print identically, but one is 5 bytes and the other 6. Compared byte by byte they are different texts. Deciding that they mean the same thing is called normalisation, and it is a separate step.Try it yourself
- Count the family emoji 👨👩👧, written as
"\u{1F468}\u{200D}\u{1F469}\u{200D}\u{1F467}". How many scalars does one picture take? - Count a waving hand with a skin tone,
"\u{1F44B}\u{1F3FD}". - Count
"한국", then the same first syllable written as three separate jamo,"\u{1112}\u{1161}\u{11AB}". - Write a function that reports whether a name fits in 20 graphemes.
Learn more
- Encoding — bytes and scalars
- UTF-8 — how the bytes are decoded into scalars
- Unicode case — another Unicode operation that works on scalars
- Format — what a placeholder's width counts