Text · Lesson 14.7

Unicode

Source
Count the same text three ways — bytes, scalars and graphemes — and find text where all three counts differ.
You'll need: Encoding, UTF-8

Encoding found two ways to count the same text: code units and scalars. There is a third, and it is the one a reader means. Ask a person how many characters are in café and they will say four — however the computer happens to store it. This lesson counts the same text all three ways, and finds text where all three answers differ.

Three levels of text

LevelWhat it isCounted here with
byteshow UTF-8 stores the textutf8.length
scalarsthe characters Unicode numbers, one char32 eachutf32.length
graphemeswhat a person would point at and call one characterCountGraphemes

The program passes each text twice — as UTF-8 for the byte count, and as UTF-32, where one unit is one scalar:

func Count(utf8: char8[..], utf32: char32[..]) {
    PrintLine("{}  bytes {}, scalars {}, graphemes {}",
        utf8, utf8.length, utf32.length, CountGraphemes(utf32));
}

CountGraphemes comes from the Unicode package, which knows where graphemes begin and end. Like everything in that package, it works on char32[..].

When the counts differ

Count("cafe", c32"cafe");
Count("caf\u{E9}", c32"caf\u{E9}");
Count("cafe\u{301}", c32"cafe\u{301}");
Count("\u{1F1FA}\u{1F1E6}", c32"\u{1F1FA}\u{1F1E6}");
TextSpelled asBytesScalarsGraphemes
cafefour ASCII letters444
caféé as one scalar, U+00E9544
cafée + U+0301, a combining accent654
🇺🇦two regional indicators821

Plain ASCII — all three counts agree, which is why the difference is so easy to miss.

é as one scalar — two bytes, but one scalar and one grapheme.

e plus a combining accent — U+0301 is a scalar of its own that attaches to the letter before it. On screen it looks exactly like the line above, yet every count differs: six bytes, five scalars, four graphemes.

A flag — the Ukrainian flag is two "regional indicator" letters, U and A, four bytes each. A screen draws the pair as one picture, so it is one grapheme.

flowchart LR
    t(["café, written with<br/>a combining accent"]) --> g["4 graphemes<br/>c · a · f · é"]
    g --> s["5 scalars<br/>c · a · f · e · U+0301"]
    s --> b["6 bytes<br/>63 61 66 65 CC 81"]

Which count is right?

It depends on the question:

QuestionCount
How much room does it take in a file or a buffer?bytes
What does an encoder or a case mapping work through?scalars
Where does a cursor step? Does it fit "maximum 20 characters"?graphemes

Bytes measure storage, scalars are what an encoding works with, and graphemes are what a person sees. A text field that limits a name to 20 "characters" by counting bytes would reject a perfectly short name written in Ukrainian, and a cursor that stepped by scalars would stop between a letter and its accent.

The program

The whole lesson is one package in the Examples repository. Its comments explain every step.

Src/Main.rux
// The Encoding lesson found two ways to count the same text: code units and scalars. There is a
// third, and it is the one a reader means. Text has three levels:
//
//     bytes       how UTF-8 stores it
//     scalars     the characters Unicode numbers, one `char32` each
//     graphemes   what a person would point at and call one character
//
// A grapheme can be made of several scalars. An accent may be its own scalar that attaches to
// the letter before it, and a flag is two "regional indicator" letters that a screen draws as
// one picture. The Unicode package knows where graphemes begin and end.
import Io::PrintLine;
import Unicode::CountGraphemes;

// The same text twice: as UTF-8 for the byte count, and as UTF-32, where one unit is one scalar.
func Count(utf8: char8[..], utf32: char32[..]) {
    PrintLine("{}  bytes {}, scalars {}, graphemes {}",
        utf8, utf8.length, utf32.length, CountGraphemes(utf32));
}

func Main() -> int {
    // Plain ASCII: all three counts agree, which is why the difference is easy to miss.
    Count("cafe", c32"cafe");

    // `é` as one scalar: two bytes, but one scalar and one grapheme.
    Count("caf\u{E9}", c32"caf\u{E9}");

    // `e` followed by U+0301, a combining acute accent. It looks the same as the line above,
    // yet every count differs: six bytes, five scalars, four graphemes.
    Count("cafe\u{301}", c32"cafe\u{301}");

    // The Ukrainian flag: two regional indicators, four bytes each, drawn as one grapheme.
    Count("\u{1F1FA}\u{1F1E6}", c32"\u{1F1FA}\u{1F1E6}");

    // Which count is right depends on the question. Bytes measure storage, scalars are what an
    // encoding works with, and graphemes are what a cursor steps over or a "maximum 20
    // characters" rule should count.
    return 0;
}

Besides Io, its Rux.toml lists Unicode under [Dependencies].

Run it

cd Examples/Text/Unicode
rux run
cafe  bytes 4, scalars 4, graphemes 4
café  bytes 5, scalars 4, graphemes 4
café  bytes 6, scalars 5, graphemes 4
🇺🇦  bytes 8, scalars 2, graphemes 1

Common mistakes

Counting graphemes in UTF-8.
CountGraphemes takes char32[..]. CountGraphemes(utf8) fails with error: argument 1 to 'CountGraphemes' has type 'char8[..]', but parameter 'text' requires 'char32[..]'.
Assuming text that looks the same is the same.
The two café lines print identically, but one is 5 bytes and the other 6. Compared byte by byte they are different texts. Deciding that they mean the same thing is called normalisation, and it is a separate step.

Try it yourself

  1. Count the family emoji 👨‍👩‍👧, written as "\u{1F468}\u{200D}\u{1F469}\u{200D}\u{1F467}". How many scalars does one picture take?
  2. Count a waving hand with a skin tone, "\u{1F44B}\u{1F3FD}".
  3. Count "한국", then the same first syllable written as three separate jamo, "\u{1112}\u{1161}\u{11AB}".
  4. Write a function that reports whether a name fits in 20 graphemes.

Learn more

  • Encoding — bytes and scalars
  • UTF-8 — how the bytes are decoded into scalars
  • Unicode case — another Unicode operation that works on scalars
  • Format — what a placeholder's width counts