Text · Lesson 14.2

Encoding

Source
Write the same text as UTF-8, UTF-16 and UTF-32 with the c8, c16 and c32 prefixes, and see that lengths count code units, not characters.

Every character has a number, its scalar value: A is 65, € is 8364, the rocket emoji 🚀 is 128640. An encoding decides how that number is stored as a run of fixed-size code units. Rux can write a literal in three encodings, and the same text takes a different number of units in each. That is why String literal found the euro sign to be three "long": .length counts code units, not characters.

Three encodings, three literal types

A prefix on the literal picks the encoding, and with it the element type of the slice:

LiteralTypeEncodingCode unit
"text"char8[..]UTF-81 byte
c8"text"char8[..]UTF-8, spelled out1 byte
c16"text"char16[..]UTF-162 bytes
c32"text"char32[..]UTF-324 bytes

The same prefixes pick a character literal's width: c8'A' is one char8 code unit, as in the previous lesson.

The program's Show takes all three slice types, so the same text can be measured in each encoding side by side:

func Show(eight: char8[..], sixteen: char16[..], thirtyTwo: char32[..]) {
    PrintLine("{}  UTF-8 {}, UTF-16 {}, UTF-32 {}",
        thirtyTwo, eight.length, sixteen.length, thirtyTwo.length);
}

How the counts drift apart

TextScalar valuesUTF-8 unitsUTF-16 unitsUTF-32 units
Rux82, 117, 120333
café…, 233 for é544
€8364311
🚀128640421
  • ASCII — every encoding uses one unit per character, and the counts agree. That is why the difference is so easy to miss.
  • é — two UTF-8 bytes, but one unit in the wider encodings.
  • € — three UTF-8 bytes, still one UTF-16 unit.
  • The rocket lies beyond the 65,536 values one UTF-16 unit can hold, so UTF-16 splits it into two units called a surrogate pair. Only UTF-32 holds it whole.
flowchart LR
    r(["🚀 scalar value 128640"]) --> u8["UTF-8<br/>4 units of 1 byte"]
    r --> u16["UTF-16<br/>2 units of 2 bytes<br/>(a surrogate pair)"]
    r --> u32["UTF-32<br/>1 unit of 4 bytes"]

A UTF-32 unit is wide enough for any scalar, so in a char32[..] the count of units is the count of scalars. The narrower encodings spend several units on one scalar once the text leaves plain ASCII.

Indexing returns a code unit

[i] gives the i-th code unit of the slice's own encoding — not the i-th character:

let rocket16 = c16"🚀";
let rocket32 = c32"🚀";
PrintLine("rocket16[0] is {}", rocket16[0] as uint32);
PrintLine("rocket32[0] is {}", rocket32[0] as uint32);

rocket32[0] is the whole rocket, 128640. rocket16[0] is 55357: the first half of a surrogate pair, a number but no character at all on its own. The as uint32 conversion prints the units as numbers, which is the honest way to look at half a character.

Which encoding to use? UTF-8 is the default for a reason: it is what files, terminals and the network speak, and the Text package works in it. The wider literals exist for the places that need them — Windows APIs take UTF-16, and the Unicode package works on char32[..], where one unit is one scalar.

The program

The whole lesson is one package in the Examples repository. Its comments explain every step.

Src/Main.rux
// A character has a number, its scalar value: `A` is 65, `€` is 8364, the rocket emoji is 128640.
// An encoding decides how that number is stored as a run of fixed-size code units, and Rux
// literals can be written in three of them:
//
//     "text"      char8[..]    UTF-8, one-byte units — the default
//     c8"text"    char8[..]    the same, with the encoding spelled out
//     c16"text"   char16[..]   UTF-16, two-byte units
//     c32"text"   char32[..]   UTF-32, four-byte units
//
// The same prefixes pick a character literal's width: `c8'A'` is one `char8` code unit.
//
// `.length` and indexing always count code units of the slice's own encoding. A UTF-32 unit is
// wide enough for any scalar, so there the count of units is the count of scalars. The narrower
// encodings spend several units on one scalar once the text leaves plain ASCII.
import Io::PrintLine;

// The three parameter types are the three literal types, so this function shows the same text
// in each encoding side by side.
func Show(eight: char8[..], sixteen: char16[..], thirtyTwo: char32[..]) {
    PrintLine("{}  UTF-8 {}, UTF-16 {}, UTF-32 {}",
        thirtyTwo, eight.length, sixteen.length, thirtyTwo.length);
}

func Main() -> int {
    // ASCII: every encoding uses one unit per character, and the counts agree.
    Show("Rux", c16"Rux", c32"Rux");

    // `é` takes two UTF-8 bytes but one unit in the wider encodings.
    Show("café", c16"café", c32"café");

    // The euro sign takes three UTF-8 bytes.
    Show("€", c16"€", c32"€");

    // The rocket lies beyond the 65,536 values one UTF-16 unit can hold, so UTF-16 splits it
    // into two units called a surrogate pair. Only UTF-32 holds it whole.
    Show("🚀", c16"🚀", c32"🚀");

    // Indexing returns a code unit, not a character. The first unit of the UTF-16 rocket is
    // half a surrogate pair: a number, but no character at all on its own.
    let rocket16 = c16"🚀";
    let rocket32 = c32"🚀";
    PrintLine("rocket16[0] is {}", rocket16[0] as uint32);
    PrintLine("rocket32[0] is {}", rocket32[0] as uint32);
    return 0;
}

Run it

cd Examples/Text/Encoding
rux run
Rux  UTF-8 3, UTF-16 3, UTF-32 3
café  UTF-8 5, UTF-16 4, UTF-32 4
€  UTF-8 3, UTF-16 1, UTF-32 1
🚀  UTF-8 4, UTF-16 2, UTF-32 1
rocket16[0] is 55357
rocket32[0] is 128640

Common mistakes

Passing one encoding where another is expected.
The three slice types are different types, and nothing converts between them silently. Show("Rux", "Rux", c32"Rux") fails with error: argument 2 to 'Show' has type 'char8[..]', but parameter 'sixteen' requires 'char16[..]'. Write the prefix: c16"Rux".
A character literal too wide for its unit.
c8'€' fails with error: character '€' (U+20AC) does not fit one 'char8' code unit, and c16'🚀' likewise for char16. The help suggests a string literal instead, such as c8"€", which may use several units.
Printing half a surrogate pair.
Without as uint32, rocket16[0] prints as �, a replacement mark: half a pair is not a character, so there is nothing real to show.

Try it yourself

  1. Call Show with naïve, 日本 and Ελλάδα. Which encoding is shortest for each?
  2. Print every unit of c16"🚀" as a number with a for loop.
  3. Find a character other than an emoji that needs a surrogate pair in UTF-16. Hint: its scalar value must be above 65,535.

Learn more

  • char8, char16 and char32 in the Rux Reference
  • Character — the character types from Part 1
  • UTF-8 — checking and decoding UTF-8 bytes one scalar at a time
  • Unicode — a third way to count, the one a reader means