Text · Lesson 14.1

String literal

Source
See that a string literal is a read-only char8[..] slice: its length counts bytes, escapes write awkward characters, and changing text means copying it first.

Text has been in every lesson so far — every PrintLine starts with one — and it was never a special string type. A literal such as "hello" is a char8[..]: the read-only slice from the Sequences part, viewing UTF-8 bytes that the compiler stored inside the program. Everything a slice can do, a literal can do. This lesson looks at what that means — and at the two surprises that follow from it.

A literal is a slice

let word = "hello";
PrintLine("{} has {} bytes", word, word.length);
PrintLine("first {}, last {}", word[0], word[word.length - 1]);
PrintLine("word[1..4] is {}", word[1..4]);
Slice operationOn wordResult
.lengthword.length5
indexingword[0]h, one char8
a rangeword[1..4]ell, another slice
passing it onPrintLine("{}", word)any char8[..] parameter
a for loopfor c in wordeach char8 in turn

Because a literal is just a slice, a function that takes char8[..] accepts text without any conversion, and a range of a literal is text too.

The length counts bytes

.length counts the slice's elements, and the elements of a char8[..] are bytes — UTF-8 code units, not characters. For plain English letters each character is one byte, so the difference is invisible. It shows up as soon as the text leaves ASCII:

let euro = "\u{20AC}";
PrintLine("{} is {} bytes", euro, euro.length);

A reader sees one character, the euro sign, but UTF-8 stores it in three bytes, so its length is 3. The next lesson, Encoding, is about exactly that gap.

Escapes

A backslash starts an escape: a character that is awkward or impossible to type between quotes. Each escape is written with two or more characters but stores fewer bytes — "\n" is one byte, a newline.

PrintLine("quote     \"to be\"");
PrintLine("backslash C:\Rux\Bin");
PrintLine("tab       one\ttwo");
PrintLine("newline   first\n          second");
EscapeStores
\"a double quote, without ending the text
\'a single quote
\one backslash
\na newline
\ta tab
\ra carriage return
\0a zero byte
\u{20AC}the Unicode character with that number

The compiler also knows \a, \b, \f and \v, the old terminal control characters. Any other letter after a backslash is an error.

Read-only: copy to change

The bytes of a literal live in a read-only part of the program, and a char8[..] is a read-only view of them. So a literal can be read, but never changed in place. To change text, copy its bytes into storage the program owns — here an array — and change the copy:

var letters: char8[5];
for i in 0..word.length {
    letters[i] = word[i];
}
letters[0] = c8'j';
PrintLine("{} became {}", word, letters[..]);

c8'j' is a character literal of type char8, one UTF-8 byte. letters[..] is a slice of the whole array, which {} prints as text — jello, while word is still hello.

flowchart LR
    ro["Read-only program data<br/>h e l l o"] -- "viewed by" --> w["word: char8[..]<br/>read only"]
    ro -- "copied byte by byte" --> a["letters: char8[5]<br/>owned by Main"]
    a -- "letters[0] = c8'j'" --> j["j e l l o"]

The program

The whole lesson is one package in the Examples repository. Its comments explain every step.

Src/Main.rux
// Text has been in every lesson so far, and it was never a special string type. A literal such
// as "hello" is a `char8[..]` — the read-only slice from the Sequences part, viewing UTF-8 bytes
// that the compiler stored inside the program. Everything a slice can do, a literal can do:
// `.length`, indexing, ranges, and passing to a function that takes a `char8[..]`.
//
// Two things follow from that. The length counts bytes (the encoding's code units), not
// characters. And the bytes live in a read-only part of the program, so a literal can be read
// but never changed in place.
import Io::PrintLine;

func Main() -> int {
    let word = "hello";
    PrintLine("{} has {} bytes", word, word.length);
    PrintLine("first {}, last {}", word[0], word[word.length - 1]);
    PrintLine("word[1..4] is {}", word[1..4]);

    // A backslash starts an escape: a character that is awkward or impossible to type between
    // quotes. Each escape is written with two or more characters but stores fewer bytes.
    PrintLine("quote     \"to be\"");
    PrintLine("backslash C:\\Rux\\Bin");
    PrintLine("tab       one\ttwo");
    PrintLine("newline   first\n          second");
    PrintLine("\"\\n\" is {} byte", "\n".length);

    // `\u{...}` names any Unicode character by its number. The euro sign takes three bytes
    // in UTF-8, so its length is 3 even though a reader sees one character. The next lesson
    // is about exactly that gap.
    let euro = "\u{20AC}";
    PrintLine("{} is {} bytes", euro, euro.length);

    // The literal itself cannot change: `word[0] = c8'j';` is rejected with "cannot modify
    // elements through read-only slice 'char8[..]'". To change text, copy its bytes into storage
    // the program owns — here an array — and change the copy.
    var letters: char8[5];
    for i in 0..word.length {
        letters[i] = word[i];
    }
    letters[0] = c8'j';
    PrintLine("{} became {}", word, letters[..]);
    return 0;
}

Run it

cd Examples/Text/StringLiteral
rux run
hello has 5 bytes
first h, last o
word[1..4] is ell
quote     "to be"
backslash C:\Rux\Bin
tab       one   two
newline   first
          second
"\n" is 1 byte
€ is 3 bytes
hello became jello

Common mistakes

Changing a literal in place.
word[0] = c8'j'; fails with error: cannot modify elements through read-only slice 'char8[..]'. Copy the bytes into an array (or, later, a String builder) and change the copy.
A single backslash in a path.
"C:\Rux\Bin" fails with error: escape sequence '\R' is not recognized (and again for \B). Double each backslash: "C:\Rux\Bin". Worse, a path such as "C:\new" compiles without complaint — \n is a real escape, so the text quietly contains a newline.
Indexing one past the end.
The last byte is word[word.length - 1]. word[word.length] compiles, but stops the program with Panic: index out of range.

Try it yourself

  1. Count how many times l appears in "hello" with a for loop over the literal.
  2. Print the length of "naïve". Is it what you expected? Which character costs the extra byte?
  3. Copy "stressed" into an array backwards and print it. (The String lesson does the same thing and explains what can go wrong with non-ASCII text.)
  4. Print a line that contains a tab, a quote and a backslash, using only escapes.

Learn more