Text
Rux has no built-in string type. Text is a slice of character code units, and a string literal is a read-only slice of code units stored in the program:
| Type | Encoding | Literal |
|---|---|---|
char8[..] | UTF-8 | "Hello", c8"Hello" |
char16[..] | UTF-16 | c16"Hello" |
char32[..] | UTF-32 | c32"Hello" |
char8[..] is the everyday text type: source files are UTF-8, unprefixed literals are UTF-8, and the standard packages take and return char8[..]. The wider encodings exist for foreign APIs that use them, such as UTF-16 on Windows.
Owned, growable text — String, StringBuilder and the validated StringView — comes from the Text package, not from the language.
Code units, not characters
Everything a slice does, it does in code units of its encoding: .length counts them, [i] returns one, and for visits each one. A character outside ASCII takes several UTF-8 units, or two UTF-16 units outside the Basic Multilingual Plane:
| Text | char8[..] length | char16[..] length | char32[..] length |
|---|---|---|---|
Hello | 5 | 5 | 5 |
€ (U+20AC) | 3 | 1 | 1 |
🚀 (U+1F680) | 4 | 2 | 1 |
let euro = "€uro";
let units = euro.length; // 6
let first = euro[0] as uint8; // 226, the first UTF-8 byte of €
let tail = euro[3..]; // "uro"
let wide = c16"\u{1F680}";
let pair = wide.length; // 2, a surrogate pair
A sub-slice can therefore cut a character in half. The Text and Unicode packages provide validated text, iteration by character and grapheme boundaries when an operation needs them.
Members and operations
| Form | Type | Meaning |
|---|---|---|
text.length | uint | Number of code units |
text.data | *char8 (and so on) | Pointer to the first code unit |
text[i] | char8 (and so on) | The code unit at index i |
text[a..b] | char8[..] (and so on) | A view of units a up to b |
for u in text | — | Visits each code unit in order |
These members need no import. An index past the end stops the program with Panic: index out of range. Slices describes indexing and sub-slicing in full.
func CountSpaces(text: char8[..]) -> uint {
var spaces: uint = 0;
for unit in text {
if unit == ' ' {
spaces += 1;
}
}
return spaces;
}
Comparing text
== is not defined on slices, because comparing two views would compare where they point rather than what they hold:
error: operator '==' is not defined for slice type 'char8[..]'
note: a slice is a view, so comparing the views would compare addresses rather than elements
help: declare '==' on 'char8[..]', or compare the elements one at a time
Compare the units, or use StringView::Equals from the Text package:
func SameText(a: char8[..], b: char8[..]) -> bool {
if a.length != b.length {
return false;
}
for i in 0..a.length {
if a[i] != b[i] {
return false;
}
}
return true;
}
Storage
A literal's code units are stored in read-only data, followed by a NUL code unit that .length does not count. "Hello".data can therefore be passed to a C function that expects a NUL-terminated string. A slice value itself is a pointer and a length — 16 bytes — so passing or copying text never copies its units.
A literal is read-only, and so is any char8[..]:
let word = "word";
word[0] = c8'W';
error: cannot modify elements through read-only slice 'char8[..]'
help: use a 'var T[..]' view to write through the sequence
To change text, copy it into storage the program owns and take a writable view var char8[..] of that:
let word = "word";
var letters: char8[4];
for i in 0..word.length {
letters[i] = word[i];
}
letters[0] = c8'W';
let changed: var char8[..] = letters[..]; // "Word"
Constants and parameters
A string literal can initialise a constant, a field, or a parameter of the matching slice type:
const Greeting: char8[..] = "Hello";
struct Message {
text: char8[..];
tag: int32;
}
func Shout(text: char8[..]) -> uint {
return text.length;
}
Escapes such as \n, \" and \u{1F680} are encoded in the literal's own encoding. Literals lists them.
See also
- Characters — the code unit and scalar value types
- Slices — the type text is made of
- Text package —
String,StringBuilderand friends - String literal, Encoding, UTF-8, String view — lessons