Characters
A character type holds one unit of text. Rux distinguishes two kinds, and the difference decides which values are valid:
- A code unit is one piece of an encoding.
char8is one byte of UTF-8 andchar16one 16-bit unit of UTF-16. Every value that fits the width is a valid code unit, but a code unit is a whole character only when the character is small enough to be encoded in one. - A scalar value is one Unicode character: a code point from U+0000 to U+10FFFF, excluding the surrogates U+D800 to U+DFFF.
char32andchar64hold scalar values.
The character types
| Type | Bits | Bytes | Holds | Range | Literal prefix |
|---|---|---|---|---|---|
char8 | 8 | 1 | A UTF-8 code unit | 0 to 255 | c8 |
char16 | 16 | 2 | A UTF-16 code unit | 0 to 65,535 | c16 |
char32 | 32 | 4 | A Unicode scalar value | U+0000 to U+10FFFF, without surrogates | c32, or none |
char64 | 64 | 8 | A Unicode scalar value | U+0000 to U+10FFFF, without surrogates | c64 |
char is a built-in alias of char32, and the type of an unprefixed character literal. Use char for characters, and char8 or char16 when working with encoded text one unit at a time — indexing a string gives code units, not characters.
Reserved widths
char128, char256 and char512 (16, 32 and 64 bytes) are named by the language but not implemented in rux 0.4.0. Using one is error: primitive type 'char128' is reserved but is not implemented in this compiler version.
Associated constants
Imported from Core:
| Constant | char8 | char16 | char32 | char64 |
|---|---|---|---|---|
Bits | 8 | 16 | 32 | 64 |
Bytes | 1 | 2 | 4 | 8 |
Min | 0 | 0 | U+0000 | U+0000 |
Max | 255 | 65,535 | U+10FFFF | U+10FFFF |
Min and Max have the character type itself; char::Max as uint32 is 1114111.
Literals
A character literal is one character or escape in single quotes. The prefix selects the type, and the character must be exactly one unit of it:
| Literal | Type | Accepts |
|---|---|---|
c8'A' | char8 | U+0000 to U+007F — one UTF-8 byte |
c16'Ж' | char16 | U+0000 to U+FFFF, not a surrogate — one UTF-16 unit |
'😀', c32'😀' | char32 | Any scalar value |
c64'😀' | char64 | Any scalar value |
An unprefixed literal initialising or assigned to a char8 or char16 takes that type by the same rule, so let initial: char8 = 'R'; works. A character that does not fit is refused, never truncated:
let accent = c8'é';
error: character 'é' (U+00E9) does not fit one 'char8' code unit
help: write a string literal such as c8"é", or a byte such as 0xE9u8
let emoji: char16 = '😀';
error: character '😀' (U+1F600) does not fit one 'char16' code unit
é takes two bytes of UTF-8 and 😀 two units of UTF-16; text that needs several units is a string.
let letter = 'A';
let emoji = '😀';
let unit = c8'R';
let word = c16'Ж';
let wide = c64'\u{1F600}';
let initial: char8 = 'R';
Operations
Characters compare by their numeric value — code-point order for scalar values, unit value for code units — with ==, !=, <, <=, > and >=. The order is not alphabetical in any language's sense: 'Z' < 'a', because U+005A comes before U+0061.
let c = c8'7';
let isDigit = c >= c8'0' && c <= c8'9'; // true
Arithmetic on characters goes through an integer: convert with as, compute, and convert back.
let lower = c8'a';
let upper = ((lower as uint8) - 32) as char8; // 'A'
let value = (c8'7' as int) - (c8'0' as int); // 7
let next = ('a' as uint32 + 1) as char; // 'b'
Conversions
Between character types
char32 widens implicitly to char64, since both hold the same scalar values. No other character conversion is implicit: char8 and char16 are units of different encodings, and a UTF-8 byte above 0x7F is not the UTF-16 unit with the same number.
let unit: char8 = 'a';
let crossed: char16 = unit;
error: cannot assign 'char8' to 'char16'
let scalar: char32 = 'a';
let narrowed: char16 = scalar;
error: cannot assign 'char32' to 'char16'
With as, a conversion to a wider character keeps the value, and a conversion to a narrower one keeps the low bits. It never re-encodes: for a face: char32 holding U+1F600, face as char16 is the unit 0xF600, not a surrogate pair. A constant that does not fit is refused instead, as '😀' as char16 is with error: constant cast from 'char32' to 'char16' is outside the target type's range. To transcode text, use the Unicode and Text packages.
To and from integers
as converts a character to any integer type, giving its numeric value, and an integer to any character type:
let code = 'A' as uint32; // 65
let byte = c8'A' as uint8; // 65
let smile = 0x263A as char; // '☺'
When the integer is a constant, the compiler checks that the character type can hold it:
let past: char8 = 0x100 as char8;
error: constant cast from 'int' to 'char8' is outside the target type's range
let surrogate: char32 = 0xD800 as char32;
error: cast from 'int' to 'char32' uses invalid surrogate code point U+D800
A conversion at run time is not checked: it keeps the low bits of the integer, so a computed number can produce a surrogate or a value above U+10FFFF in a char32. Validate such values before treating them as text.
Other conversions
Characters convert with as to floating-point types ('A' as float64 is 65.0) and to booleans (false only for U+0000). Without as, a character is never a number: let n: uint8 = c8'a'; is error: cannot assign 'char8' to 'uint8'.