Characters

A character type holds one unit of text. Rux distinguishes two kinds, and the difference decides which values are valid:

  • A code unit is one piece of an encoding. char8 is one byte of UTF-8 and char16 one 16-bit unit of UTF-16. Every value that fits the width is a valid code unit, but a code unit is a whole character only when the character is small enough to be encoded in one.
  • A scalar value is one Unicode character: a code point from U+0000 to U+10FFFF, excluding the surrogates U+D800 to U+DFFF. char32 and char64 hold scalar values.

The character types

TypeBitsBytesHoldsRangeLiteral prefix
char881A UTF-8 code unit0 to 255c8
char16162A UTF-16 code unit0 to 65,535c16
char32324A Unicode scalar valueU+0000 to U+10FFFF, without surrogatesc32, or none
char64648A Unicode scalar valueU+0000 to U+10FFFF, without surrogatesc64

char is a built-in alias of char32, and the type of an unprefixed character literal. Use char for characters, and char8 or char16 when working with encoded text one unit at a time — indexing a string gives code units, not characters.

Reserved widths

char128, char256 and char512 (16, 32 and 64 bytes) are named by the language but not implemented in rux 0.4.0. Using one is error: primitive type 'char128' is reserved but is not implemented in this compiler version.

Associated constants

Imported from Core:

Constantchar8char16char32char64
Bits8163264
Bytes1248
Min00U+0000U+0000
Max25565,535U+10FFFFU+10FFFF

Min and Max have the character type itself; char::Max as uint32 is 1114111.

Literals

A character literal is one character or escape in single quotes. The prefix selects the type, and the character must be exactly one unit of it:

LiteralTypeAccepts
c8'A'char8U+0000 to U+007F — one UTF-8 byte
c16'Ж'char16U+0000 to U+FFFF, not a surrogate — one UTF-16 unit
'😀', c32'😀'char32Any scalar value
c64'😀'char64Any scalar value

An unprefixed literal initialising or assigned to a char8 or char16 takes that type by the same rule, so let initial: char8 = 'R'; works. A character that does not fit is refused, never truncated:

let accent = c8'é';
error: character 'é' (U+00E9) does not fit one 'char8' code unit
help: write a string literal such as c8"é", or a byte such as 0xE9u8

let emoji: char16 = '😀';
error: character '😀' (U+1F600) does not fit one 'char16' code unit

é takes two bytes of UTF-8 and 😀 two units of UTF-16; text that needs several units is a string.

let letter = 'A';
let emoji = '😀';
let unit = c8'R';
let word = c16'Ж';
let wide = c64'\u{1F600}';
let initial: char8 = 'R';

Operations

Characters compare by their numeric value — code-point order for scalar values, unit value for code units — with ==, !=, <, <=, > and >=. The order is not alphabetical in any language's sense: 'Z' < 'a', because U+005A comes before U+0061.

let c = c8'7';
let isDigit = c >= c8'0' && c <= c8'9';   // true

Arithmetic on characters goes through an integer: convert with as, compute, and convert back.

let lower = c8'a';
let upper = ((lower as uint8) - 32) as char8;   // 'A'
let value = (c8'7' as int) - (c8'0' as int);    // 7
let next = ('a' as uint32 + 1) as char;         // 'b'

Conversions

Between character types

char32 widens implicitly to char64, since both hold the same scalar values. No other character conversion is implicit: char8 and char16 are units of different encodings, and a UTF-8 byte above 0x7F is not the UTF-16 unit with the same number.

let unit: char8 = 'a';
let crossed: char16 = unit;
error: cannot assign 'char8' to 'char16'

let scalar: char32 = 'a';
let narrowed: char16 = scalar;
error: cannot assign 'char32' to 'char16'

With as, a conversion to a wider character keeps the value, and a conversion to a narrower one keeps the low bits. It never re-encodes: for a face: char32 holding U+1F600, face as char16 is the unit 0xF600, not a surrogate pair. A constant that does not fit is refused instead, as '😀' as char16 is with error: constant cast from 'char32' to 'char16' is outside the target type's range. To transcode text, use the Unicode and Text packages.

To and from integers

as converts a character to any integer type, giving its numeric value, and an integer to any character type:

let code = 'A' as uint32;           // 65
let byte = c8'A' as uint8;          // 65
let smile = 0x263A as char;         // '☺'

When the integer is a constant, the compiler checks that the character type can hold it:

let past: char8 = 0x100 as char8;
error: constant cast from 'int' to 'char8' is outside the target type's range

let surrogate: char32 = 0xD800 as char32;
error: cast from 'int' to 'char32' uses invalid surrogate code point U+D800

A conversion at run time is not checked: it keeps the low bits of the integer, so a computed number can produce a surrogate or a value above U+10FFFF in a char32. Validate such values before treating them as text.

Other conversions

Characters convert with as to floating-point types ('A' as float64 is 65.0) and to booleans (false only for U+0000). Without as, a character is never a number: let n: uint8 = c8'a'; is error: cannot assign 'char8' to 'uint8'.

See also