Floating-Point
A floating-point type holds a binary fraction with an exponent, covering a vast range at a fixed number of significant digits. Rux implements the two IEEE 754 formats every supported machine has in hardware.
The floating-point types
| Type | Bits | Bytes | Format | Significant digits | Largest finite value | Suffix |
|---|---|---|---|---|---|---|
float32 | 32 | 4 | IEEE 754 binary32 | about 7 | ≈ 3.40 × 1038 | f32 |
float64 | 64 | 8 | IEEE 754 binary64 | about 16 | ≈ 1.80 × 10308 | f64 |
float is a built-in alias of float64, and float64 is the type of an unsuffixed floating-point literal. Use float32 where storage or bandwidth matters more than precision, such as large arrays of samples, and float64 otherwise.
Reserved widths
These widths are named by the language but not implemented in rux 0.4.0:
| Type | Bits | Bytes | Intended format | Suffix |
|---|---|---|---|---|
float8 | 8 | 1 | E4M3 (4 exponent, 3 significand bits) | f8 |
float16 | 16 | 2 | IEEE 754 binary16 | f16 |
float80 | 80 | 16 | x87 extended precision | f80 |
float128 | 128 | 16 | IEEE 754 binary128 | f128 |
float256 | 256 | 32 | IEEE 754 binary256 interchange | f256 |
float512 | 512 | 64 | IEEE 754 binary512 interchange | f512 |
Using one, as a type or through its suffix, is an error:
error: primitive type 'float16' is reserved but is not implemented in this compiler version
Associated constants
Imported from Core like the integer constants:
| Constant | Meaning | float32 | float64 |
|---|---|---|---|
Bits | Width in bits (uint) | 32 | 64 |
Bytes | Storage size in bytes (uint) | 4 | 8 |
Max | Largest finite value | 3.4028235e+38 | 1.7976931348623157e+308 |
Lowest | Most negative finite value, -Max | -3.4028235e+38 | -1.7976931348623157e+308 |
MinPositive | Smallest positive normal value | 1.1754944e-38 | 2.2250738585072014e-308 |
Epsilon | Gap between 1.0 and the next value up | 1.1920929e-07 | 2.220446049250313e-16 |
Infinity | Positive infinity | Inf | Inf |
NaN | A quiet not-a-number | NaN | NaN |
import Core::float64;
import Io::PrintLine;
func Main() -> int {
PrintLine("{} {}", float64::Max, float64::Epsilon);
PrintLine("{} {}", float64::Infinity, -float64::Infinity); // Inf -Inf
return 0;
}
Literals
A floating-point literal has digits on both sides of a point, an exponent, or a float suffix: 3.14, 2.998e8, 1.6e-19, 0.5f32, 1f32. Without a suffix it is a float64, whatever the context:
let single: float32 = 0.75;
error: cannot assign 'float64' to 'float32'
Write 0.75f32. Literals gives the full grammar.
Operations
| Operators | Result | Rules |
|---|---|---|
+ - * / | The operand type | IEEE 754, rounded to nearest |
% | The operand type | Remainder of truncated division: 7.5 % 2.0 is 1.5 |
prefix - | The operand type | Flips the sign, including of zero and infinity |
== != < <= > >= | bool | IEEE 754 comparison |
Floating-point arithmetic never panics. Results follow IEEE 754:
| Expression | Result |
|---|---|
1.0 / 0.0 | Inf |
-1.0 / 0.0 | -Inf |
0.0 / 0.0 | NaN |
A NaN is unordered: every comparison with it is false except !=, so x != x is true exactly when x is a NaN. The Core functions IsNaN, IsInfinite, IsFinite, IsZero and IsNegativeZero test for special values.
The bitwise and shift operators do not apply to floats: a & 1.0 is error: operator '&' requires an integer, bool, or character left operand, but found 'float64'.
Most decimal fractions have no exact binary form, so arithmetic on them rounds: 0.1 + 0.2 == 0.3 is false. Compare against a tolerance, or count in integer units such as cents.
Mixed widths
A float32 operand meets a float64 operand at float64. Because an unsuffixed literal is a float64, single * 2.0 is a float64 too; write single * 2.0f32 to stay in float32.
Conversions
float32 converts to float64 implicitly; every float32 value is exactly a float64 value. Everything else is written with as:
| Conversion | Result |
|---|---|
float64 to float32 | Rounded to the nearest float32; a value beyond its range becomes an infinity |
| To an integer type | Truncated toward zero; beyond the range, saturated at Min or Max; NaN becomes 0 |
| From an integer type | The nearest representable value |
| To a boolean type | false for zero, true for anything else, NaN included |
let ratio = 3.9;
let whole = ratio as int32; // 3
let down = -3.9 as int32; // -3
let capped = 1e10 as int32; // 2147483647
let floor = -1.0 as uint8; // 0
let single = 3.141592653589793 as float32; // 3.1415927
let back = 7 as float64; // 7.0
Float-to-integer conversion is the same on every target and whether it runs or is folded at compile time. It never wraps: 1e20 as int32 is int32::Max, not its low bits. Round first, with the Math package, when rounding rather than truncation is meant.
See also
- Integers — the other numeric family
- Arithmetic and Casts
- Float and Float special — lessons on precision, infinities and NaN
- Math — rounding, powers and roots