Floating-Point

A floating-point type holds a binary fraction with an exponent, covering a vast range at a fixed number of significant digits. Rux implements the two IEEE 754 formats every supported machine has in hardware.

The floating-point types

TypeBitsBytesFormatSignificant digitsLargest finite valueSuffix
float32324IEEE 754 binary32about 7≈ 3.40 × 1038f32
float64648IEEE 754 binary64about 16≈ 1.80 × 10308f64

float is a built-in alias of float64, and float64 is the type of an unsuffixed floating-point literal. Use float32 where storage or bandwidth matters more than precision, such as large arrays of samples, and float64 otherwise.

Reserved widths

These widths are named by the language but not implemented in rux 0.4.0:

TypeBitsBytesIntended formatSuffix
float881E4M3 (4 exponent, 3 significand bits)f8
float16162IEEE 754 binary16f16
float808016x87 extended precisionf80
float12812816IEEE 754 binary128f128
float25625632IEEE 754 binary256 interchangef256
float51251264IEEE 754 binary512 interchangef512

Using one, as a type or through its suffix, is an error:

error: primitive type 'float16' is reserved but is not implemented in this compiler version

Associated constants

Imported from Core like the integer constants:

ConstantMeaningfloat32float64
BitsWidth in bits (uint)3264
BytesStorage size in bytes (uint)48
MaxLargest finite value3.4028235e+381.7976931348623157e+308
LowestMost negative finite value, -Max-3.4028235e+38-1.7976931348623157e+308
MinPositiveSmallest positive normal value1.1754944e-382.2250738585072014e-308
EpsilonGap between 1.0 and the next value up1.1920929e-072.220446049250313e-16
InfinityPositive infinityInfInf
NaNA quiet not-a-numberNaNNaN
import Core::float64;
import Io::PrintLine;

func Main() -> int {
    PrintLine("{} {}", float64::Max, float64::Epsilon);
    PrintLine("{} {}", float64::Infinity, -float64::Infinity);   // Inf -Inf
    return 0;
}

Literals

A floating-point literal has digits on both sides of a point, an exponent, or a float suffix: 3.14, 2.998e8, 1.6e-19, 0.5f32, 1f32. Without a suffix it is a float64, whatever the context:

let single: float32 = 0.75;
error: cannot assign 'float64' to 'float32'

Write 0.75f32. Literals gives the full grammar.

Operations

OperatorsResultRules
+ - * /The operand typeIEEE 754, rounded to nearest
%The operand typeRemainder of truncated division: 7.5 % 2.0 is 1.5
prefix -The operand typeFlips the sign, including of zero and infinity
== != < <= > >=boolIEEE 754 comparison

Floating-point arithmetic never panics. Results follow IEEE 754:

ExpressionResult
1.0 / 0.0Inf
-1.0 / 0.0-Inf
0.0 / 0.0NaN

A NaN is unordered: every comparison with it is false except !=, so x != x is true exactly when x is a NaN. The Core functions IsNaN, IsInfinite, IsFinite, IsZero and IsNegativeZero test for special values.

The bitwise and shift operators do not apply to floats: a & 1.0 is error: operator '&' requires an integer, bool, or character left operand, but found 'float64'.

Most decimal fractions have no exact binary form, so arithmetic on them rounds: 0.1 + 0.2 == 0.3 is false. Compare against a tolerance, or count in integer units such as cents.

Mixed widths

A float32 operand meets a float64 operand at float64. Because an unsuffixed literal is a float64, single * 2.0 is a float64 too; write single * 2.0f32 to stay in float32.

Conversions

float32 converts to float64 implicitly; every float32 value is exactly a float64 value. Everything else is written with as:

ConversionResult
float64 to float32Rounded to the nearest float32; a value beyond its range becomes an infinity
To an integer typeTruncated toward zero; beyond the range, saturated at Min or Max; NaN becomes 0
From an integer typeThe nearest representable value
To a boolean typefalse for zero, true for anything else, NaN included
let ratio = 3.9;
let whole = ratio as int32;              // 3
let down = -3.9 as int32;                // -3
let capped = 1e10 as int32;              // 2147483647
let floor = -1.0 as uint8;               // 0
let single = 3.141592653589793 as float32;   // 3.1415927
let back = 7 as float64;                 // 7.0

Float-to-integer conversion is the same on every target and whether it runs or is folded at compile time. It never wraps: 1e20 as int32 is int32::Max, not its low bits. Round first, with the Math package, when rounding rather than truncation is meant.

See also