update design/requirements/tasks for arbitrary arithmetic based on rational numbers

This commit is contained in:
Emil Lerch 2026-07-27 18:25:23 -07:00
parent 53cf1bf827
commit 690fa18e90
Signed by: lobo
GPG key ID: A7B62D657EF764F8
3 changed files with 417 additions and 1 deletions

View file

@ -240,6 +240,335 @@ All frontends:
---
### 2.7 Numeric Model: Exact and Inexact
#### 2.7.1 The problem
Standard mode currently evaluates everything as `f64`. That produces four
distinct classes of wrong answer, all reproducible today:
```
0.1 + 0.2 -> 0.30000000000000004 binary cannot represent 0.1
1.1 + 2.2 -> 3.3000000000000003 same
9007199254740993 -> 9.007199254740992e15 SILENTLY A DIFFERENT INTEGER
2^53 + 1 -> 9.007199254740992e15 same, past f64's integer limit
12 in in ft -> 0.9999999999999998 conversion factors are decimal
factorial(171) -> error: unknown function f64 overflow, wrong error too
```
The third case is the serious one. The user typed an exact integer and got a
different integer back. Meanwhile programmer mode, one Tab away, is exact to 128
bits. That is an internal inconsistency, not a rounding preference.
#### 2.7.2 The representation
A tagged number with an exact tier and an inexact tier:
```zig
/// Exact rational: p/q with arbitrary-size integers, always reduced, q > 0.
pub const Rational = struct {
num: std.math.big.int.Managed,
den: std.math.big.int.Managed,
};
pub const Number = union(enum) {
/// Exact value. Closed under + - * / and integer powers.
exact: Rational,
/// Inexact value, produced once a computation escapes the rationals.
inexact: f64,
};
```
**Contagion rule:** any operation with an `inexact` operand yields `inexact`.
A value never silently becomes exact again. This is what keeps the guarantee
honest: `exact` means "no rounding has occurred anywhere in this value's
history", not "happens to look clean right now".
#### 2.7.3 Why rational and not decimal128
Both fix `0.1 + 0.2`. They are not otherwise equivalent:
| | decimal128 | arbitrary rational |
|---|---|---|
| Form | +/- 34-digit significand x 10^exp, 128 bits fixed | p/q, arbitrary-size integers |
| `0.1`, `0.0254` | exact | exact (`1/10`, `127/5000`) |
| `1/3` | 34-digit approximation | **exact** |
| `(1/3)*3` | `0.999...9` | **exactly 1** |
| `9007199254740993` | exact | exact |
| `sqrt(2)`, `pi` | approximation | **cannot represent** |
| Rounds | every operation | never |
| Memory | fixed 16 bytes | grows with denominators |
| Trailing zeros (`2.50`) | **preserved** (scale is part of the value) | lost |
Two conclusions:
1. **Every decimal128 value is a rational.** Rationals strictly subsume them.
What decimal128 adds is fixed cost, standard rounding modes, and scale
semantics, not additional expressiveness.
2. **decimal128 does not deliver correctness, it delivers familiarity.**
`1/3 * 3` is still not 1. Since we are prioritizing correctness, rational
wins outright.
Decimal's genuine advantage is scale preservation, which matters for money
(`2.50`, not `2.5`). We recover that in financial mode with an explicit display
scale plus banker's rounding at the commit boundary, rather than a second
numeric type. See section 5.
#### 2.7.4 The exact/inexact boundary
The trigger is **whether the operation escapes the rationals**, evaluated per
call rather than per function. Notably it is *not* the input base: `0xdeadbeef`
is exactly 3735928559, and base is a lexical and display property only.
| Stays exact | Becomes inexact |
|---|---|
| `+ - * /` on exact operands | `sin cos tan asin acos atan` |
| `^` with integer exponent | `^` with non-integer exponent |
| `sqrt` of a perfect square | `sqrt` of a non-square |
| `%`, `abs`, `ceil`, `floor`, `round` | `log ln log2 log10 exp cbrt` |
| `factorial` (unbounded) | constants `pi`, `e`, `tau` |
| integer and decimal literals | any operand already inexact |
`sqrt(4)` is 2 exactly; `sqrt(2)` is not. We deliberately stop short of a CAS:
Qalculate simplifies `sqrt(8)` to `2*sqrt(2)` symbolically, we will not.
#### 2.7.5 Denominator growth and the demotion policy
Rationals grow. Chained division with coprime denominators multiplies them, so a
pathological input can consume unbounded memory. This must be designed in from
the start, not discovered later:
- Cap the denominator at a bit-width budget (starting proposal: 4096 bits).
- On exceeding it, **demote to `inexact`** rather than erroring or growing.
- Demotion is silent in the result value but visible in the tag, so a frontend
can indicate that exactness was lost.
Interactive single calculations will not approach this. The cap exists so the
failure mode is graceful degradation to today's behavior.
#### 2.7.6 Display
Exact values still have to be rendered at finite width, so display precision
remains a separate decision from compute precision:
- An exact value with a **terminating** decimal expansion prints exactly:
`1/10` -> `0.1`, and `12 in in ft` -> `1`.
- An exact value with a **non-terminating** expansion is rendered by long
division to a digit budget with correct rounding, and marked as approximate:
`100 km in mi` is exactly `781250/12573`, displayed as `62.137119224`.
- Because the exact form is retained, we can additionally offer the fraction
(`781250/12573`) as a display option. This is a feature, not a workaround.
This supersedes the unimplemented "exceeds 15 significant digits" clause of
NFR-7, which was never a precision decision and, read literally, would have
rendered `0.9999999999999998` as `9.999999999999998e-1`. NFR-7 needs a real
answer to a better-posed question: how many digits to show for a non-terminating
exact value.
#### 2.7.7 What is unaffected
- **Programmer mode** stays fixed-width `u128`. Wrapping and masking at a chosen
bit width is the entire point of that mode; exact arithmetic would break it.
- **The IEEE 754 float view** stays binary `f32`/`f64`. It exists to show binary
encodings.
- **The formatter's f64 paths** remain, since the inexact tier still needs them.
#### 2.7.8 Fallback precision: why f64 and not f128
The obvious objection to `f64` for the inexact tier is that the usual
justification (hardware instructions) does not apply to a calculator, which
evaluates one expression per human action. That objection is correct, and
performance is *not* the reason. The reason is that Zig 0.16's f128 support is
partial and, worse, silently inconsistent. Probed directly:
| Operation | f128 status |
|---|---|
| `@sqrt` | compiles, accurate to full f128 (34 digits verified) |
| `@sin`, `@cos` | compiles, accurate to full f128 |
| `@log`, `@log10`, `@exp` | compiles, but **only f64-accurate** |
| `std.math.pow` | **does not compile** ("pow not implemented for f128") |
| `std.math.cbrt` | **does not compile** |
| `std.math.atan`, `asin` | compiles |
| `parseFloat` / formatting | works, 112 mantissa bits (~34 decimal digits) |
The middle row is the disqualifying one. `@log(2.0)` as f128 returns
`0.6931471805599452862267639829951804`, which is the *f64* value widened: it
diverges from the true value at the 17th significant digit. So f128 storage
would hand us 34 digits of which roughly 18 are noise, with nothing to
distinguish them. For a calculator that then displays those digits, this is
actively worse than knowingly using f64.
The general principle: **the precision of the inexact tier is bounded by the
accuracy of the underlying transcendental implementations, not by the width of
the storage type.** Widening storage without widening the algorithms buys
nothing but false confidence.
##### Upstream status (checked, not assumed)
The f64-accuracy above is not a misconfiguration or a version quirk. It is
visible in Zig's own source, with its own TODO markers. `lib/compiler_rt/log.zig`:
```zig
pub fn logq(a: f128) callconv(.c) f128 {
// TODO: more correct implementation
return log(@floatCast(a));
}
```
A systematic scan of `lib/compiler_rt/` in Zig 0.16.0 splits the f128 routines
cleanly in two:
- **Downcast to f64, all five marked TODO:** `expq`, `exp2q`, `logq`, `log2q`,
`log10q`
- **Genuine f128 implementations:** `sinq`, `cosq`, `tanq`, `sqrtq`, `fabsq`,
`floorq`, `ceilq`, `roundq`, `truncq`
That is an exact match for the probe results: `@sin`/`@cos`/`@sqrt` accurate,
`@log`/`@log10`/`@exp` only f64-accurate.
Tracker status, as of this writing. Note that Zig has moved to Codeberg and
GitHub is effectively frozen (GitHub last pushed 2025-11-27; Codeberg shows ~742
open issues against GitHub's ~2876, so most GitHub issues were **not** migrated):
| Gap | Tracked? |
|---|---|
| `std.math.pow` f128 | [GH #23602](https://github.com/ziglang/zig/issues/23602) open, labeled "contributor friendly". [PR #23631](https://github.com/ziglang/zig/pull/23631) implemented it but was **closed unmerged** 2026-01-05 ("no CI on GitHub anymore... reopen on Codeberg"). Nobody has. No Codeberg issue exists. |
| `std.math.cbrt` f128 | **No issue on either tracker.** Unreported. `cbrt.zig` is ported from musl, which only ships `cbrtf`/`cbrt`. |
| `logq`/`expq` etc. f64 accuracy | **No issue on either tracker.** Only the in-source TODOs. |
The nearest historical artifacts are [GH #4026](https://github.com/ziglang/zig/issues/4026)
("implement all the math functions for all the floating point types", closed
2022-12-28 under milestone 0.11.0 with its f128 column still unchecked) and
[GH #15711](https://github.com/ziglang/zig/issues/15711) (a reference table
created and closed in the same second, which lists the f128 symbols as
*existing* - existence, not accuracy).
Adjacent open work suggesting f128 remains a soft spot generally:
[GH #23865](https://github.com/ziglang/zig/pull/23865) "Better @sqrt for f32, f64
and f128", [GH #23173](https://github.com/ziglang/zig/issues/23173) "introduce new
float types", and Codeberg #36093 "compiler_rt: min(f128) / max(f128) causes
integer overflow".
**Conclusion: "wait for Zig to fix f128" is not a plan.** Two of the three gaps
are not even reported, and the one with a working patch lost it to a tracker
migration. Obtaining true f128 accuracy would mean writing or porting the
routines ourselves, which is the same category and scale of work as going
arbitrary-precision - so it should be justified on its own terms, not smuggled in
as "just use a wider type".
(Optional, off the critical path: reviving the `pow` PR on Codeberg is a
well-scoped upstream contribution, and the issue is explicitly labeled
contributor friendly.)
So `f64` is chosen because ~15-16 digits is the precision we can actually
justify, and we display no more than that. The upgrade path is therefore *not*
"switch the storage type to f128"; it is "implement or bind correctly-rounded
arbitrary-precision transcendentals" (the problem MPFR exists to solve), which is
a separate project with its own justification.
Incidentally, the probe also confirms that width alone never fixes the decimal
problem: in f128, `0.1 + 0.2` is
`0.30000000000000000000000000000000004`. Only exactness fixes it.
#### 2.7.9 Implementation sequencing
The change must not land as one diff mixing structural churn (signature changes
rippling through hundreds of test assertions) with behavioral change (results
becoming correct). Reviewing that combination is impractical: the interesting
lines are buried in mechanical ones.
Two orderings were considered.
**Rejected: structure first.** Introduce `Number`, rewrite every test and the
display path to consume it, review, then change behavior. This fails because the
structural rewrite is not behavior-neutral: rewriting a test to consume `Number`
requires asserting whether each result is `exact` or `inexact`, and that
classification *is* the behavioral design. Every such assertion would be written
against the old f64 behavior and then flipped in the later step. Double churn,
and the intermediate commit is a large diff that improves nothing.
**Chosen: behavior first, behind the existing signature.**
1. **Add `rational.zig` and `number.zig`, wired to nothing.** Pure addition with
their own tests. Reviewable in isolation, zero risk to existing behavior.
2. **Switch the evaluator's internals to the exact tier, keeping `evalString`
returning `f64`.** All existing tests must still pass. Review and commit.
3. **Change `evalString` and the display path to return and consume `Number`**,
with the resulting mechanical churn in test signatures. Review and commit.
4. **Add the exactness tests** that only the `Number` API can express.
What makes step 2 worthwhile rather than invisible is that converting an exact
value to `f64` *once, at the boundary* already fixes most of the payoff. Today's
errors are accumulated: `0.1` and `0.2` are each rounded before being added. An
exact `3/10` rounded a single time to the nearest f64 prints as `0.3`.
Bug-by-bug, where the fix lands:
| Bug | Fixed by step 2 (f64 boundary) | Needs step 3 (`Number`) |
|---|---|---|
| `0.1 + 0.2` -> `0.3` | yes | - |
| `1.1 + 2.2` -> `3.3` | yes | - |
| `0.1 * 3` -> `0.3` | yes | - |
| `12 in in ft` -> `1` | yes | - |
| `factorial(171)` | partly (becomes `inf`, not an error) | yes, for the exact value |
| `9007199254740993` | **no** (f64 cannot hold it) | yes |
| `2^53 + 1` | **no** | yes |
| `1/3` retained exactly | no (only observable as exact) | yes |
So step 2 is independently verifiable through the current API, and the
large-integer cases become the motivating tests for step 3 rather than
speculative ones.
#### 2.7.10 Migration and test impact
The dominant risk is **compile-time** breakage, not behavioral. An audit of the
402-test suite found **no test that asserts the broken behavior** - nothing
pins `0.30000000000000004`, the 2^53 cliff, or the `factorial` limit. So the
target is "no test changes its expected value", not "no test file is touched".
Unaffected by construction: `programmer.zig` (47), `float_interp.zig` (21),
`formatter.zig` (40), `types.zig` and `tui.zig` (16).
Kept passing by making the change additive:
- `tokenizer.zig` (39): `NumberValue` **gains** an exact field, does not replace
`float`/`int_value`.
- `parser.zig` (32, 10 of which read `number.float_value`): the AST number node
likewise gains a field.
- `evaluator.zig` (62, 40 asserting `@as(f64, ...)`): keep `evalString`
returning `f64` as a wrapper over the exact core. Every one of those
assertions involves small values where exact-to-f64 is bit-identical.
- `units.zig` (106, 95 routed through the `expectConvert` helper): change the
helper once, leave the call sites alone.
Three existing items must change because we are invalidating their premise:
1. `tokenizer.zig` test `"parseNumber huge decimal falls back to float"` -
`99999999999999999999999999` becomes exactly representable, so the test still
passes but its intent is now false. Rewrite it.
2. `formatter.zig` `is_integer`, gated on `< 2^53` - that guard is the display
half of the `9007199254740993` bug.
3. `evaluator.zig` `factorial`'s `x > 170` rejection, which is also the source of
the misleading `unknown function` error.
**Compatibility wrappers are scaffolding with a deadline, not an end state.** If
the CLI and TUI keep calling the `f64` API, the exactness never reaches the user
and the work is invisible. The display path must move to `Number` (step 3 of
2.7.9), and that is where `main.zig`'s ~20 output assertions get reviewed.
Literal zero test churn across the whole sequence would be evidence we had not
delivered anything.
Windows Calculator draws the same exact/inexact line we are drawing: its
arbitrary-precision rational engine covers basic arithmetic only, not
transcendentals.
New tests required: rational arithmetic and reduction, exact/inexact contagion,
the demotion policy at the denominator cap, rational-to-decimal rendering and
rounding, and a regression test per bug in 2.7.1 - including one asserting
`9007199254740993` round-trips, which nothing currently guards.
---
## 3. Struct Layout Engine
### 3.1 Mini-DSL Grammar

View file

@ -207,7 +207,7 @@ still required (FR-7.7): the mouse never becomes the only way to do something.
### NFR-7: Number Display & Formatting
- Decimal numbers must use comma grouping for display (e.g., `4,294,967,295`).
- Scientific notation only when value exceeds 15 significant digits or absolute value > 10^15 / < 10^-15. Never jump to scientific notation for values that fit in a readable decimal.
- Scientific notation only when absolute value > 10^15 / < 10^-15. Never jump to scientific notation for values that fit in a readable decimal. (An earlier draft also triggered scientific notation past 15 significant digits. That clause was never implemented and was wrong: read literally it renders `0.9999999999999998` as `9.999999999999998e-1`, which is worse. Superseded by NFR-9.)
- Programmer mode hex values display with space-separated bytes (e.g., `FF FF FF FF`).
- Programmer mode binary values display grouped by nibble with spaces (e.g., `1111 1111`).
- All frontends must distinguish between "display format" (with separators) and "clipboard format" (raw, no separators).
@ -220,3 +220,21 @@ still required (FR-7.7): the mouse never becomes the only way to do something.
- C API must not use null-terminated strings. All string inputs take `(pointer, length)` pairs to prevent buffer overread.
- Engine must validate all input lengths before processing.
- No unbounded allocations from untrusted input (enforce maximum expression length, maximum struct field count).
### NFR-9: Numeric Correctness (Exact Arithmetic)
Correctness is prioritized over evaluation speed. A calculator evaluates one
expression at a time in response to a human, so there is no hot-loop or
real-time constraint that justifies accepting wrong answers.
- **NFR-9.1**: Integer arithmetic must be exact and unbounded in standard mode. `9007199254740993` must round-trip, and `2^53 + 1` must not collapse to `2^53`. Silently returning a different integer than the user typed is a correctness bug, not a display artifact.
- **NFR-9.2**: Arithmetic on terminating decimal literals must be exact: `0.1 + 0.2` is `0.3`, and `1.1 + 2.2` is `3.3`.
- **NFR-9.3**: Division must be exact where the result is rational: `1/3` retains its exact value and `(1/3) * 3` is exactly `1`.
- **NFR-9.4**: Unit conversions between units whose factors are exact decimals must be exact: `12 in to ft` is exactly `1`.
- **NFR-9.5**: Exactness is tracked, not assumed. A value is exact only if no rounding occurred anywhere in its computation. Any operation involving an inexact operand yields an inexact result, and inexact values never silently become exact again.
- **NFR-9.6**: Operations that cannot produce a rational result (trigonometric, logarithmic, exponential, non-integer powers, roots of non-perfect-powers, and the constants `pi`/`e`/`tau`) fall back to floating point. This is a per-evaluation property: `sqrt(4)` stays exact, `sqrt(2)` does not.
- **NFR-9.7**: Input base must not affect the numeric path. `0xdeadbeef + 10` is exact integer arithmetic; base is a lexical and display concern only.
- **NFR-9.8**: Exact arithmetic must degrade gracefully rather than exhaust memory. Denominator size is capped, and exceeding the cap demotes the value to inexact rather than erroring.
- **NFR-9.9**: Display precision is a separate decision from compute precision. Exact values with terminating decimal expansions print exactly; non-terminating ones are rendered to a digit budget with correct rounding and marked approximate, with the exact fraction available.
- **NFR-9.10**: Programmer mode retains fixed-width `u128` wrapping semantics, and the IEEE 754 float view retains binary `f32`/`f64` semantics. Neither is affected by exact arithmetic.
- **NFR-9.11**: The inexact tier must not claim more precision than its operations can deliver. Displayed digits must be digits we can justify, so the storage width is chosen to match the accuracy of the available transcendental implementations rather than exceeding it (see design.md 2.7.8).

View file

@ -89,6 +89,75 @@ function arg commas).
## Phase 2: Struct Layout & Financial Engines
### Task 2.0: Exact numeric model (Rational + Number union) [NOT STARTED]
Design: see design.md 2.7. Requirements: NFR-9.
Motivation (all reproducible today): `0.1 + 0.2` = `0.30000000000000004`,
`9007199254740993` silently becomes `9.007199254740992e15`, `2^53 + 1` collapses,
`12 in in ft` = `0.9999999999999998`, `factorial(171)` reports "unknown function".
Sequenced as four reviewable commits so structural churn never mixes with
behavioral change (rationale in design.md 2.7.9):
#### 2.0a: Add the types, wired to nothing [DONE]
- `engine/src/rational.zig`: `Rational` over `std.math.big.int.Managed`, always
reduced, `q > 0`. add, sub, mul, div, negate, abs, `powInt` (negative exponents
take the exact reciprocal), `sqrtExact` (null for non-perfect-squares so the
caller can fall back), order/eql, `parseDecimal` (exact: `0.1` -> `1/10`),
`isTerminating`, `denBitCount`, `toFloat`, `toFractionString`,
`toDecimalString`. 55 tests.
- `engine/src/number.zig`: `Number = union(enum) { exact: Rational, inexact: f64 }`
with the contagion rule and the `max_denominator_bits = 4096` demotion cap.
30 tests.
- Exposed via `engine.zig` (two lines) so the tests run under `zig build test`.
Nothing else is wired: the evaluator is untouched.
- Notable implementation points:
- `toFloat` scales the numerator to ~72 significant bits before dividing, so
the result is a single correct rounding rather than a quotient of two
separately-rounded floats (which would lose precision twice and overflow for
large values).
- `toDecimalString` does long division, prints terminating expansions exactly,
and rounds half-up at the digit budget while reporting `exact = false`.
- `big.int`'s `sqrt` truncates, so `sqrtExact` squares the root back and
compares to decide exactness.
- Test gotcha worth remembering: Zig folds float literals at comptime as
`comptime_float`, so `0.1 + 0.2 != 0.3` is FALSE at comptime. Comparisons meant
to demonstrate f64 error must use runtime-known values (`var x: f64 = 0.1; _ = &x;`).
- Verify: 487 tests pass (was 402), zig fmt and zlint clean, existing tests
untouched.
#### 2.0b: Evaluator internals become exact, `evalString` still returns f64
- Exact/inexact boundary per design.md 2.7.4. Per-evaluation, not per-function:
`sqrt(4)` exact, `sqrt(2)` inexact. Deliberately not a CAS.
- `NumberValue` and the AST number node GAIN an exact field rather than replacing
`float`/`int_value`, so tokenizer and parser tests keep compiling.
- ALL existing tests must still pass unchanged.
- Payoff is visible through the f64 boundary because a single final rounding
avoids today's accumulated error: `0.1 + 0.2` -> `0.3`, `1.1 + 2.2` -> `3.3`,
`0.1 * 3` -> `0.3`, `12 in in ft` -> `1`.
- Verify: existing suite green, plus new tests for the four fixes above.
#### 2.0c: `evalString` and the display path move to `Number`
- Mechanical signature churn through evaluator/formatter/CLI/TUI tests. This is
the commit where `main.zig`'s output assertions get reviewed.
- Lift the limits this removes: `formatter.is_integer`'s `< 2^53` gate and
`evaluator.factorial`'s `x > 170` rejection (also the source of the misleading
"unknown function" error).
- Rewrite `tokenizer` test "parseNumber huge decimal falls back to float", whose
premise becomes false once big integers are exact.
- Verify: existing suite green after signature updates, no expected-value changes
beyond the three items above.
#### 2.0d: Exactness tests only the `Number` API can express
- `9007199254740993` round-trips (currently unguarded, and the one bug the f64
boundary cannot fix), `2^53 + 1`, `1/3` retained exactly, `(1/3) * 3` = 1,
unbounded `factorial`, contagion, demotion at the denominator cap.
Out of scope: arbitrary-precision transcendentals (the inexact tier is `f64`
because that is the precision we can justify; see design.md 2.7.8 for why f128 is
NOT the upgrade path), and any change to programmer mode's `u128` semantics or the
IEEE 754 float view.
### Task 2.1: Implement struct DSL tokenizer and parser
- Create `engine/src/struct_layout.zig`
- Tokenize: `struct`, `{`, `}`, `;`, type names, identifiers, `[`, `]`, numbers (for arrays)