Files
wickra/BENCHMARKS.md
T
kingchencandGitHub 3ebcb3f758 Per-binding throughput benchmarks + test-coverage gaps (#246)
Adds a `throughput` benchmark to every target and closes two small
test-coverage documentation/QA gaps. One PR, no merge of binding code beyond
the additive benchmarks and one C test.

## 1. Per-binding throughput benchmarks (all 9 targets)

Each benchmark feeds a deterministic synthetic OHLCV series through three
indicators chosen by **FFI call-signature archetype** (not algorithm — the same
Rust core runs underneath all bindings):

- `SMA(20)` — 1-in → 1-out (baseline boundary cost)
- `ATR(14)` — multi-in → 1-out (input marshalling)
- `MACD(12,26,9)` — 1-in → multi-out (output marshalling)

Streaming is timed for all three; batch for the single-output SMA and ATR
(median of 3 runs, after a warmup pass).

New: Python (PyO3), WASM, C (CMake), C# (Stopwatch), Go, Java (FFM), R, and the
Rust core baseline (`examples/rust/.../throughput.rs`, **no FFI** — the ceiling
the bindings are measured against and the value their batch paths converge
towards). Node already had `throughput.js`.

**Not a speed claim:** there is no comparable streaming TA library for C, C#,
Go, Java, R or WASM to compare against, so these are raw per-binding throughput
numbers documenting each language's FFI overhead — see BENCHMARKS.md §3. The
"Wickra is fast" claim still lives in §1/§2 (Rust core + the Python/Rust
cross-library runs).

## 2. README `## Testing`: C# and C bullets

The section listed every layer except C# and C, even though both have suites.
Adds the two missing bullets.

## 3. C archetype ctest

`examples/c/archetypes.c` drives one indicator per FFI archetype through the
real C boundary (scalar + batch==streaming, multi-output, bars, profile, array
input) plus reset, invalid-parameter and NULL-safety — the C counterpart of the
Go/R/Java archetype suites. Runs on three OSes via the existing CMake/ctest.

## Notes

- Benchmarks are not CI-gated (manual-run scripts, like the existing
  `throughput.js`); no `ci.yml`/`release.yml` changes.
- Docs: BENCHMARKS.md §3, a `## Benchmark` section in every binding README, a
  CHANGELOG entry.
- Verified locally by running: Rust, Python, C, C#, Go, Java (real numbers); the
  C archetype ctest with `-Wall -Wextra -Wpedantic -Werror`. WASM and R are
  API-correct and syntax-checked but need their own toolchains to run.
2026-06-10 03:46:38 +02:00

169 lines
9.0 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Benchmarks
Read these as **relative** speedups on identical input — absolute µs depend on
CPU, memory clock and OS scheduler, not a universal contract. **Streaming is the
headline**: it is where Wickra's design pays off and where the gap is measured in
orders of magnitude, not percent. The batch numbers come second and are shown
honestly — the leanest crates edge Wickra out on the simple recurrences, and that
is a deliberate trade for warmup/NaN semantics, not a ceiling.
- **Reproduced on:** Windows 11 Pro 26200, AMD Ryzen 9 9950X, 64 GB DDR5,
Rust 1.92 (release: `lto = "fat"`, `codegen-units = 1`), Python 3.12.
- **Reproduce yourself:**
- Rust core vs Rust crates: `cargo bench -p wickra-bench`
- Python vs Python libs: `pip install -e bindings/python[bench]` then
`python -m benchmarks.compare_libraries` (auto-detects installed peers).
## 1. Streaming — the structural win
Live trading feeds one tick at a time. Wickra updates every indicator in **O(1)**;
batch-only libraries (TA-Lib, tulipy, finta, pandas-ta) have no incremental API
and must recompute the whole history on every tick. Only `talipp` (Python) and
`ta-rs` / `yata` (Rust) carry real per-tick state. This is the gap the library
was built to expose.
**Python — per-tick latency** (seed 5 000 bars, then feed ticks one at a time):
| Indicator | **★ Wickra** | talipp | TA-Lib (recompute) |
|------------------|------------------:|------------------|-----------------------|
| SMA(20) | **0.089 µs ★** | 0.96 µs (11×) | 422 µs (4 700×) |
| EMA(20) | **0.111 µs ★** | 1.19 µs (11×) | 430 µs (3 900×) |
| RSI(14) | **0.061 µs ★** | 0.95 µs (16×) | 298 µs (4 900×) |
| MACD(12, 26, 9) | **0.079 µs ★** | 3.30 µs (42×) | 327 µs (4 100×) |
| Bollinger(20, 2) | **0.089 µs ★** | 4.97 µs (56×) | 296 µs (3 300×) |
Against the only other incremental Python peer Wickra is **1156× faster**;
against the recompute-on-every-tick libraries it is **2 80019 000× faster**
(`finta` RSI hits 19 000×). tulipy / pandas-ta land in the same recompute band
as TA-Lib.
**Rust — per-tick latency** (whole 50 000-bar series, lower = faster):
| Indicator | **★ Wickra** | kand | ta-rs | yata |
|------------------|------------------:|-----:|------:|-----:|
| SMA(20) | 50 | 38 | 47 | 38 |
| EMA(20) | 154 | 69 | 56 | 69 |
| RSI(14) | 164 | 216 | 74 | — |
| MACD(12, 26, 9) | 275 | 143 | 66 | — |
| Bollinger(20, 2) | **128 ★** | 248 | 168 | — |
| ATR(14) | 152 | 166 | 61 | — |
`ta-rs` hands back a bare `f64` from the first tick with no warmup and no
validation; it leads several rows by giving those guarantees up. Against `kand`,
Wickra wins streaming RSI, Bollinger and ATR. `yata` exposes only SMA/EMA as
raw-value methods, so its other rows are omitted rather than faked.
## 2. Batch — competitive, not the headline
Whole series in one call. Here hand-tuned C (`tulipy`, TA-Lib) and the leanest
Rust crate (`kand`) win the simple recurrences — Wickra trades a few µs per pass
for the `None`-warmup, NaN-safety and bit-exact `batch == streaming` guarantees
none of them keep. It still wins several rows outright and beats the rest of the
field everywhere.
**Python** (20 000-bar pass, µs/op, lower = faster):
| Indicator | Wickra | TA-Lib | tulipy | pandas-ta | finta |
|------------------|---------:|---------:|---------:|----------:|---------:|
| SMA(20) | 22.2 | **15.6** | 15.9 | 32.7 | 290.1 |
| EMA(20) | 30.5 | **30.4** | 30.9 | 46.7 | 198.5 |
| RSI(14) | 52.3 | 72.0 | **34.2** | 88.8 | 812.3 |
| MACD(12, 26, 9) | 129.8 | 111.1 | **38.4** | 286.8 | 716.7 |
| Bollinger(20, 2) | 87.2 | 74.6 | **37.9** | 474.3 | 1255.5 |
| ATR(14) | 74.7 | 87.3 | **35.5** | — | 3496.4 |
Wickra beats pandas-ta and finta on every row and TA-Lib on RSI and ATR;
tulipy's SIMD C (and TA-Lib on SMA/EMA) lead the remaining rows.
**Rust** (50 000-bar pass, µs, lower = faster). Only Wickra and `kand` expose a
batch API; `ta-rs` and `yata` are streaming-only:
| Indicator | **★ Wickra** | kand |
|------------------|------------------:|-------:|
| SMA(20) | 53 | **41** |
| EMA(20) | 111 | **71** |
| RSI(14) | **221 ★** | 259 |
| MACD(12, 26, 9) | 533 | **327** |
| Bollinger(20, 2) | **404 ★** | 460 |
| ATR(14) | **122 ★** | 169 |
Run the suite yourself:
```bash
cargo bench -p wickra-bench # Rust core vs kand / ta-rs / yata
pip install -e bindings/python[bench] # Python peers
python -m benchmarks.compare_libraries
```
## 3. Per-binding throughput — the cost of the boundary
The sections above compare Wickra against other libraries, which only exists for
Python and Rust (there is no comparable streaming TA library for C, C#, Go, Java,
R or WebAssembly to benchmark against). Every binding calls the **same** Rust
core, so these per-binding benchmarks are **not** a speed claim and **not** a
cross-library ratio — they document the raw cost of crossing each language's FFI
boundary, in million updates per second (Mupd/s).
Each binding ships a small `throughput` benchmark that feeds a synthetic OHLCV
series through three indicators chosen by call-signature archetype — `SMA(20)`
(1-in → 1-out), `ATR(14)` (multi-in → 1-out) and `MACD(12,26,9)` (1-in →
multi-out). Two things fall out of the numbers:
- **Batch converges.** A `batch` call crosses the boundary once and the Rust core
computes the whole series internally, so batch throughput is roughly the same
in every binding — close to the core speed.
- **Streaming reveals the boundary.** A per-tick `update` crosses the boundary
once per value, so streaming throughput is where the bindings differ: the raw C
ABI and P/Invoke-style calls are nearly free, while managed or interpreted
per-call marshalling (cgo, FFM, the R/WASM boundary) costs more per tick.
The Rust core ships the same benchmark with **no** FFI boundary
(`examples/rust/.../throughput.rs`) — it is the ceiling each binding is measured
against and the value the batch paths converge towards.
`SMA(20)`, 200 000 bars, median of 3 runs, on the reference machine (Windows 11,
AMD Ryzen 9 9950X):
| Target | streaming (Mupd/s) | batch (Mupd/s) |
|----------------------|-------------------:|---------------:|
| Rust core (no FFI) | 391 | 500 |
| C | 383 | 330 |
| C# / .NET | 337 | 244 |
| Python | 33 | 488 |
| Java | 28 | 175 |
| Go | 24 | 400 |
| WebAssembly | 19 | 167 |
| Node.js | 17 | 10 |
| R | 0.1 | 193 |
Streaming spans more than three orders of magnitude — the raw C ABI (383) is
nearly the FFI-free Rust ceiling (391), while R's per-call interpreter overhead
(0.1) makes streaming ~2000× slower than its own batch. Batch converges near the
core speed for the zero-copy bindings (numpy, slices, typed arrays); the two
outliers are Node — whose napi `batch` boxes every element into a JS `Array`
and R. These are machine-dependent and reflect FFI overhead, not algorithm
speed.
These are throughput numbers, not competitive numbers — the "Wickra is fast"
claim lives in sections 1 and 2 (Rust core + the Python/Rust cross-library runs).
Run any target's benchmark (build the C ABI library first where it links one):
```bash
cargo run -p wickra-examples --release --bin throughput # Rust core baseline (no FFI)
node bindings/node/benchmarks/throughput.js # native napi-rs
( cd bindings/python && python -m benchmarks.throughput ) # native PyO3
( cd bindings/wasm && wasm-pack build --target nodejs --out-dir pkg-node --release ) \
&& node bindings/wasm/benchmarks/throughput.mjs # wasm boundary
cargo build -p wickra-c --release # the C ABI hub
cmake -S bindings/c/benchmarks -B build/cbench && cmake --build build/cbench \
&& ./build/cbench/throughput # raw C ABI
dotnet run -c Release --project bindings/csharp/benchmarks # C# / .NET (P/Invoke)
( cd bindings/go/benchmarks && go run . ) # Go (cgo)
mvn -q -f bindings/java install -DskipTests \
&& mvn -q -f bindings/java/benchmarks exec:exec -Dexec.mainClass=org.wickra.benchmarks.Throughput
Rscript bindings/r/benchmarks/throughput.R # R (.Call)
```