Per-binding throughput benchmarks + test-coverage gaps (#246)

Adds a `throughput` benchmark to every target and closes two small
test-coverage documentation/QA gaps. One PR, no merge of binding code beyond
the additive benchmarks and one C test.

## 1. Per-binding throughput benchmarks (all 9 targets)

Each benchmark feeds a deterministic synthetic OHLCV series through three
indicators chosen by **FFI call-signature archetype** (not algorithm — the same
Rust core runs underneath all bindings):

- `SMA(20)` — 1-in → 1-out (baseline boundary cost)
- `ATR(14)` — multi-in → 1-out (input marshalling)
- `MACD(12,26,9)` — 1-in → multi-out (output marshalling)

Streaming is timed for all three; batch for the single-output SMA and ATR
(median of 3 runs, after a warmup pass).

New: Python (PyO3), WASM, C (CMake), C# (Stopwatch), Go, Java (FFM), R, and the
Rust core baseline (`examples/rust/.../throughput.rs`, **no FFI** — the ceiling
the bindings are measured against and the value their batch paths converge
towards). Node already had `throughput.js`.

**Not a speed claim:** there is no comparable streaming TA library for C, C#,
Go, Java, R or WASM to compare against, so these are raw per-binding throughput
numbers documenting each language's FFI overhead — see BENCHMARKS.md §3. The
"Wickra is fast" claim still lives in §1/§2 (Rust core + the Python/Rust
cross-library runs).

## 2. README `## Testing`: C# and C bullets

The section listed every layer except C# and C, even though both have suites.
Adds the two missing bullets.

## 3. C archetype ctest

`examples/c/archetypes.c` drives one indicator per FFI archetype through the
real C boundary (scalar + batch==streaming, multi-output, bars, profile, array
input) plus reset, invalid-parameter and NULL-safety — the C counterpart of the
Go/R/Java archetype suites. Runs on three OSes via the existing CMake/ctest.

## Notes

- Benchmarks are not CI-gated (manual-run scripts, like the existing
  `throughput.js`); no `ci.yml`/`release.yml` changes.
- Docs: BENCHMARKS.md §3, a `## Benchmark` section in every binding README, a
  CHANGELOG entry.
- Verified locally by running: Rust, Python, C, C#, Go, Java (real numbers); the
  C archetype ctest with `-Wall -Wextra -Wpedantic -Werror`. WASM and R are
  API-correct and syntax-checked but need their own toolchains to run.
This commit is contained in:
kingchenc
2026-06-10 03:46:38 +02:00
committed by GitHub
parent b31a9a3624
commit 3ebcb3f758
25 changed files with 1529 additions and 0 deletions
+13
View File
@@ -76,6 +76,19 @@ values — the equivalence is enforced by the test suite. Multi-output indicator
indicator owns a native handle freed by a `Cleaner`; `close()` releases it
eagerly (use try-with-resources).
## Benchmark
`benchmarks/` reports streaming and batch updates-per-second for `SMA`, `ATR`
and `MACD`. It measures this binding's FFI overhead, not a cross-library ratio
(the same Rust core runs under every binding) — see the repository
[BENCHMARKS.md](https://github.com/wickra-lib/wickra/blob/main/BENCHMARKS.md) §3.
```bash
cargo build -p wickra-c --release
mvn -q install -DskipTests
mvn -q -f benchmarks exec:exec -Dexec.mainClass=org.wickra.benchmarks.Throughput
```
## Documentation
The full indicator catalogue, guides, quickstarts, and API reference live in
+57
View File
@@ -0,0 +1,57 @@
<?xml version="1.0" encoding="UTF-8"?>
<project xmlns="http://maven.apache.org/POM/4.0.0"
xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"
xsi:schemaLocation="http://maven.apache.org/POM/4.0.0 http://maven.apache.org/xsd/maven-4.0.0.xsd">
<modelVersion>4.0.0</modelVersion>
<groupId>org.wickra.benchmarks</groupId>
<artifactId>wickra-benchmarks</artifactId>
<version>0.8.2</version>
<packaging>jar</packaging>
<name>Wickra Java benchmarks</name>
<description>Throughput benchmark for the Wickra Java binding.</description>
<properties>
<project.build.sourceEncoding>UTF-8</project.build.sourceEncoding>
<maven.compiler.release>22</maven.compiler.release>
</properties>
<dependencies>
<dependency>
<groupId>org.wickra</groupId>
<artifactId>wickra</artifactId>
<version>0.8.2</version>
</dependency>
</dependencies>
<build>
<plugins>
<plugin>
<groupId>org.apache.maven.plugins</groupId>
<artifactId>maven-compiler-plugin</artifactId>
<version>3.13.0</version>
</plugin>
<!--
Run the benchmark in a forked JVM with native access granted:
mvn exec:exec -Dexec.mainClass=org.wickra.benchmarks.Throughput
Requires the C ABI library and the installed binding; see Throughput.java.
-->
<plugin>
<groupId>org.codehaus.mojo</groupId>
<artifactId>exec-maven-plugin</artifactId>
<version>3.2.0</version>
<configuration>
<executable>${java.home}/bin/java</executable>
<arguments>
<argument>--enable-native-access=ALL-UNNAMED</argument>
<argument>-classpath</argument>
<classpath/>
<argument>${exec.mainClass}</argument>
</arguments>
</configuration>
</plugin>
</plugins>
</build>
</project>
@@ -0,0 +1,149 @@
package org.wickra.benchmarks;
import java.util.Arrays;
import java.util.Locale;
import org.wickra.Atr;
import org.wickra.MacdIndicator;
import org.wickra.Sma;
/**
* Throughput benchmark for the Wickra Java binding.
*
* <p>Measures how many indicator updates per second the binding sustains, both
* per-tick (streaming {@code update}) and bulk ({@code batch}), over a synthetic
* OHLCV series. It is the Java counterpart of the Node {@code throughput.js} and
* the Rust criterion benches: it benchmarks Wickra's own O(1) streaming engine
* across the Java FFM &lt;-&gt; C-ABI boundary (there is no comparable streaming
* TA library on Maven Central to compare against), so the headline number is raw
* per-binding throughput / FFI overhead, not a cross-library ratio.
*
* <p>Three indicators are timed, chosen by FFI call-signature archetype rather
* than algorithm: SMA (1-in -&gt; 1-out), ATR (multi-in -&gt; 1-out), and MACD
* (1-in -&gt; multi-out). Streaming is timed for all three; batch only for the
* single-output SMA and ATR (multi-output batch is not exposed uniformly).
*
* <p>Install the binding and build the C ABI library first, then run from the
* repo root:
*
* <pre>
* cargo build -p wickra-c --release
* mvn -q -f bindings/java install -DskipTests
* mvn -q -f bindings/java/benchmarks exec:exec -Dexec.mainClass=org.wickra.benchmarks.Throughput
* </pre>
*/
public final class Throughput {
private Throughput() {}
public static void main(String[] args) {
int bars = 200_000;
for (int i = 0; i < args.length - 1; i++) {
if (args[i].equals("--bars")) {
try {
int n = Integer.parseInt(args[i + 1]);
if (n >= 1000) {
bars = n;
}
} catch (NumberFormatException ignored) {
// keep default
}
}
}
// Deterministic synthetic OHLCV (no RNG, so runs are comparable).
double[] open = new double[bars];
double[] high = new double[bars];
double[] low = new double[bars];
double[] close = new double[bars];
double[] volume = new double[bars];
double[] timestamp = new double[bars];
for (int i = 0; i < bars; i++) {
double mid = 100 + Math.sin(i * 0.001) * 20 + i * 1e-4;
double c = mid + Math.sin(i * 0.05) * 2;
close[i] = c;
open[i] = mid;
high[i] = Math.max(c, mid) + 1.5;
low[i] = Math.min(c, mid) - 1.5;
volume[i] = 1000 + (i % 97) * 13;
timestamp[i] = i;
}
final int n = bars;
// SMA (scalar 1-in/1-out), ATR (multi-in/1-out), MACD (1-in/multi-out).
Indicator[] indicators = {
new Indicator("SMA(20)",
() -> {
try (Sma ind = new Sma(20)) {
for (int i = 0; i < n; i++) {
ind.update(close[i]);
}
}
},
() -> {
try (Sma ind = new Sma(20)) {
ind.batch(close);
}
}),
new Indicator("ATR(14)",
() -> {
try (Atr ind = new Atr(14)) {
for (int i = 0; i < n; i++) {
ind.update(open[i], high[i], low[i], close[i], volume[i], (long) timestamp[i]);
}
}
},
() -> {
try (Atr ind = new Atr(14)) {
ind.batch(open, high, low, close, volume, timestamp);
}
}),
new Indicator("MACD(12,26,9)",
() -> {
try (MacdIndicator ind = new MacdIndicator(12, 26, 9)) {
for (int i = 0; i < n; i++) {
ind.update(close[i]);
}
}
},
null), // multi-output: streaming only
};
System.out.printf(Locale.ROOT, "Wickra Java throughput - %,d bars (median of 3 runs)%n%n", bars);
System.out.printf(Locale.ROOT, "%-22s%20s%18s%n", "Indicator", "streaming (Mupd/s)", "batch (Mupd/s)");
System.out.println("------------------------------------------------------------");
for (Indicator ind : indicators) {
String streamMups = String.format(Locale.ROOT, "%.1f", mups(bars, timeNs(ind.stream)));
String batchMups = ind.batch == null
? "-"
: String.format(Locale.ROOT, "%.1f", mups(bars, timeNs(ind.batch)));
System.out.printf(Locale.ROOT, "%-22s%20s%18s%n", ind.name, streamMups, batchMups);
}
System.out.println(
"\nMupd/s = million indicator updates per second. Streaming is the per-tick\n"
+ "update path crossing the Java FFM<->C-ABI boundary once per value; batch is\n"
+ "the bulk array path (one boundary crossing). Higher is better. Numbers are\n"
+ "machine-dependent - use them for relative comparison, not as a speed claim.");
}
private static double mups(int bars, double ns) {
return bars / (ns / 1e9) / 1e6;
}
// Median elapsed-ns over a few repetitions, after one warmup pass.
private static double timeNs(Runnable fn) {
fn.run(); // warmup (JIT + cache)
final int reps = 3;
double[] samples = new double[reps];
for (int r = 0; r < reps; r++) {
long t0 = System.nanoTime();
fn.run();
samples[r] = System.nanoTime() - t0;
}
Arrays.sort(samples);
return samples[reps / 2];
}
private record Indicator(String name, Runnable stream, Runnable batch) {}
}