11 Commits
Author SHA1 Message Date
Exocet92andGitHub 6c2f4eda9e bench: publish the results where they can be read (#12)
An artifact is not a publication. It needs a token to download, expires
after ninety days, and nothing outside GitHub can link to it, so a number
that only lives in an artifact is a number nobody can check.

A third job merges a green run's two payloads into
benchmarks/vs_vectorbt/results/latest.json and commits it. The two are
stored side by side rather than folded into one table: they run on two
runners, and timings from two machines are not rows of the same table.

Only a run where both measuring jobs came back green is published, and a
dispatch that pins an old version measures and reports without becoming
the published number.
2026-08-20 21:55:08 +02:00
Exocet92andGitHub 4a5b740ecb bench: stop manual runs from cancelling each other (#11)
`cancel-in-progress: true` is right for a CI triggered by pushes, where a newer
commit makes an older run pointless. It is wrong here. Three times in one
afternoon a second dispatch killed a run that was mid-measurement, once eight
minutes in, and each left a cancelled entry in a run list whose whole job is to
be readable by someone checking the numbers.

The fix people reach for then is deleting runs, which breaks the run_url the
published data links to, and with it the only reason to believe the numbers.

A manual run is never superseded: nobody dispatches a benchmark to invalidate
the one already running. A release or the weekly cron still supersedes, because
there an older run really is measuring a version nobody asks about any more.
2026-08-20 20:19:48 +02:00
Exocet92andGitHub 6cae686283 bench: add a cost workload and a multi-asset one, and go to three repetitions (#10)
Two gaps a reader could name without running anything: costs appeared on one
workload out of four, and nothing in the suite was a portfolio.

Costs could not simply be switched on across the board, and the reason is
measured. On FractionOfEquity sizing a 5 bps fee puts the engines 1.3e-4 of
capital apart and 2 bps of slippage 2.1e-5, against a 1e-9 tolerance, while the
round-trip counts stay identical: the trading agrees, the cost arithmetic does
not, because one charges the fee on top of the notional and the other reserves
it out of cash first. In fixed units both land exactly, to 1e-15. So
`sma_cross_costs` carries a fee and slippage on the headline signal, sized in
units, and the price of that is visible rather than hidden: x48.0 against x50.6
at 100k bars.

`multi_asset` runs five independent series in one shared book. It is the
workload manifoldbt does worst on, and it is here for that reason: going from
one asset to five costs it 6.1x and vectorbt 1.4x, so the ratio falls from x36.7
to x8.8 at a million bars. Broadcasting a column per asset is close to free;
walking five books is not. A portfolio is also what people actually run, and a
suite that only measures where it wins is not evidence.

Both are capped where a materialised five-column simulation would stop measuring
the engine and start measuring the swap file, and `ema_rsi_fees` keeps the
ceiling it got for going bankrupt.

Repetitions go from two to three: the floor at which a median is a median rather
than the mean of two.
2026-08-20 18:45:05 +02:00
Exocet92andGitHub 9ddefc64df bench: bigger sweep points, and call them sweeps (#9)
The three points were sized before this runner had ever run one. It has now,
so they are sized from what it measured: 87.5 us per combination for manifoldbt
at 20,000 bars, 1.16 ms for vectorbt, 1.34 ms for raptorbt, and 15.0 ms for
raptorbt at 200,000 bars.

20,000 x 5,000 stays, because it is the only one of the three vectorbt can hold:
it materialises 1.57 MB per combination at that length, so 5,000 already costs
it 2.5 GB. The other two grow to 20,000 and 10,000 combinations and put it out
of scope, which is where a sweep stops being a speed comparison and becomes a
capability one.

raptorbt sets the budget, not manifoldbt. With no fan-out API its sweep is a
Python loop costing a full backtest per cell, so the large point goes deep in
combinations on a short series rather than the reverse: 20,000 combinations on
20,000 bars costs it 27 s a call, where 5,000 combinations on a million bars
would cost it 25 minutes.

Also: sweeps, not grids. `run_sweep`, `run_sweep_lite` and `--sweep` are what
the product calls this, and a second word for the same thing is a second thing
to learn. `grid` is kept only where it means the parameter space itself.
2026-08-20 17:14:13 +02:00
Exocet92andGitHub df13224efc bench: run on Linux only, and name the jobs for what they measure (#8)
Windows and macOS were carried on an argument that does not survive
examination. They were the only place in the whole chain that installed the
published wheel and ran it, which made this benchmark an install smoke test by
accident. That check is worth having, and worth forty seconds next to the build
in release.yml rather than thirteen minutes inside a performance measurement:
nobody reads a benchmark to find out whether a package imports. release.yml
already builds on all four targets, it just never executes what it built.

What is lost is a per-platform timing that was never quoted; the numbers that
get published are the Linux ones. Adding a platform back is one matrix entry.

The jobs are also named for what they measure rather than for the runner they
landed on, which the row already says: `backtests (ubuntu-latest)` and `grids
(ubuntu-latest)` instead of a raw label next to a hand-written one.
2026-08-20 17:08:43 +02:00
Exocet92andGitHub a4040375e2 bench: fix the red runs, trim the matrix, add a grid job (#7)
Five consecutive red runs, two unrelated causes.

Four of them never reached an engine: the workflow installs the tag it is
handed, but 0.18.0rc1 and rc2 were previews that never reached PyPI.
Pre-releases are now skipped, and the report step checks for its input file
instead of dying on a missing one and reporting the wrong cause twice.

The fifth came from adding 10M bars, which broke a workload whose validity
depended on the ladder stopping at 5M. ema_rsi_fees sizes in fixed units and
pays 5 bps a side, so over 10M one-minute bars the fees compound into the whole
account: -15% of capital at 1M, -74% at 5M, exactly -100% at 10M, where fees
reach 99,611 of the 100,000 it started with. Both engines then sit at zero and
disagree by 9,085 round-trips about how many worthless trades to book on a dead
account. Workloads can now declare a ceiling, and the runner skips past it out
loud.

Fewer points per axis: three series lengths instead of five, a decade apart
each step. 10k measured the clock rather than the work, and 5M sat between two
points that already bracketed it. sma_cross crosses on 30/150 rather than
10/50, worth about 15% on the ratio because it books a third of the trades.

Grids get their own job, licensed through ci_activate.py, which refuses to run
unlicensed rather than time a wait. Three points, not a matrix: across the
plane the four-core ratio moves only between x32 and x38.
2026-08-20 17:00:41 +02:00
Exocet92andGitHub 52cbe1ba54 bench: add raptorbt as a third engine, and a 10M-bar point (#6)
The harness compared two engines everywhere; it now compares N against a
reference. manifoldbt is the reference: every parity check and every ratio is
a challenger against it, never two challengers against each other.

raptorbt 0.9.0 joins on three of the four workloads. Its sma_cross comes back
bit-identical to the reference's final equity, and its rsi matches to the last
bit; its ema seeds on a different warmup and it has no fixed-quantity sizing,
so the fee workload records it as unsupported with the reason rather than
leaving a blank cell. On the bracket it diverges in its own documented way: it
never re-arms while the entry level holds, so it books exactly the reference's
round-trips minus the ones that re-enter on the exit bar.

Python moves to 3.12, which raptorbt pins rather than we do: it is built
against pyo3 0.20.3, whose maximum supported CPython is 3.12. Timings from runs
before this change are therefore not directly comparable.

The bar matrix gains 10M and the repetition default drops from 7 to 2. Measured,
those two almost cancel: the job stays around 16 minutes. macOS keeps its old
ceiling, since 10M bars adds 1.55 GB on vectorbt's side alone and that runner
has 7 GB.
2026-08-20 16:16:33 +02:00
Exocet92andGitHub d9f1862fd9 bench: run the vectorbt comparison on public runners (#5)
A speed claim a reader cannot reproduce is a screenshot. This harness
installs manifoldbt from PyPI like any user would, generates its own data,
and gates every timing behind a parity check: a workload where the two
engines disagree publishes nothing and fails the run.

It lives here rather than in the engine repository because it benchmarks the
published wheel, not the source. Anyone can fork this repository and press
"Run workflow" to get the same table on their own runner.

The workflow runs on demand, weekly, and on every published release, so a
version that gets slower says so in public.
2026-08-18 02:49:20 +02:00
Exocet92 bb86eaf84c ci: consolidate workflow 2026-07-04 05:42:55 +02:00
Exocet92 fa67614ad8 ci: verify published package layout 2026-07-04 05:42:54 +02:00
Exocet92 3b0cf6c981 ci: fail on any Rust source in the public repo (source-leak safety net) 2026-07-04 04:34:44 +02:00