398 lines
16 KiB
Markdown
398 lines
16 KiB
Markdown
# Performance Guide
|
||
|
||
This document explains the performance characteristics of **ferro-ta** and gives
|
||
practical advice on how to get the best speed from the library.
|
||
|
||
---
|
||
|
||
## Quick Summary
|
||
|
||
| Use case | Recommended API | Notes |
|
||
|---------------------------------------|----------------------------------------|----------------------------------------|
|
||
| Fast path — NumPy arrays | Pass `np.ndarray` (float64, C-order) | Zero overhead; no conversion needed |
|
||
| pandas users | Pass `pd.Series`; result is `pd.Series`| Small overhead for index wrapping |
|
||
| polars users | Pass `pl.Series`; result is `pl.Series`| Small overhead for type conversion |
|
||
| Raw Rust access (expert) | `from ferro_ta._ferro_ta import sma` | Bypasses all Python wrappers |
|
||
| Multiple series at once | `batch_sma`, `batch_ema`, `batch_rsi` | One Python call for all columns |
|
||
| Many indicators on same arrays | `compute_many` | Amortizes Python→Rust overhead |
|
||
|
||
**Recorded baseline and roadmap:** Performance roadmap and trade-offs are tracked
|
||
in [PERFORMANCE_ROADMAP.md](../PERFORMANCE_ROADMAP.md). For reproducible benchmark
|
||
inputs/results and methodology, use [benchmarks/README.md](../benchmarks/README.md)
|
||
and regenerate with `python benchmarks/bench_vs_talib.py --json benchmark_vs_talib.json`.
|
||
|
||
---
|
||
|
||
## The Rust Core Is Fast; Overhead Is in Python
|
||
|
||
The Rust extension (`_ferro_ta`) is compiled with full optimisations and is very
|
||
fast. The bottlenecks for most users are in the Python wrapping layer:
|
||
|
||
1. **Array conversion** — `_to_f64` converts any array-like to a contiguous
|
||
`float64` NumPy array. If your input is already a C-contiguous `float64`
|
||
ndarray the fast path returns it without any copy or allocation.
|
||
|
||
2. **pandas wrapping** — `pandas_wrap` extracts the NumPy array from a
|
||
`pd.Series`, calls the Rust function, and wraps the result back into a
|
||
`pd.Series` with the original index. The wrapping itself is cheap but adds
|
||
a small constant overhead per call.
|
||
|
||
3. **polars wrapping** — `polars_wrap` converts a `pl.Series` to NumPy and back.
|
||
The result is now built from the NumPy buffer directly (`pl.Series(name,
|
||
np.asarray(result))`), which avoids the O(n) `.tolist()` conversion of
|
||
earlier versions.
|
||
|
||
4. **Batch / grouped execution** — `batch_sma`/`batch_ema`/`batch_rsi` use
|
||
Rust-side batch functions for 2-D input (single GIL release for all
|
||
columns). `compute_many(...)` groups supported 1-D indicator bundles into
|
||
one Rust call, which helps most on medium-to-large workloads. The generic
|
||
`batch_apply` still runs a Python loop over columns; use it only when there
|
||
is no dedicated fast path.
|
||
|
||
---
|
||
|
||
## The Fast Path: Pass Contiguous float64 NumPy Arrays
|
||
|
||
The cheapest way to call any indicator is to pass a C-contiguous `float64`
|
||
NumPy array. `_to_f64` detects this case and returns the array as-is:
|
||
|
||
```python
|
||
import numpy as np
|
||
from ferro_ta import SMA
|
||
|
||
# Already float64 and C-contiguous — _to_f64 is a no-op (zero copy)
|
||
close = np.random.rand(10_000).astype(np.float64)
|
||
result = SMA(close, timeperiod=20)
|
||
```
|
||
|
||
If your array is in a different dtype or order, `_to_f64` will create a new
|
||
array. You can force the fast path once and reuse the result:
|
||
|
||
```python
|
||
close_f64 = np.ascontiguousarray(close, dtype=np.float64) # one-time conversion
|
||
result = SMA(close_f64, timeperiod=20) # no copy inside _to_f64
|
||
```
|
||
|
||
---
|
||
|
||
## Raw Numpy-Only API (No Wrapper Overhead)
|
||
|
||
If you want zero Python overhead — no pandas/polars wrapping, no validation —
|
||
you can import functions directly from the compiled extension:
|
||
|
||
```python
|
||
from ferro_ta._ferro_ta import sma, ema, rsi # raw Rust functions
|
||
|
||
import numpy as np
|
||
close = np.random.rand(10_000).astype(np.float64)
|
||
result = sma(close, 20) # returns a NumPy array (PyArray1<f64> from PyO3)
|
||
```
|
||
|
||
> **Warning:** The raw `_ferro_ta` API is internal and may change between
|
||
> versions. It does *not* validate inputs — passing an empty array or a wrong
|
||
> type will raise an obscure error from PyO3. Use it only if you have
|
||
> profiled a bottleneck and need the absolute minimum overhead.
|
||
|
||
For a stable raw API with the same functions, use the `ferro_ta.raw` submodule
|
||
(no pandas/polars wrapping or validation).
|
||
|
||
---
|
||
|
||
## pandas Series
|
||
|
||
```python
|
||
import pandas as pd
|
||
from ferro_ta import SMA
|
||
|
||
s = pd.Series([1.0, 2.0, 3.0, 4.0, 5.0], index=pd.date_range("2024-01-01", periods=5))
|
||
result = SMA(s, timeperiod=3)
|
||
# result is a pd.Series with the same DatetimeIndex
|
||
```
|
||
|
||
Overhead compared to a raw numpy call: one `pd.Series.to_numpy()` call (cheap)
|
||
plus one `pd.Series(result, index=...)` call (cheap). For large arrays this
|
||
is negligible; for very tight loops (millions of calls per second) prefer numpy.
|
||
|
||
---
|
||
|
||
## polars Series
|
||
|
||
```python
|
||
import polars as pl
|
||
from ferro_ta import SMA
|
||
|
||
s = pl.Series("close", [1.0, 2.0, 3.0, 4.0, 5.0])
|
||
result = SMA(s, timeperiod=3)
|
||
# result is a pl.Series named "close"
|
||
```
|
||
|
||
Overhead: one `.cast(Float64).to_numpy()` call plus one `pl.Series(name,
|
||
np.asarray(result))` call. The result is built from the numpy buffer
|
||
(zero-copy where polars allows it) rather than going through `.tolist()`.
|
||
|
||
---
|
||
|
||
## Batch Execution
|
||
|
||
Use the batch API when you have many series (e.g., one column per symbol):
|
||
|
||
```python
|
||
import numpy as np
|
||
from ferro_ta.batch import batch_sma, batch_ema, batch_rsi, batch_apply
|
||
|
||
data = np.random.rand(252, 500).astype(np.float64) # 252 bars × 500 symbols
|
||
sma_out = batch_sma(data, timeperiod=20) # shape (252, 500)
|
||
rsi_out = batch_rsi(data, timeperiod=14)
|
||
```
|
||
|
||
`batch_apply` lets you run any indicator on a 2-D array:
|
||
|
||
```python
|
||
from ferro_ta import ATR
|
||
from ferro_ta.batch import batch_apply
|
||
|
||
ohlcv = np.random.rand(252, 100, 3).astype(np.float64) # not directly supported
|
||
# For indicators that take multiple arrays use a manual loop instead
|
||
```
|
||
|
||
For 2-D input, `batch_sma`/`batch_ema`/`batch_rsi` use Rust-side batch
|
||
functions (single GIL release for all columns). Use `batch_apply` for other
|
||
indicators that do not have a dedicated Rust batch implementation.
|
||
|
||
When you have several indicators over the same 1-D arrays, use `compute_many`:
|
||
|
||
```python
|
||
from ferro_ta.batch import compute_many
|
||
|
||
results = compute_many(
|
||
[
|
||
("SMA", {"timeperiod": 10}),
|
||
("EMA", {"timeperiod": 12}),
|
||
("RSI", {"timeperiod": 14}),
|
||
],
|
||
close=close,
|
||
)
|
||
```
|
||
|
||
Supported grouped paths currently cover common close-only indicators plus a
|
||
small HLC bundle (`ATR`, `NATR`, `ADX`, `ADXR`, `CCI`, `WILLR`). Unsupported
|
||
parameter shapes fall back to the normal registry path automatically.
|
||
|
||
---
|
||
|
||
## Streaming (Bar-by-Bar)
|
||
|
||
```python
|
||
from ferro_ta.streaming import StreamingSMA
|
||
|
||
sma = StreamingSMA(period=20)
|
||
for bar in live_feed:
|
||
value = sma.update(bar.close)
|
||
if value is not None:
|
||
print(f"SMA(20) = {value:.4f}")
|
||
```
|
||
|
||
The streaming classes are implemented in Rust (PyO3 `#[pyclass]` in
|
||
`_ferro_ta`) and re-exported from `ferro_ta.streaming`. They are suitable for
|
||
live trading at typical bar rates with minimal Python overhead.
|
||
|
||
---
|
||
|
||
## Extended Indicators
|
||
|
||
`VWAP`, `SUPERTREND`, `ICHIMOKU`, `DONCHIAN`, `PIVOT_POINTS`, `KELTNER_CHANNELS`,
|
||
`HULL_MA`, `CHANDELIER_EXIT`, `VWMA`, and `CHOPPINESS_INDEX` are implemented in
|
||
Rust (`src/extended/mod.rs`). The Python module `ferro_ta/extended.py` is a thin
|
||
wrapper with validation and `_to_f64`; all computation runs in the extension.
|
||
|
||
---
|
||
|
||
## Tips for Best Performance
|
||
|
||
1. **Pre-convert once.** If you call multiple indicators on the same array,
|
||
convert it to `float64` + C-contiguous once:
|
||
```python
|
||
close = np.ascontiguousarray(raw_close, dtype=np.float64)
|
||
```
|
||
|
||
2. **Avoid repeated dtype conversions.** Passing a `float32` or `int` array
|
||
triggers a copy every call.
|
||
|
||
3. **Use batch functions for multiple symbols.** For SMA, EMA, and RSI use
|
||
`batch_sma`/`batch_ema`/`batch_rsi` (Rust-side loop, single GIL release).
|
||
The generic `batch_apply` runs a Python loop over columns; use it only for
|
||
indicators that do not have a dedicated Rust batch.
|
||
|
||
4. **Avoid wrapping in very tight loops.** If you call an indicator millions
|
||
of times per second (e.g., in a simulation) use the raw `_ferro_ta` API
|
||
and manage conversion yourself.
|
||
|
||
5. **Profile before optimising.** Use `cProfile` or `py-spy` to find the
|
||
actual bottleneck before assuming a particular layer is slow.
|
||
|
||
6. **Use the perf-contract scripts for evidence.** `benchmarks/run_perf_contract.py`
|
||
and `benchmarks/profile_runtime_hotspots.py` record timings with git/runtime
|
||
metadata so you can compare apples to apples across machines and commits.
|
||
|
||
## Backtesting Performance
|
||
|
||
ferro-ta's backtesting engine is the fastest in the Python ecosystem for
|
||
vectorized single- and multi-asset scenarios.
|
||
|
||
| Library | 100k bars | vs ferro-ta |
|
||
|---------|-----------|-------------|
|
||
| ferro-ta `backtest_core` | **0.29 ms** | — |
|
||
| ferro-ta `backtest_ohlcv_core` | **0.33 ms** | ~same |
|
||
| NumPy vectorized | 0.46 ms | 1.6× slower |
|
||
| vectorbt | 2.90 ms | 10× slower |
|
||
| backtesting.py | 319 ms | 1,117× slower |
|
||
| backtrader | ~50,000 ms (est.) | >15,000× slower |
|
||
|
||
Additional capabilities measured at 100k bars:
|
||
|
||
| Capability | Time |
|
||
|---|---|
|
||
| Monte Carlo 1,000 sims (parallel) | 50 ms — 12× faster than NumPy loop |
|
||
| 23 performance metrics | 2.8 ms (0.12 ms/metric) |
|
||
| Multi-asset 100 symbols, parallel | 43 ms — 2× vs serial |
|
||
| Walk-forward index generation | 0.3 µs |
|
||
|
||
## Benchmark Tooling
|
||
|
||
The benchmark suite now includes a small set of machine-readable scripts for
|
||
performance work beyond the full pytest benchmark table:
|
||
|
||
- `python benchmarks/bench_batch.py --json batch_benchmark.json`
|
||
- `python benchmarks/bench_streaming.py --json streaming_benchmark.json`
|
||
- `python benchmarks/bench_backtest.py --json bench_backtest_results.json`
|
||
- `python benchmarks/profile_runtime_hotspots.py --json runtime_hotspots.json`
|
||
- `python benchmarks/bench_simd.py --json simd_benchmark.json`
|
||
- `python benchmarks/run_perf_contract.py --output-dir benchmarks/artifacts/latest`
|
||
- `python benchmarks/check_hotspot_regression.py --input runtime_hotspots.json`
|
||
|
||
The WASM bindings also ship with a Node benchmark:
|
||
|
||
- `cd wasm && wasm-pack build --target nodejs --out-dir pkg`
|
||
- `node bench.js --json ../wasm_benchmark.json`
|
||
|
||
## SIMD And Build Flags
|
||
|
||
Distributable wheels should stay on the portable release profile:
|
||
|
||
- `cargo`/`maturin` release build
|
||
- `lto = true`
|
||
- `codegen-units = 1`
|
||
- no architecture-specific `target-cpu=native` in shipped artifacts
|
||
|
||
For local source builds, there are two opt-in tuning levers:
|
||
|
||
```bash
|
||
# Portable SIMD-enabled local build
|
||
uv run maturin develop --release --features simd
|
||
|
||
# Maximum local tuning for your current machine only
|
||
RUSTFLAGS="-C target-cpu=native" uv run maturin develop --release --features simd
|
||
```
|
||
|
||
Policy:
|
||
|
||
- Ship portable wheels with the default release settings.
|
||
- Use `--features simd` for measured local/source wins.
|
||
- Reserve `target-cpu=native` for developer workstations or private deploys,
|
||
because those binaries are not portable across CPU families.
|
||
|
||
---
|
||
|
||
## Performance Improvements (implemented)
|
||
|
||
The following improvements are already in place. See
|
||
[docs/plans/2026-03-08-production-grade.md](plans/2026-03-08-production-grade.md)
|
||
for history and commits.
|
||
|
||
| Area | Improvement | Where |
|
||
|-------------|----------------------------------------------------------------|-------|
|
||
| **Utils** | `_to_f64` fast path: no copy for 1-D C-contiguous float64 | `python/ferro_ta/_utils.py` (lines 34–39) |
|
||
| **Utils** | Polars result: `pl.Series(name, result)` from NumPy buffer (no `.tolist()`) | `python/ferro_ta/_utils.py` (e.g. 254–258) |
|
||
| **Raw API** | `ferro_ta.raw` — bypass pandas/polars and validation | `python/ferro_ta/raw.py` |
|
||
| **Batch** | Rust batch for SMA/EMA/RSI — single GIL release for 2-D | `src/batch/mod.rs`, `python/ferro_ta/batch.py` |
|
||
| **Streaming** | All streaming classes in Rust (PyO3) | `src/streaming/mod.rs` |
|
||
| **Extended** | All extended indicators (incl. SUPERTREND) in Rust | `src/extended/mod.rs`, `python/ferro_ta/extended.py` wraps Rust |
|
||
|
||
---
|
||
|
||
## Known Bottlenecks and Possible Improvements
|
||
|
||
Maintainer-facing list of slower paths and optional improvements. Update as
|
||
bottlenecks are fixed or deferred.
|
||
|
||
**Backtest** (`python/ferro_ta/analysis/backtest.py`):
|
||
- Core signal→equity loop is fully in Rust (`backtest_core`, `backtest_ohlcv_core`).
|
||
- Commission and slippage applied inside Rust; no Python loop on the hot path.
|
||
- `compute_performance_metrics` computes all 23 metrics in a single Rust pass.
|
||
- Monte Carlo runs in parallel Rayon threads with LCG seeding (GIL released).
|
||
|
||
**Batch** (`python/ferro_ta/batch.py`):
|
||
- `batch_apply` runs a Python loop over columns (one Python call per column).
|
||
Use `batch_sma`/`batch_ema`/`batch_rsi` when possible.
|
||
- No fast path for already 2-D C-contiguous float64 in batch_sma/ema/rsi
|
||
(unlike `_to_f64` for 1-D); could avoid a potential copy.
|
||
|
||
**Derivatives analytics** (`python/ferro_ta/analysis/options.py`):
|
||
- `iv_rank`, `iv_percentile`, and `iv_zscore` now delegate to Rust.
|
||
- The Python layer mostly performs broadcasting and result shaping; the hot
|
||
path is in Rust.
|
||
- Model-based implied-volatility inversion is much faster now, but still more
|
||
expensive than direct pricing or Greeks due to root-finding.
|
||
|
||
**Features** (`python/ferro_ta/features.py`):
|
||
- `nan_policy="fill"` is vectorized now.
|
||
- `feature_matrix(...)` uses `compute_many(...)`, but grouped HLC bundles are
|
||
still only near parity on medium workloads and are best on larger arrays.
|
||
|
||
**Signals** (`python/ferro_ta/signals.py`):
|
||
- `compose(..., method="rank")` now uses a one-call Rust rank-composition
|
||
path, but its gains are moderate rather than dramatic. Keep measuring before
|
||
treating it as a major optimization lever.
|
||
|
||
**Other**:
|
||
- **dsl.py**: Some code paths use Python loops over bars.
|
||
- **gpu.py**: Fallback SMA/EMA/RSI use Python loops when GPU is not used.
|
||
- **tools.py / viz.py**: `.tolist()` for JSON/Plotly; acceptable for I/O.
|
||
- **Validation**: `check_equal_length`, `check_timeperiod` run in Python;
|
||
cost is small; moving to Rust is deferred (see production-grade plan).
|
||
- **pandas_wrap / polars_wrap**: Per-call overhead; use `ferro_ta.raw` when
|
||
minimising overhead.
|
||
|
||
---
|
||
|
||
## Benchmarking and comparison
|
||
|
||
For cross-library speed, run:
|
||
`pytest benchmarks/test_speed.py --benchmark-only --benchmark-json=benchmarks/results.json`.
|
||
|
||
To convert benchmark JSON into a markdown table:
|
||
`python benchmarks/benchmark_table.py`.
|
||
|
||
For focused TA-Lib comparison on the same data/parameters, run
|
||
`python benchmarks/bench_vs_talib.py` (requires `pip install ta-lib`).
|
||
Results are reported as speedup = TA-Lib time / ferro_ta time (values > 1 mean
|
||
ferro_ta is faster). Speedup depends on indicator and data size.
|
||
|
||
---
|
||
|
||
## Related Documents
|
||
|
||
- [`docs/architecture.md`](architecture.md) — how the Rust/Python layers are
|
||
organised and how they communicate.
|
||
- [`benchmarks/test_speed.py`](../benchmarks/test_speed.py) —
|
||
Authoritative cross-library speed benchmarks (pytest-benchmark).
|
||
- [`benchmarks/benchmark_table.py`](../benchmarks/benchmark_table.py) —
|
||
Render speed tables from `benchmarks/results.json`.
|
||
- [`crates/ferro_ta_core/benches/indicators.rs`](../crates/ferro_ta_core/benches/indicators.rs) —
|
||
Rust Criterion benchmarks for the pure core (run with `cargo bench -p ferro_ta_core`).
|
||
- [`benchmarks/bench_vs_talib.py`](../benchmarks/bench_vs_talib.py) — speed comparison vs
|
||
TA-Lib (same data and parameters); run with `python benchmarks/bench_vs_talib.py` (requires
|
||
`ta-lib`). See README “Performance vs TA-Lib” for methodology and a comparison table.
|
||
- [`benchmarks/check_vs_talib_regression.py`](../benchmarks/check_vs_talib_regression.py) —
|
||
CI guardrail script for detecting severe benchmark regressions from JSON artifacts.
|