Files
Pratik Bhadane 436954138f chore: prepare v1.1.0 release
Update version numbers across Rust, Python, and documentation files to 1.1.0. Enhance the .gitignore to include macOS dSYM files and plans directory. Introduce new dependencies in the Rust core library and update the README to reflect recent performance benchmarks and backtesting engine capabilities. Add new artifacts to the benchmarks manifest and improve documentation for the backtesting engine API.
2026-03-30 12:45:52 +05:30

398 lines
16 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Performance Guide
This document explains the performance characteristics of **ferro-ta** and gives
practical advice on how to get the best speed from the library.
---
## Quick Summary
| Use case | Recommended API | Notes |
|---------------------------------------|----------------------------------------|----------------------------------------|
| Fast path — NumPy arrays | Pass `np.ndarray` (float64, C-order) | Zero overhead; no conversion needed |
| pandas users | Pass `pd.Series`; result is `pd.Series`| Small overhead for index wrapping |
| polars users | Pass `pl.Series`; result is `pl.Series`| Small overhead for type conversion |
| Raw Rust access (expert) | `from ferro_ta._ferro_ta import sma` | Bypasses all Python wrappers |
| Multiple series at once | `batch_sma`, `batch_ema`, `batch_rsi` | One Python call for all columns |
| Many indicators on same arrays | `compute_many` | Amortizes Python→Rust overhead |
**Recorded baseline and roadmap:** Performance roadmap and trade-offs are tracked
in [PERFORMANCE_ROADMAP.md](../PERFORMANCE_ROADMAP.md). For reproducible benchmark
inputs/results and methodology, use [benchmarks/README.md](../benchmarks/README.md)
and regenerate with `python benchmarks/bench_vs_talib.py --json benchmark_vs_talib.json`.
---
## The Rust Core Is Fast; Overhead Is in Python
The Rust extension (`_ferro_ta`) is compiled with full optimisations and is very
fast. The bottlenecks for most users are in the Python wrapping layer:
1. **Array conversion**`_to_f64` converts any array-like to a contiguous
`float64` NumPy array. If your input is already a C-contiguous `float64`
ndarray the fast path returns it without any copy or allocation.
2. **pandas wrapping**`pandas_wrap` extracts the NumPy array from a
`pd.Series`, calls the Rust function, and wraps the result back into a
`pd.Series` with the original index. The wrapping itself is cheap but adds
a small constant overhead per call.
3. **polars wrapping**`polars_wrap` converts a `pl.Series` to NumPy and back.
The result is now built from the NumPy buffer directly (`pl.Series(name,
np.asarray(result))`), which avoids the O(n) `.tolist()` conversion of
earlier versions.
4. **Batch / grouped execution**`batch_sma`/`batch_ema`/`batch_rsi` use
Rust-side batch functions for 2-D input (single GIL release for all
columns). `compute_many(...)` groups supported 1-D indicator bundles into
one Rust call, which helps most on medium-to-large workloads. The generic
`batch_apply` still runs a Python loop over columns; use it only when there
is no dedicated fast path.
---
## The Fast Path: Pass Contiguous float64 NumPy Arrays
The cheapest way to call any indicator is to pass a C-contiguous `float64`
NumPy array. `_to_f64` detects this case and returns the array as-is:
```python
import numpy as np
from ferro_ta import SMA
# Already float64 and C-contiguous — _to_f64 is a no-op (zero copy)
close = np.random.rand(10_000).astype(np.float64)
result = SMA(close, timeperiod=20)
```
If your array is in a different dtype or order, `_to_f64` will create a new
array. You can force the fast path once and reuse the result:
```python
close_f64 = np.ascontiguousarray(close, dtype=np.float64) # one-time conversion
result = SMA(close_f64, timeperiod=20) # no copy inside _to_f64
```
---
## Raw Numpy-Only API (No Wrapper Overhead)
If you want zero Python overhead — no pandas/polars wrapping, no validation —
you can import functions directly from the compiled extension:
```python
from ferro_ta._ferro_ta import sma, ema, rsi # raw Rust functions
import numpy as np
close = np.random.rand(10_000).astype(np.float64)
result = sma(close, 20) # returns a NumPy array (PyArray1<f64> from PyO3)
```
> **Warning:** The raw `_ferro_ta` API is internal and may change between
> versions. It does *not* validate inputs — passing an empty array or a wrong
> type will raise an obscure error from PyO3. Use it only if you have
> profiled a bottleneck and need the absolute minimum overhead.
For a stable raw API with the same functions, use the `ferro_ta.raw` submodule
(no pandas/polars wrapping or validation).
---
## pandas Series
```python
import pandas as pd
from ferro_ta import SMA
s = pd.Series([1.0, 2.0, 3.0, 4.0, 5.0], index=pd.date_range("2024-01-01", periods=5))
result = SMA(s, timeperiod=3)
# result is a pd.Series with the same DatetimeIndex
```
Overhead compared to a raw numpy call: one `pd.Series.to_numpy()` call (cheap)
plus one `pd.Series(result, index=...)` call (cheap). For large arrays this
is negligible; for very tight loops (millions of calls per second) prefer numpy.
---
## polars Series
```python
import polars as pl
from ferro_ta import SMA
s = pl.Series("close", [1.0, 2.0, 3.0, 4.0, 5.0])
result = SMA(s, timeperiod=3)
# result is a pl.Series named "close"
```
Overhead: one `.cast(Float64).to_numpy()` call plus one `pl.Series(name,
np.asarray(result))` call. The result is built from the numpy buffer
(zero-copy where polars allows it) rather than going through `.tolist()`.
---
## Batch Execution
Use the batch API when you have many series (e.g., one column per symbol):
```python
import numpy as np
from ferro_ta.batch import batch_sma, batch_ema, batch_rsi, batch_apply
data = np.random.rand(252, 500).astype(np.float64) # 252 bars × 500 symbols
sma_out = batch_sma(data, timeperiod=20) # shape (252, 500)
rsi_out = batch_rsi(data, timeperiod=14)
```
`batch_apply` lets you run any indicator on a 2-D array:
```python
from ferro_ta import ATR
from ferro_ta.batch import batch_apply
ohlcv = np.random.rand(252, 100, 3).astype(np.float64) # not directly supported
# For indicators that take multiple arrays use a manual loop instead
```
For 2-D input, `batch_sma`/`batch_ema`/`batch_rsi` use Rust-side batch
functions (single GIL release for all columns). Use `batch_apply` for other
indicators that do not have a dedicated Rust batch implementation.
When you have several indicators over the same 1-D arrays, use `compute_many`:
```python
from ferro_ta.batch import compute_many
results = compute_many(
[
("SMA", {"timeperiod": 10}),
("EMA", {"timeperiod": 12}),
("RSI", {"timeperiod": 14}),
],
close=close,
)
```
Supported grouped paths currently cover common close-only indicators plus a
small HLC bundle (`ATR`, `NATR`, `ADX`, `ADXR`, `CCI`, `WILLR`). Unsupported
parameter shapes fall back to the normal registry path automatically.
---
## Streaming (Bar-by-Bar)
```python
from ferro_ta.streaming import StreamingSMA
sma = StreamingSMA(period=20)
for bar in live_feed:
value = sma.update(bar.close)
if value is not None:
print(f"SMA(20) = {value:.4f}")
```
The streaming classes are implemented in Rust (PyO3 `#[pyclass]` in
`_ferro_ta`) and re-exported from `ferro_ta.streaming`. They are suitable for
live trading at typical bar rates with minimal Python overhead.
---
## Extended Indicators
`VWAP`, `SUPERTREND`, `ICHIMOKU`, `DONCHIAN`, `PIVOT_POINTS`, `KELTNER_CHANNELS`,
`HULL_MA`, `CHANDELIER_EXIT`, `VWMA`, and `CHOPPINESS_INDEX` are implemented in
Rust (`src/extended/mod.rs`). The Python module `ferro_ta/extended.py` is a thin
wrapper with validation and `_to_f64`; all computation runs in the extension.
---
## Tips for Best Performance
1. **Pre-convert once.** If you call multiple indicators on the same array,
convert it to `float64` + C-contiguous once:
```python
close = np.ascontiguousarray(raw_close, dtype=np.float64)
```
2. **Avoid repeated dtype conversions.** Passing a `float32` or `int` array
triggers a copy every call.
3. **Use batch functions for multiple symbols.** For SMA, EMA, and RSI use
`batch_sma`/`batch_ema`/`batch_rsi` (Rust-side loop, single GIL release).
The generic `batch_apply` runs a Python loop over columns; use it only for
indicators that do not have a dedicated Rust batch.
4. **Avoid wrapping in very tight loops.** If you call an indicator millions
of times per second (e.g., in a simulation) use the raw `_ferro_ta` API
and manage conversion yourself.
5. **Profile before optimising.** Use `cProfile` or `py-spy` to find the
actual bottleneck before assuming a particular layer is slow.
6. **Use the perf-contract scripts for evidence.** `benchmarks/run_perf_contract.py`
and `benchmarks/profile_runtime_hotspots.py` record timings with git/runtime
metadata so you can compare apples to apples across machines and commits.
## Backtesting Performance
ferro-ta's backtesting engine is the fastest in the Python ecosystem for
vectorized single- and multi-asset scenarios.
| Library | 100k bars | vs ferro-ta |
|---------|-----------|-------------|
| ferro-ta `backtest_core` | **0.29 ms** | — |
| ferro-ta `backtest_ohlcv_core` | **0.33 ms** | ~same |
| NumPy vectorized | 0.46 ms | 1.6× slower |
| vectorbt | 2.90 ms | 10× slower |
| backtesting.py | 319 ms | 1,117× slower |
| backtrader | ~50,000 ms (est.) | >15,000× slower |
Additional capabilities measured at 100k bars:
| Capability | Time |
|---|---|
| Monte Carlo 1,000 sims (parallel) | 50 ms — 12× faster than NumPy loop |
| 23 performance metrics | 2.8 ms (0.12 ms/metric) |
| Multi-asset 100 symbols, parallel | 43 ms — 2× vs serial |
| Walk-forward index generation | 0.3 µs |
## Benchmark Tooling
The benchmark suite now includes a small set of machine-readable scripts for
performance work beyond the full pytest benchmark table:
- `python benchmarks/bench_batch.py --json batch_benchmark.json`
- `python benchmarks/bench_streaming.py --json streaming_benchmark.json`
- `python benchmarks/bench_backtest.py --json bench_backtest_results.json`
- `python benchmarks/profile_runtime_hotspots.py --json runtime_hotspots.json`
- `python benchmarks/bench_simd.py --json simd_benchmark.json`
- `python benchmarks/run_perf_contract.py --output-dir benchmarks/artifacts/latest`
- `python benchmarks/check_hotspot_regression.py --input runtime_hotspots.json`
The WASM bindings also ship with a Node benchmark:
- `cd wasm && wasm-pack build --target nodejs --out-dir pkg`
- `node bench.js --json ../wasm_benchmark.json`
## SIMD And Build Flags
Distributable wheels should stay on the portable release profile:
- `cargo`/`maturin` release build
- `lto = true`
- `codegen-units = 1`
- no architecture-specific `target-cpu=native` in shipped artifacts
For local source builds, there are two opt-in tuning levers:
```bash
# Portable SIMD-enabled local build
uv run maturin develop --release --features simd
# Maximum local tuning for your current machine only
RUSTFLAGS="-C target-cpu=native" uv run maturin develop --release --features simd
```
Policy:
- Ship portable wheels with the default release settings.
- Use `--features simd` for measured local/source wins.
- Reserve `target-cpu=native` for developer workstations or private deploys,
because those binaries are not portable across CPU families.
---
## Performance Improvements (implemented)
The following improvements are already in place. See
[docs/plans/2026-03-08-production-grade.md](plans/2026-03-08-production-grade.md)
for history and commits.
| Area | Improvement | Where |
|-------------|----------------------------------------------------------------|-------|
| **Utils** | `_to_f64` fast path: no copy for 1-D C-contiguous float64 | `python/ferro_ta/_utils.py` (lines 3439) |
| **Utils** | Polars result: `pl.Series(name, result)` from NumPy buffer (no `.tolist()`) | `python/ferro_ta/_utils.py` (e.g. 254258) |
| **Raw API** | `ferro_ta.raw` — bypass pandas/polars and validation | `python/ferro_ta/raw.py` |
| **Batch** | Rust batch for SMA/EMA/RSI — single GIL release for 2-D | `src/batch/mod.rs`, `python/ferro_ta/batch.py` |
| **Streaming** | All streaming classes in Rust (PyO3) | `src/streaming/mod.rs` |
| **Extended** | All extended indicators (incl. SUPERTREND) in Rust | `src/extended/mod.rs`, `python/ferro_ta/extended.py` wraps Rust |
---
## Known Bottlenecks and Possible Improvements
Maintainer-facing list of slower paths and optional improvements. Update as
bottlenecks are fixed or deferred.
**Backtest** (`python/ferro_ta/analysis/backtest.py`):
- Core signal→equity loop is fully in Rust (`backtest_core`, `backtest_ohlcv_core`).
- Commission and slippage applied inside Rust; no Python loop on the hot path.
- `compute_performance_metrics` computes all 23 metrics in a single Rust pass.
- Monte Carlo runs in parallel Rayon threads with LCG seeding (GIL released).
**Batch** (`python/ferro_ta/batch.py`):
- `batch_apply` runs a Python loop over columns (one Python call per column).
Use `batch_sma`/`batch_ema`/`batch_rsi` when possible.
- No fast path for already 2-D C-contiguous float64 in batch_sma/ema/rsi
(unlike `_to_f64` for 1-D); could avoid a potential copy.
**Derivatives analytics** (`python/ferro_ta/analysis/options.py`):
- `iv_rank`, `iv_percentile`, and `iv_zscore` now delegate to Rust.
- The Python layer mostly performs broadcasting and result shaping; the hot
path is in Rust.
- Model-based implied-volatility inversion is much faster now, but still more
expensive than direct pricing or Greeks due to root-finding.
**Features** (`python/ferro_ta/features.py`):
- `nan_policy="fill"` is vectorized now.
- `feature_matrix(...)` uses `compute_many(...)`, but grouped HLC bundles are
still only near parity on medium workloads and are best on larger arrays.
**Signals** (`python/ferro_ta/signals.py`):
- `compose(..., method="rank")` now uses a one-call Rust rank-composition
path, but its gains are moderate rather than dramatic. Keep measuring before
treating it as a major optimization lever.
**Other**:
- **dsl.py**: Some code paths use Python loops over bars.
- **gpu.py**: Fallback SMA/EMA/RSI use Python loops when GPU is not used.
- **tools.py / viz.py**: `.tolist()` for JSON/Plotly; acceptable for I/O.
- **Validation**: `check_equal_length`, `check_timeperiod` run in Python;
cost is small; moving to Rust is deferred (see production-grade plan).
- **pandas_wrap / polars_wrap**: Per-call overhead; use `ferro_ta.raw` when
minimising overhead.
---
## Benchmarking and comparison
For cross-library speed, run:
`pytest benchmarks/test_speed.py --benchmark-only --benchmark-json=benchmarks/results.json`.
To convert benchmark JSON into a markdown table:
`python benchmarks/benchmark_table.py`.
For focused TA-Lib comparison on the same data/parameters, run
`python benchmarks/bench_vs_talib.py` (requires `pip install ta-lib`).
Results are reported as speedup = TA-Lib time / ferro_ta time (values &gt; 1 mean
ferro_ta is faster). Speedup depends on indicator and data size.
---
## Related Documents
- [`docs/architecture.md`](architecture.md) — how the Rust/Python layers are
organised and how they communicate.
- [`benchmarks/test_speed.py`](../benchmarks/test_speed.py) —
Authoritative cross-library speed benchmarks (pytest-benchmark).
- [`benchmarks/benchmark_table.py`](../benchmarks/benchmark_table.py) —
Render speed tables from `benchmarks/results.json`.
- [`crates/ferro_ta_core/benches/indicators.rs`](../crates/ferro_ta_core/benches/indicators.rs) —
Rust Criterion benchmarks for the pure core (run with `cargo bench -p ferro_ta_core`).
- [`benchmarks/bench_vs_talib.py`](../benchmarks/bench_vs_talib.py) — speed comparison vs
TA-Lib (same data and parameters); run with `python benchmarks/bench_vs_talib.py` (requires
`ta-lib`). See README “Performance vs TA-Lib” for methodology and a comparison table.
- [`benchmarks/check_vs_talib_regression.py`](../benchmarks/check_vs_talib_regression.py) —
CI guardrail script for detecting severe benchmark regressions from JSON artifacts.