Update version numbers across Rust, Python, and documentation files to 1.1.0. Enhance the .gitignore to include macOS dSYM files and plans directory. Introduce new dependencies in the Rust core library and update the README to reflect recent performance benchmarks and backtesting engine capabilities. Add new artifacts to the benchmarks manifest and improve documentation for the backtesting engine API.
16 KiB
Performance Guide
This document explains the performance characteristics of ferro-ta and gives practical advice on how to get the best speed from the library.
Quick Summary
| Use case | Recommended API | Notes |
|---|---|---|
| Fast path — NumPy arrays | Pass np.ndarray (float64, C-order) |
Zero overhead; no conversion needed |
| pandas users | Pass pd.Series; result is pd.Series |
Small overhead for index wrapping |
| polars users | Pass pl.Series; result is pl.Series |
Small overhead for type conversion |
| Raw Rust access (expert) | from ferro_ta._ferro_ta import sma |
Bypasses all Python wrappers |
| Multiple series at once | batch_sma, batch_ema, batch_rsi |
One Python call for all columns |
| Many indicators on same arrays | compute_many |
Amortizes Python→Rust overhead |
Recorded baseline and roadmap: Performance roadmap and trade-offs are tracked
in PERFORMANCE_ROADMAP.md. For reproducible benchmark
inputs/results and methodology, use benchmarks/README.md
and regenerate with python benchmarks/bench_vs_talib.py --json benchmark_vs_talib.json.
The Rust Core Is Fast; Overhead Is in Python
The Rust extension (_ferro_ta) is compiled with full optimisations and is very
fast. The bottlenecks for most users are in the Python wrapping layer:
-
Array conversion —
_to_f64converts any array-like to a contiguousfloat64NumPy array. If your input is already a C-contiguousfloat64ndarray the fast path returns it without any copy or allocation. -
pandas wrapping —
pandas_wrapextracts the NumPy array from apd.Series, calls the Rust function, and wraps the result back into apd.Serieswith the original index. The wrapping itself is cheap but adds a small constant overhead per call. -
polars wrapping —
polars_wrapconverts apl.Seriesto NumPy and back. The result is now built from the NumPy buffer directly (pl.Series(name, np.asarray(result))), which avoids the O(n).tolist()conversion of earlier versions. -
Batch / grouped execution —
batch_sma/batch_ema/batch_rsiuse Rust-side batch functions for 2-D input (single GIL release for all columns).compute_many(...)groups supported 1-D indicator bundles into one Rust call, which helps most on medium-to-large workloads. The genericbatch_applystill runs a Python loop over columns; use it only when there is no dedicated fast path.
The Fast Path: Pass Contiguous float64 NumPy Arrays
The cheapest way to call any indicator is to pass a C-contiguous float64
NumPy array. _to_f64 detects this case and returns the array as-is:
import numpy as np
from ferro_ta import SMA
# Already float64 and C-contiguous — _to_f64 is a no-op (zero copy)
close = np.random.rand(10_000).astype(np.float64)
result = SMA(close, timeperiod=20)
If your array is in a different dtype or order, _to_f64 will create a new
array. You can force the fast path once and reuse the result:
close_f64 = np.ascontiguousarray(close, dtype=np.float64) # one-time conversion
result = SMA(close_f64, timeperiod=20) # no copy inside _to_f64
Raw Numpy-Only API (No Wrapper Overhead)
If you want zero Python overhead — no pandas/polars wrapping, no validation — you can import functions directly from the compiled extension:
from ferro_ta._ferro_ta import sma, ema, rsi # raw Rust functions
import numpy as np
close = np.random.rand(10_000).astype(np.float64)
result = sma(close, 20) # returns a NumPy array (PyArray1<f64> from PyO3)
Warning: The raw
_ferro_taAPI is internal and may change between versions. It does not validate inputs — passing an empty array or a wrong type will raise an obscure error from PyO3. Use it only if you have profiled a bottleneck and need the absolute minimum overhead.
For a stable raw API with the same functions, use the ferro_ta.raw submodule
(no pandas/polars wrapping or validation).
pandas Series
import pandas as pd
from ferro_ta import SMA
s = pd.Series([1.0, 2.0, 3.0, 4.0, 5.0], index=pd.date_range("2024-01-01", periods=5))
result = SMA(s, timeperiod=3)
# result is a pd.Series with the same DatetimeIndex
Overhead compared to a raw numpy call: one pd.Series.to_numpy() call (cheap)
plus one pd.Series(result, index=...) call (cheap). For large arrays this
is negligible; for very tight loops (millions of calls per second) prefer numpy.
polars Series
import polars as pl
from ferro_ta import SMA
s = pl.Series("close", [1.0, 2.0, 3.0, 4.0, 5.0])
result = SMA(s, timeperiod=3)
# result is a pl.Series named "close"
Overhead: one .cast(Float64).to_numpy() call plus one pl.Series(name, np.asarray(result)) call. The result is built from the numpy buffer
(zero-copy where polars allows it) rather than going through .tolist().
Batch Execution
Use the batch API when you have many series (e.g., one column per symbol):
import numpy as np
from ferro_ta.batch import batch_sma, batch_ema, batch_rsi, batch_apply
data = np.random.rand(252, 500).astype(np.float64) # 252 bars × 500 symbols
sma_out = batch_sma(data, timeperiod=20) # shape (252, 500)
rsi_out = batch_rsi(data, timeperiod=14)
batch_apply lets you run any indicator on a 2-D array:
from ferro_ta import ATR
from ferro_ta.batch import batch_apply
ohlcv = np.random.rand(252, 100, 3).astype(np.float64) # not directly supported
# For indicators that take multiple arrays use a manual loop instead
For 2-D input, batch_sma/batch_ema/batch_rsi use Rust-side batch
functions (single GIL release for all columns). Use batch_apply for other
indicators that do not have a dedicated Rust batch implementation.
When you have several indicators over the same 1-D arrays, use compute_many:
from ferro_ta.batch import compute_many
results = compute_many(
[
("SMA", {"timeperiod": 10}),
("EMA", {"timeperiod": 12}),
("RSI", {"timeperiod": 14}),
],
close=close,
)
Supported grouped paths currently cover common close-only indicators plus a
small HLC bundle (ATR, NATR, ADX, ADXR, CCI, WILLR). Unsupported
parameter shapes fall back to the normal registry path automatically.
Streaming (Bar-by-Bar)
from ferro_ta.streaming import StreamingSMA
sma = StreamingSMA(period=20)
for bar in live_feed:
value = sma.update(bar.close)
if value is not None:
print(f"SMA(20) = {value:.4f}")
The streaming classes are implemented in Rust (PyO3 #[pyclass] in
_ferro_ta) and re-exported from ferro_ta.streaming. They are suitable for
live trading at typical bar rates with minimal Python overhead.
Extended Indicators
VWAP, SUPERTREND, ICHIMOKU, DONCHIAN, PIVOT_POINTS, KELTNER_CHANNELS,
HULL_MA, CHANDELIER_EXIT, VWMA, and CHOPPINESS_INDEX are implemented in
Rust (src/extended/mod.rs). The Python module ferro_ta/extended.py is a thin
wrapper with validation and _to_f64; all computation runs in the extension.
Tips for Best Performance
-
Pre-convert once. If you call multiple indicators on the same array, convert it to
float64+ C-contiguous once:close = np.ascontiguousarray(raw_close, dtype=np.float64) -
Avoid repeated dtype conversions. Passing a
float32orintarray triggers a copy every call. -
Use batch functions for multiple symbols. For SMA, EMA, and RSI use
batch_sma/batch_ema/batch_rsi(Rust-side loop, single GIL release). The genericbatch_applyruns a Python loop over columns; use it only for indicators that do not have a dedicated Rust batch. -
Avoid wrapping in very tight loops. If you call an indicator millions of times per second (e.g., in a simulation) use the raw
_ferro_taAPI and manage conversion yourself. -
Profile before optimising. Use
cProfileorpy-spyto find the actual bottleneck before assuming a particular layer is slow. -
Use the perf-contract scripts for evidence.
benchmarks/run_perf_contract.pyandbenchmarks/profile_runtime_hotspots.pyrecord timings with git/runtime metadata so you can compare apples to apples across machines and commits.
Backtesting Performance
ferro-ta's backtesting engine is the fastest in the Python ecosystem for vectorized single- and multi-asset scenarios.
| Library | 100k bars | vs ferro-ta |
|---|---|---|
ferro-ta backtest_core |
0.29 ms | — |
ferro-ta backtest_ohlcv_core |
0.33 ms | ~same |
| NumPy vectorized | 0.46 ms | 1.6× slower |
| vectorbt | 2.90 ms | 10× slower |
| backtesting.py | 319 ms | 1,117× slower |
| backtrader | ~50,000 ms (est.) | >15,000× slower |
Additional capabilities measured at 100k bars:
| Capability | Time |
|---|---|
| Monte Carlo 1,000 sims (parallel) | 50 ms — 12× faster than NumPy loop |
| 23 performance metrics | 2.8 ms (0.12 ms/metric) |
| Multi-asset 100 symbols, parallel | 43 ms — 2× vs serial |
| Walk-forward index generation | 0.3 µs |
Benchmark Tooling
The benchmark suite now includes a small set of machine-readable scripts for performance work beyond the full pytest benchmark table:
python benchmarks/bench_batch.py --json batch_benchmark.jsonpython benchmarks/bench_streaming.py --json streaming_benchmark.jsonpython benchmarks/bench_backtest.py --json bench_backtest_results.jsonpython benchmarks/profile_runtime_hotspots.py --json runtime_hotspots.jsonpython benchmarks/bench_simd.py --json simd_benchmark.jsonpython benchmarks/run_perf_contract.py --output-dir benchmarks/artifacts/latestpython benchmarks/check_hotspot_regression.py --input runtime_hotspots.json
The WASM bindings also ship with a Node benchmark:
cd wasm && wasm-pack build --target nodejs --out-dir pkgnode bench.js --json ../wasm_benchmark.json
SIMD And Build Flags
Distributable wheels should stay on the portable release profile:
cargo/maturinrelease buildlto = truecodegen-units = 1- no architecture-specific
target-cpu=nativein shipped artifacts
For local source builds, there are two opt-in tuning levers:
# Portable SIMD-enabled local build
uv run maturin develop --release --features simd
# Maximum local tuning for your current machine only
RUSTFLAGS="-C target-cpu=native" uv run maturin develop --release --features simd
Policy:
- Ship portable wheels with the default release settings.
- Use
--features simdfor measured local/source wins. - Reserve
target-cpu=nativefor developer workstations or private deploys, because those binaries are not portable across CPU families.
Performance Improvements (implemented)
The following improvements are already in place. See docs/plans/2026-03-08-production-grade.md for history and commits.
| Area | Improvement | Where |
|---|---|---|
| Utils | _to_f64 fast path: no copy for 1-D C-contiguous float64 |
python/ferro_ta/_utils.py (lines 34–39) |
| Utils | Polars result: pl.Series(name, result) from NumPy buffer (no .tolist()) |
python/ferro_ta/_utils.py (e.g. 254–258) |
| Raw API | ferro_ta.raw — bypass pandas/polars and validation |
python/ferro_ta/raw.py |
| Batch | Rust batch for SMA/EMA/RSI — single GIL release for 2-D | src/batch/mod.rs, python/ferro_ta/batch.py |
| Streaming | All streaming classes in Rust (PyO3) | src/streaming/mod.rs |
| Extended | All extended indicators (incl. SUPERTREND) in Rust | src/extended/mod.rs, python/ferro_ta/extended.py wraps Rust |
Known Bottlenecks and Possible Improvements
Maintainer-facing list of slower paths and optional improvements. Update as bottlenecks are fixed or deferred.
Backtest (python/ferro_ta/analysis/backtest.py):
- Core signal→equity loop is fully in Rust (
backtest_core,backtest_ohlcv_core). - Commission and slippage applied inside Rust; no Python loop on the hot path.
compute_performance_metricscomputes all 23 metrics in a single Rust pass.- Monte Carlo runs in parallel Rayon threads with LCG seeding (GIL released).
Batch (python/ferro_ta/batch.py):
batch_applyruns a Python loop over columns (one Python call per column). Usebatch_sma/batch_ema/batch_rsiwhen possible.- No fast path for already 2-D C-contiguous float64 in batch_sma/ema/rsi
(unlike
_to_f64for 1-D); could avoid a potential copy.
Derivatives analytics (python/ferro_ta/analysis/options.py):
iv_rank,iv_percentile, andiv_zscorenow delegate to Rust.- The Python layer mostly performs broadcasting and result shaping; the hot path is in Rust.
- Model-based implied-volatility inversion is much faster now, but still more expensive than direct pricing or Greeks due to root-finding.
Features (python/ferro_ta/features.py):
nan_policy="fill"is vectorized now.feature_matrix(...)usescompute_many(...), but grouped HLC bundles are still only near parity on medium workloads and are best on larger arrays.
Signals (python/ferro_ta/signals.py):
compose(..., method="rank")now uses a one-call Rust rank-composition path, but its gains are moderate rather than dramatic. Keep measuring before treating it as a major optimization lever.
Other:
- dsl.py: Some code paths use Python loops over bars.
- gpu.py: Fallback SMA/EMA/RSI use Python loops when GPU is not used.
- tools.py / viz.py:
.tolist()for JSON/Plotly; acceptable for I/O. - Validation:
check_equal_length,check_timeperiodrun in Python; cost is small; moving to Rust is deferred (see production-grade plan). - pandas_wrap / polars_wrap: Per-call overhead; use
ferro_ta.rawwhen minimising overhead.
Benchmarking and comparison
For cross-library speed, run:
pytest benchmarks/test_speed.py --benchmark-only --benchmark-json=benchmarks/results.json.
To convert benchmark JSON into a markdown table:
python benchmarks/benchmark_table.py.
For focused TA-Lib comparison on the same data/parameters, run
python benchmarks/bench_vs_talib.py (requires pip install ta-lib).
Results are reported as speedup = TA-Lib time / ferro_ta time (values > 1 mean
ferro_ta is faster). Speedup depends on indicator and data size.
Related Documents
docs/architecture.md— how the Rust/Python layers are organised and how they communicate.benchmarks/test_speed.py— Authoritative cross-library speed benchmarks (pytest-benchmark).benchmarks/benchmark_table.py— Render speed tables frombenchmarks/results.json.crates/ferro_ta_core/benches/indicators.rs— Rust Criterion benchmarks for the pure core (run withcargo bench -p ferro_ta_core).benchmarks/bench_vs_talib.py— speed comparison vs TA-Lib (same data and parameters); run withpython benchmarks/bench_vs_talib.py(requiresta-lib). See README “Performance vs TA-Lib” for methodology and a comparison table.benchmarks/check_vs_talib_regression.py— CI guardrail script for detecting severe benchmark regressions from JSON artifacts.