2026-06-26 18:47:35 +08:00
# 02 — Architecture & Data Flow
The whole lab is a set of **layers with one-directional dependencies** . Higher layers call lower
ones; lower layers never know about higher ones. This is what lets you freeze the engine, swap
instruments, and add strategies without anything leaking sideways.
---
## 1. The layer stack
```
┌─────────────────────────────────────────────────────────────────────┐
│ STRATEGY / CALLER one folder per strategy │
│ loads data, computes signals + stops, calls the engine, scores │
└───────────────┬──────────────────────────────┬──────────────────────┘
│ │
┌───────▼────────┐ ┌─────────▼──────────┐
│ OPTIMIZER │ │ MT5 BRIDGE │
│ objective, │ │ compile, run, │
│ Optuna study, │ │ parse, compare │
│ diverse top-N │ └─────────┬──────────┘
└───────┬─────────┘ │
│ │ (gold-standard verify)
┌───────▼─────────┐ │
│ ROBUSTNESS │ anti-overfit checks over results (read-only)
└───────┬─────────┘
│
┌───────▼─────────────────────────────────────────────┐
│ ENGINE (frozen) bar-by-bar fill simulator │
│ knows: bars, signals, stop/target prices, instrument │
│ knows NOT: your strategy, indicators, broker │
└───┬─────────────┬────────────────┬───────────────────┘
│ │ │
┌────────▼───┐ ┌──────▼──────┐ ┌──────▼───────┐ ┌──────────────┐
│ INDICATORS │ │ INSTRUMENTS │ │ GATES │ │ DATA │
│ RSI/ATR/.. │ │ per-symbol │ │ entry-filter │ │ loaders + │
│ pure fns │ │ config объ. │ │ masks │ │ MT5 parser │
└────────────┘ └─────────────┘ └──────────────┘ └──────────────┘
```
### Layer responsibilities
| Layer | Folder | Owns | Must NOT |
|-------|--------|------|----------|
| **Data** | `shared/data/` | load bars from Parquet; fetch from MT5; parse MT5 HTML reports | contain strategy logic |
| **Instruments** | `shared/instruments/` | per-symbol/broker config objects (tick value, spread, swap, lot steps) | know any strategy |
| **Indicators** | `shared/indicators/` | pure functions: RSI, ATR, EMA, SMA, your custom range filter, etc. | hold state across calls |
| **Gates** | `shared/gates/` | reusable boolean masks that *filter* entries (regime, time-of-day, exhaustion) | open/close trades |
| **Engine** | `shared/core/` | the bar-by-bar fill/exit simulator; **frozen** once validated | compute signals or stops |
| **Robustness** | `shared/robustness/` | read-only anti-overfit analyses over a result | mutate the engine or results |
| **Optimizer** | `shared/optimizer/` | objective fn, Optuna wiring, diverse top-N selection | know broker specifics |
| **MT5 bridge** | `shared/mt5_pipeline/` | generate `.set` /`.ini` , run the tester, pull & parse the report, compare | compute strategy |
| **Strategy / caller** | `strategies/<name>/` | glue: data → signals → stops → engine → metrics; the only layer that knows the *whole* picture | be reused by another strategy (copy, don't import) |
| **Registry** | `registry/` | approved, locked results — the source of truth | be edited casually |
---
## 2. The core contract: engine knows nothing about strategy
This is the single most important design decision. The engine's input is **pre-computed** :
```
engine.run(
bars, # DataFrame [timestamp, open, high, low, close, spread]
signals_long, # boolean array — True on the bar an entry is triggered
signals_short, # boolean array
sl_prices, # array — the stop price for an entry on that bar (NaN if none)
tp_prices, # array — the target price
instrument, # InstrumentConfig — all symbol mechanics
2026-06-26 20:50:07 +08:00
sizing, # SizingInputs — position sizing inputs
2026-06-26 18:47:35 +08:00
initial_deposit,
2026-06-26 20:50:07 +08:00
*, # keyword-only from here
m1_bars=None, # OPTIONAL: M1 bars for tick-level exit simulation
# REQUIRED if the EA moves its SL intra-trade (BE / trailing /
# basket trailing) — bar-level mode is untrustworthy for that
# class (doc 03 §7 failure mode, doc 03 §8 target gates).
2026-06-26 18:47:35 +08:00
) -> Result
```
> **Crucial:** the stop and target arrive as **finished prices** computed by the caller. The engine
> never decides *where* a stop goes — only *whether* price touched it. This is the seam that
> separates "the strategy" from "the simulator". Change your strategy → you change the caller and the
> arrays you hand in; the engine is untouched.
2026-06-26 20:50:07 +08:00
>
> **The `m1_bars` parameter is not optional for trailing/BE strategies.** When the EA updates its SL
> during a trade, the bar-level engine can produce a − 40% to − 50% net gap vs MT5 (doc 03 §7 measured
> failure mode). Pass M1 bars and the engine switches to tick-level exit simulation, dropping the
> gap to ~− 5%. Bar-level mode remains the right (and faster) choice for clean-directional setups that
> don't move the SL.
2026-06-26 18:47:35 +08:00
The engine's output is equally generic:
```
Result:
trades # list of closed trades (entry/exit time, price, lots, pnl, swap)
equity_curve # DataFrame [timestamp, balance, equity], sampled (e.g. hourly)
final_balance
initial_deposit
```
A separate `compute_metrics(result)` turns that into the numbers you optimize on: Net Profit, Profit
Factor, Win Rate, max Balance/Equity Drawdown, Sharpe, APR, trade count.
---
## 3. Data flow of one backtest
```
2026-06-26 20:50:07 +08:00
data/<symbol>/<SYM>_M1_<years>.parquet data/<symbol>/<SYM>_M5_<years>.parquet
│ load_bars() → M1 DataFrame │ load_bars() → signal-TF DataFrame
│ │ (resampled from M1 if needed)
│ ▼
│ caller: indicators.* on the signal timeframe
│ │
│ ▼
│ caller: edge-detect signals (True only on the bar
│ │ the condition first flips, not every bar after)
│ │ + optional gates (regime/time filters AND-ed in)
│ ▼
│ caller: compute SL/TP price arrays from params
│ │ (ATR stop, % stop, indicator band, …)
└──────────────┐ ┌──────────────────────────┘
▼ ▼
engine.run(signal_bars, signals, sl/tp, instrument, sizing, deposit,
m1_bars=m1_bars) ← M1 is passed back to the engine
│ for tick-level exit simulation when the EA
│ moves its SL intra-trade (BE / trailing). Without
│ it, the bar-level engine over-credits BE exits
│ (doc 03 §7 failure mode).
▼
Result → compute_metrics → dict
│
▼ (finalists only)
MT5 bridge: build .set from the same params → run real tester → parse report → compare
2026-06-26 18:47:35 +08:00
```
2026-06-26 20:50:07 +08:00
Three subtleties that cause most bugs if missed (the first two explained in doc 03):
2026-06-26 18:47:35 +08:00
- **Edge detection.** Signals must be `True` only on the *transition* bar, not forward-filled, or the
engine re-enters every bar.
- **Timeframe alignment.** Indicators computed on a higher timeframe must be mapped back onto the M1
bars correctly (no look-ahead — a daily value is only known after that day closes).
2026-06-26 20:50:07 +08:00
- **M1 must reach the engine for trailing/BE EAs.** Downloading M1 only to resample it up to the signal
timeframe is **not** enough if the EA moves its SL during a trade. Pass the M1 bars to `engine.run`
(the `m1_bars=` kwarg); otherwise the bar-level exit simulation produces a − 40% to − 50% net gap
(doc 03 §7) that looks like a "fidelity issue" but is actually a missing-input bug.
2026-06-26 18:47:35 +08:00
---
## 4. Why "shared" vs "strategies"
- **`shared/` ** is infrastructure you write **once** and treat as a library: engine, indicators,
instruments, optimizer, robustness, bridge. It is import-stable. Breaking changes here ripple
everywhere, so it changes rarely and deliberately.
- **`strategies/<name>/` ** is where experimentation happens. Each strategy is **self-contained** : its
EA source, its presets, and its `iterations/` (each iteration = one research attempt with its own
`optimize.py` , search space, results, and verification). Strategies **copy** glue code between
iterations rather than importing it, so an old iteration always reproduces even after you change
how you do things.
```
strategies/<strategy>/
├── source/ # your EA .mq5/.ex5 + custom indicators (read-only after compile)
├── presets/ # canonical baseline .set files (frozen reference)
├── production/ # live-ready presets
├── iterations/<base>-<approach>-<YYYY-MM-DD>/
│ ├── README.md # context, status, result summary
│ ├── parameter-space.md # the search space for this attempt
│ ├── optimize.py # the run script (a snapshot — self-contained)
│ ├── wizard-answers.yaml # the run settings (reproducibility)
│ ├── top-1/ top-2/ # the diverse finalists: params, preset, py+mt5 metrics
│ └── auto-verification.md# Python-vs-MT5 comparison table
└── archive/ # discarded attempts
```
---
## 5. The registry is the source of truth
Nothing is "real" until it is in `registry/` . A registry entry is a result that passed **all three**
gates: Python search → MT5 verification → human approval. Once written it is **locked** — you never
silently edit a registry entry; you create a *new* one. This is what stops a research lab from
drifting: the registry is the curated, trustworthy subset of everything you ever tried.
Doc 04 covers the isolation rules that keep these boundaries honest. Doc 03 is next — how to actually
build the engine.
Next: [`03-engine-design.md` ](03-engine-design.md ).