Files
mymt5opp/04-isolation-rules.md
2026-06-26 18:47:35 +08:00

144 lines
6.7 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# 04 — Isolation Rules
The rules that keep a research lab trustworthy as it grows. Every one exists because the alternative
quietly corrupts results. Treat them as non-negotiable; they cost a little friction and save you from
optimizing on a silently-broken engine or shipping a number you can't reproduce.
---
## Rule 1 — The validated engine is FROZEN
Once an engine reproduces your EA within the expected fidelity gap (doc 03 §8), it becomes the
**satisfactory baseline** and is frozen. Frozen means: **you do not edit it to test an idea.**
Why so strict: the engine's validation was expensive (trade-by-trade reconciliation against MT5).
A one-line "quick tweak" to test a hypothesis can silently shift fills across millions of bars, and
now every result the engine produces is suspect — including the ones you already trusted. Freezing
protects the validation from regressions you won't notice until much later.
---
## Rule 2 — Engine changes happen in an ISOLATED FORK
Any hypothesis that needs the engine to behave differently is implemented as a **separate engine
file** — a copy of the frozen engine plus the minimal hook for the experimental change. Never by
editing the frozen file.
```
shared/core/grid_engine.py ← frozen, never touched
shared/core/grid_engine_<idea>.py ← fork: copy + one experimental hook
```
The process is fixed:
1. **Fork** = exact copy of the frozen engine + the experimental change, **defaulting to OFF**.
2. **Regression-verify:** run the fork with the change *disabled* and confirm it reproduces the frozen
engine **1:1** — identical Net, drawdown, and trade count — on identical data. If it doesn't, your
copy isn't clean; fix that before testing anything.
3. **A/B** the hypothesis: fork-with-change vs frozen-baseline on the same data.
4. **Decide:** significant, robust improvement → consider promotion (Rule 3). Otherwise **delete the
fork** — revert is just removing a file, the core was never touched.
The regression-verify in step 2 is the linchpin. It proves your fork differs from the baseline *only*
in the one thing you're testing, so the A/B measures the hypothesis and nothing else.
---
## Rule 3 — Promoting a fork to the new baseline is a deliberate, rare event
A fork stays a fork until it has passed a **full round**: A/B + optimization + MT5 verification + an
explicit human decision. Only then does it replace the frozen engine as the new baseline — and at that
moment **every preset/EA built on that engine must be re-validated**, because the ground shifted.
This is never automatic. No script and no agent swaps the baseline engine on its own. It is a manual
procedure under human supervision, precisely because it invalidates prior numbers and forces
re-verification.
---
## Rule 4 — Instruments are isolated as data, not code
Everything symbol- or broker-specific lives in an `InstrumentConfig` object (doc 05), never in the
engine and never hard-coded in a strategy. The engine pulls tick value, spread model, swap, and lot
steps *from the config*.
Consequences:
- Testing the same strategy on a different symbol = a different config, **zero engine changes**.
- Cost-stress testing = `real` / `worst_case` / `best_case` variants of the same config.
- A broker change = edit one file.
If you ever find yourself writing `if symbol == ...` in the engine, that logic belongs in the config.
---
## Rule 5 — Strategies are isolated folders; iterations copy, not import
Each strategy is a self-contained folder (doc 02 §4). Within it, each **iteration** (one research
attempt) carries its **own snapshot** of the run script, search space, and answers.
- Iterations **copy** glue code rather than importing a shared "current" version. That way an
iteration from months ago still runs and reproduces exactly, even after you've changed how you do
things. Reproducibility beats DRY here — these are lab notebooks, not production code.
- Shared *infrastructure* (engine, indicators, instruments, optimizer) is imported normally; it is
import-stable by Rule 1.
---
## Rule 6 — Gates and robustness sit OVER the frozen engine, read-only
Two kinds of add-on never modify the engine:
- **Gates** (`shared/gates/`) are boolean masks AND-ed into the signal *before* it reaches the engine:
`allow = directional_signal & ~block_condition`. A regime filter, a time-of-day filter, an
exhaustion filter — all are caller-side masks. The engine still just consumes a signal array.
- **Robustness layers** (`shared/robustness/`, doc 06) are read-only analyses *over* a finished result
(the study, the trade list, the equity curve). They never touch the engine or mutate results.
Keeping both outside the engine means you can add or remove a filter or a check without re-validating
the engine.
---
## Rule 7 — The registry is curated, locked, and append-only
`registry/` holds only results that passed **all three** gates: Python search → MT5 verification →
human approval. Rules for it:
- **Append-only in spirit:** you don't edit an approved entry; a new finding is a new entry.
- **Locked:** registry entries are not casual scratch space. Treat them like committed releases.
- **Self-documenting:** each entry records the params, the Python metrics, the MT5 report, and the
context (period, instrument, why it was approved).
The registry is the *trustworthy* slice of everything you ever tried. Scratch experiments you don't
intend to keep go in a `results/` scratchpad, never the registry.
---
## Rule 8 — Heavy compute runs detached; never relaunch a running job
A full-history A/B (millions of bars) or a full Optuna study is **minutes** of compute and will block
a foreground shell or hit a timeout. Discipline:
- **Always run heavy jobs in the background**, then poll (read the study's trial count, tail the
output file). Don't sit blocked online.
- **Smoke-test first:** a tiny run (e.g. 3040 trials, a short period) validates the script end-to-end
before you commit to one full run.
- **Dedup guard:** before launching, check whether the same job is already running. Relaunching on a
perceived timeout spawns orphan duplicates that thrash the CPU and corrupt nothing but waste
everything. One study with N workers beats N competing scripts.
---
## Why these rules pay off
Individually each rule is a small constraint. Together they guarantee three properties that a pile of
ad-hoc scripts never has:
1. **Trust** — the engine that produced a number is the same validated engine, every time.
2. **Reproducibility** — any past result can be re-run from its own snapshot.
3. **Fast, safe revert** — a failed idea is a deleted fork, not a half-removed change festering in the
core.
Next: [`05-config-and-inputs.md`](05-config-and-inputs.md) — how the test inputs themselves are kept
separate and declared.