feat: multi-TF SMC scalping pipeline + critical leakage fixes
Add M1+M15 multi-timeframe SMC scalping training pipeline (GPU XGBoost), then fix data-leakage and non-stationarity issues found in a skeptical audit. Pipeline: - src/triple_barrier.py: TP/SL/time labeling (ATR-scaled, asymmetric RR) - src/multi_tf_dataset.py: M1 base + M15 HTF context, point-in-time join_asof (only CLOSED M15 candles visible to each M1 bar - proven no leakage) - src/economic_calendar.py: point-in-time forecast/actual/surprise provider - src/smc_polars.py: add premium/discount + displacement SMC features - scripts/train_multitf_scalper.py: GPU (device=cuda) training + walk-forward - scripts/download_training_data.py: 1y data downloader Leakage / robustness fixes (audit): - CRITICAL: order block signal was written to the ORIGIN bar (future info); now assigned at the CONFIRMATION bar -> matches live conditions - replace non-stationary absolute features (ema_9/21, macd*) with scale-free forms (ema*_dist_atr, ema_spread_atr, macd_*_bps) -> valid at any price level - drop constant-zero calendar features from defaults (recurring provider has no real values); re-add when a real calendar CSV is configured - walk-forward + train/test now embargo the max_holding label horizon and drop warmup rows (NaN->0 artifacts) - news calendar features remain point-in-time (actual only at/after release) Honest result: after fixes the spurious +2.35% edge collapses to ~random (AUC 0.49). The prior edge was caused by the order-block look-ahead. Pipeline is now leakage-free; a real edge still needs more M1 history / better features. Also: test infra (pytest.ini asyncio, hmmlearn), TRAIN_BARS, cleanup of dead modules. 14 tests pass.
This commit is contained in:
+6
-2
@@ -230,12 +230,16 @@ def main():
|
||||
return
|
||||
|
||||
try:
|
||||
# Fetch data - MORE DATA for better generalization
|
||||
# Fetch data - ~1 year of history for better generalization.
|
||||
# M15: ~96 bars/day * ~260 trading days ≈ 25k bars per year.
|
||||
# MT5 returns as many bars as the broker has (up to this cap), so a
|
||||
# higher cap simply grabs the maximum available. Override via TRAIN_BARS.
|
||||
train_bars = int(os.getenv("TRAIN_BARS", "35000"))
|
||||
df = fetch_training_data(
|
||||
connector,
|
||||
config.symbol,
|
||||
config.execution_timeframe,
|
||||
bars=15000, # Increased for better HMM regime separation
|
||||
bars=train_bars,
|
||||
)
|
||||
|
||||
# Prepare features
|
||||
|
||||
Reference in New Issue
Block a user