Files
LSTM-PPO-Reinforcement-Lear…/README.md
T

158 lines
2.9 KiB
Markdown
Raw Normal View History

2026-06-19 16:59:27 +02:00
# MT5 XAUUSD LSTM PPO Trading Bot
2026-06-19 16:59:27 +02:00
A MetaTrader 5 reinforcement learning trading bot for XAUUSD (Gold).
2026-06-19 16:59:27 +02:00
The bot combines an LSTM neural network for sequence learning with Proximal Policy Optimization (PPO) to learn trading decisions directly from historical market data.
2026-06-19 16:59:27 +02:00
Unlike traditional bots that rely on fixed rules, the PPO agent learns when to Buy, Sell or Hold from thousands of market examples.
2026-06-19 17:01:42 +02:00
The program should be in a folder or the desktop where the 5m csv file with OHLC data from https://www.kaggle.com/datasets/novandraanugrah/xauusd-gold-price-historical-data-2004-2024 is, in training it will then make a LSTM-PPO-saves folder where the training is saved (used for making decisions, also in testing).
---
## Features
2026-06-19 16:59:27 +02:00
### Reinforcement Learning
- LSTM policy network
- PPO training
- Continuous online training
- Live inference on MT5
- Automatic checkpoint saving/loading
### Technical Indicators
- EMA 7
- EMA 21
- EMA Difference (Momentum)
- ADX
- +DI
- -DI
- Stochastic
- VWAP
- VWAP Bands
- VWAP Slope
- Volume Moving Average
### Smart Money Concepts
- Bullish Order Blocks
- Bearish Order Blocks
- Bullish Fair Value Gaps
- Bearish Fair Value Gaps
- Bullish Rejection Blocks
- Bearish Rejection Blocks
- Equal Highs
- Equal Lows
- Market Breaks
- Indecision Candles
### PPO State Features
Current state contains:
- OHLC
- EMA trend
- Momentum
- ADX trend strength
- DI Direction
- Stochastic
- VWAP
- VWAP Bands
- VWAP Position
- VWAP Slope
- Volume MA
- Order Blocks
- Fair Value Gaps
- Rejection Blocks
- Equal Highs/Lows
- Buy Score
- Sell Score
---
2026-06-19 16:59:27 +02:00
## Current Strategy
2026-06-19 16:59:27 +02:00
Current reward structure:
2026-06-19 16:59:27 +02:00
- Take Profit: 20 pips
- Stop Loss: 40 pips
- Risk/Reward: 1 : 0.5
2026-06-19 16:59:27 +02:00
The current focus is high-probability momentum trades rather than large swing trades.
---
2026-06-19 16:59:27 +02:00
## Training
2026-06-19 16:59:27 +02:00
The PPO agent trains continuously over historical MT5 data.
2026-06-19 16:59:27 +02:00
Example metrics during training:
2026-06-19 16:59:27 +02:00
- Win rate: 7085%
- Profit Factor: 1.53+
- Weekly performance: typically 1030R during training (varies by market conditions)
2026-06-19 16:59:27 +02:00
These figures are training statistics only and are not guarantees of future performance.
---
## Requirements
2026-06-19 16:59:27 +02:00
- Python 3.11+
- MetaTrader 5
- MetaTrader5
- pandas
- numpy
- torch
2026-06-19 16:59:27 +02:00
Install dependencies:
2026-06-19 16:59:27 +02:00
```bash
pip install MetaTrader5 pandas numpy torch
```
---
## Running
Train:
```bash
2026-06-21 08:21:33 +02:00
python mt5-xau-lstm-ppo-bot.py --train
2026-06-19 16:59:27 +02:00
```
Live trading:
```bash
2026-06-21 08:21:33 +02:00
python mt5-xau-lstm-ppo-bot.py --test
2026-06-19 16:59:27 +02:00
```
Train and trade simultaneously:
```bash
2026-06-21 08:21:33 +02:00
python mt5-xau-lstm-ppo-bot.py --train --test
2026-06-19 16:59:27 +02:00
```
---
## Project Goals
Current work focuses on:
- Improving PPO policy learning
- Better feature engineering
- Dynamic trade management
- VWAP and volume analysis
- Smart Money Concept detection
- MFE/MAE prediction research
- Higher timeframe context
---
## Disclaimer
2026-06-19 16:59:27 +02:00
This project is for educational and research purposes only.
2026-06-19 16:59:27 +02:00
Trading leveraged products involves substantial risk. Always test thoroughly on historical data and demo accounts before risking real capital.