96680cde63f8c3970105782bdeb9e573fe0541bd
7
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
20fd86f065 |
fix: dashboard goes silent during Phase 1 (no AI thinking, empty Param Changes panel)
Three issues caused the dashboard to look frozen during the long Phase 1 exploration phase, even when 2+ runs had completed: 1. AI Thinking feed only got ONE entry (the phase-start banner) and then went silent until Phase 2. Phase 1 is LHS sampling so there are no AI calls — but we can still narrate observations rule-based. Added _narrate_phase1_run() that emits a per-result message tuned to what actually happened: 'PF 1.8, DD 14% — strong region, worth refining' (success), '$-734 loss with 191 trades — skipping' (warning), 'only 3 trades — params too restrictive' (info), etc. Mediocre runs are throttled to 1-in-5 to avoid feed spam 2. Parameter Changes panel showed 'No iterations yet — AI-driven param edits appear per iteration' which is misleading during Phase 1 (where there are no iterations YET, by design). Empty-state message now phase-aware: shows 'Exploration phase — changes appear in Phase 2' during phase1, 'Waiting for first AI iteration' if autonomous, or 'Random-neighbor refinement — enable Autonomous AI for AI changes' if not autonomous 3. AI Summary still said 'Waiting for optimization to start...' even while the optimization was actively running. New refreshAIWaitingState() updates the message based on (a) whether a run is active and (b) whether the API key is set. When running with no key, shows 'AI insights are off — no Anthropic API key set' with a 'Configure API key in Settings →' link Frontend now also tracks state.phase and state.autonomous from phase_start events so all empty-state messages stay in sync. |
||
|
|
51a73c953b |
docs: Cerebral Valley Opus 4.7 hackathon submission packet
- AIReasoner default model now claude-opus-4-7 (was sonnet-4-6); pipeline reads ai.model from config and passes to constructor. The reasoner class attribute is also Opus 4.7 so any caller without explicit model defaults to the hackathon model - README hero gains a "Built with Claude Opus 4.7" badge + a line linking to the Cerebral Valley × Anthropic hackathon page - New SUBMISSION.md — copy-paste blocks for every field of the submission form: name, taglines, short / long description, "how Claude is used" detail, GitHub URL, demo instructions, tags, tech stack, team, license, and a final pre-submission checklist |
||
|
|
a584f46891 |
feat: 6 user-facing upgrades + realistic demo metrics + animated hero
Demo + assets
- Synthetic backtests now occasionally fail (regime failures, OOS degradation)
so verdicts span RECOMMENDED / RISKY / NOT_RELIABLE realistically. Phase 1
has ~35% failure rate, Phase 2 AI loop ~10%, OOS has 30% chance of severe
degradation — matches what real markets look like
- screenshots/dashboard.png + best_result_modal.png regenerated against the
current UI; screenshots/apex_demo.gif (6-frame autonomous-run timelapse)
embedded in the README
FEATURE 1 — Live AI token streaming
- AIReasoner._call_claude() now streams via SSE when a callback is
registered. Each text delta forwards to the dashboard as
`ai_thinking_chunk` events
- The Live AI Thinking Feed renders a single growing bubble with a blinking
cursor while text streams in, finalising on `end`. Looks and feels like
watching the AI type
FEATURE 2 — Pre-flight check on /setup
- New /api/preflight endpoint runs 5–7 probes: config readable, API key
set, MT5 paths exist (skipped in demo), EA registered, reports folder
writable. Returns {ok, blocking_count, checks[]}
- Setup page renders a colour-coded checklist on load and refocus.
Replaces "click Start, wait 5s, see generic error"
FEATURE 3 — Hot-reload settings into the running pipeline
- pipeline.reload_config() applies AI model / timeout / API-key swaps to
the live reasoner mid-run. Threshold changes surface for next run
- /api/settings POST detects a running pipeline and calls reload_config(),
returning the changed keys plus a "hot-reloaded into the running
optimization" note
FEATURE 4 — Replay scrubber on Best Result
- Evolution path now renders as an interactive scrubber: range slider +
prev/next/play buttons. Each step shows the run ID, phase, score, full
metrics grid, parameter changes for that step, and the AI's analysis
text — auto-plays at 700ms/step
FEATURE 5 — Compare runs on /reports
- Each card has a checkbox; selecting 2–4 reveals a floating Compare bar.
Compare modal renders a side-by-side table with metric winners
highlighted (Calmar / PF / profit favour higher; DD favours lower)
and a parameter-diff section showing changed values
FEATURE 6 — Discord / Slack / generic webhook on completion
- New `notifications.webhook_url` + `webhook_style` config keys
- Auto-detects Discord vs Slack from the URL host. Posts a one-line
summary on `optimization_complete`: verdict + best run + PF/Calmar/DD/
profit/trades/elapsed
|
||
|
|
c42345ea1e |
fix: live-activity persistence on refresh + reports page rebuild
Refresh-during-run no longer wipes the dashboard - pipeline now keeps in-memory rolling logs of ai_thinking, param_changes, validation events, and the current early-termination state. Logs reset at the start of each new run (capped: 200 thinking / 50 param-change records / 40 validation events) - new GET /api/live_activity returns those logs + current phase + running state in one shot — the dashboard hits it on page load - restoreHistory() in dashboard.js now replays each event into addThinking / addParamChanges / validationRunStart-Complete-Done / showEarlyTermination, and re-applies the active phase via setPhaseActive. Reload F5 mid-run no longer shows a fresh empty dashboard - addThinking() preserves the original ts on replay (was using nowStr() so every replayed entry got the refresh time) Reports page (/reports) rebuilt - previous template crashed with 500 on legacy summary.json files that pre-date the win_rate / drawdown_pct fields. Server now backfills sane defaults for every metric the template touches - runs now sorted by ts (was filesystem-iterdir order) - new template: search bar + filter chips (All / Exploration / AI Iteration / Validation / AI insight / Has .set), phase tags, AI/.set badges, three action buttons per card (View Params / Download .set / Full Report) - per-card detail modal shows metrics grid + full parameters table + AI reasoning + Download .set button — the missing "click to see params + download" path the user reported - has_set detected per-run by globbing run_dir/*.set; AI insight tag shown when ai_insight.json exists on disk |
||
|
|
746ab8fb11 |
fix: hands-on bugs found during pre-submission audit
API + data integrity - /api/settings GET no longer leaks the active Anthropic API key — returns a masked preview (sk-ant-XX…YYYY) plus a boolean `anthropic_api_key_set` flag - /api/settings POST won't overwrite a real key with the masked placeholder the client receives back on GET (length<30 / "…" / "..." / "***" markers trigger a preserve-existing path) - /api/best_result, /api/status, _make_run_dict, _result_to_dict, all AI-loop emits, validation_run_complete, optimization_complete, ai_iteration_complete, ai_targets_met, run_complete: max_drawdown and win_rate are now consistently emitted as PERCENTAGES (0–100), matching the dashboard's existing display formatters. They were previously emitted as fractions (0.13 = 13%) so the UI rendered "0.13%" instead of "13%" - ResultRanker.make_result() now sets `passing` and `raw_score` on every result it produces. Previously these were only set during a full ranker.rank() pass, so individual Phase 2 / Phase 3 runs hit _make_run_dict with passing=False even when they cleared all gates (history showed "0 passing" when 22/22 actually passed) |
||
|
|
6caafdb794 |
feat: open-source release — AI-driven autonomous optimization with live visibility
Major upgrade making the AI loop visible and the project ready for public release. UI / UX - Live AI Thinking Feed: streams reasoning, decisions, and outcomes per iteration - Parameter Changes panel: prev → new + reason for every AI-driven edit - Validation Activity panel: out-of-sample + sensitivity runs with live metrics - Early Termination banner: surfaces why optimization stopped (targets met, no profit, budget, stuck, user stop) - 3-phase tracker renamed Exploration / Iteration / Validation with live N/total - Best Result modal exposes Evolution Path showing how the AI arrived at the winner - Run-detail modal accessible from every recent run row - Setup form validation (dates, walk-forward order, params selection, AI targets) - Pause button removed; misleading sidebar nav consolidated to Dashboard / New Run / Reports / Source Backend - AIGuidedLoop streams ai_thinking, param_changes, ai_targets_met, ai_stuck - Pipeline emits validation_start / validation_run_start / validation_run_complete / validation_done - Pipeline emits early_termination on every early-stop path - /api/best_result returns best run + full evolution chain - /api/run/<id> + /api/runs sorted by ts - AIReasoner falls back to ANTHROPIC_API_KEY env var when config is a placeholder - Demo mode (APEX_DEMO_MODE=1) generates deterministic synthetic backtests so judges can run end-to-end without MT5 Open-source readiness - README.md with pitch, demo flow, architecture diagram, quickstart, event reference - LICENSE (MIT) - config.example.yaml template (config.yaml now git-ignored) - requirements.txt: added anthropic / requests / psutil / beautifulsoup4, capped majors - .gitignore: secrets, *.set, scratch screenshots, ea_registry.yaml - demo/run_demo.py: one-command offline demo runner - 10 polished screenshots for README + judge review |
||
|
|
47f6012eb2 |
feat: Smart Autonomous 3-Phase Optimizer — complete redesign
PROBLEM: Old system repeated identical params, all scores flat at 0.2500,
user had zero control over symbol/TF/dates. Not smart, not dynamic.
NEW ARCHITECTURE:
optimizer/ (NEW package)
├── __init__.py
├── session_config.py User choices (EA, symbol, TF, dates, budget, objective)
├── lhs_sampler.py Latin Hypercube Sampling — diverse exploration
├── result_ranker.py Relative scoring (best in session=1.0, worst=0.0)
├── budget.py Time budget tracker
└── pipeline.py 3-phase orchestrator
Phase 1 — Broad Discovery (LHS, 20-26 runs):
Samples FULL parameter space, not just defaults±tiny step
LHS guarantees coverage: all 17 optimizable LEGSTECH params explored
Relative ranking: profitable configs float top, losers score 0
Phase 2 — Refinement (9 runs):
Neighbor search around top 3 configs at ±20% range (not ±0.5 step)
Keeps best of Phase1 vs Phase2 — never regresses
Phase 3 — Validation (5 runs):
OOS backtest on unseen data period
Sensitivity test: nudge params ±20%, detect fragility
Verdict: RECOMMENDED / RISKY / NOT_RELIABLE
Output: Clean downloadable .set file via /download_set/<run_id>
UI REDESIGN:
ui/templates/landing.html New / homepage (was old dashboard)
ui/templates/setup.html New /setup — EA, symbol, TF, dates, budget, objective
ui/templates/dashboard.html Updated /dashboard with:
- 5-step phase indicator
- Real progress bar per run
- Phase 1 results table (top 5 after phase1)
- Verdict banner with download button
- No-profitable-config warning
ui/static/js/dashboard.js Handles 8 new pipeline SocketIO events
app.py New routes: /, /setup, /dashboard
/api/start accepts full SessionConfig JSON
/download_set/<id> serves optimized .set
ea/registry.py +list_all() for setup page dropdown
ui/templates/reports_index.html Back to Dashboard → /dashboard (was /)
ui/static/css/style.css +dot-warn, dot-done, profit-pos/neg, aliases
FIXES:
Score no longer flat 0.2500 (was: absolute thresholds on losing EA)
User now controls: symbol, timeframe, dates, budget, objective
Parameters now span full range (was: tiny step from defaults)
Verdict is actionable: RECOMMENDED / RISKY / NOT_RELIABLE with reason
TESTED:
8/8 pre-flight checks pass
Browser test: landing ✓, setup form ✓, /dashboard ✓,
phase indicator active ✓, /reports ✓, back link ✓
|