docs: lock in Study C #22 + Study D #23 state (wwf-lagbot deployed, shakedown findings, batch-two verdicts)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
jaxperro
2026-07-23 17:09:09 -04:00
parent d4053ee2b9
commit 41812b63f6
4 changed files with 88 additions and 26 deletions
+25 -2
View File
@@ -683,6 +683,27 @@ kills to its survivors: **speed is the moat where repricing is fast, and
every edge we can actually reach lives where repricing is slow — stale
siblings, absorbed inventory, and the sharps' own entry prices.**
**Stage-2 the same night: the lead-lag edge gets its instrument
(2026-07-23).** T9's +$9.73/$100 carried a stated optimism — entries at
stale prints nobody may still be quoting. Per the fill-model lesson we
didn't argue with the number, we built the instrument: **wwf-lagbot**
(Study D, #23, deployed 2026-07-23 20:56Z) buys the lagging sibling's
book for real — paper $100 FAKs, premium cap stale+4¢, and down-moves
routed through the sibling's *complement* token (chasing a crashing
sibling's own asks would measure the mirage, not the lag). Both arms of
the question are instrumented: every attempt logs the standing ask's
premium over the stale print, so the **observational kill-switch (median
premium ≥ +8¢ over 3 days = the sim's entries never existed)** can kill
the study without waiting for a paper sample. Two shakedown hours taught
what the tape couldn't: the first premium sample's median sat exactly AT
that bar (+8.5¢ — mirages concentrate in handicap/spread siblings wearing
ancient prints under 0.920.98 asks), and a stale-but-cheap ask can be
stale-but-*empty* ($0.55 of $100 filled on a +1.2¢ "bargain") — so the
grader reports EV per episode **and** per dollar staked, and partials
can't flatter the verdict. Either outcome is a finding: the premium
median kills the mirage in three days, or the paper ledger prices the
first edge that survives its own execution.
## Repo layout
- `insider.py` — the detector: z-score/p-value, timing/freshness/sizing signals,
@@ -699,8 +720,10 @@ siblings, absorbed inventory, and the sharps' own entry prices.**
in-progress system — this finder is selection + tracking only.*
- `research/` — the tape-era edge factory (SILO'd from the bots): read-only
RTDS loaders, execution sim calibrated on the live ledger, pre-registered
studies (#16 surge momentum, #17 oracle fair value), nightly forward
ledger. Verdicts come from `research/forward_ledger.jsonl` only. See
studies (#16 surge and #17 oracle — both killed chain-true; #22
lean-follow and #23 lead-lag — windows open), three measurement harnesses
(wwf-surgebot A2, wwf-oraclebot, wwf-lagbot), nightly forward ledger.
Verdicts come from the chain-graded ledgers only. See
`research/README.md`.
- `wide/` — bulk subgraph→DuckDB scanner (survivorship-bias-free, all wallets);
public subgraph frozen at Jan 2026, so historical-only. See `wide/README.md`.
+35 -14
View File
@@ -29,9 +29,13 @@ momentum **CLOSED 2026-07-23** (kill executed as pre-registered — both
instruments chain-true; A2 runs through Friday's #19 read), #17 oracle fair
value (E0.04 killed; higher tiers ledger-positive but harness-vetoed —
taker arm evidence-dead, decision Friday), #19 surge sprint plan (Friday:
disposition + #20/#21 flip decisions). #18 (empty-cond copies
unsettleable) closed same day: RTDS seed enrichment + falsy-cond repair
pass + 1h alarm.
disposition), #20/#21 execution flips **FLIPPED LIVE 2026-07-23 18:36Z**
(windows open, bars at n≥30 each), #22 Study C lean-follow
(pre-registered, frozen 052eda0; window opens with the first nightly
lean rows), #23 Study D sibling lead-lag (pre-registered;
**wwf-lagbot DEPLOYED 2026-07-23 20:56Z**, shakedown numbers on the
issue). #18 (empty-cond copies unsettleable) closed same day: RTDS seed
enrichment + falsy-cond repair pass + 1h alarm.
## Operating boundary (user, 2026-07-13 — standing)
**Full autonomy on the bots**; the real-money bot **stays ARMED**. Never
@@ -99,7 +103,20 @@ push through.
sim ledger reads positive. Both instruments share the scorer now, so the
divergence is the sim's 6.7s-lag FILL model flattering the tiers (round
3's lesson recurring in the fill model). Taker arm is evidence-dead;
formal tier bars keep accruing; maker pivot (T1 sim) is the successor.**
formal tier bars keep accruing; maker pivot (T1 sim) KILLED at Stage 1
same night — staleness IS the adverse selection.**
- **wwf-lagbot (PAPER, ~$3/mo, deployed 2026-07-23 20:56Z)**: Study D #23
— T9's sibling lead-lag edge (+$9.73/$100 tape, n=2,028, stale-print
optimism stated) at real books. Leader bursts ≥10¢/120s → paper-FAK
$100 on ≤2 lagging same-outcome siblings at ≤ stale+4¢; down-moves buy
the sibling's COMPLEMENT token; cooldown 600s/event. Attempts log every
standing ask premium — the observational kill-switch (median ≥+8¢ over
3 days = mirage) can kill without a paper sample. Shakedown 07-23: first
8 attempts' median +8.5¢ (AT the bar — mirages concentrate in
handicap/spread siblings with ancient prints); the 3 fills were +1-2¢,
one $0.55-of-$100 partial → grade_lag reports $/episode AND %-of-staked.
PASS ≥+$4/$100 @ n≥400/≥5d · KILL ≤0 @ n≥300. Nightly grade_lag.py →
lag_paper_ledger.jsonl.
- **Data moat (2026-07-22, DATA LAW in research/README)**: all raw streams
append-only and Mac-independent (Fly volumes + daily snapshots; recorder
has ~3+ weeks offline headroom); forward.py backfills ledger-missing
@@ -111,21 +128,25 @@ push through.
Friday's combined read (#13 bench + #16 formal close [A2 chain grade
$7.54/fill × 1,344 independently confirms the kill; virtual book's
+26% on n=46 is the variance footnote, not a signal] + #17 taker-arm
decision + #19 disposition + **#20/#21 flip decisions** — every number
chain-true). Five tandem tests 2026-07-23 (research/: copy_edge_slices,
copy_maker_entry, maker_sharps, sibling_sum_scan, sell_mirror_study):
T3 maker entries +$17.45 vs +$12.86 taker → #20; sells anti-signal →
#21; 673-wallet maker-sharp species (+$5.9M, 86% invisible to taker
screens) → inventory-lean follow-on; sibling-sum = print artifact,
parked; T1 crypto maker-quote sim needs its re-run.
- Dashboards: jaxperro.com/{trading,live,test,value} — /test = both paper
decision + #19 disposition + early #20/#21 window check — every number
chain-true) · **#22 lean-follow** (PASS ≥+$2/lean & hit≥.56 @ n≥1,500;
KILL ≤0 @ n≥1,000) · **#23 lead-lag** (bars above). Ten tandem tests
2026-07-23 (scripts + verdicts in research/): batch one — T3 maker
entries +$17.45 vs +$12.86 taker → #20; sells anti-signal → #21;
673-wallet maker-sharp species → inventory-lean line; sibling-sum =
print artifact; T5 esports concentration → #14 tension. Batch two —
T6 lean-follow → #22 (fade arm FAILED its concentration gate,
report-only); T9 lead-lag → #23; T1 maker-quote Stage-1 KILL; T7
settlement-discount industrialized, parked; T8 crater-rejects were
good misses; T10 age gradient needs sample.
- Dashboards: jaxperro.com/{trading,live,test,value} — /test = all four
studies on one page (old /surge + /oracle URLs redirect) · daily pipeline
on the Mac at 08:00 (launchd, lockfile) — floors, bench forward table,
edge row, tape sync, Discord digest · tape mirror every 15 min
(com.jaxperro.tape-sync → sync_tape.py, sftp + base64-console fallback)
· research nightly fires 09:15 then WAITS for fresh tape
(com.jaxperro.research-nightly; self-commits ledger, informed set,
surge/oracle grades, virtual book, meta snapshot). All launchd agents
(com.jaxperro.research-nightly; self-commits ledger + lean rows, informed
set, surge/oracle/lag grades, virtual book, meta snapshot). All launchd agents
removable with `launchctl unload ~/Library/LaunchAgents/<label>.plist`.
## Ops quick-reference
+2 -2
View File
@@ -29,7 +29,7 @@ Three deployed pieces + one static dashboard:
| **VALUE bot** (CLOSED 2026-07-19 — [`archive/value/README.md`](archive/value/README.md)) | ~~Fly app `wwf-valuebot`~~ destroyed after the paper verdict: 994 tickets, 1W/993L, 0.075x — today's sub-2¢ asks are fully informed; the 2025-era calibration edge is gone — a hard SILO: own code (`value/valuebot.py`, zero copybot imports), own state/feed, own image, no shared wallet ever | Systematically buys sub-2¢ contracts (the calibration study's one underpriced bucket — [`archive/value/PLAN.md`](archive/value/PLAN.md)); honest FAK fill model, chain-truth settles, $1 tickets, event cap 1. Verdict at ~2k resolved tickets vs a ~1.05% break-even hit rate |
| **RTDS tape recorder** ([`recorder/README.md`](recorder/README.md)) | **Fly.io app `wwf-recorder`, `arn`, 24/7** — third silo: no repo clone (code in image), no keys, 25GB volume | Records the FULL firehose (~8M events/day): every trade, order MATCH (maker side), comment, crypto tick — dual-socket capture (one socket alone measured 92.9% coverage; the twin covers per-conn silences). **Stage-0 warehouse (2026-07-21): the box folds its own segments into row-verified Parquet on the volume; the Mac mirrors + appends into `live/rtds.duckdb` every 15 min** (`com.jaxperro.tape-sync`). Nothing deletes without a verified second copy (gz needs its parquet; parquet needs the Mac's ack) |
| **Discord digest** (`live/discord_daily.py`) | end of the daily pipeline | one message/day: the sharp list with profile links + 30-day conviction stats (per-trade pings retired 2026-07-04 for PAPER; the LIVE book pings every real placement/exit/settle) |
| **dashboard** | [jaxperro.com/trading](https://jaxperro.com/trading) + [jaxperro.com/live](https://jaxperro.com/live) + [jaxperro.com/test](https://jaxperro.com/test) (static, in the `jaxperro` repo) | `/trading` renders the paper book, backtest book and sharp table; `/live` is the REAL MONEY page (reads `live/copybot_live_real.json` — live since 2026-07-10); `/test` is BOTH paper studies on one page (A2 surge measurement arm + oracle harness, each reading its bot's /feed + nightly chain-graded ledger; old /surge and /oracle URLs redirect). Wallet cards everywhere show LIFETIME numbers server-side (window sums lied twice — 2026-07-21/22) |
| **dashboard** | [jaxperro.com/trading](https://jaxperro.com/trading) + [jaxperro.com/live](https://jaxperro.com/live) + [jaxperro.com/test](https://jaxperro.com/test) (static, in the `jaxperro` repo) | `/trading` renders the paper book, backtest book and sharp table; `/live` is the REAL MONEY page (reads `live/copybot_live_real.json` — live since 2026-07-10); `/test` is ALL FOUR studies on one page (A2 surge + oracle + lead-lag harnesses reading their bots' /feeds + nightly chain-graded ledgers, plus the Study C lean forward window from the forward ledger; old /surge and /oracle URLs redirect). Wallet cards everywhere show LIFETIME numbers server-side (window sums lied twice — 2026-07-21/22) |
**The calibration experiment (running now):** a fresh $1,000 paper book,
**reset 2026-07-08** so it measures exactly one thing — the follow set in
@@ -137,7 +137,7 @@ does).
| `insider.py` | the original detector: z-score, pre-resolution timing, fresh-wallet flags, funding-cluster rings |
| `smart_money.py` | shared HTTP helper + survivorship-corrected win-rate dashboard (`:8899`) |
| `archive/webhook_receiver.py` | retired 2026-07-04: Alchemy webhook → per-trade Discord pings (replaced by `live/discord_daily.py`'s daily digest) |
| `research/` | tape-era edge factory, **SILO'd from the bots** ([research/README](research/README.md)): read-only RTDS loaders, calibrated execution sim, pre-registered studies scored through a MANDATORY chain-truth overlay (the 2026-07-22 scorer law — tape-resolution timing is win-biased and killed Study A #16 at $6/fill once corrected, independently confirmed by the A2 measurement arm at $7.54/fill chain-true; oracle #17: E0.04 killed, higher tiers ledger-positive but vetoed by the harness's real-latency chain grade — the 2026-07-23 fill-model lesson), TWO measurement harnesses (wwf-surgebot A2 + wwf-oraclebot: every-trigger $100 paper FAKs, attempts/markouts/settles append-only, bankroll specs replayed offline), nightly forward ledger — the only source of study verdicts |
| `research/` | tape-era edge factory, **SILO'd from the bots** ([research/README](research/README.md)): read-only RTDS loaders, calibrated execution sim, pre-registered studies scored through a MANDATORY chain-truth overlay (the 2026-07-22 scorer law — tape-resolution timing is win-biased and killed Study A #16 at $6/fill once corrected, independently confirmed by the A2 measurement arm at $7.54/fill chain-true; oracle #17: E0.04 killed, higher tiers ledger-positive but vetoed by the harness's real-latency chain grade — the 2026-07-23 fill-model lesson), THREE measurement harnesses (wwf-surgebot A2 + wwf-oraclebot + wwf-lagbot: every-trigger $100 paper FAKs, attempts/markouts/settles append-only, bankroll specs replayed offline), nightly forward ledger — the only source of study verdicts. Open pre-registered windows: #20/#21 execution flips (live 2026-07-23), Study C #22 lean-follow (tape-scored forward), Study D #23 sibling lead-lag (wwf-lagbot, deployed 2026-07-23) |
| `wide/` | frozen-subgraph bulk scanner (1.76M wallets, historical only — subgraph froze Jan 2026) |
| `archive/` | everything retired, kept honest ([archive/README](archive/README.md)): the six failed strategies, earlier research sweeps (`hunt/huntwide/oos/copyback`), the superseded live selection layer (`live-research/`), the scrapped Polymarket-US venue probe (`us-venue/`), and retired infra (`retired-infra/`: Railway config, Mac launchd runner, GH-Actions cron) |
+26 -8
View File
@@ -59,13 +59,16 @@ Layout:
dominated; needs standing-book data; parked)
maker_lean.py T6 — WALK-FORWARD POSITIVE: follow small maker
inventory leans (+$2.51/lean, 59% hit, positive
all 3 days); fade $2k+ whale bags (+$13.38,
n=422). Pre-registration candidate pending
event-concentration + execution checks.
all 3 days) → STUDY C #22 PRE-REGISTERED
(params/study_lean.json frozen 052eda0; follow
arm verdict-gated, fade $2k+ FAILED its
event-concentration gate — report-only; nightly
rows via forward.py score_lean, study:"lean")
event_leadlag.py T9 — POSITIVE: same-outcome siblings reprice
slowly after leader bursts (+$9.73/$100 chain,
n=2,028; stale-print entry optimism stated
Stage-2 is execution realism)
n=2,028; stale-print entry optimism stated) →
STUDY D #23: wwf-lagbot harness (below) is the
Stage-2 execution-realism instrument
settle_discount.py T7 — settlement-discount niche: industrialized
by incumbents (+2-3%/hold for the good ones);
residual unproven (63% of vol outside
@@ -96,9 +99,21 @@ Layout:
jaxperro.com/test): fair value tick-by-tick on the venue's
own settlement feed, all E-tiers tracked, three settle
layers (own-feed tick → CLOB flags → nightly chain truth)
grade_surge.py / grade_oracle.py nightly chain-truth re-grades
surge_meas_ledger.jsonl / oracle_paper_ledger.jsonl;
also pull the raw volume streams to local .pull copies
lagbot.py Study D real-time PAPER harness (wwf-lagbot
jaxperro.com/test, deployed 2026-07-23, SEM_VER l1):
leader bursts ≥10¢/120s → paper-FAK $100 on ≤2 lagging
same-outcome siblings at ≤ stale+4¢; DOWN-moves buy the
sibling's COMPLEMENT token; cooldown 600s/event, 30-min
warmup (leader maps built live from the stream). Every
attempt logs the ask premium over the stale print — the
observational kill-switch (median ≥+8¢ over 3 days =
mirage) can kill #23 without a paper sample
grade_surge.py / grade_oracle.py / grade_lag.py nightly chain-truth
re-grades → surge_meas / oracle_paper / lag_paper
ledgers (grade_lag also prints EV per episode AND
%-of-staked — partial fills can't flatter the verdict —
plus the median-premium kill-switch read); all three
pull the raw volume streams to local .pull copies
markout_flow.py exploratory markout-exit curve (chain-truth verdict: NO
scalp inside the dead surge signal — losers bleed from
minute one; v0's res_tok version was round-3-biased)
@@ -121,6 +136,9 @@ Data streams (the moat — all append-only, per DATA LAW):
.jsonl (durable settles) · surge2_state.json
wwf-oraclebot /data oracle_attempts / oracle_markouts / oracle_settles
.jsonl + oracle_state.json (same shapes)
wwf-lagbot /data lag_attempts.jsonl (every attempt + top-5 asks +
premium vs stale + latency) · lag_settles.jsonl
(durable settles) · lag_state.json
Mac (nightly pulls + git): ledgers + params committed; raw .pull copies
and meta/ gzips local. forward.py backfills any tape-covered day the
ledger has never seen — a Mac gap > RESCORE_DAYS leaves no holes.