docs: lock in Study C #22 + Study D #23 state (wwf-lagbot deployed, shakedown findings, batch-two verdicts)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
jaxperro
2026-07-23 17:09:09 -04:00
parent d4053ee2b9
commit 41812b63f6
4 changed files with 88 additions and 26 deletions
+25 -2
View File
@@ -683,6 +683,27 @@ kills to its survivors: **speed is the moat where repricing is fast, and
every edge we can actually reach lives where repricing is slow — stale every edge we can actually reach lives where repricing is slow — stale
siblings, absorbed inventory, and the sharps' own entry prices.** siblings, absorbed inventory, and the sharps' own entry prices.**
**Stage-2 the same night: the lead-lag edge gets its instrument
(2026-07-23).** T9's +$9.73/$100 carried a stated optimism — entries at
stale prints nobody may still be quoting. Per the fill-model lesson we
didn't argue with the number, we built the instrument: **wwf-lagbot**
(Study D, #23, deployed 2026-07-23 20:56Z) buys the lagging sibling's
book for real — paper $100 FAKs, premium cap stale+4¢, and down-moves
routed through the sibling's *complement* token (chasing a crashing
sibling's own asks would measure the mirage, not the lag). Both arms of
the question are instrumented: every attempt logs the standing ask's
premium over the stale print, so the **observational kill-switch (median
premium ≥ +8¢ over 3 days = the sim's entries never existed)** can kill
the study without waiting for a paper sample. Two shakedown hours taught
what the tape couldn't: the first premium sample's median sat exactly AT
that bar (+8.5¢ — mirages concentrate in handicap/spread siblings wearing
ancient prints under 0.920.98 asks), and a stale-but-cheap ask can be
stale-but-*empty* ($0.55 of $100 filled on a +1.2¢ "bargain") — so the
grader reports EV per episode **and** per dollar staked, and partials
can't flatter the verdict. Either outcome is a finding: the premium
median kills the mirage in three days, or the paper ledger prices the
first edge that survives its own execution.
## Repo layout ## Repo layout
- `insider.py` — the detector: z-score/p-value, timing/freshness/sizing signals, - `insider.py` — the detector: z-score/p-value, timing/freshness/sizing signals,
@@ -699,8 +720,10 @@ siblings, absorbed inventory, and the sharps' own entry prices.**
in-progress system — this finder is selection + tracking only.* in-progress system — this finder is selection + tracking only.*
- `research/` — the tape-era edge factory (SILO'd from the bots): read-only - `research/` — the tape-era edge factory (SILO'd from the bots): read-only
RTDS loaders, execution sim calibrated on the live ledger, pre-registered RTDS loaders, execution sim calibrated on the live ledger, pre-registered
studies (#16 surge momentum, #17 oracle fair value), nightly forward studies (#16 surge and #17 oracle — both killed chain-true; #22
ledger. Verdicts come from `research/forward_ledger.jsonl` only. See lean-follow and #23 lead-lag — windows open), three measurement harnesses
(wwf-surgebot A2, wwf-oraclebot, wwf-lagbot), nightly forward ledger.
Verdicts come from the chain-graded ledgers only. See
`research/README.md`. `research/README.md`.
- `wide/` — bulk subgraph→DuckDB scanner (survivorship-bias-free, all wallets); - `wide/` — bulk subgraph→DuckDB scanner (survivorship-bias-free, all wallets);
public subgraph frozen at Jan 2026, so historical-only. See `wide/README.md`. public subgraph frozen at Jan 2026, so historical-only. See `wide/README.md`.
+35 -14
View File
@@ -29,9 +29,13 @@ momentum **CLOSED 2026-07-23** (kill executed as pre-registered — both
instruments chain-true; A2 runs through Friday's #19 read), #17 oracle fair instruments chain-true; A2 runs through Friday's #19 read), #17 oracle fair
value (E0.04 killed; higher tiers ledger-positive but harness-vetoed — value (E0.04 killed; higher tiers ledger-positive but harness-vetoed —
taker arm evidence-dead, decision Friday), #19 surge sprint plan (Friday: taker arm evidence-dead, decision Friday), #19 surge sprint plan (Friday:
disposition + #20/#21 flip decisions). #18 (empty-cond copies disposition), #20/#21 execution flips **FLIPPED LIVE 2026-07-23 18:36Z**
unsettleable) closed same day: RTDS seed enrichment + falsy-cond repair (windows open, bars at n≥30 each), #22 Study C lean-follow
pass + 1h alarm. (pre-registered, frozen 052eda0; window opens with the first nightly
lean rows), #23 Study D sibling lead-lag (pre-registered;
**wwf-lagbot DEPLOYED 2026-07-23 20:56Z**, shakedown numbers on the
issue). #18 (empty-cond copies unsettleable) closed same day: RTDS seed
enrichment + falsy-cond repair pass + 1h alarm.
## Operating boundary (user, 2026-07-13 — standing) ## Operating boundary (user, 2026-07-13 — standing)
**Full autonomy on the bots**; the real-money bot **stays ARMED**. Never **Full autonomy on the bots**; the real-money bot **stays ARMED**. Never
@@ -99,7 +103,20 @@ push through.
sim ledger reads positive. Both instruments share the scorer now, so the sim ledger reads positive. Both instruments share the scorer now, so the
divergence is the sim's 6.7s-lag FILL model flattering the tiers (round divergence is the sim's 6.7s-lag FILL model flattering the tiers (round
3's lesson recurring in the fill model). Taker arm is evidence-dead; 3's lesson recurring in the fill model). Taker arm is evidence-dead;
formal tier bars keep accruing; maker pivot (T1 sim) is the successor.** formal tier bars keep accruing; maker pivot (T1 sim) KILLED at Stage 1
same night — staleness IS the adverse selection.**
- **wwf-lagbot (PAPER, ~$3/mo, deployed 2026-07-23 20:56Z)**: Study D #23
— T9's sibling lead-lag edge (+$9.73/$100 tape, n=2,028, stale-print
optimism stated) at real books. Leader bursts ≥10¢/120s → paper-FAK
$100 on ≤2 lagging same-outcome siblings at ≤ stale+4¢; down-moves buy
the sibling's COMPLEMENT token; cooldown 600s/event. Attempts log every
standing ask premium — the observational kill-switch (median ≥+8¢ over
3 days = mirage) can kill without a paper sample. Shakedown 07-23: first
8 attempts' median +8.5¢ (AT the bar — mirages concentrate in
handicap/spread siblings with ancient prints); the 3 fills were +1-2¢,
one $0.55-of-$100 partial → grade_lag reports $/episode AND %-of-staked.
PASS ≥+$4/$100 @ n≥400/≥5d · KILL ≤0 @ n≥300. Nightly grade_lag.py →
lag_paper_ledger.jsonl.
- **Data moat (2026-07-22, DATA LAW in research/README)**: all raw streams - **Data moat (2026-07-22, DATA LAW in research/README)**: all raw streams
append-only and Mac-independent (Fly volumes + daily snapshots; recorder append-only and Mac-independent (Fly volumes + daily snapshots; recorder
has ~3+ weeks offline headroom); forward.py backfills ledger-missing has ~3+ weeks offline headroom); forward.py backfills ledger-missing
@@ -111,21 +128,25 @@ push through.
Friday's combined read (#13 bench + #16 formal close [A2 chain grade Friday's combined read (#13 bench + #16 formal close [A2 chain grade
$7.54/fill × 1,344 independently confirms the kill; virtual book's $7.54/fill × 1,344 independently confirms the kill; virtual book's
+26% on n=46 is the variance footnote, not a signal] + #17 taker-arm +26% on n=46 is the variance footnote, not a signal] + #17 taker-arm
decision + #19 disposition + **#20/#21 flip decisions** — every number decision + #19 disposition + early #20/#21 window check — every number
chain-true). Five tandem tests 2026-07-23 (research/: copy_edge_slices, chain-true) · **#22 lean-follow** (PASS ≥+$2/lean & hit≥.56 @ n≥1,500;
copy_maker_entry, maker_sharps, sibling_sum_scan, sell_mirror_study): KILL ≤0 @ n≥1,000) · **#23 lead-lag** (bars above). Ten tandem tests
T3 maker entries +$17.45 vs +$12.86 taker → #20; sells anti-signal → 2026-07-23 (scripts + verdicts in research/): batch one — T3 maker
#21; 673-wallet maker-sharp species (+$5.9M, 86% invisible to taker entries +$17.45 vs +$12.86 taker → #20; sells anti-signal → #21;
screens) → inventory-lean follow-on; sibling-sum = print artifact, 673-wallet maker-sharp species → inventory-lean line; sibling-sum =
parked; T1 crypto maker-quote sim needs its re-run. print artifact; T5 esports concentration → #14 tension. Batch two —
- Dashboards: jaxperro.com/{trading,live,test,value} — /test = both paper T6 lean-follow → #22 (fade arm FAILED its concentration gate,
report-only); T9 lead-lag → #23; T1 maker-quote Stage-1 KILL; T7
settlement-discount industrialized, parked; T8 crater-rejects were
good misses; T10 age gradient needs sample.
- Dashboards: jaxperro.com/{trading,live,test,value} — /test = all four
studies on one page (old /surge + /oracle URLs redirect) · daily pipeline studies on one page (old /surge + /oracle URLs redirect) · daily pipeline
on the Mac at 08:00 (launchd, lockfile) — floors, bench forward table, on the Mac at 08:00 (launchd, lockfile) — floors, bench forward table,
edge row, tape sync, Discord digest · tape mirror every 15 min edge row, tape sync, Discord digest · tape mirror every 15 min
(com.jaxperro.tape-sync → sync_tape.py, sftp + base64-console fallback) (com.jaxperro.tape-sync → sync_tape.py, sftp + base64-console fallback)
· research nightly fires 09:15 then WAITS for fresh tape · research nightly fires 09:15 then WAITS for fresh tape
(com.jaxperro.research-nightly; self-commits ledger, informed set, (com.jaxperro.research-nightly; self-commits ledger + lean rows, informed
surge/oracle grades, virtual book, meta snapshot). All launchd agents set, surge/oracle/lag grades, virtual book, meta snapshot). All launchd agents
removable with `launchctl unload ~/Library/LaunchAgents/<label>.plist`. removable with `launchctl unload ~/Library/LaunchAgents/<label>.plist`.
## Ops quick-reference ## Ops quick-reference
+2 -2
View File
@@ -29,7 +29,7 @@ Three deployed pieces + one static dashboard:
| **VALUE bot** (CLOSED 2026-07-19 — [`archive/value/README.md`](archive/value/README.md)) | ~~Fly app `wwf-valuebot`~~ destroyed after the paper verdict: 994 tickets, 1W/993L, 0.075x — today's sub-2¢ asks are fully informed; the 2025-era calibration edge is gone — a hard SILO: own code (`value/valuebot.py`, zero copybot imports), own state/feed, own image, no shared wallet ever | Systematically buys sub-2¢ contracts (the calibration study's one underpriced bucket — [`archive/value/PLAN.md`](archive/value/PLAN.md)); honest FAK fill model, chain-truth settles, $1 tickets, event cap 1. Verdict at ~2k resolved tickets vs a ~1.05% break-even hit rate | | **VALUE bot** (CLOSED 2026-07-19 — [`archive/value/README.md`](archive/value/README.md)) | ~~Fly app `wwf-valuebot`~~ destroyed after the paper verdict: 994 tickets, 1W/993L, 0.075x — today's sub-2¢ asks are fully informed; the 2025-era calibration edge is gone — a hard SILO: own code (`value/valuebot.py`, zero copybot imports), own state/feed, own image, no shared wallet ever | Systematically buys sub-2¢ contracts (the calibration study's one underpriced bucket — [`archive/value/PLAN.md`](archive/value/PLAN.md)); honest FAK fill model, chain-truth settles, $1 tickets, event cap 1. Verdict at ~2k resolved tickets vs a ~1.05% break-even hit rate |
| **RTDS tape recorder** ([`recorder/README.md`](recorder/README.md)) | **Fly.io app `wwf-recorder`, `arn`, 24/7** — third silo: no repo clone (code in image), no keys, 25GB volume | Records the FULL firehose (~8M events/day): every trade, order MATCH (maker side), comment, crypto tick — dual-socket capture (one socket alone measured 92.9% coverage; the twin covers per-conn silences). **Stage-0 warehouse (2026-07-21): the box folds its own segments into row-verified Parquet on the volume; the Mac mirrors + appends into `live/rtds.duckdb` every 15 min** (`com.jaxperro.tape-sync`). Nothing deletes without a verified second copy (gz needs its parquet; parquet needs the Mac's ack) | | **RTDS tape recorder** ([`recorder/README.md`](recorder/README.md)) | **Fly.io app `wwf-recorder`, `arn`, 24/7** — third silo: no repo clone (code in image), no keys, 25GB volume | Records the FULL firehose (~8M events/day): every trade, order MATCH (maker side), comment, crypto tick — dual-socket capture (one socket alone measured 92.9% coverage; the twin covers per-conn silences). **Stage-0 warehouse (2026-07-21): the box folds its own segments into row-verified Parquet on the volume; the Mac mirrors + appends into `live/rtds.duckdb` every 15 min** (`com.jaxperro.tape-sync`). Nothing deletes without a verified second copy (gz needs its parquet; parquet needs the Mac's ack) |
| **Discord digest** (`live/discord_daily.py`) | end of the daily pipeline | one message/day: the sharp list with profile links + 30-day conviction stats (per-trade pings retired 2026-07-04 for PAPER; the LIVE book pings every real placement/exit/settle) | | **Discord digest** (`live/discord_daily.py`) | end of the daily pipeline | one message/day: the sharp list with profile links + 30-day conviction stats (per-trade pings retired 2026-07-04 for PAPER; the LIVE book pings every real placement/exit/settle) |
| **dashboard** | [jaxperro.com/trading](https://jaxperro.com/trading) + [jaxperro.com/live](https://jaxperro.com/live) + [jaxperro.com/test](https://jaxperro.com/test) (static, in the `jaxperro` repo) | `/trading` renders the paper book, backtest book and sharp table; `/live` is the REAL MONEY page (reads `live/copybot_live_real.json` — live since 2026-07-10); `/test` is BOTH paper studies on one page (A2 surge measurement arm + oracle harness, each reading its bot's /feed + nightly chain-graded ledger; old /surge and /oracle URLs redirect). Wallet cards everywhere show LIFETIME numbers server-side (window sums lied twice — 2026-07-21/22) | | **dashboard** | [jaxperro.com/trading](https://jaxperro.com/trading) + [jaxperro.com/live](https://jaxperro.com/live) + [jaxperro.com/test](https://jaxperro.com/test) (static, in the `jaxperro` repo) | `/trading` renders the paper book, backtest book and sharp table; `/live` is the REAL MONEY page (reads `live/copybot_live_real.json` — live since 2026-07-10); `/test` is ALL FOUR studies on one page (A2 surge + oracle + lead-lag harnesses reading their bots' /feeds + nightly chain-graded ledgers, plus the Study C lean forward window from the forward ledger; old /surge and /oracle URLs redirect). Wallet cards everywhere show LIFETIME numbers server-side (window sums lied twice — 2026-07-21/22) |
**The calibration experiment (running now):** a fresh $1,000 paper book, **The calibration experiment (running now):** a fresh $1,000 paper book,
**reset 2026-07-08** so it measures exactly one thing — the follow set in **reset 2026-07-08** so it measures exactly one thing — the follow set in
@@ -137,7 +137,7 @@ does).
| `insider.py` | the original detector: z-score, pre-resolution timing, fresh-wallet flags, funding-cluster rings | | `insider.py` | the original detector: z-score, pre-resolution timing, fresh-wallet flags, funding-cluster rings |
| `smart_money.py` | shared HTTP helper + survivorship-corrected win-rate dashboard (`:8899`) | | `smart_money.py` | shared HTTP helper + survivorship-corrected win-rate dashboard (`:8899`) |
| `archive/webhook_receiver.py` | retired 2026-07-04: Alchemy webhook → per-trade Discord pings (replaced by `live/discord_daily.py`'s daily digest) | | `archive/webhook_receiver.py` | retired 2026-07-04: Alchemy webhook → per-trade Discord pings (replaced by `live/discord_daily.py`'s daily digest) |
| `research/` | tape-era edge factory, **SILO'd from the bots** ([research/README](research/README.md)): read-only RTDS loaders, calibrated execution sim, pre-registered studies scored through a MANDATORY chain-truth overlay (the 2026-07-22 scorer law — tape-resolution timing is win-biased and killed Study A #16 at $6/fill once corrected, independently confirmed by the A2 measurement arm at $7.54/fill chain-true; oracle #17: E0.04 killed, higher tiers ledger-positive but vetoed by the harness's real-latency chain grade — the 2026-07-23 fill-model lesson), TWO measurement harnesses (wwf-surgebot A2 + wwf-oraclebot: every-trigger $100 paper FAKs, attempts/markouts/settles append-only, bankroll specs replayed offline), nightly forward ledger — the only source of study verdicts | | `research/` | tape-era edge factory, **SILO'd from the bots** ([research/README](research/README.md)): read-only RTDS loaders, calibrated execution sim, pre-registered studies scored through a MANDATORY chain-truth overlay (the 2026-07-22 scorer law — tape-resolution timing is win-biased and killed Study A #16 at $6/fill once corrected, independently confirmed by the A2 measurement arm at $7.54/fill chain-true; oracle #17: E0.04 killed, higher tiers ledger-positive but vetoed by the harness's real-latency chain grade — the 2026-07-23 fill-model lesson), THREE measurement harnesses (wwf-surgebot A2 + wwf-oraclebot + wwf-lagbot: every-trigger $100 paper FAKs, attempts/markouts/settles append-only, bankroll specs replayed offline), nightly forward ledger — the only source of study verdicts. Open pre-registered windows: #20/#21 execution flips (live 2026-07-23), Study C #22 lean-follow (tape-scored forward), Study D #23 sibling lead-lag (wwf-lagbot, deployed 2026-07-23) |
| `wide/` | frozen-subgraph bulk scanner (1.76M wallets, historical only — subgraph froze Jan 2026) | | `wide/` | frozen-subgraph bulk scanner (1.76M wallets, historical only — subgraph froze Jan 2026) |
| `archive/` | everything retired, kept honest ([archive/README](archive/README.md)): the six failed strategies, earlier research sweeps (`hunt/huntwide/oos/copyback`), the superseded live selection layer (`live-research/`), the scrapped Polymarket-US venue probe (`us-venue/`), and retired infra (`retired-infra/`: Railway config, Mac launchd runner, GH-Actions cron) | | `archive/` | everything retired, kept honest ([archive/README](archive/README.md)): the six failed strategies, earlier research sweeps (`hunt/huntwide/oos/copyback`), the superseded live selection layer (`live-research/`), the scrapped Polymarket-US venue probe (`us-venue/`), and retired infra (`retired-infra/`: Railway config, Mac launchd runner, GH-Actions cron) |
+26 -8
View File
@@ -59,13 +59,16 @@ Layout:
dominated; needs standing-book data; parked) dominated; needs standing-book data; parked)
maker_lean.py T6 — WALK-FORWARD POSITIVE: follow small maker maker_lean.py T6 — WALK-FORWARD POSITIVE: follow small maker
inventory leans (+$2.51/lean, 59% hit, positive inventory leans (+$2.51/lean, 59% hit, positive
all 3 days); fade $2k+ whale bags (+$13.38, all 3 days) → STUDY C #22 PRE-REGISTERED
n=422). Pre-registration candidate pending (params/study_lean.json frozen 052eda0; follow
event-concentration + execution checks. arm verdict-gated, fade $2k+ FAILED its
event-concentration gate — report-only; nightly
rows via forward.py score_lean, study:"lean")
event_leadlag.py T9 — POSITIVE: same-outcome siblings reprice event_leadlag.py T9 — POSITIVE: same-outcome siblings reprice
slowly after leader bursts (+$9.73/$100 chain, slowly after leader bursts (+$9.73/$100 chain,
n=2,028; stale-print entry optimism stated n=2,028; stale-print entry optimism stated) →
Stage-2 is execution realism) STUDY D #23: wwf-lagbot harness (below) is the
Stage-2 execution-realism instrument
settle_discount.py T7 — settlement-discount niche: industrialized settle_discount.py T7 — settlement-discount niche: industrialized
by incumbents (+2-3%/hold for the good ones); by incumbents (+2-3%/hold for the good ones);
residual unproven (63% of vol outside residual unproven (63% of vol outside
@@ -96,9 +99,21 @@ Layout:
jaxperro.com/test): fair value tick-by-tick on the venue's jaxperro.com/test): fair value tick-by-tick on the venue's
own settlement feed, all E-tiers tracked, three settle own settlement feed, all E-tiers tracked, three settle
layers (own-feed tick → CLOB flags → nightly chain truth) layers (own-feed tick → CLOB flags → nightly chain truth)
grade_surge.py / grade_oracle.py nightly chain-truth re-grades lagbot.py Study D real-time PAPER harness (wwf-lagbot
surge_meas_ledger.jsonl / oracle_paper_ledger.jsonl; jaxperro.com/test, deployed 2026-07-23, SEM_VER l1):
also pull the raw volume streams to local .pull copies leader bursts ≥10¢/120s → paper-FAK $100 on ≤2 lagging
same-outcome siblings at ≤ stale+4¢; DOWN-moves buy the
sibling's COMPLEMENT token; cooldown 600s/event, 30-min
warmup (leader maps built live from the stream). Every
attempt logs the ask premium over the stale print — the
observational kill-switch (median ≥+8¢ over 3 days =
mirage) can kill #23 without a paper sample
grade_surge.py / grade_oracle.py / grade_lag.py nightly chain-truth
re-grades → surge_meas / oracle_paper / lag_paper
ledgers (grade_lag also prints EV per episode AND
%-of-staked — partial fills can't flatter the verdict —
plus the median-premium kill-switch read); all three
pull the raw volume streams to local .pull copies
markout_flow.py exploratory markout-exit curve (chain-truth verdict: NO markout_flow.py exploratory markout-exit curve (chain-truth verdict: NO
scalp inside the dead surge signal — losers bleed from scalp inside the dead surge signal — losers bleed from
minute one; v0's res_tok version was round-3-biased) minute one; v0's res_tok version was round-3-biased)
@@ -121,6 +136,9 @@ Data streams (the moat — all append-only, per DATA LAW):
.jsonl (durable settles) · surge2_state.json .jsonl (durable settles) · surge2_state.json
wwf-oraclebot /data oracle_attempts / oracle_markouts / oracle_settles wwf-oraclebot /data oracle_attempts / oracle_markouts / oracle_settles
.jsonl + oracle_state.json (same shapes) .jsonl + oracle_state.json (same shapes)
wwf-lagbot /data lag_attempts.jsonl (every attempt + top-5 asks +
premium vs stale + latency) · lag_settles.jsonl
(durable settles) · lag_state.json
Mac (nightly pulls + git): ledgers + params committed; raw .pull copies Mac (nightly pulls + git): ledgers + params committed; raw .pull copies
and meta/ gzips local. forward.py backfills any tape-covered day the and meta/ gzips local. forward.py backfills any tape-covered day the
ledger has never seen — a Mac gap > RESCORE_DAYS leaves no holes. ledger has never seen — a Mac gap > RESCORE_DAYS leaves no holes.