- surgebot/oraclebot: settles now append-only to /data/*_settles.jsonl
(SETTLED_TRIM rotation can never lose a settle; graders pull them)
- forward.py: backfills any tape-covered day missing from the ledger —
a Mac offline gap > RESCORE_DAYS no longer leaves verdict-evidence holes
- meta_snap.py: nightly gzipped snapshot of all active markets (12k, 2MB/d,
local-only) — end dates make tau knowable at trigger for every tape
trigger; token->outcome maps kill the label-gap scorer artifact class
- .gitignore: pulled raw streams + meta stay local (re-fetchable, lean repo)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
v0 (res_tok only) said 'hold wins everywhere, +$43/fill' — that was
resolution-timing survivorship (round 3). Chain-graded, 1,146 forward
fills: hold -$5.92/fill and EVERY exit horizon negative too (best -$4.18
at +30m). Cohort split: tape-resolved winners drift to +$44 held; the
hidden-loss cohort bleeds monotonically from minute one (-$6@60s ->
-$44@2h). No scalp, no rescue — surge moments are symmetric information
events; net of fees + worst-print entry the taker case is closed.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Forward days, 559 resolved fills: exit +60s = -$3.43/fill (burst-top entry
is briefly underwater), +30m = +$13.63, +2h = +$35.51 vs hold-to-resolution
+$43.56. Monotone accrual to resolution in every slice — the under-reaction
converges over hours, not seconds. v1's >3h 'bleed bucket' is POSITIVE on
the full signal (+$44.66/fill): another confirmation the v1 book's losses
were cash-gate selection, not hold-time. All entry bands positive forward.
No A2-scalp pre-registration warranted; markout re-reads stay on to verify
with real bids.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
No signal-path change (SEM_VER unchanged): records more, alters nothing.
Purpose: real exit marks for the markout-exit study (prints can't show the
bid) + book depth at mispricing moments (maker-study stage-1 groundwork).
oraclebot gains the same append-only attempts log as A2. Nightly pulls the
new files alongside state.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
v1's cash-gated book halted at its pre-registered -50% line; post-mortem
showed the ~2% cash-gated subsample was adversely selected (-$9/fill vs
+$41/fill full-signal, same day). A2 samples every trigger and replays
bankroll specs offline (surge_book_replay.py -> surge_book.json). Signal
semantics verbatim; grades to surge_meas_ledger.jsonl. See #19.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
FINDINGS: full round-3 section (mechanism, the 81%-vs-26% audit, how the
paper harness caught it, corrected verdicts, three standing rules) +
scorecard updates (surge dead, sub5c dead 0/38, oracle revived at E>=0.07)
+ correction banner on the 07-20 tape-era section. HANDOFF: queue +
snapshot reflect the kill, surgebot's instrument role, scorer law, Friday's
combined agenda. research/README: SCORER LAW + independent-instrument rule
+ full current layout. recorder/README: sync transport fallback. README:
/surge dashboard + research row truth.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
THE SCORER BUG (found 2026-07-22 by the surge paper book divergence):
tape-resolution timing is win-biased. When our side LOSES, the winning
sibling keeps trading at 99c until close, the sibling-veto keeps the
market 'alive', and the loss sits in the ignored pending bucket — while
wins tape-resolve within hours and score. Jul-21 audit: tape-resolved
fills 81% hit (+$46/fill); chain-resolving the 329 'pending' fills: 26%
hit, -$49.61/fill; combined truth 53% ≈ the paper book's 57.5%. The
paper harness was the honest instrument; every ledger arm was flattered.
payouts_for(): tape proxy first, CTF chain truth for the remainder
(payouts.py cache — immutable, so nightly incremental cost is small);
refunds (0.5) now booked as scratches. Applied to flow, controls, and
oracle (sub5c inherits via score_flow). Ledger recompute follows; #16
verdict evaluates corrected numbers only.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Mirrors forward.py::score_oracle semantics verbatim (constants frozen from
study_oracle/tape); live-only suppress-only guards; $100 FAK walks the live
asks inside p_ref*1.05; three settle layers (own-feed tick, CLOB flags,
nightly CTF re-grade via grade_oracle.py -> oracle_paper_ledger.jsonl).
Shakedown 2026-07-22: 52 events/9 fills/43 craters at p50 541ms; 30/30
event agreement vs tape scorer on the comparable set. Verdict still binds
to forward_ledger only (#17).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Fetched from CLOB end_date_iso at fill time; settle passes backfill
pre-ETA fills. Feeds the /surge Open tab's Resolves column.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
CORS-open GET /feed serving the paper book (equity, counters, open,
settled, skips, informed-set meta) — paper data only, nothing mutable.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Real-time PAPER trader of the FROZEN surge signal on its own Fly app
(wwf-surgebot, recorder-pattern image, no keys, no bot imports):
$100 paper book · 5%-of-equity daily stakes ($1 venue floor) · cash-gated
all-or-nothing · event cap 2 · paper FAK against the live CLOB book inside
p_ref*1.05 · provisional CLOB settles re-graded nightly with CTF payout
vectors (grade_surge.py -> surge_paper_ledger.jsonl). Informed set
published daily (informed_set.py -> params/informed_set.json, frozen
method). Unit-tested: fill/crater/event-cap/cash paths exact; live smoke:
dual sockets + sizing clean. Believing any of it stays gated on the #16
forward verdict.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Same frozen surge signal, band opened to 0.5-5c, all niches, F=$300 —
tracked nightly as flow_sub5c_EXPLORATORY (NOT pre-registered). First
tape scan: they live in in-play sports totals/blowout tails + esports
maps; ~70% crater; 31 resolved fills, 4 wins, EV sign owned entirely by
two lottery hits; sim depth-blind at longshot share counts. Accumulate
to ~100+ resolved fills before any pre-registration or kill.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
First run exited 1 after a successful score: the exit trap's relative
rmdir ran from the repo root (post-cd) — wrong dir, stale lock left in
research/, every future night would have skipped itself.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Replays candidate sets through the engine's mirrored mechanics (stake rule
+ DD halving + their-shares ceiling + one-market-one-stake adds + all-or-
nothing cash gate + proportional sell mirror, copytrade.py cited) with the
calibrated sim fill model and tape proxy-resolution. Per-wallet conviction
floors from the paper config's pinned p80s (candidates without pins get
tape-p80, same rule). Outputs per-set×bankroll: realized/open, deployment
stats, miss families (capital/crater/band), capital-miss hypothetical P&L,
per-wallet realized, and --loo leave-one-out marginals at $1k.
Validated against the real paper book on the same window: 33 replay opens
vs 26 real (backfill bias documented — pre-tape positions' adds replay as
opens), capital misses 0 vs 0, peak deploy 62% vs the era's 74%, mean
deployed $297 vs ~$360. SEARCH TOOL ONLY per the silo README — verdicts
stay with forward_ledger.jsonl.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The nightly Mac-coupled bulk ingest was the fragile link (today: Mac slept
through 08:00, the pipeline started 09:31, and the sftp bulk pull wedged in
timeout cascades holding the db lock — digest blocked behind it). Stage 0
moves the fold onto the box and reduces the Mac to an incremental mirror:
- recorder/fold.py (sidecar, capture stays PID 1): closed gz segments ->
zstd parquet /data/parquet/<fam>/date=*/segment.parquet, row-parity
verified, manifest-logged, THEN gz deleted. Deletion invariant STRONGER:
raw needs a verified parquet; parquet needs a Mac ACK before the disk
guard may prune it. duckdb capped 384MB; VM 256MB -> 1GB.
- recorder/sync_tape.py (com.jaxperro.tape-sync every 15 min + daily.sh):
manifest-driven incremental sftp pull, per-file row verify, append into
rtds.duckdb NATIVE tables (views-over-parquet rejected: sim's per-asset
point queries would crawl), segment-keyed idempotence shared with the
legacy ingest.py (kept as fallback), ack back to the box.
- recorder/bootstrap_parquet.py: pre-fold history exported to the mirror,
parity OK (13.84M trades + 3.07M aux). live/parquet/ is now the complete
durable layer Stage 1 (MotherDuck/ClickHouse) would consume.
- research/tape.py connect(): brief retry — the 15-min sync holds the
write lock for seconds.
First run: backlog folded in <60s on the box, 36/36 files mirrored+
verified, rtds.duckdb 13.8M -> 18.2M trades, tape age 24h+ -> ~15-45 min.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-21 first firing failed two ways: launchd's bare python3 has no
duckdb (framework python does), and 09:15 preceded the sleep-delayed daily
pipeline's tape ingest, so even a clean run would have scored a stale
tape. nightly.sh now uses the framework python and polls tape max-ts until
it is <6h old (15min polls, 8h deadline, then scores anyway and logs the
staleness).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>