Align adaptation with audited outcomes

This commit is contained in:
Theodore Song
2026-08-18 10:45:58 -04:00
parent 7b81dbfa4b
commit ad11de39ad
5 changed files with 128 additions and 55 deletions
+37 -21
View File
@@ -15,7 +15,7 @@ https://polymarket-site-eta.vercel.app/personal.html
The site fetches live Polymarket markets, generates agent suggestions, lets you
run frequent paper cycles, and syncs the shared arena state through Neon or
Vercel Blob. Build 44 also installs an offline app shell and caches timestamped
Vercel Blob. Build 45 also installs an offline app shell and caches timestamped
market snapshots. During an outage, cycles continue locally; cached entries are
allowed for 90 minutes, older snapshots become mark-only, and all cached data
expires after 24 hours.
@@ -26,7 +26,7 @@ small samples toward neutral, caps sizing changes to 0.68x-1.30x, and reserves
15% of candidates for deterministic exploration so a stale regime cannot become
permanent.
Strategy 42 treats each binary stake as capable of falling to zero even when the
Strategy 43 treats each binary stake as capable of falling to zero even when the
18% stop cannot fill. New core positions are capped at 2.5%-4% of equity and
aggressive positions at 3%-5%, with lower limits for near-term, extreme-price,
reversal, and fast-moving setups. Oversized positions inherited from older
@@ -35,9 +35,11 @@ The two-agent overlap guard counts only positions worth at least 1.25% of an
agent's equity, so tiny profit-lock runners do not block a new material trade.
A separate walk-forward ledger records each trade-ready signal before its future
price is known, grades it at least 12 hours later, and combines that broad market
price is known, grades it at least 24 hours later, and combines that broad market
calibration with each agent's personal outcomes. This expands the learning sample
without backfilling future information into old decisions.
without backfilling future information into old decisions. The 24-hour horizon
matches the engine's minimum ordinary holding policy; stops and profit locks still
act immediately from fresh prices.
The initial seven-day chart seed is an approximate replay, not a live return.
It uses only prices available on each simulated date, computes daily and weekly
@@ -52,13 +54,18 @@ applies a conservative half-cent cost estimate, and reports a chronological
70/30 split plus three consecutive time segments. Results are also clustered by
market so repeated observations from one contract cannot masquerade as broad
evidence. Set `EVAL_MARKETS`, `EVAL_CONCURRENCY`, `EVAL_HORIZONS`, or
`EVAL_COST_CENTS` to change the audit.
The first 80-market audit found that reversal signals lost 4.34% on average in
both chronological partitions, while crypto and longshot samples were also
negative overall. Strategy 42 therefore blocks reversal and sports-trend entries outside the fixed
15% exploration lane and applies modest sizing penalties to crypto and longshots.
It does not boost any rule from this audit because no positive rule was robust
across the chronological split.
`EVAL_COST_CENTS` to change the audit. Set `EVAL_SUMMARY=1` for the compact,
decision-focused report.
The latest 120-active-market audit produced 1,241 twelve-hour observations from
41 markets with no fetch failures. The broad rule averaged -1.35% net and was
negative in all three chronological segments. Reversals averaged -3.69%, with a
market-clustered 90% interval entirely below zero. Crypto and Sports were also
negative but covered only three and five markets. The 24-hour cohort improved to
-0.82% row mean and +1.31% market mean, with no rule robustly negative across all
segments. Strategy 43 therefore disables reversal entries, retains their signals
for paper grading, and evaluates adaptation at 24 hours. It does not promote any
rule because no positive cohort passed the same robustness checks.
A corrected 200-market audit paged through 197 markets with usable history and
1,912 twelve-hour outcomes. Reversals remained negative in every chronological
@@ -66,23 +73,23 @@ segment and averaged -4.13%. Sports trends were negative in train and test and
averaged -3.53% at 72 hours. Politics trends were the sole cohort with positive
row-level returns in all three 72-hour segments, but its market-cluster interval
still crossed zero; that supports a longer hold test, not a larger entry bet.
Strategy 42 gives Politics trend positions that 72-hour observation window before
Strategy 43 gives Politics trend positions that 72-hour observation window before
ordinary signal exits. Stops, profit locks, settlement handling, and risk-budget
reductions remain immediate.
Strategy 42 also subtracts a half-cent round-trip cost when grading each live
Strategy 43 also subtracts a half-cent round-trip cost when grading each live
walk-forward signal. Confidence uses the largest independent matching bucket,
not the sum of five overlapping feature buckets, and evidence from older engine
versions is down-weighted. This prevents a handful of duplicated observations
from authorizing larger positions or hiding a modest negative regime.
Strategy 42 adds uncertainty-aware promotion and demotion. A matching setup must
Strategy 43 adds uncertainty-aware promotion and demotion. A matching setup must
accumulate at least eight effective observations and agree across at least two
feature views before repeatable positive evidence can increase size or repeatable
negative evidence can block a new entry. Mixed evidence stays close to neutral
instead of being mistaken for an edge.
Build 44 enforces the documented offline boundary end to end. Cached snapshots
Build 45 enforces the documented offline boundary end to end. Cached snapshots
under 90 minutes old may continue paper execution. Older snapshots remain usable
for valuation and chart snapshots for up to 24 hours, but cannot trigger entries,
stop-losses, gain-stops, risk rebalances, settlements, or policy exits. Network
@@ -95,16 +102,17 @@ adaptive baselines, pending signal grades, and trade evidence remain in one stra
lineage until the actual entry, sizing, or exit logic changes. Legacy build 40 and 41
records are migrated into the same strategy lineage without losing evidence.
Build 44 independently refreshes markets for matured pending signals that have
Build 45 independently refreshes markets for matured pending signals that have
left the current top-500 activity scan. Unavailable markets remain queued for a
bounded retry window. This prevents activity-rank survivorship from deciding
which wins and losses reach the adaptive calibration ledger.
Strategy 42 coordinates high-risk exploration globally. Crypto, reversal,
near-term, extreme-price, and other gap-prone positions may be held materially by
Strategy 43 coordinates high-risk exploration globally. Crypto, near-term,
extreme-price, legacy reversal, and other gap-prone positions may be held materially by
only one agent, while ordinary independently confirmed markets retain the
two-agent cap. A historically blocked setup can enter the exploration lane for
only one designated agent, preventing duplicated speculative losses.
two-agent cap. A historically blocked Sports trend can enter the exploration lane
for only one designated agent, preventing duplicated speculative losses. Reversal
signals cannot enter that lane and remain observation-only.
Run `npm run evaluate:settlements` to evaluate fixed decisions made 1, 3, 7,
14, 30, and 90 days before known binary settlements. The audit uses one
@@ -113,7 +121,8 @@ applies the same half-cent cost assumption, clusters related contracts by event,
and requires positive event-clustered confidence bounds in train and test plus
positive results in three chronological segments before it calls a settlement
cohort robust. Environment variables beginning with
`SETTLEMENT_` control its market count, concurrency, horizons, and cost.
`SETTLEMENT_` control its market count, concurrency, horizons, and cost. Set
`SETTLEMENT_SUMMARY=1` for the compact report.
The first event-clustered run loaded 498 of the 500 highest-volume resolved
markets. No side, price band, category, or 1-90 day holding rule passed the
@@ -121,6 +130,13 @@ required train/test confidence checks. In particular, older YES/underdog gains
reversed in the recent test segment. The engine therefore does not install a
static settlement-direction boost from this audit.
The latest 200-resolved-market audit loaded history for 199 markets with no
fetch failures. No positive rule passed the robustness gate. Buying NO with
7-14 days remaining was robustly negative in pooled, train, test, and
event-clustered results: the 14-day cohort averaged -37.33% by observation and
-46.84% by event. Strategy 43 therefore blocks new NO entries with 21 days or
less to resolution while continuing to record their signals for future evidence.
Paper accounts created with a password are also saved through the backend, so a
user can log in from another device and see the same paper portfolio, activity,
and value history. Passwordless paper accounts remain local-only.