Files
winning-wallet-finder_github/live/validate_timing.py
T
jaxperro 4941818d51 scoring: chain-truth resolutions overlay — refunds no longer count as wins
The cache's won (curPrice>=0.5 at pull) counts 50/50 refunds as wins for
BOTH sides — 521 of 2,128 chain-checked follow-set markets (24%) were
refunds, which is where the whales' 92-100% displayed win rates came from.

- live/payouts.py: resolutions table in cache.duckdb, filled from the CTF
  contract's payout vectors (batched+paced JSON-RPC, ~12 calls/s free tier;
  resolved rows immutable, unresolved recheck 6h, RPC failures never cached).
  truth(cond, asset) -> 1/0/0.5/None; refunds need no asset side.
- validate_timing: every displayed stat (conv/conv30/all-time/realized/copy
  replay) settles at truth; refunds count as neither W nor L, P&L is
  size*(wp-p)/p; new conv_ref/conv30_ref/all_ref feed fields.
- trust.conviction_record: optional truthfn — the selection gates
  (trust_wr/trust_roi) no longer select on refund inflation.
- portfolio.py: replay pays wp (refunds 0.5/share, was 1.0).
- conviction_scan: documented as the (refund-inflated) candidate layer;
  final selection re-judges against truth downstream.

Validation: 0x4bFb-whale conv 174-16 91.6% $1.61M -> 34-16 +140ref 68%
$214k, and truth-adjusted all-time P&L now sits within ~16% of lb-api's
PM P&L (was 7x apart); LSB1 (0 refunds) byte-identical, its P&L matches
PM P&L to 0.07%.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-07 01:55:50 -04:00

355 lines
18 KiB
Python
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
#!/usr/bin/env python3
"""Select the COPYABLE conviction wallets — by what a copier actually earns.
The earlier version gated on entry->resolution lead time (a proxy for "can we
mirror it"). That was too blunt: it kept scalpers whose position win% looks great
but lose when copied, and dropped fast-resolving holders that are perfect for a
small fast-recycling bankroll. The fix: run a full flat-$50 copy replay on every
conviction wallet and SELECT on copyability directly —
* copy_pnl > 0 — copying them actually makes money, AND
* held_pnl > 0 over >= MIN_HELD — their hold-to-resolution edge is real (the
latency-robust leg), not just scalp-sell timing
* active in 30d, median lead >= MIN_LEAD_H (light guard vs true sub-hour snipers)
This keeps Kruto (sells often but profitably) and surfaces copy-positive holders
the lead gate used to discard; it drops scalper-traps like a wallet that's only
positive via sells while its held bets lose.
2026-07-03 holder fix: the held-edge gates no longer use the replay's held leg.
That leg only counts bets entered AND resolved inside the Jun-1->now window, so
a ~7-day-lead holder always showed `held 0-0, ~20 unresolved` and failed
held_n>=8 — the filter structurally rejected the most copyable wallets (whale
0x73afc816: 100% fwd win in conviction_scan, "held 0-0" here; and pre-Jul-2 the
winner=False bug booked those unresolved bets as LOSSES, which is where the
iohihoo $749 / ArbTrader $790 "scalper trap" numbers came from). The held-edge
gate now reads the wallet's trailing TRUSTED conviction record from the cache
(trust.py: consensus-resolution rows, outcome observed post-resolution), which
includes bets entered before the window that resolved inside it. The replay's
copy_pnl (fees, mirror exits) remains the other selection leg, and held stats
are still computed for display.
"""
import json
import os
import ssl
import statistics as st
import time
import urllib.request
from concurrent.futures import ThreadPoolExecutor
import cache
import payouts
import smart_money as sm
import trust
HERE = os.path.dirname(__file__)
COPYABLE_MED_LEAD = 24.0 # median lead (h) on winning conviction bets to count as copyable
JUN1 = time.mktime(time.strptime("2026-06-01", "%Y-%m-%d")) # portfolio copy-start
STAKE = 50.0 # flat $/trade the copy portfolio uses
# Polymarket taker fee (since 2026-03-30): fee = shares·rate·p·(1p), paid on
# marketable entries AND mirror exits; redeeming at resolution is free. 0.03 is
# the sports rate (the follow set's category). Making copy_pnl fee-aware makes
# the SELECTION fee-aware — a wallet only counts as a copyable sharp if copying
# it clears the fees a real copier pays.
FEE_RATE = 0.03
_SSL = ssl._create_unverified_context()
_CLOB = {} # conditionId -> {token_id: winner-price 1/0/None}
def _clob_winner(cond, token):
"""Authoritative resolution for a token: 1 if it won, 0 if it lost, None if the
market hasn't resolved. Matched by token_id (exact, no outcome-name guessing).
NB: the CLOB reports winner=False on EVERY token of an UNRESOLVED market —
only a present True winner means resolved. Treating False as "lost" counted
every unresolved held bet as a loss, biasing copy_pnl (the selection metric)
downward."""
if cond not in _CLOB:
try:
req = urllib.request.Request("https://clob.polymarket.com/markets/" + cond,
headers={"User-Agent": "Mozilla/5.0"})
m = json.loads(urllib.request.urlopen(req, timeout=20, context=_SSL).read())
toks = m.get("tokens") or []
resolved = any(t.get("winner") is True for t in toks)
_CLOB[cond] = {str(t.get("token_id")):
((1 if t.get("winner") is True else 0) if resolved else None)
for t in toks}
except Exception:
_CLOB[cond] = {}
return _CLOB[cond].get(str(token))
def _pm_profit(w):
"""The wallet's own all-time account P&L as Polymarket reports it
(lb-api /profit): cash-flow truth including early sells. The sanity anchor
next to the hold-to-resolution columns — a ~1x gap means true holder, a
huge gap means scalper (ArbTraderRookie: +$8.6k real vs +$462k held, 53x —
a 0.5% margin on $1.7M volume)."""
try:
req = urllib.request.Request(
"https://lb-api.polymarket.com/profit?window=all&limit=1&address=" + w,
headers={"User-Agent": "Mozilla/5.0"})
r = json.loads(urllib.request.urlopen(req, timeout=15, context=_SSL).read())
return round(r[0]["amount"]) if r else None
except Exception:
return None
def _wp(cond, asset, won):
"""Chain-truth payout for a bet (1/0/0.5), falling back to the cache's
`won` mark when the chain can't say (unresolved, legacy NULL-asset rows on
decided markets, RPC gaps). The fallback keeps old behavior; the truth
path kills the two cache lies: 50/50 refunds counted as wins for BOTH
sides (28% of the follow set's resolved markets!) and stale both-sides-won
marks on operator-resolved markets."""
wp = payouts.truth(cond, asset)
return (1.0 if won else 0.0) if wp is None else wp
def _bet_pnl(b, wp=None):
"""Resolved P&L of one cache bet at payout wp: a $size stake at avg price p
returns size·(wpp)/p — wp=1 win, 0 loss, 0.5 refund ($0.50/share, NOT
money-back: flat near coin-flip entries, ruinous for favorites)."""
p = max(0.001, min(0.999, b["p"] or 0))
if wp is None:
wp = _wp(b.get("cond"), b.get("asset"), b["won"])
return b["size"] * (wp - p) / p
def display_stats(w):
"""Everything the dashboard's sharp table renders, precomputed so the page makes
ZERO per-wallet data-api calls.
conv win%/record/P&L : over the wallet's conviction (top-20%-stake) bets — a
POSITION stat from the cache (large 180d sample)
realized P&L : reconstructed P&L over the last 500 resolved bets
copy P&L : the TRUTH for a copier — what a flat-$50 copy of their
conviction bets ACTUALLY realizes since Jun 1: replays
their entries, mirrors their exits, settles held bets at
AUTHORITATIVE clob resolution (by token id). This exposes
scalpers whose position win% looks great but don't copy
(e.g. ArbTrader: ~100% conv win but $790 copy P&L).
name / last-bet : from the /activity pull
"""
# ---- position win%/record/P&L from the cache (large, survivorship-corrected).
# res_t <= now: the cache stores early-sold positions in UNRESOLVED markets with
# a future res_t and won = current price — a mark, not an outcome; skip them. ----
now = time.time()
bets = [b for b in cache.get_bets(w)
if (b["size"] or 0) > 0 and (b["res_t"] or 0) <= now]
# chain-truth payouts for everything this wallet's stats touch (cached
# in the resolutions table — incremental after the first backfill)
trows = trust.trusted_wallet_rows(cache.query, w)
payouts.ensure({b["cond"] for b in bets} | {r[0] for r in trows})
# ---- ALL-TIME stats over EVERY trusted bet (any size): the dashboard's
# "of every bet placed" columns. Trusted rows only, deduped one-per-market,
# truth-adjusted: refunds (wp=0.5) count as neither won nor lost. ----
tbest = {}
for cond, asset, won, p, res_t, size in trows:
if cond not in tbest or size > tbest[cond][3]:
tbest[cond] = (cond, asset, won, p, size)
def tally(rows):
"""(won, lost, refunds, pnl) over (cond, asset, won, p, size) rows."""
w_ = l_ = r_ = 0
pnl = 0.0
for cond, asset, won, p, size in rows:
wp = _wp(cond, asset, won)
pc = max(0.001, min(0.999, p or 0))
pnl += size * (wp - pc) / pc
if wp > 0.5:
w_ += 1
elif wp < 0.5:
l_ += 1
else:
r_ += 1
return w_, l_, r_, pnl
all_won, all_lost, all_ref, all_pnl = tally(tbest.values())
thr = cache.conv_cutoff(b["size"] for b in bets)
conv = [b for b in bets if b["size"] >= thr]
recent = sorted(bets, key=lambda b: b["res_t"] or 0, reverse=True)[:500]
cut30 = time.time() - 30 * 86400
conv30 = [b for b in conv if (b["res_t"] or 0) >= cut30]
brow = lambda bs: [(b["cond"], b.get("asset"), b["won"], b["p"], b["size"]) for b in bs]
cw, cl, cr, cpnl = tally(brow(conv))
c3w, c3l, c3r, c3pnl = tally(brow(conv30))
out = {
"conv_win": round(100 * cw / (cw + cl), 1) if (cw + cl) else None,
"conv_won": cw, "conv_lost": cl, "conv_ref": cr,
"conv_pnl": round(cpnl),
"conv30_win": round(100 * c3w / (c3w + c3l), 1) if (c3w + c3l) else None,
"conv30_won": c3w, "conv30_lost": c3l, "conv30_ref": c3r,
"conv30_pnl": round(c3pnl),
"realized_pnl": round(tally(brow(recent))[3]),
"all_win": round(100 * all_won / (all_won + all_lost), 1) if (all_won + all_lost) else None,
"all_won": all_won, "all_lost": all_lost, "all_ref": all_ref, "all_pnl": round(all_pnl),
"pm_pnl": _pm_profit(w),
"avg_bet": round(sum(b["size"] for b in conv) / len(conv)) if conv else 0,
"copy_pnl": 0, "held_pnl": 0, "held_won": 0, "held_lost": 0, "sold": 0,
"name": None, "last_trade": 0, "last_conv_bet": 0,
}
# ---- resolution map from a FRESH positions pull (curPrice extreme = resolved);
# cheap, so the copy replay can run on every conviction wallet. clob fills gaps.
resmap = {}
for p in (sm.get_json("/closed-positions", {"user": w, "limit": 500,
"sortBy": "TIMESTAMP", "sortDirection": "DESC"}) or []) + \
(sm.get_json("/positions", {"user": w, "limit": 500, "sizeThreshold": 0}) or []):
cp = p.get("curPrice", 0) or 0
if (cp <= 0.001 or cp >= 0.999) and p.get("asset") and p["asset"] not in resmap:
resmap[p["asset"]] = 1 if cp >= 0.5 else 0
# ---- activity: name, last-bet, and the flat-$50 copy replay ----
a = []
for off in range(0, 4000, 500):
pg = sm.get_json("/activity", {"user": w, "type": "TRADE", "limit": 500, "offset": off}) or []
a += pg
if len(pg) < 500 or (pg and (pg[-1].get("timestamp", 0) < JUN1)):
break
if a:
out["last_trade"] = a[0].get("timestamp", 0)
out["name"] = next((t.get("name") for t in a if t.get("name")), None)
# position-level conviction: each market's TOTAL buy stake, top-20% (p80)
mkt = {}
for t in a:
if t.get("side") == "BUY" and t.get("conditionId"):
mkt[t["conditionId"]] = mkt.get(t["conditionId"], 0) + (t.get("usdcSize", 0) or 0)
cthr = cache.conv_cutoff(mkt.values())
for t in a:
if t.get("side") == "BUY" and mkt.get(t.get("conditionId"), 0) >= cthr:
out["last_conv_bet"] = t.get("timestamp", 0)
break
# replay a flat-$50 copy of their conviction markets since Jun 1. Split P&L into
# the SOLD (scalp) leg and the HELD-to-resolution leg — the held leg is the
# latency-robust edge; a wallet whose copy P&L is positive only via scalp sells
# (held leg negative) isn't a reliable copy target.
ev = sorted([t for t in a if t.get("timestamp", 0) >= JUN1], key=lambda t: t.get("timestamp", 0))
openp, entered, scalp, held = {}, set(), 0.0, 0.0
hw = hl = sold = 0
for t in ev:
c, pr, asset = t.get("conditionId"), t.get("price", 0) or 0, t.get("asset")
if not c or pr <= 0:
continue
if t.get("side") == "BUY":
if mkt.get(c, 0) < cthr or c in entered or c in openp:
continue
fee_in = STAKE * FEE_RATE * (1 - pr) # taker fee on the entry
entered.add(c); openp[c] = {"sh": STAKE / pr, "a": asset, "fee": fee_in}
elif c in openp: # mirror their exit (scalp)
sh = openp[c]["sh"]
fee_out = sh * FEE_RATE * pr * (1 - pr) # taker fee on the exit too
scalp += sh * pr - STAKE - openp[c]["fee"] - fee_out
sold += 1; del openp[c]
for c, p in openp.items(): # settle held bets at resolution
wv = payouts.truth(c, p["a"]) # chain first: refunds pay 0.5
if wv is None:
wv = resmap.get(p["a"])
if wv is None:
wv = _clob_winner(c, p["a"]) # clob fallback for out-of-pull markets
if wv is None:
continue # not resolved yet -> exclude
held += p["sh"] * wv - STAKE - p["fee"] # redeem itself is fee-free
if wv > 0.5:
hw += 1
elif wv < 0.5:
hl += 1
out.update(copy_pnl=round(scalp + held), held_pnl=round(held),
held_won=hw, held_lost=hl, sold=sold)
return out
def lead_profile(w):
ent = cache.get_entries(w)
now = time.time()
bets = [b for b in cache.get_bets(w) if (b["res_t"] or 0) <= now] # resolved only
cut = cache.conv_cutoff(b["size"] for b in bets) # this wallet's top-20% stake cutoff
leads = [(b["res_t"] - ent[b["cond"]]) / 3600.0 for b in bets
if b["won"] and (b["size"] or 0) >= cut and b["cond"] in ent
and b["res_t"] and b["res_t"] >= ent[b["cond"]]]
if not leads:
return None
med = st.median(leads)
u6 = sum(1 for l in leads if l < 6) / len(leads)
verdict = ("last-minute" if (med < 6 or sum(1 for l in leads if l < 1) / len(leads) > 0.5)
else "borderline" if med < COPYABLE_MED_LEAD else "sharp")
return dict(n=len(leads), med=med, u6=u6, verdict=verdict)
MIN_HELD = 8 # need this many trailing trusted conviction bets to trust the held edge
MIN_HELD_WR = 0.55 # they must WIN a clear majority — excludes longshot-variance
# players (+EV but ~34% win) that don't fit the high-win-rate thesis
MIN_LEAD_H = 1.0 # light sniper guard: drop wallets whose median winning lead < 1h
TRUST_DAYS = 90 # trailing window for the trusted conviction record (long enough
# that week-lead holders have real resolved sample in it)
def main():
conv = json.load(open(os.path.join(HERE, "conviction_wallets.json")))
print(f"copy-testing {len(conv)} conviction wallets…\n", flush=True)
# run the full copy replay on EVERY conviction wallet (cheap now: fresh-positions
# resolution, clob only fills gaps), then select on copyability — not lead time.
# Per-wallet guard: one wallet's unexpected error must not kill the whole
# selection run (a single RemoteDisconnected once took out the nightly refresh);
# a failed wallet is retried once, then excluded from this run and logged.
def safe_stats(c):
for attempt in (1, 2):
try:
return display_stats(c["wallet"])
except Exception as e:
if attempt == 2:
print(f" ⚠ {c['wallet'][:10]}… stats failed ({e}) — excluded this run",
flush=True)
return None
time.sleep(2)
trust.ensure_cons(cache.query)
with ThreadPoolExecutor(max_workers=8) as ex:
stats = list(ex.map(safe_stats, conv))
cut30 = time.time() - 30 * 86400
sharps = []
for c, ds in zip(conv, stats):
if ds is None:
continue
c.update(ds)
if ds.get("name"):
c["name"] = ds["name"]
lp = lead_profile(c["wallet"])
c["med_lead_h"] = round(lp["med"], 1) if lp else None
# held-to-resolution edge from the trailing TRUSTED cache record — includes
# bets entered before the replay window that resolved inside it, so long-lead
# holders are judged on their real resolved sample (the replay's own held leg
# is mostly "unresolved" for them and only reported for display).
tr = trust.conviction_record(cache.query, c["wallet"], days=TRUST_DAYS,
pctile=cache.CONV_PCTILE, truthfn=payouts.truth)
c["trust_n"], c["trust_wr"], c["trust_roi"] = tr["n"], round(tr["wr"], 3), round(tr["roi"], 3)
c["trust_refunds"] = tr.get("refunds", 0)
# SELECT a copyable sharp: active, copy-positive (fee-aware replay), and a
# genuine hold-to-resolution edge — trailing trusted conviction record wins a
# clear majority with positive flat-stake ROI on a real sample, so the edge
# survives live latency and isn't longshot variance or all sell-timing. A
# light lead floor drops true sub-hour snipers.
if ((ds["last_trade"] or 0) >= cut30 and ds["copy_pnl"] > 0
and tr["n"] >= MIN_HELD and tr["wr"] >= MIN_HELD_WR and tr["roi"] > 0
and (c["med_lead_h"] is None or c["med_lead_h"] >= MIN_LEAD_H)):
sharps.append(c)
sharps.sort(key=lambda c: c["copy_pnl"], reverse=True)
print(f"copy-positive holders (copy>0, trust_n>={MIN_HELD}, trust_wr>={MIN_HELD_WR:.0%}, "
f"trust_roi>0 over {TRUST_DAYS}d, active, lead>={MIN_LEAD_H}h): "
f"{len(sharps)} of {len(conv)}\n")
h = (f"{'copyP&L':>8}{'trustRec':>10}{'trustROI':>9}{'heldP&L':>8}{'held':>9}"
f"{'sold%':>6}{'medLeadH':>9} wallet")
print(h); print("-" * len(h))
for c in sharps[:35]:
n = c["held_won"] + c["held_lost"]
sp = 100 * c["sold"] / (c["sold"] + n) if (c["sold"] + n) else 0
ld = f"{c['med_lead_h']:.0f}" if c["med_lead_h"] is not None else "—"
rec = f"{round(c['trust_wr']*c['trust_n'])}-{round((1-c['trust_wr'])*c['trust_n'])}"
print(f"{c['copy_pnl']:>+8}{rec:>10}{c['trust_roi']:>+9.0%}{c['held_pnl']:>+8}"
f"{(str(c['held_won'])+'-'+str(c['held_lost'])):>9}"
f"{sp:>5.0f}%{ld:>9} {(c.get('name') or c['wallet'][:10])}")
json.dump(sharps, open(os.path.join(HERE, "watch_sharps.json"), "w"), indent=2)
print(f"\n-> watch_sharps.json ({len(sharps)} copy-positive holders)")
if __name__ == "__main__":
main()