id: 5bf5676aed9e439f9934f2d603144744
parent_id: a0cc3cf6624541ba9b5b61aaa75ace97
item_type: 1
item_id: 47031623e2c6451fb36730ffe93380ea
item_updated_time: 1783318408786
title_diff: "[]"
body_diff: "[{\"diffs\":[[0,\"d 2026-0\"],[-1,\"6-30\"],[1,\"7-06\"],[0,\".\\\n\\\n---\\\n\\\n\"]],\"start1\":19,\"start2\":19,\"length1\":20,\"length2\":20},{\"diffs\":[[0,\"ion, **2\"],[-1,\"0\"],[1,\"2\"],[0,\" registe\"]],\"start1\":184,\"start2\":184,\"length1\":17,\"length2\":17},{\"diffs\":[[0,\"gression\"],[1,\"\\\n- **PEM2 preflop equity matrix** — 169×169 distribution matrix (10 metrics × 8 bins)\"],[0,\"\\\n\\\n---\\\n\\\n#\"]],\"start1\":477,\"start2\":477,\"length1\":16,\"length2\":101},{\"diffs\":[[0,\" Gen 2-4\"],[1,\".7\"],[0,\" Strateg\"]],\"start1\":588,\"start2\":588,\"length1\":16,\"length2\":18},{\"diffs\":[[0,\"en 4\"],[1,\".5\"],[0,\": P\"],[-1,\"rofile persistence to disk\"],[1,\"layability-enhanced preflop + adaptive postflop\"],[0,\"\\\n- ✅ \"],[1,\"**\"],[0,\"Gen 4\"],[-1,\": Preflop looseness wired (VPIP-based threshold scaling)\\\n- ✅ Gen 4: VPIP tracking fixed (survivorship bias + to_call computation)\\\n- ✅ Gen 4: `range_tightness_scale` discount removed (double-counting fix)\\\n- ✅ Gen 4: Post-discount sweep applied (`bet_turn_base=0.40`, `bet_flop_base=0.15`)\\\n- ✅ Gen 4: Opponent features bridge for Gen 5 data col\"],[1,\".7: Matrix-based equity preflop (PEM2) — CURRENT PRODUCTION**\\\n  - PFR-based raiser range with bet-sizing-aware narrowing\\\n  - Hard equity floor for 3-bets (min_reraise_equity=0.42)\\\n  - Milder looseness on reraise (looseness^0.3)\\\n  - Ppot-based call decisions\\\n  - Short-stack strategy removed (folds defer to G47)\\\n  - Tuned via 2-round multi-field sweep (10+8 params)\\\n  - **Performance: +10.6 BB/hand vs G45 field**\\\n\\\n### Engine Fixes (rusty-marvin)\\\n- ✅ Crash fix: seat restoration + chip refund on `run_to_comp\"],[0,\"le\"],[-1,\"c\"],[0,\"tion\"],[-1,\"\\\n- ✅ **Production**: `cash_nl_g4_live.toml` — validated +132.9/hand vs chumps, +0.95/hand vs gen field\\\n\\\n### Open Tuning Items\\\n- [ ] Passive postflop — bot could value-bet more first-in (needs Gen 5 for full fix)\\\n- [ ] Automated sweep → apply pipeline (currently manual winner selection)\\\n- [ ] Range prediction misses (88% hit rate — \"],[1,\"()` error\\\n- ✅ `NotEnoughPlayers` now logs warning (was silent)\\\n- ✅ `NoLimit { max_raises: Option<usize> }` field added\\\n\\\n### Open Tuning Items\\\n- [ ] Wire per-opponent PFR into G47 equity lookup (currently table-level only)\\\n- [ ] Postflop equity matrix (extend PEM2 to common board t\"],[0,\"ext\"],[1,\"u\"],[0,\"re\"],[-1,\"me opponents cause 12% miss)\"],[1,\"s)\\\n- [ ] Chump_bot `RaiseBelowMinimum` root cause\"],[0,\"\\\n\\\n--\"]],\"start1\":875,\"start2\":875,\"length1\":762,\"length2\":922},{\"diffs\":[[0,\")\\\n- \"],[-1,\"Call mismatch fix (blind reconciliation before non-blind actions)\\\n- BB-check-via-call fix (zero-cost calls detected as checks)\\\n- Raise-by → raise-to fix (parser sends target amount directly)\\\n- Graceful shutdown/restart endpoints (`/api/shutdown`, `/api/restart`)\\\n- AllIn executor confirmed working\\\n- DOM debug logging (toggleable via `config.ts`\"],[1,\"Graceful shutdown/restart endpoints\\\n- **Default bot switched to G47** (2026-07-06\"],[0,\")\\\n\\\n#\"]],\"start1\":2135,\"start2\":2135,\"length1\":353,\"length2\":89},{\"diffs\":[[0,\"d\\\n- \"],[-1,\"28\"],[1,\"**29\"],[0,\"0-dim\"],[1,\"**\"],[0,\" fea\"]],\"start1\":2476,\"start2\":2476,\"length1\":15,\"length2\":19},{\"diffs\":[[0,\" actions\"],[1,\", range-weighted + multiway\"],[0,\")\\\n- 11-d\"]],\"start1\":2546,\"start2\":2546,\"length1\":16,\"length2\":43},{\"diffs\":[[0,\"twork (2\"],[-1,\"8\"],[1,\"9\"],[0,\"0→512→LN\"]],\"start1\":2705,\"start2\":2705,\"length1\":17,\"length2\":17},{\"diffs\":[[0,\"ed loss,\"],[1,\" focal loss,\"],[0,\" gradien\"]],\"start1\":2782,\"start2\":2782,\"length1\":16,\"length2\":28},{\"diffs\":[[0,\"ping\"],[-1,\", cosine LR\\\n- DAgger strategy (teacher fallback + model inference)\\\n- Opponent features bridge from Gen 4 observ\"],[1,\"\\\n- Multi-head policy (action + sizing heads, class weights)\\\n- DAgger strategy (teacher fallback + model inference)\\\n- PPO self-play loop (v11 = best model, +27.9 avg chips/hand real performance)\\\n\\\n### Data Collection\\\n- **3.79M transitions** collected overnight 2026-07-05 (G4 teach\"],[0,\"er\"],[1,\")\"],[0,\"\\\n- \"],[-1,\"Self-play\"],[1,\"Total accumulated: ~10M+ transitions across multiple\"],[0,\" col\"]],\"start1\":2816,\"start2\":2816,\"length1\":133,\"length2\":345},{\"diffs\":[[0,\"tion\"],[-1,\" +\"],[1,\"s\\\n-\"],[0,\" DAgger \"],[-1,\"loop scripts\\\n\\\n### Bugs Found & Fixed (v1–v4 all degenerate)\\\n- ✅ LayerNorm was documented but not implemented → added\\\n- ✅ `engine_action_to_discrete` broken (all bets→BetHalf, all raises→Raise25x) → fixed\\\n- ✅ Opponent features not bridged (60/280 dims zeroed) → bridged\\\n- ✅ Class imbalance → model collapsed to majority class → c\"],[1,\"round 1: ~1.43M transitions\\\n\\\n### Models\\\n| Version | Training | Eval Acc | Status |\\\n|---------|----------|----------|--------|\\\n| v1-v4 | Various | Degenerate | Fixed (LayerNorm, action mapping, features) |\\\n| v5-v10 | Imitation | 45-46% | Production baseline |\\\n| **v11** | 3.15M transitions, 290-dim, GPU | **45.2%** | **Current best** (+27.9 chips/hand) |\\\n| v12 | C\"],[0,\"lass\"],[-1,\"-\"],[1,\" \"],[0,\"weight\"],[-1,\"ed loss\\\n- ✅ Per-batch optimizer recreation destroyed AdamW state → epoch-level LR schedule\\\n\\\n### Next Steps\\\n1. ⬜ Rebuild `train_gen5` with CUDA, retrain v5 on fresh data (~8M transitions)\\\n2. ⬜ Verify v5 non-degeneracy (per-class accuracy, action distribu\"],[1,\"s | 45.5% | Marginal improvement |\\\n\\\n### Next Steps\\\n1. ⬜ Update `gen5_collect_full` to use **G47 as teacher** (currently G4)\\\n2. ⬜ Train v13 on 3.79M G47-collected transi\"],[0,\"tion\"],[-1,\")\"],[1,\"s\"],[0,\"\\\n3. \"]],\"start1\":3164,\"start2\":3164,\"length1\":615,\"length2\":567},{\"diffs\":[[0,\"aluate v\"],[-1,\"5\"],[1,\"13\"],[0,\" vs G\"],[-1,\"en \"],[0,\"4\"],[1,\"7\"],[0,\" in simu\"]],\"start1\":3735,\"start2\":3735,\"length1\":26,\"length2\":25},{\"diffs\":[[0,\"tion\"],[-1,\" (validation gate)\\\n4. ⬜ Implement reward assignment (for PPO Phase 2)\\\n5. ⬜ Fix `gen5_pipeline.sh` bash arithmetic bug for DAgger automation\\\n6. ⬜ PPO self-play training loop (curriculum Gen 1→2→3→4)\"],[1,\"\\\n4. ⬜ PPO fine-tuning (if imitation plateaus)\\\n5. ⬜ Postflop equity matrix for better postflop features\"],[0,\"\\\n\\\n--\"]],\"start1\":3762,\"start2\":3762,\"length1\":205,\"length2\":110},{\"diffs\":[[0,\"26-0\"],[-1,\"6-30\"],[1,\"7-06\"],[0,\")\\\n\\\n|\"]],\"start1\":3893,\"start2\":3893,\"length1\":12,\"length2\":12},{\"diffs\":[[0,\"gies | 2\"],[-1,\"0\"],[1,\"2\"],[0,\" |\\\n| Pro\"]],\"start1\":3961,\"start2\":3961,\"length1\":17,\"length2\":17},{\"diffs\":[[0,\"n bot | \"],[1,\"**\"],[0,\"Gen 4\"],[1,\".7** (g47_equity_preflop +\"],[0,\" \"],[-1,\"(\"],[0,\"g4_adapt\"]],\"start1\":3984,\"start2\":3984,\"length1\":23,\"length2\":50},{\"diffs\":[[0,\"onal\"],[-1,\", playing profitably |\\\n| Range prediction accuracy | ~88% hit rate |\\\n| g4_live vs chump | +132.9/hand (60k-hand validation) |\\\n| g4_live vs gen field | +0.95/hand (60k-hand validation) |\\\n| Gen 5 models trained | 4 (v1–v4, all degenerate — fixes in code\"],[1,\" |\\\n| G47 vs G45 field | **+10.6 BB/hand** (3 seeds × 10K hands) |\\\n| G47 vs Flock field | **+10.6 BB/hand** |\\\n| Gen 5 best model | v11, eval_acc=45.2%, +27.9 chips/hand |\\\n| Gen 5 data collected | 3.79M transitions (overnight G4 teacher) |\\\n| PEM2 matrix | 8.7 MB, 10 metrics × 8 bins × 169² pairs |\\\n| Regression tests | 28 hand replay tests (all pass\"],[0,\") |\\\n\"]],\"start1\":4084,\"start2\":4084,\"length1\":259,\"length2\":356},{\"diffs\":[[0,\". **\"],[-1,\"Rebuild `t\"],[1,\"T\"],[0,\"rain\"],[-1,\"_gen5` with CUDA** and retrain v5 on fresh collection data\\\n2. **Verify v5 non-degeneracy** — check per-class accuracy improves for minority a\"],[1,\" Gen5 v13** on the 3.79M overnight transitions with G47 preflop teacher\\\n2. **Update gen5_collect_full** to use G47 as teacher for future colle\"],[0,\"ction\"],[-1,\"s\"],[0,\"\\\n3. **\"],[-1,\"Evaluate v5 vs Gen 4** — if model loses, keep Gen 4 as production\\\n4. **Fix `gen5_pipeline.sh`** for automated DAgger iterations\\\n5. **Continue live play** with Gen 4 (profitable, stable)\"],[1,\"Live monitoring** of G47 performance (watch for preflop leaks)\\\n4. **Postflop exploration** — consider postflop equity matrix or learned postflop\\\n5. **Per-opponent PFR wiring** — use observed PFR per seat, not table average\"]],\"start1\":4485,\"start2\":4485,\"length1\":356,\"length2\":384}]"
metadata_diff: {"new":{},"deleted":[]}
encryption_cipher_text: 
encryption_applied: 0
updated_time: 2026-07-06T06:18:07.597Z
created_time: 2026-07-06T06:18:07.597Z
type_: 13