bf43f1b2026-07-25DW Suite index page (/suite, /dw) — categorized landing for the 20 agent-built DW builds (Tools/Games/Social/Art), reads /api/arcade; +category tags on featured entries + API exposes them
01935282026-07-255x sweep 1: fix Color Story modal — visibility:hidden when closed so the × isn't in the tab/hit-test/a11y tree (was timing out clickthrough)
682fd462026-07-25Run on Claude Max subscription only: drop paid API models (kimi/gpt/grok) from arena roster — claude-code (Max plan) + 5 local models = $0. Agent-driven builds already on Max plan.
7527c242026-07-25Agent-driven pivot: DW Color Story Studio (frontend-developer, plan+execute, 72KB self-contained luxury tool) — stripped Google Fonts for CSP-safety, registered as featured arcade game
f631a812026-07-25Cost: drop paid models (kimi/gpt/grok) from loop roster per Steve — now /bin/zsh local/CLI-only (6 free models). Was ~$0.17/battle; total night spend $25.13.
86330ef2026-07-25night-loop: cycle 04:48 — judged=64da14ce1509 · fired 1 →; FIRED idx=14/16 id=4497fc709c11 title=Designer Wallcoverings — Colorway Swatch Match Game
1ffc3f92026-07-25Shell: +no-external-images (picsum breaks under CSP) +instantiate-wordmark-in-DOM (weak models defined class but never used it). Log 53ec (2.5 model-exec fail) + polish
447e1f12026-07-25night-loop: cycle 04:36 — judged=45271386f3f4 · fired 2 →; FIRED idx=12/16 id=035b7009e366 title=Designer Wallcoverings — Pattern Runner Game; FIRED idx=13/16 id=64da14ce1509 title=Designer Wallcoverings — This or That Style Poll
f4b407f2026-07-25ALL-MUST-WORK: local models (qwen3-14b, hermes3-8b) use design PACK not the live tool loop — 8B/14B returned 0 chars through function-calling; API models keep the tool loop
5092bf12026-07-25night-loop: cycle 01:45 — judged=71a0784e71c8 · fired 1 →; FIRED idx=5/16 id=5410cf08e245 title=Designer Wallcoverings — Match the Motif Memory Game
49020512026-07-25Fix run-next payload: designTools True (Python bool, not JS true) — the last fires were erroring
f9fd5f62026-07-25night-loop: cycle 00:49 — judged=d7f4f8ca5f8d · fired 2 →; NameError: name 'true' is not defined. Did you mean: 'True'?; NameError: name 'true' is not defined. Did you mean: 'True'?
3e64ec72026-07-25Beauty gate: auto-remix winners scoring <7 (make it more beautiful) via arena referee+remix lineage; MAXP-aware, one remix/battle
e6af3f82026-07-25Enable 🎨 design-tools belt (opendesign/hyperframes/figma) on all loop battles — models fetch real palettes/type/motion for beautiful UI
5f0ddc02026-07-25Repoint loop to games + DW social-media idea pool (16 ideas: IG carousel/reel/shoppable/calendar/link-in-bio + swatch/memory/runner/roulette games); pointer reset
0e921bc2026-07-24Add Games Arcade: /arcade gallery + /api/arcade auto-discovery of crowned game challenges, /game/:slug standalone champions, DW+AA idea pool + all-night loop harness
19eb62d2026-07-23gemini image-gen fallback backend for the photoshop tool (Adobe preferred when creds land); fix .env load order — design-tools registered before env was loaded, so key-gated tools never appeared
7bd9ccc2026-07-23photoshop design tool: Adobe Firefly Services — IMS auth, Firefly text-to-image (jpeg-compressed assets), Photoshop cutout via unguessable auth-exempt asset GET + single-use PUT-back routes, {{PS_ASSET:id}} placeholders inlined as data URIs at artifact save; tool self-registers only when ADOBE_CLIENT_ID/SECRET are set
8a62aa02026-07-23qwen3: disable thinking (think:false) and raise num_predict to 10240 — the <think> block was eating the whole token budget on creative challenges, returning 0 chars of HTML
afeb0922026-07-23README: document the design-tools belt
7df27982026-07-23design tools for arena models: figma / opendesign / hyperframes tool belt — OpenAI-style function-calling loop for GPT/Grok/Kimi + tool-capable Ollama models (qwen3, hermes3), design-pack prompt injection for the rest; per-battle 🎨 toggle, per-run toolCalls recorded and badged in UI
ca4eed62026-07-23watchdog: age queued runs from their own queued_at, not challenge created_at — retries/resumes on battles older than 20 min were being falsely killed as 'never started'; bump queue ceiling to 45 min for deep serialized queues
bde0e0c2026-07-23add 'Re-run entire challenge' button (re-queues every model in a battle)
0dbc4862026-07-23auto-resume runs interrupted by a server restart: on boot, re-queue running/queued/mid-run-error runs (cap 4 resumes per run) instead of abandoning them as errors; heal stuck judging flag and re-judge affected battles
41ff05b2026-07-23arena: harden with unhandledRejection/uncaughtException guards (a stray async abort under a heavy batch no longer kills the server mid-run)
b25a81b2026-07-23resilience: global unhandledRejection/uncaughtException handlers so a stray async failure under batch load doesn't kill the server mid-run
1e4746a2026-07-23arena: move Build-in-iTerm from whole-challenge to per-MODEL-CHIP — pick which model's attempt to build; brief embeds that model's actual artifact so Claude builds it forward. Per-pane ⚒ in detail view too.
584400e2026-07-23arena: Build-in-iTerm button per challenge (+ detail view) — writes the challenge brief to a temp file and opens a real interactive Claude session in a new iTerm2 window to build it; safe file-based prompt escaping
e93e6412026-07-23arena: Claude via Max plan (Claude Code CLI adapter, $0 subscription not API) + HuggingFace top-coder auto-register (bartowski Qwen2.5-Coder-32B GGUF pull, real ollama-tag availability check, HF badge + downloading state)
341ac472026-07-23roster: remove Claude Fable 5 (Anthropic account has no credit) — 7 live models, no dead option erroring every battle
c7f13162026-07-23yolo: RUN COMPLETE — 14 wild ideas across 3 iterations, final summary; 6-way clean, all local models live, one gated draft, Claude pending credit
89557232026-07-23yolo idea E: tournament bracket — seeded single-elim visualization over a battle's AI-referee consensus scores (Semis/Final/Champion), $0 client-side
0c31fff2026-07-23yolo: regression note (documents the b-cat hidden-view clickthrough false-positive, proven working) + README iteration-2 features
6fea94d2026-07-23yolo idea D: head-to-head record matrix — pairwise W-L (row vs col) from crowned battles, color-coded, under the leaderboard
8264d562026-07-23yolo idea C: model profile page — clickable leaderboard rows open a model's artifact gallery across all battles + aggregate stats
2a631a52026-07-23yolo idea B: export a battle as a self-contained standalone HTML (artifacts inlined via sandboxed srcdoc + scoreboard), downloadable, renders offline
dfbd21a2026-07-23yolo idea A: per-category leaderboards — auto-tag challenges (Games/Real Work/Custom), category-filtered ELO ledger; 'which model is best at real work?' is now answerable
d2bb0692026-07-22yolo idea 9: multi-vision CONSENSUS referee — panel of local vision models (qwen2.5vl + minicpm-v) vote per artifact; consensus score + per-panelist breakdown + spread/disagreement flag; consensus now agrees with human crown on the smoke test
0c1d6912026-07-22yolo: log backlog + mid-run status (8 shipped, continuing)
0688aaa2026-07-22yolo idea 8: artifact source view (</> per pane) + README documents all 8 yolo features
9944dae2026-07-22yolo idea 7: remix mode — feed the winning/AI-pick artifact's HTML back to models to improve on, spawning an evolutionary round with remixOf lineage; $0 local default, verified on product page
b15a98b2026-07-22yolo idea 6: daily auto-challenge generator (accrues real judged data, $0 local) + AI-judged; launchd install drafted to pending-approval (gated, not installed)
38540992026-07-22yolo idea 5: real-workload preset pack (Strategist's fix) — DW-relevant challenges (product page, sortable catalog, marketing email, room visualizer, JSON→grid) grouped under Real Work vs Games; verified gen+render+vision-judge end-to-end on a product page
7ec7ada2026-07-22yolo idea 4: blind judging mode — anonymize models to Model A/B/C during voting (Chatbot-Arena style), reveal on crown; hides AI pick + names/alts/titles
f109f272026-07-22yolo idea 3: ELO ranking (K=24) + cost-per-win + avg AI-referee score + human/AI agreement % — leaderboard is now a real ranking, not a flat win-rate
dc3d7322026-07-22yolo idea 2: AI auto-referee — local vision model (qwen2.5vl:7b) LOOKS at each artifact screenshot, scores 0-10 + one-line reason, auto-judges on battle settle, surfaces an AI PICK alongside the human crown. Directly answers contrarian 'one builder vote = theater'. $0 local.
37d69932026-07-22yolo idea 1: artifact auto-screenshots — headless-render each artifact to a thumbnail; visual card strip + click-to-run-live poster panes; startup backfill + /thumb route
84ddac12026-07-22mac1: revive Ollama (unstuck pinned qwen3:14b); restore arena 4th-local slot to Mac1 host (OLLAMA_MAC1=1) for 2-host parallelism — verified concurrent Mac1+Mac2 run
Wraps the 7:15am daily battle: persistent log (not /tmp) + CNCP parking-lot
card + best-effort George email on any nonzero exit or missing fresh daily-log
entry. Also un-silences the fs.appendFile catch. Closes the silent-death gap
from the /yolo contrarian gate. $0 (local only). TK-00107.
6c375f2 · 2026-07-23 · Pin the Claude arena competitor to Opus (CLAUDE_MODEL, default opus)
The claude-code entrant was spawned with bare 'claude -p', so it raced on
whatever the CLI defaulted to (could drift to Sonnet), silently benchmarking
the wrong Claude and skewing the win-rate ledger. Both the headless -p run and
the interactive build-forward launch now pass --model opus, env-overridable.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
File tree
2959 files tracked. Click any to browse the source at HEAD.