← back to Model Arena
BRIEF.md
26 lines
# model-arena — brief
- **Seed:** Julian Goldie SEO (@JulianGoldieSEO) X post, Jul 20 2026
(https://x.com/JulianGoldieSEO/status/2079008392811544994) — head-to-head
real-world testing of Kimi K3 vs Claude Fable 5 vs GPT 5.6 on build
challenges (Skyrim-style open world, Dragon Realm, Neon City driving).
Takeaway: *"Don't marry one AI model. Build a workflow where every model
does the job it performs best at — trust real-world testing over
leaderboard screenshots."*
- **Client:** internal (Steve) first; productizable later.
- **Outcome / "high value":** a Model Arena — fan ONE real-world build
challenge out to multiple AI models, render each model's single-file HTML
artifact side-by-side in sandboxed iframes, crown a winner per challenge,
and accumulate a **real-world win-rate ledger per model** over time.
First-party benchmark data nobody else has (Codex panel: "defensible moat").
- **Direction locked by DTD panel 2026-07-22: unanimous 5/5** for the arena
(vs an AI-SEO tool or a static SEO landing page). Panel cost ~$0.006.
- **Constraints:**
- $0 by default — local Ollama models (Mac2 localhost + Mac1 192.168.1.133).
- Paid model calls (Claude / Kimi / Codex / Grok) are OPT-IN, per-run,
cost-estimated up front and cost-logged after.
- Demo behind basic auth `admin / DW2024!` (env-overridable `BASIC_AUTH`).
- House rules: grid gets sort + density slider; admin cards show created
date + time; gitified; zero-dep Node server.
- Launch is GATED — nothing deploys or goes public without Steve's go.