← back to Model Arena

BRIEF.md

26 lines

# model-arena — brief

- **Seed:** Julian Goldie SEO (@JulianGoldieSEO) X post, Jul 20 2026
  (https://x.com/JulianGoldieSEO/status/2079008392811544994) — head-to-head
  real-world testing of Kimi K3 vs Claude Fable 5 vs GPT 5.6 on build
  challenges (Skyrim-style open world, Dragon Realm, Neon City driving).
  Takeaway: *"Don't marry one AI model. Build a workflow where every model
  does the job it performs best at — trust real-world testing over
  leaderboard screenshots."*
- **Client:** internal (Steve) first; productizable later.
- **Outcome / "high value":** a Model Arena — fan ONE real-world build
  challenge out to multiple AI models, render each model's single-file HTML
  artifact side-by-side in sandboxed iframes, crown a winner per challenge,
  and accumulate a **real-world win-rate ledger per model** over time.
  First-party benchmark data nobody else has (Codex panel: "defensible moat").
- **Direction locked by DTD panel 2026-07-22: unanimous 5/5** for the arena
  (vs an AI-SEO tool or a static SEO landing page). Panel cost ~$0.006.
- **Constraints:**
  - $0 by default — local Ollama models (Mac2 localhost + Mac1 192.168.1.133).
  - Paid model calls (Claude / Kimi / Codex / Grok) are OPT-IN, per-run,
    cost-estimated up front and cost-logged after.
  - Demo behind basic auth `admin / DW2024!` (env-overridable `BASIC_AUTH`).
  - House rules: grid gets sort + density slider; admin cards show created
    date + time; gitified; zero-dep Node server.
  - Launch is GATED — nothing deploys or goes public without Steve's go.