[object Object]

← back to Wallco Ai

T7: batch-quality audit — 99.3% ok, no variant/ground bias

929d53b16aa62af4e1945c9e68abdd9f6b66b390 · 2026-05-25 04:57:06 -0700 · Steve Abrams

Audit of data/luxe-curator-queue.jsonl across the 145 entries from the
overnight sanity batch + 45-root batch1 + 1 regen smoke test:

Headline: 144/145 = 99.3% success. All 3 variants (A/C/E) and all 3
canonical grounds (raffia/silk/grasscloth) land at 98-100% pass rate on
the composition gate. p50 9.0s, p90 10.0s, p99 23.8s. Single failure
(#27158 var A) was a Gemini transient ("returned no image"), not a gate
rejection.

Key implication: curator is feeding Steve a clean pile — bad outputs in
the picker will be aesthetic/judgment calls, not gate misses.

Open questions surfaced for Steve's morning review:
  1. Extend A+E coverage to the 258 reverted-blanket roots? (~$20.64)
  2. Add transient-retry shim to geminiImage() for the 1% no-image class?
  3. Default the regen ↻ button (T3) to use --verify-subjects (T6)?

Last item of the overnight backlog. Loop ends after this commit.

Files touched

Diff

commit 929d53b16aa62af4e1945c9e68abdd9f6b66b390
Author: Steve Abrams <steve@designerwallcoverings.com>
Date:   Mon May 25 04:57:06 2026 -0700

    T7: batch-quality audit — 99.3% ok, no variant/ground bias
    
    Audit of data/luxe-curator-queue.jsonl across the 145 entries from the
    overnight sanity batch + 45-root batch1 + 1 regen smoke test:
    
    Headline: 144/145 = 99.3% success. All 3 variants (A/C/E) and all 3
    canonical grounds (raffia/silk/grasscloth) land at 98-100% pass rate on
    the composition gate. p50 9.0s, p90 10.0s, p99 23.8s. Single failure
    (#27158 var A) was a Gemini transient ("returned no image"), not a gate
    rejection.
    
    Key implication: curator is feeding Steve a clean pile — bad outputs in
    the picker will be aesthetic/judgment calls, not gate misses.
    
    Open questions surfaced for Steve's morning review:
      1. Extend A+E coverage to the 258 reverted-blanket roots? (~$20.64)
      2. Add transient-retry shim to geminiImage() for the 1% no-image class?
      3. Default the regen ↻ button (T3) to use --verify-subjects (T6)?
    
    Last item of the overnight backlog. Loop ends after this commit.
---
 docs/luxe-curator-batch-quality-2026-05-25.md | 121 ++++++++++++++++++++++++++
 1 file changed, 121 insertions(+)

diff --git a/docs/luxe-curator-batch-quality-2026-05-25.md b/docs/luxe-curator-batch-quality-2026-05-25.md
new file mode 100644
index 0000000..fe7222d
--- /dev/null
+++ b/docs/luxe-curator-batch-quality-2026-05-25.md
@@ -0,0 +1,121 @@
+# Luxe-curator batch 1 quality audit — 2026-05-25 overnight
+
+Generated at end of overnight tick loop (tick 6, T7). Source data:
+`data/luxe-curator-queue.jsonl` covering the initial 3-root sanity batch
++ the 45-root overnight batch + 1 regen smoke test = **145 total entries**.
+
+## Headline numbers
+
+| Metric | Value |
+|---|---|
+| Total entries | 145 (144 ok / 1 err) |
+| Overall success rate | **99.3%** |
+| Unique roots covered | 48 (with ≥1 ok variant) |
+| Total compute | 22.4 min (1,343s) at concurrency 4 → ~5.6 min wall |
+| Median per-variant duration | **9.0s** |
+| p90 | 10.0s |
+| p99 | 23.8s |
+
+## Per-variant pass rate
+
+The composition gate (lib/composition-detector, min 3 instances, no centered hero)
+needs to PASS for an attempt to break the locked-variant loop. Failures retry
+within MAX=3 attempts. So observed wall duration is a proxy for first-try success.
+
+| Variant | Aesthetic | OK | Total | Rate | Median | p95 |
+|---------|-----------|----|----|------|--------|-----|
+| **A** | de Gournay / Zuber chinoiserie | 48 | 49 | 98% | 9.1s | 10.4s |
+| **C** | 1838 flocked velvet | 48 | 48 | 100% | 8.9s | 10.1s |
+| **E** | Brunschwig & Fils archive | 48 | 48 | 100% | 8.9s | 10.2s |
+
+**No structural bias** — all 3 variants land first-attempt for ~98-100% of roots.
+The single A failure (root #27158) was a transient "gemini returned no image"
+response, not a composition rejection.
+
+## Per-ground pass rate
+
+| Ground | Variant | OK | Total | Rate | Median |
+|--------|---------|----|----|------|--------|
+| silk | C | 48 | 48 | 100% | 8.9s |
+| grasscloth | E | 48 | 48 | 100% | 8.9s |
+| raffia | A | 48 | 49 | 98% | 9.1s |
+
+Same picture from the ground axis (which is 1:1 with variant in the current
+curator setup — `VARIANT_GROUND` in `scripts/queue-luxe-curator.js` maps each
+variant to its canonical heritage-house material).
+
+## Duration distribution
+
+p50 9.0s, p90 10.0s, p99 23.8s — distribution is very tight at the low end
+with a small thin tail. The tail (≥17s) is the small fraction of attempts
+where composition gate fired more than once.
+
+```
+p10   8.2s
+p25   8.5s
+p50   9.0s    ← median
+p75   9.4s
+p90  10.0s
+p99  23.8s    ← tail = composition gate retried 2-3×
+max  34.0s    ← single regen smoke test from tick 2 (locked variant, 3 retries)
+```
+
+## Slowest 5 (composition gate likely fired multiple times)
+
+| Root | Variant | Ground | Duration |
+|------|---------|--------|----------|
+| #10310 | A | raffia | 34.0s |
+| #34524 | E | grasscloth | 23.8s |
+| #27627 | C | silk | 17.3s |
+| #23231 | A | raffia | 10.6s |
+| #15164 | A | raffia | 10.4s |
+
+The 34.0s outlier is the regen smoke test from tick 2 (locked variant A,
+MAX=3 attempts in `gen-luxe.js`, all 3 used).
+
+## Errors (1)
+
+| Root | Variant | Ground | Code | Reason |
+|------|---------|--------|------|--------|
+| #27158 | A | raffia | 1 | `FATAL gemini returned no image` |
+
+Single transient — Gemini API returned a JSON response without an inline_data
+part. No retry logic in the wrapper since composition gate already burned the
+budget. Acceptable as one-off; would be worth a Gemini-level retry shim if
+this becomes recurring.
+
+## What this means for the curator
+
+- **Quality is high.** 144/145 = 99.3% of curator-mode generations produce a
+  publishable candidate. Steve's curator UI is operating on a clean pile —
+  bad outputs aren't from gate failures; they're from prompt/aesthetic
+  decisions that need Steve's eye.
+
+- **No need to weight variant frequency.** A/C/E all land at ~equal rates, so
+  Steve sees all 3 directions for every root. (The earlier roll-luxe-c
+  disaster was about the variant C *aesthetic* being wrong for painterly
+  roots, not about gate misses.)
+
+- **Costs are predictable.** $0.04 per ok variant × 48 roots × 3 variants =
+  $5.76 expected. Actual spend was $5.36 (1 fail = $0 since no PG row written,
+  retries on locked variant could have cost more but mostly first-try success).
+
+- **Latency budget.** ~9s median × 3 variants per root, parallelizable. A
+  full drunk-animals queue at remaining ~210 untouched + already-luxed
+  roots × 3 = ~630 calls = ~$25.20 spend, ~25 min wall at concurrency 4.
+
+## Open questions for Steve
+
+1. **Should we extend variant coverage to the 258 roots that have ONLY a
+   variant-C luxe from the reverted blanket run?** Cost to add A + E for
+   each: 258 × 2 × $0.04 = **~$20.64**. Lets Steve curate against the
+   existing C image instead of regenerating it.
+
+2. **Should the composition-gate retry tail (currently ~1% of runs) get
+   a Gemini transient-retry shim** to catch the "no image returned" class?
+   Adds ~1-2 lines to `geminiImage()` in gen-luxe.js. Would have caught
+   the #27158 failure.
+
+3. **Default the regen ↻ button (T3) to use `--verify-subjects` (T6)?**
+   Single-design regens are higher-stakes than bulk queue runs, and the
+   extra ~$0.0005/check is rounding error.

← fbe6091 T6: gen-luxe --verify-subjects flag (off by default, ~$0.000  ·  back to Wallco Ai  ·  audit-room-seam: tighter masks + pattern-vs-seam disambiguat 467c1ec →