← back to Exo Helper

verification/model-placement-TK11324.md

66 lines

# Model placement audit — TK-11324

Read-only measurements taken September 9, 2026, approximately 10:59–11:03 AM America/Los_Angeles. No model files were transferred or deleted; no inference placement was changed.

## Recommendation

Keep the working Exo helper and active llama-server model on the M3 Ultra. Do not move the large model caches to the other Macs under their current load. Investigate archiving the two unloaded Exo caches below to the attached Henry drive, subject to confirming future use and validating an archive and restore procedure. This is a conditional recommendation, not an approved migration.

## Current capacity

| Machine | Total RAM | Available RAM | Disk free | Swap used |
| --- | ---: | ---: | ---: | ---: |
| M3 Ultra, this Mac | 96 GiB | 13.2 GiB | 25.4 GiB | 4.9 GiB |
| M1 Max, 192.168.1.133 | 32 GiB | 1.7 GiB | 61.5 GiB | 10.2 GiB |
| M2 Max, ms2 | 32 GiB | 8.4 GiB | 2.7 GiB | 8.3 GiB |

Available memory is a snapshot, not a permanent limit on possible workloads. Model file size does not include all inference memory requirements. Henry is a locally attached external disk mounted at `/Volumes/Henry`, with approximately 346 GiB free. Capacity and readability were checked; archive integrity and restore performance were not tested.

## Local models

The Exo cache occupies approximately 81 GiB; the Ollama-format cache occupies approximately 42 GiB. Files in the latter are also used by llama.cpp, independently of the Ollama runtime.

| Model | Disk size | Finding and recommendation |
| --- | ---: | --- |
| Exo Qwen3.6-35B-A3B-8bit | 35.6 GiB | Complete cache; no active Exo instance. Candidate for archival after use audit. |
| Exo Qwen3.8-27B-8bit | 27.5 GiB | Complete cache; no active Exo instance. Candidate for archival after use audit. |
| Exo Shiftedx Qwen3.8-27B-Abliterated-MLX-MXFP4 | 14.2 GiB | Expected weights present; no active Exo instance. Broader use not established. |
| Exo Qwen3-VL-4B-Instruct-4bit | 2.9 GiB | Active Claude Code/Codex helper. Keep local. |
| Exo Qwen3-0.6B-8bit | 0.6 GiB | Small possible future peer helper; moving it provides little storage relief. |
| Ollama-format qwen3:14b | 8.64 GiB | Active llama-server has its weight blob open. Keep in place. |
| Ollama-format qwen3.8-27b-heretic | 16.0 GiB | Present; future use not audited. |
| Ollama-format gemma3:12b | 7.59 GiB | Present; future use not audited. |
| Ollama-format hermes3:8b | 4.34 GiB | Present; future use not audited. |
| Ollama-format qwen2.5vl:7b | 5.56 GiB | Present; future use not audited. |

The two principal archival candidates total approximately **63.1 GiB**. Copying them alone would not reclaim local disk space or RAM. Local space would only be reclaimed after a verified archive and separately authorized removal of the originals.

Partial Exo downloads include Meta-Llama-3.1-70B-Instruct-4bit (~0.23 GiB) and Qwen3-Coder-Next-6bit (~0.28 GiB). The Shiftedx MTP variant contains metadata only. These are not complete transferable inference models.

## Peer and archive findings

- M1 Max: approximately 4.1 GiB of Exo files, predominantly partial downloads, and 24 GiB of Ollama-format files across its user accounts.
- M2 Max: approximately 21 GiB of Exo files, including a partial Qwen3.8-27B-8bit download, and 33 GiB of Ollama-format files. Its 2.7 GiB free disk makes it unsuitable for another large copy today.
- Henry already contains an Exo offload directory with Qwen3.6-27B-4bit (~15 GiB) and gpt-oss-20b-MXFP4-Q8 (~11 GiB), plus an Ollama-format model directory. Their integrity was not established by this audit.

## Evidence and limits

Measurements used the live Exo `/state` endpoint, local and SSH `df`/`du` inspections, cache manifests and shard inventories, and `lsof`/`ps`. The active llama-server process had the Qwen3:14b blob open; that digest was matched to its manifest. A targeted scan of project source and LaunchAgents found no direct static references to the two principal archival candidates. That does not exclude dynamic callers or future scheduled use. An unloaded model is not necessarily unused.

## DTD decision record

The canonical cost mode was ZERO_COST_REQUIRED. The panel ran without Ollama or paid API voters.

| Participant | Actual execution | Result |
| --- | --- | --- |
| Codex | Signed-in CLI; log identifies gpt-6-astra | B |
| Qwen | Local Exo Qwen3-VL-4B-Instruct-4bit | B |
| Claude | Disabled under zero-cost mode | Abstained |
| Grok | Unavailable under zero-cost mode | Abstained |
| Kimi | Unavailable under zero-cost mode | Abstained |
| Muse | Unavailable without Ollama | Abstained |

Options: A, move large models to peers now; B, preserve active models and investigate conditional archival; C, make no recommendation. Result: B, 2/2 valid votes, 2/6 participants available; medium confidence. The post-decision Codex adversarial review returned **FINAL: KEEP**, emphasizing that B remains conditional and authorizes no migration. Its concerns—snapshot RAM, unknown future usage, untested restore integrity, and copying not reclaiming space—are retained above.

Panel evidence remains in `/private/tmp/dtd-model-placement-TK11324/`; sampled Exo state is `/private/tmp/exo-model-placement-state.json`. Temporary evidence may not survive reboot. This report preserves the findings.