{"slug":"exo","total":2078,"limit":100,"offset":0,"since":null,"commits":[{"hash":"27b85c8f","date":"2026-07-02 19:27:52 -0700","author":"Steve Abrams","subject":"local: darwin mlx from stock PyPI wheel 0.31.2 (no Metal toolchain needed; LAN cluster, TB-RDMA fork not applicable)","body":""},{"hash":"6abccd1b","date":"2026-07-02 19:26:32 -0700","author":"Steve Abrams","subject":"local: darwin mlx from PyPI wheel (no Metal toolchain; LAN cluster, no TB-RDMA)","body":""},{"hash":"b5375f8c","date":"2026-06-22 15:38:54 +0200","author":"aidiffuser","subject":"Add Kimi K2.7-Code model card (official INT4 weights + vision) (#2167)","body":"Adds a model card for\n[moonshotai/Kimi-K2.7-Code](https://huggingface.co/moonshotai/Kimi-K2.7-Code),\nreleased 2026-06-12.\n\nSame architecture as Kimi K2.6 (`kimi_k25`, 61 layers, official INT4),\nso the card mirrors the existing `moonshotai--Kimi-K2.6.toml`. Sampling\ndefaults per the model card (temperature 1.0 / top_p 0.95 for thinking\nmode).\n\n**Vision:** the official repo ships MoonViT weights inline, so I\nextracted the 335 `vision_tower.*` / `mm_projector.*` tensors\n(unmodified bf16) into\n[aidiffuser/Kimi-K2.7-Code-vision](https://huggingface.co/aidiffuser/Kimi-K2.7-Code-vision),\nfollowing the `exolabs/Kimi-K2.6-vision` format. The vision config is\nbyte-identical to K2.6's; the extraction script is included in the repo\nfor verification. Happy to have this re-hosted under the exolabs org if\nyou prefer — it's a one-line change to the card.\n\n**Tested:** distributed serving on 2× Mac Studio M3 Ultra (512 GB),\ntensor parallelism, text + thinking + image understanding all confirmed\nworking.\n\nCo-authored-by: aidiffuser <your-noreply-email@users.noreply.github.com>\nCo-authored-by: Claude Fable 5 <noreply@anthropic.com>"},{"hash":"cdf1add8","date":"2026-06-22 18:59:15 +0530","author":"OrbisAI Security","subject":"fix: upgrade devalue to 5.6.2 (CVE-2026-22774) (#2150)","body":"## Summary\nUpgrade devalue from 5.5.0 to 5.6.2 to fix CVE-2026-22774.\n\n## Vulnerability\n| Field | Value |\n|-------|-------|\n| **ID** | CVE-2026-22774 |\n| **Severity** | HIGH |\n| **Scanner** | trivy |\n| **Rule** | `CVE-2026-22774` |\n| **File** | `dashboard/package-lock.json` |\n| **Assessment** | Likely exploitable |\n\n**Description**: devalue: devalue: Denial of Service due to excessive\nresource consumption from untrusted input\n\n## Evidence\n\n**Scanner confirmation**: trivy rule `CVE-2026-22774` flagged this\npattern.\n\n**Production code**: This file is in the production codebase, not\ntest-only code.\n\n## Threat Model Context\n\nThis is a web service - vulnerabilities in request handlers are directly\nexploitable by remote attackers.\n\n## Changes\n- `dashboard/package.json`\n- `dashboard/package-lock.json`\n\n## Verification\n- [x] Build passes\n- [x] Scanner re-scan confirms fix\n- [x] LLM code review passed\n\n---\n*This change addresses a pattern flagged by static analysis. The code\npath handles user-influenced input and the fix reduces the attack\nsurface against both manual and automated exploitation.*\n\n---\n*Automated security fix by [OrbisAI Security](https://orbisappsec.com)*"},{"hash":"09f9ea31","date":"2026-06-03 16:31:56 +0100","author":"Evan Quiney","subject":"libp2p -> zenoh (#2132)","body":"supercedes #2076 and #2073\n\n---------\n\nCo-authored-by: Andrei Cravtov <the.andrei.cravtov@gmail.com>"},{"hash":"81d7cb0f","date":"2026-06-03 00:35:18 +0900","author":"Sakutaro","subject":"docs: add Homebrew cask install instructions (#2140)","body":"## Motivation\n\nexo is now available as a Homebrew cask, so the README should show the\nsimplest macOS installation path alongside the existing DMG download.\n\nFixes https://github.com/exo-explore/exo/issues/2105\nhttps://github.com/exo-explore/exo/issues/176\n\n## Changes\n\n- Added `brew install --cask exo` to the macOS App section of\n`README.md`\n- Kept the existing DMG download link as the first installation option\n\n## Why It Works\n\nAdding the Homebrew cask command gives macOS users a\npackage-manager-managed installation path while preserving the existing\nDMG download option.\n\n## Test Plan\n\n### Manual Testing\n\n- Reviewed the rendered Markdown structure in `README.md`\n\n### Automated Testing\n\n- Not run. Documentation-only change.\n\n## Related\n\n- https://github.com/Homebrew/homebrew-cask/pull/265956"},{"hash":"629c55d6","date":"2026-05-31 19:23:41 +0100","author":"Andrei Cravtov","subject":"Rename exo_pyo3_bindings to exo_rs (#2131)","body":"## Motivation\n\n(I think it) Makes Evan's massive PR easier to merge later on\n\n## Changes\n\n- Renamed exo_pyo3_bindings to exo_rs\n- Upgraded versions of pyo3-based dependencies\n- Renamed PyFromSwarm to just FromSwarm, and PyNetworkingHandle to just\nNetworkingHandle"},{"hash":"f9f8cbb3","date":"2026-05-29 18:37:47 +0100","author":"Andrei Cravtov","subject":"fix: make app builds work again (#2127)","body":"## Motivation\n\nThey didn't\n\n## Changes\n\nThey now do\n\n## Why It Works\n\nI changed an env flag, and added a keyword"},{"hash":"051a64e3","date":"2026-05-28 14:42:36 -0700","author":"ciaranbor","subject":"Capture energy in prefill and ageneration separately (#2124)","body":"## Motivation\n\nEnergy was reported as a single aggregate. Split into prefill vs.\ngeneration so each phase can be analysed independently.\n\n## Changes\n\n- `PowerSampler`: `mark_prefill_done()` + `trapezoidal_energy_range()`\nhelper; `result()` now emits per-phase splits.\n- `PowerUsage` / `NodePowerStats`: optional `prefill_*` / `generation_*`\nfields (back-compat: `None` if unmarked).\n- API marks the boundary on the first non-`PrefillProgressChunk`.\n- `bench/exo_bench.py` surfaces the split in the log line and persists\n`power_usage` to JSON.\n- METHODOLOGY: one sentence + one bullet.\n\n## Why It Works\n\nFirst non-prefill chunk *is* the boundary. Anchoring a sample there and\ninterpolating power at the boundary makes phase energies sum exactly to\nthe unsplit total.\n\n## Test Plan\n\n### Manual Testing\n\n`eco`-reserved nodes:\n- M3 Ultra, Qwen3-VL-4B, pp=8192/tg=1024: server 1940 J vs client 1931 J\n(+0.5 %)\n- M4 Pro, Qwen3.6-27B, pp=16384/tg=2048: server 20,292 J vs client\n20,221 J (+0.35 %)\n\n### Automated Testing\n\n5 new tests in `test_power_sampler.py` (range integrator,\nsplits-sum-to-total, `None`-when-unmarked, idempotency). 14/14 pass."},{"hash":"a8602ea6","date":"2026-05-26 14:42:39 +0100","author":"Andrei Cravtov","subject":"fix(bug): no longer repeated _trigger_notify_user_to_download_model (#2114)","body":"## Motivation\n\nPartially fixes [this](https://github.com/exo-explore/exo/issues/2098)\nissue. Removed erroneous logic for telling user to download when they\nalready downloaded.\n\nCould not figure out about the \"spontaneous crashes\" in that issue,\nauthor should consolidate more logs and open a new issue dedicated to\nthat. I believe\n[this](https://github.com/exo-explore/exo/commit/74e9fe15e62fe189dc7e019db86e75c83eca2721)\ncommit solved some EventRouter-related crashes, which was mentioned in\n[this](https://github.com/exo-explore/exo/issues/2098) issue, so it may\nhave already been solved. If not, should be re-submitted as a new issue.\n\n## Changes\n\n- Consolidated _resolve_and_validate_text_model and\n_validate_image_model into one function: _validate_model_has_instance;\n- + They already had virtually identical logic, it being different seems\nto be an artifact of history\n- + Added logic to ensure that _trigger_notify_user_to_download_model is\nonly called when no such model is downloaded, not just if there is no\ninstance of it\n- Added a new `/instance/await` SSE streaming endpoint to wait for when\na model has an instance available. Complements instance-placement API,\nso we can wait till that is done without client-side polling.\n- Updated docs and a /tmp script to reflect some of the changes\n- Updated dashboard `getModelForRequest` to only return model ID if an\ninstance exists for it, and updated bits to use `handleChatSend` instead\nof `sendMessage` because that checks for if a model instance exists\nfirst.\n\n## Why It Works\n\nThe problem was that there was erroneous logging for model not\ndownloaded. I fixed that logic. The rest is extra."},{"hash":"a1a22b5f","date":"2026-05-25 20:42:47 +0100","author":"Andrei Cravtov","subject":"feat: added background/daemon support (#2106)","body":"## Motivation\n\nAddresses [this](https://github.com/exo-explore/exo/issues/1931) issue.\n\n## Changes\n\nYou can now launch Exo as a legacy SysV-style daemin (in the background)\nwith `--legacy-daemon` flag.\nNOTE: don't use it if you're managing Exo with systemd or launchd\n\nSIDE FIX: the macmon process not found trace is no longer displayed on\nprocess shutdown via ctrl+c, that error is supressed.\n\n## Why It Works\n\nBecause I used a daemonization library and tweaked it not to break\nmultiprocessing.\n\n## Test Plan\n\nI ran it in daemon mode, non daemon mode, etc., and pid locking +\ninference + everything else works just fine.\n\nAlso ran it `ssh user@host -t 'cd exo && nohup nix run .#exo --\n--legacy-daemon'` on a 4-node TB mac-mini cluster and the mDNS didn't\ndie"},{"hash":"74e9fe15","date":"2026-05-22 14:20:04 +0100","author":"Andrei Cravtov","subject":"fix(bug): EventRouter lifetime-handling fixed, no more process crashes (#2102)","body":"## Motivation\n\nTrying to (partially) fix\n[this](https://github.com/exo-explore/exo/issues/2101) issue.\n\n## Changes\n\nChanged channels (in channels.py) to support exception overriding.\n\nMade EventRouter channels throw a subclass of the resource closed/broken\nerrors.\n\nThe current lifetime logic of EventRouter in event loop no longer blows\nup because components that use channels from EventRouter now catch the\nsubclass exceptions in the run method: Worker, Master,\nDownloadCoordinator, RunnerSupervisor.\n\nAdded logic to throw when API server exits without being asked to shut\ndown - this kill the sleep-forever in the task-group."},{"hash":"90f24bef","date":"2026-05-15 16:17:35 +0100","author":"Evan Quiney","subject":"fix model cards not validating properly after #2071 (#2096)","body":""},{"hash":"5097b266","date":"2026-05-15 14:04:50 +0100","author":"Andrei Cravtov","subject":"Tweaked workspace settings (#2095)","body":"workspace settings"},{"hash":"bc6661e6","date":"2026-05-15 13:52:12 +0100","author":"rltakashige","subject":"Add node backends to model cards (#2071)","body":"Co-authored-by: Evan <evanev7@gmail.com>"},{"hash":"14aab356","date":"2026-05-15 13:40:59 +0100","author":"Andrei Cravtov","subject":"Runner error handling (#2093)","body":"# Runner error handling\n\n## Motivation\n\nRunner failures were mostly surfaced as plain shutdown messages, which\nmade root cause hard to spot from API errors or runner status.\n\nThis adds a MVP path for preserving runner crash context and attaching\nknown stderr diagnostics to failure reports.\n\n## Changes\n\n- Added `RunnerTerminationError` for Python exceptions raised inside\nrunner bootstrap\n- Changed runner bootstrap to send `Event | RunnerTerminationError` over\nthe private runner channel\n- Moved public `RunnerFailed` emission back into supervisor\n- Added stderr-only `RunnerDiagnosticCollector`\n- + Added known diagnostics for Metal GPU timeout, ring socket receive\nerrno, and ring transport abort\n- Added diagnostics to `RunnerFailed` and `ErrorChunk`\n- Tweaked async process termination to join briefly before\nterminate/kill\n- Updated tests/fixtures for new failure payload shape\n- Added Ruff VS Code formatter settings\n\n## Why It Works\n\nRunner child now reports raw-ish failure context to supervisor instead\nof publishing failed status directly.\n\nSupervisor still owns process lifecycle, exit code/signal handling,\nin-flight task error chunks, and final runner status. Stderr diagnostics\nstay best effort and only known root-cause variants are surfaced.\n\n## Test Plan\n\n### Manual Testing\n\nHardware: remote runner logs from e16/e11/e4/e2\n\nWhat you did:\n- inspected live runner stderr logs\n- used observed Metal GPU timeout and ring socket errors as initial\ndiagnostic targets\n\n### Automated Testing\n\n- `nix flake check`\n- supervisor test covers error chunk + failed status emission\n- plan lifecycle test updated for failed runner diagnostics\n- type/lint checks cover new runner channel union\n\n---------\n\nCo-authored-by: Evan Quiney <evanev7@gmail.com>"},{"hash":"88d46d46","date":"2026-05-14 17:32:54 +0100","author":"Heidar","subject":"fix: omit null delta fields in streaming chat completions (issue #2082) (#2092)","body":"## Motivation\n\nStreaming /v1/chat/completions responses emitted null for tool_calls,\nfunction_call, name, and tool_call_id in every delta chunk. The OpenAI\nstreaming spec marks these fields as non-nullable — they must either\ncarry a\n  real value or be absent entirely. Spec-correct clients doing\ndelta.get(\"tool_calls\", []) receive None and crash with 'NoneType'\nobject is\n  not iterable.\n\nRoot cause: the streaming serialisation path called model_dump_json()\nwithout\nexclude_none=True, while the request-parsing path already used it\ncorrectly.\nThree call sites in chat_completions.py and two in responses.py were\naffected.\n\n## Testing\n\nBefore — every delta carries explicit nulls:\n\n  $ curl -sN -X POST http://localhost:52415/v1/chat/completions \\\n    -H 'Content-Type: application/json' \\\n-d\n'{\"model\":\"mlx-community/Qwen3.5-2B-MLX-8bit\",\"messages\":[{\"role\":\"user\",\"\n  content\":\"hi\"}],\"max_tokens\":3,\"stream\":true}' \\\n    | grep \"^data: \"\ndata:\n{\"id\":\"7c4dae10-...\",\"choices\":[{\"index\":0,\"delta\":{\"role\":\"assistant\",\"c\n\nontent\":null,\"reasoning_content\":\"Okay\",\"name\":null,\"tool_calls\":null,\"tool_cal\n\nl_id\":null,\"function_call\":null},\"logprobs\":null,\"finish_reason\":null,\"usage\":n\n  ull}],\"usage\":null,\"service_tier\":null}\ndata:\n{\"id\":\"7c4dae10-...\",\"choices\":[{\"index\":0,\"delta\":{\"role\":\"assistant\",\"c\n\nontent\":null,\"reasoning_content\":\",\",\"name\":null,\"tool_calls\":null,\"tool_call_i\n\nd\":null,\"function_call\":null},\"logprobs\":null,\"finish_reason\":null,\"usage\":null\n  }],\"usage\":null,\"service_tier\":null}\ndata:\n{\"id\":\"7c4dae10-...\",\"choices\":[{\"index\":0,\"delta\":{\"role\":\"assistant\",\"c\nontent\":\"\nthe\",\"reasoning_content\":null,\"name\":null,\"tool_calls\":null,\"tool_cal\n\nl_id\":null,\"function_call\":null},\"logprobs\":null,\"finish_reason\":\"length\",\"usag\n  e\":{\"prompt_tokens\":11,...}}],\"usage\":null,\"service_tier\":null}\n  data: [DONE]\n\n  After — only populated fields are emitted:\ndata:\n{\"id\":\"demo\",\"object\":\"chat.completion\",\"created\":...,\"model\":\"mlx-commun\n\nity/Qwen3.5-2B-MLX-8bit\",\"choices\":[{\"index\":0,\"delta\":{\"role\":\"assistant\",\"rea\n  soning_content\":\"Okay\"}}]}\ndata:\n{\"id\":\"demo\",\"object\":\"chat.completion\",\"created\":...,\"model\":\"mlx-commun\n\nity/Qwen3.5-2B-MLX-8bit\",\"choices\":[{\"index\":0,\"delta\":{\"role\":\"assistant\",\"rea\n  soning_content\":\",\"}}]}\ndata:\n{\"id\":\"demo\",\"object\":\"chat.completion\",\"created\":...,\"model\":\"mlx-commun\n\nity/Qwen3.5-2B-MLX-8bit\",\"choices\":[{\"index\":0,\"delta\":{\"role\":\"assistant\",\"con\ntent\":\"\nthe\"},\"finish_reason\":\"length\"}],\"usage\":{\"prompt_tokens\":11,\"completio\n  n_tokens\":3,\"total_tokens\":14,...}}\n  data: [DONE]"},{"hash":"e8ec8d50","date":"2026-05-14 17:12:58 +0100","author":"Heidar","subject":"fix ollama API compatibility for VS Code Copilot (#2091)","body":"Ollama adapter fixes for VS Code Copilot (#2042):\n\n  - /api/version: bare semver \"1.0.0\" - Copilot parseInts each segment.\n- /api/show: populate model_info + capabilities - Copilot crashes on\nnull model_info and filters by `tools`.\n- Add POST /ollama/v1/chat/completions - ollama serves the OpenAI-compat\nroute here, BYOK clients 405 without it.\n\n\nBefore:\n<img width=\"1380\" height=\"144\" alt=\"image\"\nsrc=\"https://github.com/user-attachments/assets/99d5464f-187d-4432-9a31-8229c55aa209\"\n/>\n\nAfter:\n<img width=\"1362\" height=\"181\" alt=\"image\"\nsrc=\"https://github.com/user-attachments/assets/361dc006-d8df-435f-8d8b-4fa4f44a8c23\"\n/>\n<img width=\"279\" height=\"909\" alt=\"image\"\nsrc=\"https://github.com/user-attachments/assets/4621aba7-bd57-4762-8568-34a3383a6025\"\n/>\n\nCo-authored-by: Claude Opus 4.7 <noreply@anthropic.com>"},{"hash":"1fd15d59","date":"2026-05-14 17:03:51 +0100","author":"Heidar","subject":"create directory on startup (#2089)","body":"## Motivation\n\n<!-- Why is this change needed? What problem does it solve? -->\n<!-- If it fixes an open issue, please link to the issue here -->\n\nWhen you first run `uv run exo` you get an error like :\n\n`FileNotFoundError: [Errno 2] No such file or directory:\n'/Users/heidar/.exo/models'`\n\nManually tested on Macbook Pro M1 32GB\n\nFixes issue - https://github.com/exo-explore/exo/issues/2090"},{"hash":"4466cd53","date":"2026-05-13 10:45:11 +0100","author":"Evan Quiney","subject":"use custom mlx sources for linux (#2087)","body":"switch to hosting mlx sources on github & cachix instead of using a\nbroken version of mlx. closes #2043."},{"hash":"ed2d10bd","date":"2026-05-12 11:48:08 +0100","author":"Andrei Cravtov","subject":"Redirect runner stdout/stderr to file logs  (#2084)","body":"## Motivation\n\nWe want to use log mining tools like\n[Drain3](https://github.com/logpai/Drain3) to get standardized error\nformats, but for that we should record runner stdout/stderr in a massive\nappend-only log to gather training data for such tools. Also useful for\nfuture opt-in telemetry.\n\n## Changes\n\nThe stdout/stderr from runner now splits into 3 tasks: \n1) raw write to dedicated runner logs \n2) sanitized line-by-line logging with log-guru \n3) stub for further error-processing (i.e. turning lines into errors)\n\n### Manual Testing\nWorks on 4x mac mini clusted connected as TB4 ring."},{"hash":"87c72fc1","date":"2026-05-11 13:15:22 +0100","author":"Andrei Cravtov","subject":"Fixes issue #2068 (#2083)","body":"## Motivation\n\nTo fix https://github.com/exo-explore/exo/issues/2068\n\n## Changes\n\nAdds queue shutdown logic & hard-timeouts for closing server.\n\n## Why It Works\n\nPrevents API from hanging more than 5 seconds."},{"hash":"b76bc301","date":"2026-05-10 18:11:46 +0100","author":"Evan Quiney","subject":"bump rust versions (#2081)","body":""},{"hash":"08ffa5f6","date":"2026-05-10 10:02:22 -0700","author":"team-wcv","subject":"Map GLM 4.7 stop tokens to GLM 4 IDs (#2061)","body":"## Motivation\n\nGLM 4.7 reuses the GLM 4 chat-template tokenizer, but the model card and\nEOS-detection path didn't have an explicit mapping for it, so\nOpenAI-compatible clients didn't see a clean stop and the runner emitted\nfollow-on role turns (e.g. \\`<|user|>\\` continuations after\n\\`<|assistant|>\\`'s output).\n\n## Changes\n\n\\`src/exo/worker/engines/mlx/utils_mlx.py\\` — add the GLM 4 stop-token\nIDs as the EOS set when the loaded model's tokenizer matches GLM 4 / 4.7\nchat templates.\n\n## Why It Works\n\nThe GLM 4 tokenizer's \\`<|user|>\\`, \\`<|observation|>\\`, and\n\\`<|endoftext|>\\` IDs are stable across the GLM 4 / 4.7 line; treating\nany of them as EOS lets the runner stop at the assistant turn boundary\nthe same way it stops at \\`</s>\\` for Llama-style models. No\nprompt-template changes — only the stop set widens.\n\n## Test Plan\n\n### Automated Testing\n\nNew unit test\n\\`src/exo/worker/tests/unittests/test_mlx/test_eos_token_ids.py\\`\ncovering: GLM 4 / 4.7 path returns the expected stop ID set; non-GLM\npath returns the standard EOS only.\n\n\\`\\`\\`\nsrc/exo/worker/tests/unittests/test_mlx/test_eos_token_ids.py ..\n=== 2 passed in 0.01s ===\n\\`\\`\\`\n\n\\`uv run basedpyright\\` and \\`uv run ruff check\\` both clean.\n\n### Manual Testing\n\nHardware: 4-node Apple Silicon cluster, M5 Max master.\n\n- Loaded \\`mlx-community/GLM-4.7-Air-mlx-4bit\\`, ran chat completion via\n\\`/v1/chat/completions\\`. Before this fix the assistant turn ran on into\na synthetic \\`<|user|>\\` continuation; after the fix the response stops\ncleanly at the assistant boundary.\n\n---------\n\nCo-authored-by: jw-wcv <101585096+jw-wcv@users.noreply.github.com>\nCo-authored-by: Evan Quiney <evanev7@gmail.com>"},{"hash":"45df74ba","date":"2026-05-09 22:45:14 +0100","author":"Andrei Cravtov","subject":"Andrei/mp capture stdio (#2056)","body":"## Motivation\n\nProcess-isolated runner crashes and C-extension failures can write\ndirectly to fd-level stdout/stderr, bypassing Python/loguru. We need to\ncapture that output per runner process without polluting the main\nprocess or other workers, and without breaking operation when the parent\nstdio is detached.\n\n## Changes\n\n- Added `AsyncProcess`, a spawn-only multiprocessing wrapper that\nredirects child stdout/stderr to pipes and exposes them as in-memory\n`Receiver[bytes]`s\n- Replaced runner-supervisor's raw `multiprocessing.Process` usage with\n`AsyncProcess`\n- Added `--no-stdio`, redirecting stdin/stdout/stderr to `/dev/null`\nafter logging is configured\n- Disabled verbose MLX\n- Added tests covering stdio capture, child crashes, repeated bad\nchildren, SIGTERM/SIGKILL shutdown escalation, stdio detachment, and\nspawning captured children from a stdio-detached parent\n\n## Why It Works\n\nThe parent can redirect its own stdio fds to `/dev/null`, while\n`AsyncProcess` installs fresh pipe fds over fd 1 and 2 inside each\nspawned child. That keeps stdio-detached parents quiet while preserving\nper-runner stdout/stderr capture. Runner shutdown is still bounded:\nSIGTERM grace first, then SIGKILL escalation if needed.\n\nNext direction: the runner supervisor currently drains captured output\nand logs it as stdout/debug and stderr/warning. This should be split\ninto more useful process-isolated error reporting instead of just log\nforwarding (regex match on errors to obtain \"reason\" string, best\neffort).\n\n## Test Plan\n\n### Manual Testing\n\nRan on 4 Mac Minis in a Thunderbolt 4 ring, can see that runner's\nstdout/stderr contents are being captured.\n\n### Automated Testing\n\n- Added async-process tests for fd-level stdout/stderr capture, Python\ntraceback capture, bounded-buffer output, child `exit`/abort, parent\nstdio preservation, fd leak checks, spawn-context mp channels, and\nSIGTERM/SIGKILL shutdown behavior\n- Added stdio-detach tests proving stdio detaches to `/dev/null`, a\nstdio-detached parent can still spawn and capture a child, and the same\nstdio-detached parent can spawn/capture multiple children sequentially\n- Updated runner-supervisor tests for the new `AsyncProcess.exitcode`\npath"},{"hash":"ce37bdce","date":"2026-05-09 15:10:22 +0300","author":"Kerollos Magdy","subject":"fix: Create directory for PID file if it doesn't exist (#2075)","body":"Ensure the directory for the PID file exists before creating it.\n\n## Motivation\n\nFixes https://github.com/exo-explore/exo/issues/2074\n\n## Changes\n\n<!-- Describe what you changed in detail -->\n\n## Why It Works\n\n<!-- Explain why your approach solves the problem -->\n\n## Test Plan\n\n### Manual Testing\n<!-- Hardware: (e.g., MacBook Pro M1 Max 32GB, Mac Mini M2 16GB,\nconnected via Thunderbolt 4) -->\n<!-- What you did: -->\n<!-- - -->\n\n### Automated Testing\n<!-- Describe changes to automated tests, or how existing tests cover\nthis change -->\n<!-- - -->"},{"hash":"e5a1e5da","date":"2026-05-08 18:50:18 +0100","author":"Andrei Cravtov","subject":"Create PID file locking for EXO (#2072)","body":"## Motivation\n\nEXO should be PID file locked, to prevent duplicate processes from\nclobbering the log, right now this isn't the case.\n\n## Changes\n\nI added a wrapper around a Rust PID file lock library, and used it to\nimplement PID locking for EXO, with the PID file being in exo cache\ndirectory.\n\n## Test Plan\n\n### Manual Testing\nTested on e11, trying to spawn duplicate EXO processes prevented."},{"hash":"fa571313","date":"2026-05-08 17:15:08 +0100","author":"ciaranbor","subject":"Integration tests infra (#1995)","body":"## Motivation\n\nNo automated integration tests exist for exo. Manual testing against\nreal hardware clusters is slow and error-prone. We need a pytest\nframework that deploys clusters via `eco`, runs inference scenarios, and\ntears down cleanly.\n\n## Changes\n\n- **`tools/src/exo_tools/`** — New workspace member shared by bench,\neval, and tests:\n- `client.py` — `ExoClient` HTTP client (extracted from\n`bench/harness.py`)\n- `harness.py` — instance lifecycle helpers (placement, wait-for-ready,\netc.)\n- `cluster.py` — `EcoSession` for eco cluster lifecycle\n(deploy/stop/start/release/logs/exec) with unique `USER=<prefix>-<uuid>`\nper session and atexit/signal cleanup\n- **`tests/integration/`** — 17 pytest tests across 5 files:\n- `test_1node.py` — place, chat, multi-turn, delete, state/models\nendpoints, cluster snapshot, download-from-scratch\n- `test_2node.py` — parametrized tensor/jaccl + pipeline/ring inference\nand multi-turn\n- `test_4node.py` — parametrized 4-node pipeline/ring inference, cluster\nstate\n- `test_resilience.py` — full disconnect/reconnect cycle (2-node →\ndisconnect → 1-node → reconnect → 2-node)\n- `test_dashboard.py` — Playwright: dashboard loads, shows node info,\nchat flow\n- `helpers.py` — placement/inference helpers, re-exports from\n`exo_tools`\n- `conftest.py` — session-scoped cluster fixtures with constraint-based\neco reservations; `--hosts` override; `EXO_REF` env var for CI\ndeployments from a GitHub branch\n- **`bench/`** — Updated imports from `exo_tools.client` /\n`exo_tools.harness`\n- **`pyproject.toml`** — Added `tools` workspace member, `playwright`\ndev dep, `--ignore=tests/integration`\n\n## Why It Works\n\nTests use `eco` for cluster lifecycle and `ExoClient` for API\ninteractions — same tools humans use. Session-scoped fixtures deploy\nonce per file. Unique eco users prevent test runs from interfering with\neach other or manual usage.\n\n## Test Plan\n\n### Automated Testing\n\n- `uv run pytest tests/integration/ -v -s` — full suite (~4-5 min, 17/17\npassing)\n- `uv run pytest tests/integration/ -v -s --hosts s4,s9,s10,s22` — pin\nspecific hosts\n- `EXO_REF=main uv run pytest tests/integration/ -v` — deploy from a\nGitHub branch (CI)\n- `uv run pytest` — confirms integration tests are excluded from default\nruns"},{"hash":"414132ae","date":"2026-05-07 03:42:14 -0700","author":"Alex Cheema","subject":"Use time-weighted power sampling (#2038)","body":"## Why\n\nThe power sampler currently averages sampled wattage values\narithmetically. That can be materially wrong when sample intervals are\nuneven: a short high-power spike gets the same weight as a long steady\ninterval. Energy should be computed by integrating power over time, and\naverage power should be derived from energy / elapsed time.\n\n## How\n\n- Store each power sample with its relative timestamp.\n- Anchor the first sample at `t=0` and take a final sample at `elapsed`\nwhen producing results.\n- Integrate per-node power using the trapezoidal rule.\n- Sum node energy for total cluster energy, then derive total average\nsystem power from total energy / elapsed.\n- Add focused unit tests for uneven sample intervals and the\nsingle-sample fallback.\n\n## Tests\n\n- `uv run pytest src/exo/utils/tests/test_power_sampler.py`\n- `uv run basedpyright`\n- `uv run ruff check src/exo/utils/power_sampler.py\nsrc/exo/utils/tests/test_power_sampler.py`\n- `nix fmt`"},{"hash":"edef8004","date":"2026-05-07 01:06:39 -0700","author":"Alex Cheema","subject":"Store custom model cards in State (#2024)","body":"## Why\n\nWorkers currently update their custom model-card cache by reacting to\n`CustomModelCardAdded` / `CustomModelCardDeleted` events directly. That\nis another snapshot footgun: a worker restored from State may never see\nthe historical add/delete event, so the durable State must include the\ndesired custom-card set.\n\n## How\n\n- Add `State.custom_model_cards`, keyed by `ModelId`.\n- Reduce `CustomModelCardAdded` into State.\n- Reduce `CustomModelCardDeleted` into State.\n- Add focused reducer tests for add and delete.\n\nThis PR only makes custom cards durable in State. A follow-up PR will\nmake workers reconcile their on-disk custom-card cache from this state\ninstead of relying on those events directly.\n\n## Tests\n\n- `uv run pytest\nsrc/exo/shared/tests/test_apply/test_apply_custom_model_cards.py\nsrc/exo/shared/tests/test_state_serialization.py`\n- `uv run pytest`\n- `uv run ruff check src/exo/shared/types/state.py\nsrc/exo/shared/apply.py\nsrc/exo/shared/tests/test_apply/test_apply_custom_model_cards.py`\n- `uv run basedpyright`\n- `nix fmt`"},{"hash":"a0c00f9d","date":"2026-05-07 00:00:15 -0700","author":"Alex Cheema","subject":"fix(placement): gate RDMA on nodeRdmaCtl.enabled at both endpoints (#2014)","body":"## Summary\n\n- Fixes a bug where `POST /place_instance` (and the dashboard UI) would\naccept an MlxJaccl/RDMA instance spanning nodes whose\n`nodeRdmaCtl.enabled` was `false`, because topology + placement\nconsulted Thunderbolt-derived RDMA edges without checking the per-node\n`rdma_ctl` status.\n- Three-layer fix: topology only emits `RDMAConnection` edges when both\nendpoints have `nodeRdmaCtl.enabled = true`; flipping a node to disabled\nimmediately purges every RDMA edge touching it; `place_instance`\nadditionally rejects RDMA cycles containing any disabled or unobserved\nnode as a defense-in-depth check on the API/master path.\n\n## Details\n\n- `src/exo/shared/apply.py`\n- `MacThunderboltConnections` case now filters out RDMA connections\nwhose source or sink lacks observed-and-enabled `rdma_ctl` status\n(missing entry → treated as disabled).\n- `RdmaCtlStatus` case now calls\n`topology.remove_all_rdma_connections_touching(node_id)` when the node\nreports disabled, so consumers don't have to wait for the next TB poll.\n- `src/exo/shared/topology.py`\n- New `Topology.remove_all_rdma_connections_touching(node_id)` removes\nevery RDMA edge incident to the node (incoming and outgoing) while\nleaving socket edges intact.\n- `src/exo/master/placement.py`\n- `place_instance` accepts `node_rdma_ctl: Mapping[NodeId,\nNodeRdmaCtlStatus] | None`. The `is_rdma_cycle` filter now also requires\n`nodeRdmaCtl.enabled` for every node in the cycle. MlxJaccl placement\nraises the existing \"no RDMA-connected cycles available\" error if no\nqualifying cycle remains.\n- `src/exo/api/main.py`, `src/exo/master/main.py`\n  - Both placement entrypoints now pass `state.node_rdma_ctl` through.\n\n## Tests\n\n- `src/exo/shared/tests/test_apply/test_apply_rdma_gating.py` (new): six\nunit tests covering enabled/disabled/missing combinations on apply, the\nimmediate-purge transition, and that purging RDMA edges leaves socket\nedges untouched.\n- `src/exo/master/tests/test_placement.py`: existing\n`test_tensor_rdma_backend_connectivity_matrix` updated to pass\n`node_rdma_ctl`. Two new tests assert MlxJaccl placement is rejected\nwhen any cycle node is `enabled=false` or has no `rdma_ctl` entry.\n\n## Test plan\n\n- [x] `uv run basedpyright` — 0 errors\n- [x] `uv run ruff check` — clean\n- [x] `nix fmt`\n- [x] `uv run pytest` — 429 passed, 1 skipped\n- [ ] On a real mixed cluster (s15/s16 disabled, s17/s18 enabled),\nconfirm:\n- [ ] `POST /place_instance` for an RDMA instance including s15 or s16\nreturns an error\n  - [ ] An RDMA instance can still be placed across {s17, s18}\n- [ ] `GET /state` shows no `sourceRdmaIface`/`sinkRdmaIface` on s15↔s16\nconnections\n- [ ] Dashboard previews don't surface RDMA-spanning options that\ninclude s15/s16\n\n🤖 Generated with [Claude Code](https://claude.com/claude-code)\n\nCo-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>"},{"hash":"89d20c18","date":"2026-05-06 07:24:58 -0500","author":"Drifter4242","subject":"fix(inference): prevent TP collective deadlock via agree_on_tasks order (#2048)","body":"If you have two machines and make two requests at the same time, it can\ncrash. This is because the tasks can sometimes end up in different\norders on different machines. We need to sort the tasks and\nmx_all_gather_tasks already sorts the tasks but the code ignores that\nordering. The fix is to make sure the sort order is preserved.\n\nThe rest is written by Sonnet (reviewed by me):\n\nTensor-parallel inference requires that every rank enqueues tasks in the\nsame order before running agree_on_tasks collectives. The old\nimplementation filtered from _maybe_queue:\n\nself._queue.extend(task for task in self._maybe_queue if task in agreed)\nself._maybe_queue = [task for task in self._maybe_queue if task in\ndifferent]\n\nBecause _maybe_queue is independently ordered per-rank (tasks arrive via\ngRPC in whatever order the API server sends them), two concurrent\nrequests could produce different _maybe_queue orderings on rank 0 vs\nrank 1. The filter then preserved those different orders into _queue, so\neach rank started processing tasks in a different sequence. The next mlx\ncollective (all_reduce, all_gather, etc.) on rank 0 corresponded to a\ndifferent task than on rank 1 → permanent deadlock.\n\nFix: extend from agreed directly. mx_all_gather_tasks returns agreed as\na list sorted by task_id on all ranks, so every rank appends the same\nsequence regardless of local arrival order.\n\nApplies to both SequentialGenerator and BatchGenerator.\n\n## Motivation\n\n`agree_on_tasks` is called on every rank after accumulating new requests\nin\n`_maybe_queue`. Its job is to run an `all_gather` collective so all\nranks agree\non which tasks to promote to `_queue` before the next inference step.\n\nThe old implementation re-imposed **local arrival order** when extending\n`_queue`:\n\n```python\nself._queue.extend(task for task in self._maybe_queue if task in agreed)\n```\n\n`mx_all_gather_tasks` already returns `agreed` sorted by `task_id` — the\nsame\ndeterministic order on every rank. But iterating `self._maybe_queue`\ninstead of\n`agreed` discarded that sort and substituted the local gRPC arrival\norder, which\ndiffers per rank under concurrent load. Two concurrent requests arriving\nin\n`[A, B]` order on rank 0 and `[B, A]` on rank 1 caused the first MLX\ncollective\nin the next step to hang permanently: each rank was executing a\ndifferent task's\ncollective and would never match.\n\n## Changes\n\n`SequentialGenerator.agree_on_tasks` and\n`BatchGenerator.agree_on_tasks`:\n\n```python\n# Before\nself._queue.extend(task for task in self._maybe_queue if task in agreed)\nself._maybe_queue = [task for task in self._maybe_queue if task in different]\n\n# After\nself._queue.extend(agreed)          # preserves mx_all_gather_tasks sort order\nself._maybe_queue = list(different) # already in local order; filter was redundant\n```\n\n## Why It Works\n\n`mx_all_gather_tasks` (in `utils_mlx.py`) computes the agreed set then\nsorts by\n`task_id`:\n\n```python\nagreed = [local_tasks[tid] for tid in sorted(agreed_ids)]\n```\n\nBecause `task_id` is a UUID and the sort is lexicographic, every rank\nproduces\nthe same `agreed` list regardless of local arrival order. Using `agreed`\ndirectly\npreserves this guarantee. The `different` list (tasks not yet seen on\nall ranks)\nis built by iterating `tasks` in local order, which is already correct.\n\n## Test Plan\n\n### Manual Testing\n\n**Hardware:** 2× Mac Studio M3 Ultra 512 GB, Thunderbolt 5 direct\nbridge,\n`MlxJaccl` RDMA tensor-parallel (`moonshotai/Kimi-K2.6`, 595 GB INT4, 61\nlayers).\n\n- Sent concurrent streaming requests; confirmed all complete without\ndeadlock.\n- This hardware configuration (sub-millisecond inter-node latency) is\nthe most\nlikely to trigger the race, as requests from separate HTTP connections\ncan\nreach rank 0 and rank 1 in opposite order before `agree_on_tasks` runs.\n\n### Automated Testing\n\nAll existing tests pass: `pytest src -m \"not slow\"\n--import-mode=importlib`\n— 422/422 passed. The existing `test_event_ordering.py` covers the\n`agree_on_tasks` call path with a mock that returns tasks in consistent\norder;\nthe race requires real distributed hardware to reproduce\ndeterministically."},{"hash":"dbcceaa5","date":"2026-05-05 17:27:57 +0100","author":"Evan Quiney","subject":"Initialise _cancelled_tasks in ImageEngine (#2051)","body":"we yielded nonsense chunks from engines; we didn't initialize the image\nengine correctly. mostly rewrite of #2049\n\n---------\n\nCo-authored-by: ciaranbor <ciaranborourke-dev@proton.me>"},{"hash":"9c6ff4ce","date":"2026-05-01 04:18:57 -0700","author":"Sam Bradbury","subject":"feat: update rdma_ctl instructions (#1977)","body":"## Motivation\n\nThe RDMA setup instructions were missing a step: after booting to\nRecovery mode, users need to open Terminal from the Utilities menu\nbefore they can run the `rdma_ctl` command. Without this step, users\nfollowing the instructions wouldn't know how to access a terminal in\nRecovery mode. This step was already in the README just not in the UI\nnotifications.\n\n## Changes\n\nAdded a missing instruction step — \"Open Terminal from the Utilities\nmenu\" — to three instances of the RDMA setup flow in\n`dashboard/src/routes/+page.svelte`.\n\n## Why It Works\n\nN/A copy change only. \n\n## Test Plan\n\n### Manual Testing\nHardware: MacBook Pro M4 Max 48GB\n\n### Automated Testing\nNo automated tests affected; this is a UI copy change only.\n\nCo-authored-by: Sam Bradbury <sam@consultbradbury.com>"},{"hash":"b26268df","date":"2026-05-01 06:41:10 -0400","author":"ecohash-co","subject":"fix(macos-app): disable URL response caching for cluster-state polling (#2005)","body":"Fixes #2004.\n\n`ClusterStateService` polls `/state` at 2 Hz via `URLSession.shared`,\nwhich keeps an on-disk `URLCache` attached by default. Every polled\nresponse body gets persisted under `~/Library/Caches/exolabs.EXO/`,\nsustaining ~500–620 KB/sec of file-backed memory dirtied — far above\nmacOS's ~25 KB/sec per-process daily-average baseline. Six\nmicrostackshot reports observed on a single Mac Studio M3 Ultra over\neight days, with one 15-hour run accumulating 34.36 GB of cache writes.\n\nHeaviest stack on every diagnostic report (96–98% of samples):\n\n```\n_dispatch_workloop_worker_thread → _dispatch_block_async_invoke2 →\n  __CFURLCache::CreateAndStoreCacheNode → write\n```\n\nFull diagnostic data and analysis in #2004.\n\n## What changed\n\n`ClusterStateService` now defaults to an ephemeral, non-caching\n`URLSession` instead of `URLSession.shared`. Cluster-state responses are\ntime-sensitive and small; nothing benefits from being cached on disk.\n\n```swift\nprivate static func makeNonCachingSession() -> URLSession {\n    let config = URLSessionConfiguration.ephemeral\n    config.urlCache = nil\n    config.requestCachePolicy = .reloadIgnoringLocalCacheData\n    return URLSession(configuration: config)\n}\n```\n\nThe existing per-request `request.cachePolicy =\n.reloadIgnoringLocalCacheData` calls are kept as defense in depth — they\nonly affect read behavior, but harmless to leave alongside the\nsession-level config.\n\n## Scope\n\n- **Behavioral**: none. Polled requests still go out at the same\ncadence; responses still parse the same; no semantic change to any API\nsurface.\n- **Test injection**: the `session:` parameter remains in `init`, so\ntests can still inject a custom mock session unchanged.\n- **`BugReportService` and other `URLSession.shared` callers**:\nuntouched. If maintainers prefer an app-wide URLCache disable instead,\nhappy to switch the approach (issue body has the alternative spelled\nout).\n\n## Verification\n\nVerified locally that compiling EXO with this change produces a working\nmenubar app and `ClusterStateService` continues to fetch state\ncorrectly. After ~30 min of idle polling, no new entries in\n`/Library/Logs/DiagnosticReports/EXO_*.diag` and no growth in\n`~/Library/Caches/exolabs.EXO/`.\n\n## Test plan\n- [ ] Build EXO from this branch on macOS 26.4\n- [ ] Launch, let cluster state polling run for 30+ min\n- [ ] Confirm no new microstackshot diagnostic reports\n- [ ] Confirm `~/Library/Caches/exolabs.EXO/Cache.db*` does not grow\n\n🤖 Generated with [Claude Code](https://claude.com/claude-code)\n\nCo-authored-by: Jordan Miller <jordan.d.miller@gmail.com>"},{"hash":"8dae3ecb","date":"2026-04-30 19:06:15 +0100","author":"ciaranbor","subject":"A few targeted tweaks to address HF rate limits (#2009)","body":"## Motivation\n\n- exo bursts ~200 HF Hub-API requests on every cold start, blowing past\nthe anonymous 500-req/5-min budget.\n- The existing retry loop catches 429 generically and gives up in ~3s —\nwell before HF's reset window.\n- `file_meta` and `_download_file` had no 429 handling at all (became\n`AssertionError`).\n- Disk file-list cache was bypassed on every process restart.\n\n## Changes\n\nAll in `src/exo/download/download_utils.py` + tests.\n\n- Parse `t=` from HF's `RateLimit` header on 429; sleep `min(t, 300s) +\njitter`.\n- Handle 429 at all three call sites (`_fetch_file_list`, `file_meta`,\n`_download_file`).\n- `n_attempts`: 3 → 5.\n- Disk cache now primary across restarts (24h mtime TTL).\n- `?recursive=true` instead of N+1 subdir walks.\n\n## Why It Works\n\n`t=<seconds>` is HF's \"wait this long and you'll be unblocked\" —\nsleeping that long lets the window reset. Disk-cache-as-primary plus\nrecursive listing cuts cold-start Hub-API traffic by ~10×.\n\n## Test Plan\n\n### Manual Testing\n\nMacBook Pro M1 Max. Tripped the real HF 429. Pre-fix: failed in 3.4s.\nPost-fix: slept (HF returned `t=158`) and recovered.\n\n### Automated Testing\n\n- New `test_rate_limit_handling.py` (19 tests) — header parsing,\nretry-loop behaviour, plus HTTP-level coverage that mocks aiohttp to\nreturn a 429 and asserts each call site raises\n`HuggingFaceRateLimitError(retry_after=52.0)`.\n- New `TestFileListCacheTTL` in `test_offline_mode.py` — fresh cache\nhits, stale cache refetches.\n- 421 tests pass; basedpyright / ruff / nix fmt clean."},{"hash":"fb12b403","date":"2026-04-30 15:10:26 +0100","author":"Alex Cheema","subject":"fix(app): tighten Share Bug Report prompt layout (#2008)","body":"## Summary\n\nFollow-ups to #2003 based on feedback that the Share Bug Report window\nfelt visually weighty: too much padding above and below, and a\ndescription editor that invited an essay rather than a one-liner.\n\n## Changes (one file)\n\n`app/EXO/EXO/Views/BugReportWindowController.swift`:\n\n- **Auto-size the window to its content.** Switched from `NSHostingView`\n+ fixed `contentRect: 480x380` + SwiftUI `frame(minHeight: 320)` to\n`NSHostingController` with `sizingOptions = [.preferredContentSize,\n.minSize]`. The fixed-min combo was centering the form in dead vertical\nspace.\n- **Smaller, lower-pressure editor.** Field is now labeled `Description\n(optional)` with a placeholder hint (`What were you doing when it\nbroke?`) inside the editor. Editor height fixed at 72pt (was 120pt min).\nReplaced the long lead-in paragraph and headline with a single one-line\ncaption between field and buttons: `Diagnostic logs will be uploaded\nwith your report.`\n- **Tighter spacing.** Outer padding 20 -> 16, root spacing 16 -> 12,\nprompting-section spacing 12 -> 8.\n- **Remove em dash from copy.**\n\n`BugReportService` and the menu wiring are unchanged.\n\n## Test plan\n\n- [ ] Click `Share Bug Report...` from the menu bar.\n- [ ] The window opens centered and sized to its content (no big empty\nbands top/bottom).\n- [ ] Description editor is visibly compact, with the placeholder hint\nshowing when empty.\n- [ ] The optional-ness is conveyed by the field label (no separate help\nparagraph).\n- [ ] Caption `Diagnostic logs will be uploaded with your report.`\nappears in `.caption` style under the editor, above the buttons.\n- [ ] Resize the window: persists across re-opens (frame autosave still\nworks).\n- [ ] Send/Cancel/Try Again/Done flows behave the same as before.\n\n🤖 Generated with [Claude Code](https://claude.com/claude-code)\n\n---------\n\nCo-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>"},{"hash":"1606e638","date":"2026-04-30 13:05:07 +0100","author":"Alex Cheema","subject":"feat(app): open Share Bug Report in a dedicated window (#2003)","body":"## Summary\n\n- Adds a top-level **Share Bug Report…** menu item to the macOS popover\n(between *Check for Updates* and *Quit*) with SF Symbol `ladybug`.\n- Clicking it opens a dedicated resizable `NSWindow` (\"Send a Bug\nReport\") that hosts the prompting / sending / success / failure flow.\n- Removes the description-less duplicate from Settings → Debug Info, and\nthe dead `debugSection` it nominally lived behind.\n\n## Why\n\nPR #1959 added a user-description prompt to the bug-report flow, but its\ntrigger lived inside `ContentView.debugSection` — a view that's defined\nbut never rendered in the body. The path users actually hit was\n`SettingsView.sendBugReportButton`, which called\n`BugReportService.sendReport(isManual: true)` without ever passing\n`userDescription`. So the description prompt was unreachable in the\nbuilt app.\n\n## Approach\n\nPer Apple HIG, an action that requires further input before completing\nshould open a dialog, not transform the menu inline. So:\n\n- Add a top-level menu entry that ends in `…` (HIG: ellipsis indicates\n\"further input required\").\n- Move the prompting/sending/success/failure state machine into a\nstandalone `BugReportWindowController` modeled after the existing\n`SettingsWindowController`.\n- Single-instance window with frame-autosave name, sensible\n`contentMinSize`, resizable, native button layout (`.cancelAction` /\n`.defaultAction` keyboard shortcuts), light/dark-mode-correct\n`.textBackgroundColor` and `.separatorColor`.\n- Auto-focus the description field on open. `Try Again` from failure,\n`Open GitHub Issue` + `Done` from success.\n\n## Files\n\n- `app/EXO/EXO/Views/BugReportWindowController.swift` (new) — controller\n+ view.\n- `app/EXO/EXO/EXOApp.swift` — wire `BugReportWindowController` as a\n`@StateObject` and inject as environment object.\n- `app/EXO/EXO/ContentView.swift` — replace inline state machine with\nmenu item that calls `bugReportWindowController.open()`. Remove\nnow-unused state, helpers, and dead `debugSection`.\n- `app/EXO/EXO/Views/SettingsView.swift` — remove duplicate\n`sendBugReportButton`, `sendBugReport()`, and related `@State`. Section\n\"Debug Info\" keeps Thunderbolt / interface / RDMA info.\n\n`BugReportService` is unchanged.\n\n## Test plan\n\n- [ ] Open the menu-bar popover → confirm **Share Bug Report…** appears\nbetween *Check for Updates* and *Quit*, with a ladybug icon.\n- [ ] Click it → a window titled \"Send a Bug Report\" appears, centered,\nwith the description editor focused.\n- [ ] Resize the window → size persists across re-opens (frame\nautosave).\n- [ ] Type a description, press Return → upload succeeds, success card\nwith **Open GitHub Issue** + **Done** appears.\n- [ ] Click **Open GitHub Issue** → browser opens with the description\npre-filled into the issue template.\n- [ ] Send with empty description → upload still succeeds.\n- [ ] Press Esc from the prompting state → window closes.\n- [ ] On failure (e.g., offline) → error card with **Try Again** +\n**Close** appears; Try Again returns to the editor with the description\npreserved.\n- [ ] Open the Settings window → Debug Info section is unchanged except\nthe Send Bug Report button is gone.\n\n🤖 Generated with [Claude Code](https://claude.com/claude-code)\n\n---------\n\nCo-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>"},{"hash":"667a3bb0","date":"2026-04-28 02:06:02 +0100","author":"Alex Cheema","subject":"feat: keep-models option when uninstalling EXO (#1997)","body":"## Summary\n\n- Adds a **Keep downloaded models (~/.exo/models)** checkbox to the\nmacOS uninstall confirmation dialog (Settings → Advanced → Danger Zone).\nThe full `~/.exo` directory is now removed on uninstall by default; if\nthe checkbox is checked, `~/.exo/models` is preserved.\n- The standalone `app/EXO/uninstall-exo.sh` gains a matching\n`--keep-models` flag and the same `~/.exo` cleanup so GUI and CLI flows\nstay in sync. Resolves the user home via `$SUDO_USER` since the script\nruns under `sudo`.\n\nPreviously, \"Uninstall EXO\" only cleaned up system-level components\n(LaunchDaemon, network location, logs, app bundle) and left the entire\n`~/.exo` directory behind. Now uninstalling actually removes EXO's user\ndata, with a one-click opt-out for the (potentially many GB) of\ndownloaded models.\n\n![Uninstall dialog with new\ncheckbox](https://raw.githubusercontent.com/exo-explore/exo/703b7fbbf13441217ad2903bb199f07e92af4490/uninstall-dialog.png)\n\n> Note: the rendered icon in the screenshot above is the generic system\nfolder icon because it was captured from a small standalone Swift binary\n(no app bundle / icon resource). When triggered from the actual EXO.app,\nthe EXO app icon is shown.\n\n## Test plan\n\n- [ ] Build EXO.app locally; open Settings → Advanced → Danger Zone →\nUninstall EXO; confirm the new \"Keep downloaded models (~/.exo/models)\"\ncheckbox is present and unchecked by default.\n- [ ] Uninstall with the checkbox **checked** → `~/.exo/models/`\nsurvives, all other entries under `~/.exo` are gone, system components\nremoved, app moved to Trash.\n- [ ] Uninstall with the checkbox **unchecked** → `~/.exo` is fully\nremoved.\n- [ ] `sudo app/EXO/uninstall-exo.sh --keep-models` → `~/.exo/models/`\nis preserved, the rest of `~/.exo` is removed.\n- [ ] `sudo app/EXO/uninstall-exo.sh` (no flag) → `~/.exo` is fully\nremoved.\n- [ ] `app/EXO/uninstall-exo.sh --help` prints usage and exits 0;\nunknown args exit 2 with a usage hint.\n\n🤖 Generated with [Claude Code](https://claude.com/claude-code)\n\n---------\n\nCo-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>\nCo-authored-by: Evan <evanev7@gmail.com>"},{"hash":"c80b10c0","date":"2026-04-28 01:58:17 +0100","author":"Evan Quiney","subject":"implement engine abstraction for mlx and mflux (#2000)","body":"refactor for future versions."},{"hash":"18ffe1df","date":"2026-04-28 01:28:12 +0100","author":"Alex Cheema","subject":"fix: uninstall-exo.sh removes both current and legacy bridge scripts (#1998)","body":"## Summary\n\nThe standalone `app/EXO/uninstall-exo.sh` only knew about the legacy\nfilename `disable_bridge_enable_dhcp.sh`. On machines installed with\nnewer EXO versions, the current `/Library/Application\nSupport/EXO/disable_bridge.sh` was left behind, and the script then\nreported `EXO support directory not empty, leaving in place`.\n\nThis PR makes the script try both filenames, removing whichever ones\nexist. Tolerates **either**, **both**, or **neither** being present\nwithout erroring.\n\nThe Swift `NetworkSetupHelper.makeUninstallScript()` already handles\nboth paths correctly, so the GUI uninstall flow is unaffected — this is\na script-only fix.\n\nCaught while running an end-to-end uninstall on a real machine for\n#1997.\n\n## Test plan\n\nVerified the new block in isolation against all four states:\n\n- [x] both `disable_bridge.sh` and `disable_bridge_enable_dhcp.sh`\npresent → both removed\n- [x] only `disable_bridge.sh` present → removed cleanly\n- [x] only `disable_bridge_enable_dhcp.sh` present → removed cleanly\n(legacy install)\n- [x] neither present → prints the existing \"already removed?\" warning,\nexits 0\n\n🤖 Generated with [Claude Code](https://claude.com/claude-code)\n\nCo-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>"},{"hash":"f0d1371d","date":"2026-04-28 01:12:42 +0100","author":"rltakashige","subject":"MLX P/D (#1993)","body":"## Motivation\n\nMLX only prefill server for Apple Silicon"},{"hash":"5d10188d","date":"2026-04-27 11:03:12 -0500","author":"Adam Durham","subject":"fix: route by in-flight tasks only — completed tasks were skewing load balance (#1989)","body":"The load balancer counted ALL tasks (Complete, Cancelled, TimedOut,\nFailed) instead of only Pending/Running ones. With 138 accumulated tasks\nand only 7 active, routing decisions were based on historical\ndistribution, causing one node to appear permanently 'busier' and\nstarving the other of work.\n\nCo-authored-by: Adam Durham <adam@example.com>\nCo-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>"},{"hash":"f2a0db4e","date":"2026-04-27 16:53:43 +0100","author":"ciaranbor","subject":"Extend bench/eval tooling (#1905)","body":"## Motivation\n\nExtend bench/eval tooling with robustness features, streaming support,\nand align model configs with vllm eval for reproducible comparisons.\n\n## Changes\n\n- **exo_eval**: Checkpoint/resume (JSONL), instance health monitoring +\nearly abort, `top_k`/`min_p`/`enable_thinking` params, LCB\n`--release-version`/`--offset`\n- **exo_bench**: Streaming SSE (`--stream`), Kimi tokenizer fix for\ntransformers 5.x\n- **Both tools**: Auto-detect running instances instead of requiring\n`--skip-instance-setup`; `--fresh-instance` to override\n- **harness**: SSE streaming client, `find_existing_instance()` shared\nhelper, removed download timeout, settle-timeout default 0→7200s\n- **models.toml**: Added `enable_thinking`, aligned `max_tokens`/temps\nwith vllm, added new models\n- **API**: Streaming SSE for `/bench/chat/completions`\n\n## Why It Works\n\n- Checkpoint/resume uses append-only JSONL + skip-on-load so interrupted\nevals resume without re-running completed questions\n- Health monitoring races an `asyncio.Event` against API calls for fast\nabort when the instance dies\n- Auto-detection queries `/state` for existing instances matching the\nmodel ID before attempting placement\n- Streaming reuses the existing `generate_chat_stream` infrastructure\nfrom the regular chat endpoint"},{"hash":"37f6f4f6","date":"2026-04-27 15:20:50 +0100","author":"rltakashige","subject":"Add DeepSeek V4 Flash/Pro (#1978)","body":"Wait for upstream merge.\n\n---------\n\nCo-authored-by: Evan <evanev7@gmail.com>"},{"hash":"48a922fd","date":"2026-04-27 03:58:59 -0500","author":"Adam Durham","subject":"fix: map presence_penalty and frequency_penalty from ChatCompletionRequest (#1991)","body":"Upstream PR #1947 added `presence_penalty` and `frequency_penalty` to\n`TextGenerationTaskParams` and the mlx-lm generator call sites, but\nmissed wiring them up in the API adapter so they were silently dropped\nfrom incoming requests. This fixes the API mapping.\n\nCo-authored-by: Adam Durham <adam@example.com>"},{"hash":"fd707de3","date":"2026-04-23 15:28:40 +0100","author":"rltakashige","subject":"Add more model cards (#1970)","body":""},{"hash":"45248c5c","date":"2026-04-23 15:08:16 +0100","author":"Alex Cheema","subject":"chore(app): hardcode bug report presigned-URL endpoint (#1971)","body":"## Motivation\n\nThe bug-report presigned-URL endpoint\n(`https://reports.exolabs.net/presigned-urls`) was injected at build\ntime from the `EXO_BUG_REPORT_PRESIGNED_URL_ENDPOINT` GitHub Actions\nsecret into `Info.plist`, then read at runtime by `BugReportService`. It\nisn't actually a secret — the POST body is just `{\"keys\":[...]}` with no\ncredential (see `app/EXO/EXO/Services/BugReportService.swift:136-142`),\nabuse prevention lives server-side on the lambda, and the URL is already\nvisible in every publicly-distributed DMG's `Info.plist`. Treating it as\na repo secret added plumbing with no security benefit and broke local\ndev builds — hitting **Send Bug Report** on an uncustomised `just\nbuild-app` raised \"Bug report endpoint is invalid\".\n\n## Changes\n\n- `app/EXO/EXO/Info.plist`: replace\n`$(EXO_BUG_REPORT_PRESIGNED_URL_ENDPOINT)` with the literal URL.\n- `.github/workflows/build-app.yml`: drop the\n`EXO_BUG_REPORT_PRESIGNED_URL_ENDPOINT` job-level env var and the\nxcodebuild build-setting passthrough. No other workflow changes.\n\nSwift code is unchanged — `BugReportService` still reads from\n`Info.plist`, which leaves an escape hatch if anyone ever needs to\noverride via `xcodebuild EXOBugReportPresignedUrlEndpoint=...` without\nrecompiling.\n\nFollow-up: the `EXO_BUG_REPORT_PRESIGNED_URL_ENDPOINT` repo secret can\nnow be deleted in the GitHub Actions settings UI.\n\n## Why It Works\n\n`Info.plist` variable substitution turns `$(FOO)` into whatever build\nsetting `FOO` resolves to. CI was setting `FOO` via xcodebuild; local\ndev wasn't, so the key resolved to an empty string, which\n`BugReportService.fetchPresignedUploadUrls` rejects via the\n`!trimmedEndpointString.isEmpty` guard at `BugReportService.swift:131`.\nHardcoding the literal string removes the substitution entirely, so\nevery build — local or CI — gets the right value.\n\n## Test Plan\n\n### Manual Testing\n<!-- Hardware: MacBook Pro (macOS app build via Xcode) -->\n- `just build-app` with no extra env vars (reproduces the failure path\non `main`).\n- `/usr/libexec/PlistBuddy -c \"Print :EXOBugReportPresignedUrlEndpoint\"\napp/EXO/build/Build/Products/Debug/EXO.app/Contents/Info.plist` →\nreturns `https://reports.exolabs.net/presigned-urls` (was empty before\nthis change).\n- `open app/EXO/build/Build/Products/Debug/EXO.app` → menubar → **Debug\nInfo** → **Send Bug Report** → type a description → **Send** → upload\nsucceeds and the **Create GitHub Issue** button appears (was failing\nwith \"Bug report endpoint is invalid\" before).\n- Cross-check on the Slack side that the uploaded `report.json` lands\nunder `reports/YYYY/MM/DD/<ts>/` as before.\n\n### Automated Testing\n<!-- Describe changes to automated tests, or how existing tests cover\nthis change -->\n- No new tests. This is a single-string change to `Info.plist` plus a\nworkflow cleanup. `nix flake check` in CI verifies formatting/lint for\nthe rest of the tree.\n\n🤖 Generated with [Claude Code](https://claude.com/claude-code)\n\nCo-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>"},{"hash":"290e3fd9","date":"2026-04-23 11:39:36 +0100","author":"rltakashige","subject":"Keep image cache fresh (temporary fix) (#1961)","body":"## Motivation\n\nWhen a new node joins, it might not have the cache.\n\n\n\nCaveat: \nThis is potentially fallible if a new node joins and updates real\ntopology, but the API topology hasn't caught up with this fact and the\nuser queues up a new text generation. In practice, there is only a split\nsecond where this is the case, and this is only for users of the\ndashboard interface. We should fix this properly after the release."},{"hash":"3894cf13","date":"2026-04-23 02:50:39 +0100","author":"rltakashige","subject":"Fix Gemma 4 E2B TP + DeepSeek V32 thinking parsing (#1967)","body":""},{"hash":"8993ccaf","date":"2026-04-22 18:39:36 +0100","author":"Alex Cheema","subject":"feat(app): add friendly context message to bug report prompt (#1959)","body":"## Motivation\n\nWhen a user clicks **Send Bug Report** in the macOS app, we already give\nthem the option to add more context via an optional text field. But the\ncurrent prompt is just a terse label — `\"What's the issue? (optional)\"`\n— which doesn't tell the user why bothering to fill it in matters. A\nfriendly one-line explanation increases the chance they'll describe what\nwent wrong, which is the single most useful signal when we triage the\nresulting diagnostic bundle.\n\n## Changes\n\n- `app/EXO/EXO/ContentView.swift`: In the `.prompting` phase of\n`sendBugReportButton`, replace the single label with a two-line\nhierarchy:\n  - Primary: `Tell us what went wrong (optional)`\n- Helper: `A quick description of what you were doing and what happened\nhelps us track down the bug for you.`\n- The helper uses `.caption2` + `.secondary` + `.opacity(0.8)` +\n`.fixedSize(horizontal: false, vertical: true)` so it stays visually\nsubordinate and wraps cleanly inside the 340pt popover.\n\nNo changes to `BugReportService`, the `user_description` payload, or any\nother flow.\n\n## Why It Works\n\nThe optional description is already plumbed end-to-end (text editor →\n`bugReportUserDescription` state → `BugReportService.sendReport(...,\nuserDescription:)` → `report.json`'s `user_description` field → GitHub\nissue pre-fill). The only gap was user-facing motivation, so this is\npurely a copy/layout tweak inside the existing `.prompting` case — no\nnew state, bindings, or service changes.\n\n## Test Plan\n\n### Manual Testing\n<!-- Hardware: MacBook Pro (macOS app build via Xcode) -->\n- Build the macOS app in Xcode (`app/EXO/EXO.xcodeproj`) and launch it.\n- Open the menubar popover → expand **Debug Info** → click **Send Bug\nReport**.\n- Verify the new primary label and helper sentence both appear above the\ntext editor and wrap cleanly within the popover width.\n- Leave the field empty → click **Send** → upload should succeed (no\n`user_description` in payload, same as before).\n- Fill in a description → click **Send** → upload succeeds and the\nsuccess card with **Create GitHub Issue** appears; clicking it opens\nGitHub with the description pre-filled.\n- Click **Cancel** from the prompting state → returns to idle.\n\n### Automated Testing\n<!-- Describe changes to automated tests, or how existing tests cover\nthis change -->\n- No new automated tests. This is a SwiftUI copy/layout change; existing\n`EXOTests` are smoke-level and don't cover `ContentView` view bodies,\nand UI snapshot tests aren't worth adding for a two-line copy tweak.\n- `nix fmt` reports 0 files changed after the edit; `nix flake check` in\nCI will verify formatting/lint for the rest of the tree.\n\n🤖 Generated with [Claude Code](https://claude.com/claude-code)\n\nCo-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>"},{"hash":"4939fbe9","date":"2026-04-22 21:29:47 +0400","author":"Nadeem Hilal Wani","subject":"feat(dashboard): add Pi integration tab (#1925)","body":"## Summary\nAdds a new **Pi** tab to the Integrations page (`/#/integrations`)\nalongside the existing Claude Code, OpenCode, Codex, OpenClaw, Open\nWebUI, n8n, and Firefox tabs.\n\n[pi](https://pi.dev) (`@mariozechner/pi-coding-agent`) is a terminal\ncoding agent that supports custom OpenAI-compatible providers via\n`~/.pi/agent/models.json`.\nThis tab gives users a copy-pasteable config to wire pi up to their exo\ncluster.\n\n## What's in the tab\n- **Model selector** (shown when multiple models are running) — picks\nthe default model for the generated shell command.\n- **Models Config card** — generates `~/.pi/agent/models.json`\nregistering `exo` as a custom provider:\n     - `baseUrl` → `<apiUrl>/v1`\n     - `api` → `openai-completions`\n     - `apiKey` → `\"exo\"` (placeholder; exo ignores it)\n- `compat.supportsDeveloperRole: false` and\n`compat.supportsReasoningEffort: false`, per pi docs recommendation for\nlocal OpenAI-compatible servers\n- Auto-populates every running model with `id`, `contextWindow` (from\n`/v1/models`), and `input: [\"text\", \"image\"]` for vision-capable models\n- **Shell Command card** — `pi --provider exo --model <model>` for quick\nlaunch.\n\nThe tab gracefully falls back to `your-model-id` when no models are\nrunning, matching the behavior of the other tabs.\n\n   ## Usage\n\n   1. `npm install -g @mariozechner/pi-coding-agent`\n   2. Paste the generated config into `~/.pi/agent/models.json`\n3. Run `pi` and pick an exo model via `/model` — or run the shell\ncommand directly\n\n   ## Changes\n\n- `dashboard/src/routes/integrations/+page.svelte` — adds `\"Pi\"` to the\n`tabs` tuple, `piModel` state, `piModelsJson` + `piShellCommand`\nderivations, and the tab content block.\n\n   Single-file, scoped change — no backend or type changes.\n\n   ## Testing\n\n   - `cd dashboard && npm run build` — ✅ builds cleanly\n   - `svelte-check` on the edited file — no new errors\n- Manually verified the tab renders, the model selector updates the\ngenerated JSON, and the config reflects `/v1/models` capabilities\n(vision → `input: [\"text\",\"image\"]`,\n `context_length` → `contextWindow`).\n\n   ## Screenshots\n\n<img width=\"1545\" height=\"1236\" alt=\"pi-tab\"\nsrc=\"https://github.com/user-attachments/assets/38aa179f-4ed9-4a1e-9783-d3baa7738263\"\n/>"},{"hash":"73782ecc","date":"2026-04-22 17:12:24 +0100","author":"rltakashige","subject":"Fix event mutation causing indexed vs event mismatch (#1964)","body":"Fixes small issue with #1957"},{"hash":"f6e418ed","date":"2026-04-22 17:05:49 +0100","author":"rltakashige","subject":"Cleanup on #1952 (#1960)","body":""},{"hash":"7a312a17","date":"2026-04-22 16:43:27 +0100","author":"rltakashige","subject":"Misc fixes: upstream JACCL all_sum, API, etc. + Add Kimi K2.6 (#1952)","body":"## Motivation\n\nThis fixes a bunch of observed model quality issues introduced upstream\nin JACCL, as well as API issues and prefix cache calculation.\n\n\n## Test Plan\n\n### Manual Testing\nTested a bunch\n\n### Automated Testing\nAdded a test, automated eval tool calls on Kimi K2.6, Minimax M2.7, GPT\nOSS and Qwen3.6 models.\n\n---------\n\nCo-authored-by: Evan <evanev7@gmail.com>"},{"hash":"0a549f88","date":"2026-04-22 14:03:31 +0100","author":"Evan Quiney","subject":"remove layer loading callback (#1890)","body":"first part of modularising the backend is simplifying some of the\ncontrol flow. more tbd."},{"hash":"df332035","date":"2026-04-22 12:49:25 +0100","author":"Evan Quiney","subject":"swap camelcasemodels for frozenmodels globally (#1957)","body":""},{"hash":"af673845","date":"2026-04-22 11:11:01 +0100","author":"ciaranbor","subject":"Ignore HF remote repo changes (temporary fix) (#1958)","body":"## Motivation\n\nFixes #1918. Downloaded model status reverts from \"completed\" to\n\"pending\" during each download scan. Reproduced with `zai-org/GLM-5.1`.\n\n## Changes\n\n- `coordinator.py`: In the periodic rescan, don't downgrade\nalready-completed models; fall back to `resolve_existing_model()`\n(safetensors weight check) when per-file size check reports incomplete\n- New `test_download_status_not_lost.py`: 3 regression tests\n\n## Why It Works\n\nThe rescan compares local file sizes against HF's `main` revision. When\nHF updates text files (README, jinja, etc.), remote sizes change but\nlocal files still match the old revision — causing a false \"incomplete\".\nThe fix uses the safetensors weight check as ground truth instead.\n\nLong-term: pin the downloaded revision SHA rather than always checking\nagainst `main`.\n\n## Test Plan\n\n### Manual Testing\n\n- Mac Studio M3 Ultra with GLM-5.1 downloaded (natural reproduction of\nthe issue)\n- Confirmed GLM-5.1 stays `DownloadCompleted` through multiple rescan\ncycles\n\n### Automated Testing\n\n- 3 new tests: completed-not-downgraded, fallback-to-resolve,\ngenuinely-incomplete-stays-pending"},{"hash":"49670c86","date":"2026-04-21 16:39:14 +0100","author":"ciaranbor","subject":"Handle missing total_size in safetensors index files (#1956)","body":"## Motivation\n\nImage models fail to load after a mid-download instance deletion and\nrecreation. The system skips the download and crashes with\n`FileNotFoundError: No safetensors files found in .../vae`.\n\n## Changes\n\n- Make `ModelSafetensorsIndexMetadata.total_size` optional (`PositiveInt\n| None = None`)\n- Add null guard in `fetch_safetensors_size`\n- Add regression test\n\n## Why It Works\n\nExolabs quantized image models have safetensors index files with mflux\nmetadata (`quantization_level`, `mflux_version`) but no `total_size`.\nThe required `PositiveInt` field caused Pydantic validation to fail,\nwhich was silently swallowed by `except Exception: continue` in\n`_scan_model_directory`. This skipped all weight map checks, making\nincomplete models appear complete.\n\n## Test Plan\n\n### Manual Testing\n\n- Hardware: Mac Studio\n- Before: `CreateRunner → LoadModel` (crash). After: `CreateRunner →\nDownloadModel` (correct).\n\n### Automated Testing\n\n- `test_safetensors_index.py`: 3 cases covering missing, valid, and null\nmetadata"},{"hash":"fcc3718e","date":"2026-04-21 07:45:33 +0100","author":"rltakashige","subject":"Add sampling defaults (#1947)","body":"## Motivation\n\nModel quality issues\n\n### Manual Testing\nTODO"},{"hash":"8ccfd7fc","date":"2026-04-21 07:41:07 +0100","author":"rltakashige","subject":"Fix some misc build issues (#1948)","body":"## Motivation\n\n<!-- Why is this change needed? What problem does it solve? -->\n<!-- If it fixes an open issue, please link to the issue here -->\n\n## Changes\n\n<!-- Describe what you changed in detail -->\n\n## Why It Works\n\n<!-- Explain why your approach solves the problem -->\n\n## Test Plan\n\n### Manual Testing\n<!-- Hardware: (e.g., MacBook Pro M1 Max 32GB, Mac Mini M2 16GB,\nconnected via Thunderbolt 4) -->\n<!-- What you did: -->\n<!-- - -->\n\n### Automated Testing\n<!-- Describe changes to automated tests, or how existing tests cover\nthis change -->\n<!-- - -->"},{"hash":"7b416155","date":"2026-04-20 12:41:40 +0100","author":"Evan Quiney","subject":"bump uv lock for linux builds (#1942)","body":"followup from a rebase issue in #1874"},{"hash":"93a24748","date":"2026-04-20 11:49:12 +0100","author":"Evan Quiney","subject":"mlx cuda 13 (dgx spark) support (#1874)","body":""},{"hash":"e32829e5","date":"2026-04-20 10:39:28 +0100","author":"Evan Quiney","subject":"chore: bump versions in line with release (#1941)","body":""},{"hash":"09e894dd","date":"2026-04-19 12:55:20 +0100","author":"rltakashige","subject":"Fix vision models on M5 Pro/Max MacBooks (#1927)","body":"## Motivation\n\nVision models don't understand images on M5 series MacBooks. The\nupstream NAX addmm fix (https://github.com/ml-explore/mlx/pull/3422)\nfixes this.\n\n## Why It Works\nSame conclusion I came to when I was debugging the issue on an M5 Max.\nIt works after this fix.\n\n## Test Plan\n\n### Manual Testing\nWorks for Qwen3.5 27B"},{"hash":"bf8aacfd","date":"2026-04-17 21:55:15 +0100","author":"rltakashige","subject":"Improve build CI (#1920)","body":"## Motivation\n\n<!-- Why is this change needed? What problem does it solve? -->\n<!-- If it fixes an open issue, please link to the issue here -->\n\n## Changes\n\n<!-- Describe what you changed in detail -->\n\n## Why It Works\n\n<!-- Explain why your approach solves the problem -->\n\n## Test Plan\n\n### Manual Testing\n<!-- Hardware: (e.g., MacBook Pro M1 Max 32GB, Mac Mini M2 16GB,\nconnected via Thunderbolt 4) -->\n<!-- What you did: -->\n<!-- - -->\n\n### Automated Testing\n<!-- Describe changes to automated tests, or how existing tests cover\nthis change -->\n<!-- - -->"},{"hash":"af9e847e","date":"2026-04-17 12:57:32 -0500","author":"Adam Durham","subject":"fix: force gc + clear_cache after KV prefix cache eviction (#1832)","body":"## Summary\n- After `KVPrefixCache` evicts LRU entries, the MLX Metal buffers stay\nallocated until Python's GC runs\n- This leaks ~3-4 GB between long-context requests, reducing the\neffective context ceiling for back-to-back requests\n- Adding `gc.collect()` + `mx.clear_cache()` after eviction frees Metal\nbuffers promptly\n\n## Test plan\n- [x] Measured on 2-node PP cluster with Qwen3.5-397B-A17B-4bit at 63K\ncontext\n- [x] Before: 108.88 GB retained after eviction (3.78 GB above baseline)\n- [x] After: 105.48 GB retained after eviction (0.38 GB above baseline —\ndraft model KV + minor overhead)\n- [x] `gc.collect()` adds ~2-3ms latency, runs once per eviction cycle\n(not per token)\n- [ ] Verify with `uv run pytest`\n\n🤖 Generated with [Claude Code](https://claude.com/claude-code)\n\n---------\n\nCo-authored-by: Adam Durham <adam@example.com>\nCo-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>\nCo-authored-by: rltakashige <rl.takashige@gmail.com>"},{"hash":"01598960","date":"2026-04-17 13:06:23 +0300","author":"mlpy0","subject":"Add model card for Qwen3.6-35B-A3B-8bit (#1917)","body":"Adds the 8bit variant missing from #1907 — the safetensors index is now\nlive on HF.\n\n- `mlx-community/Qwen3.6-35B-A3B-8bit` (~35 GB)\n\nArchitectural fields match the existing 4bit/5bit/bf16 cards.\n`storage_size.in_bytes` is taken from `metadata.total_size` of the\nupstream `model.safetensors.index.json`."},{"hash":"63b8e647","date":"2026-04-16 23:25:26 +0100","author":"Alex Cheema","subject":"Add model cards for Qwen3.6-35B-A3B variants (#1907)","body":"## Motivation\n\n`mlx-community` has just published the new **Qwen3.6-35B-A3B**\nmultimodal MoE family on HuggingFace. Without static model cards exo\ndoesn't surface these models in the dashboard picker or match its\nplacement / prefill logic, so users can't one-click launch them. This PR\nadds cards for the three quants whose safetensors indexes are already\nlive on HF (4bit / 5bit / bf16).\n\n## Changes\n\nThree new TOML files in `resources/inference_model_cards/`:\n\n- `mlx-community--Qwen3.6-35B-A3B-4bit.toml` (~19 GB)\n- `mlx-community--Qwen3.6-35B-A3B-5bit.toml` (~23 GB)\n- `mlx-community--Qwen3.6-35B-A3B-bf16.toml` (~65 GB)\n\nAll three share the same architectural fields (`n_layers = 40`,\n`hidden_size = 2048`, `num_key_value_heads = 2`, `context_length =\n262144`, capabilities `text, thinking, thinking_toggle, vision`,\n`base_model = \"Qwen3.6 35B A3B\"`) — only `model_id`, `quantization`, and\n`storage_size.in_bytes` differ between variants.\n\n## Why It Works\n\n- Qwen3.6-35B-A3B reuses the `qwen3_5_moe` architecture\n(`Qwen3_5MoeForConditionalGeneration`) — the same one already wired into\nexo's MLX runner at `src/exo/worker/engines/mlx/auto_parallel.py:47` via\n`Qwen3_5MoeModel`. The architectural fields are taken verbatim from the\nHF `config.json.text_config` and match the existing `Qwen3.5-35B-A3B-*`\ncards.\n- Storage sizes are the exact `metadata.total_size` read from each\nvariant's `model.safetensors.index.json` on HF, so download progress and\ncluster-memory-fit checks are accurate.\n- Vision support is flagged in `capabilities`; the `[vision]` block is\nauto-detected by `ModelCard._autodetect_vision` from the upstream\n`config.json`, so no hand-written vision config is required.\n- The card loader (`_refresh_card_cache` in\n`src/exo/shared/models/model_cards.py`) globs every `.toml` in\n`resources/inference_model_cards/` on startup, so nothing else needs to\nchange — the `/models` endpoint and the dashboard picker pick them up\nautomatically.\n\nThe `mxfp4` / `mxfp8` / `nvfp4` variants are still uploading upstream\n(index JSONs currently 404) and can be added in a follow-up PR once HF\ncompletes.\n\n## Test Plan\n\n### Manual Testing\n\nHardware: MacBook Pro M4 Max, 48 GB unified memory.\n\n- Built the dashboard, ran `uv run exo`, waited for the API to come up\non `http://localhost:52415`.\n- `curl -s http://localhost:52415/models` returns the three new model\nids (`mlx-community/Qwen3.6-35B-A3B-{4bit,5bit,bf16}`) alongside\nexisting models.\n- Opened the dashboard, clicked SELECT MODEL, typed \"Qwen3.6\" into the\nsearch box. A single **\"Qwen3.6 35B A3B\"** group appears showing `3\nvariants (19GB-65GB)`. Expanding it lists the `4bit` / `5bit` / `bf16`\nquants with sizes `19GB` / `23GB` / `65GB`, exactly as expected:\n\n![Qwen3.6 35B A3B in model\npicker](https://gist.githubusercontent.com/AlexCheema/68c2c02da9450b44968e6b0e0b1d255e/raw/127119f70382353c65a847188e5a2c9013db68d2/qwen36-picker.png)\n\n- Programmatically loaded each TOML via `ModelCard.load_from_path(...)`\nand confirmed the parsed fields (layers / hidden / KV heads / context /\nquant / base_model / caps / bytes) match what's written in the files.\n\n### Automated Testing\n\nNo code paths were touched — these are pure TOML data files that plug\ninto the existing model-card loader. The existing pytest suite covers\nTOML parsing and card serving; adding new TOMLs doesn't require new test\nscaffolding. `uv run ruff check` and `nix fmt` are clean.\n\n---------\n\nCo-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>\nCo-authored-by: Ryuichi Leo Takashige <rl.takashige@gmail.com>"},{"hash":"28c79784","date":"2026-04-16 11:59:33 +0100","author":"rltakashige","subject":"Update mlx and mlx lm to latest (#1906)","body":"Just bumping to the very latest upstream versions."},{"hash":"058bb082","date":"2026-04-15 23:30:12 +0100","author":"rltakashige","subject":"Allow copying on dashboard even on HTTP (#1902)","body":"## Motivation\n\n<!-- Why is this change needed? What problem does it solve? -->\n<!-- If it fixes an open issue, please link to the issue here -->\n\n## Changes\n\n<!-- Describe what you changed in detail -->\n\n## Why It Works\n\n<!-- Explain why your approach solves the problem -->\n\n## Test Plan\n\n### Manual Testing\n<!-- Hardware: (e.g., MacBook Pro M1 Max 32GB, Mac Mini M2 16GB,\nconnected via Thunderbolt 4) -->\n<!-- What you did: -->\n<!-- - -->\n\n### Automated Testing\n<!-- Describe changes to automated tests, or how existing tests cover\nthis change -->\n<!-- - -->"},{"hash":"3eead802","date":"2026-04-15 21:14:52 +0100","author":"rltakashige","subject":"Better environment variables in MacOS app (#1901)","body":"## Motivation\n\nCloses #1858 \n\n## Changes\n\n<!-- Describe what you changed in detail -->\n\n## Why It Works\n\n<!-- Explain why your approach solves the problem -->\n\n## Test Plan\n\n### Manual Testing\n<!-- Hardware: (e.g., MacBook Pro M1 Max 32GB, Mac Mini M2 16GB,\nconnected via Thunderbolt 4) -->\n<!-- What you did: -->\n<!-- - -->\n\n### Automated Testing\n<!-- Describe changes to automated tests, or how existing tests cover\nthis change -->\n<!-- - -->"},{"hash":"87329c80","date":"2026-04-15 19:40:00 +0100","author":"rltakashige","subject":"Add usage stats to tool calls and handle multiple tool calls correctly (#1899)","body":"## Motivation\n\nTool calls are usually not end tokens, so they didn't have usage stats."},{"hash":"8cdc8338","date":"2026-04-15 15:23:07 +0100","author":"rltakashige","subject":"Drain tokens silently skipped in thinking parsing   (#1898)","body":"## Motivation\nCloses #1882"},{"hash":"2cd66ae4","date":"2026-04-15 09:46:02 +0100","author":"rltakashige","subject":"Fix out of order event idx causing fatal crashes (#1894)","body":"## Motivation\n\n<img width=\"828\" height=\"373\" alt=\"Screenshot 2026-04-14 at 22 56 52\"\nsrc=\"https://github.com/user-attachments/assets/f8f48c1d-68c5-4acc-a6de-9d180672da9d\"\n/>\n\nif is_new_master=True, _elect_loop creates a new EventRouter before the\nworker has receivers. Then, event router runs _run_ext_in and\nbuf.drain_indexed() will pick off events, even though\nself.internal_outbound is not populated fully.\n\nFinally, when the worker does try requesting events, the next event it\nreceives is not the first event, meaning the worker crashes.\n\n## Changes\n\nStart the event router after all the receivers are registered\n\n## Why It Works\n\nself.internal_outbound is populated before the loop begins.\n\n## Test Plan\n\n### Manual Testing\nNo more crashes observed in testing (it's actually quite easy to\nreproduce the issue if you have one node with this fix but the other\nnode on main).\n\nI'm convinced this is a fix, at least."},{"hash":"2ecefa0c","date":"2026-04-14 23:05:55 +0100","author":"rltakashige","subject":"Fix Qwen3-VL and autodetect vision config (#1893)","body":"## Motivation\n\nQwen3 VL TP doesn't work atm, and vision is not behaving.\n\n## Test Plan\n\n### Manual Testing\nWorks now."},{"hash":"b8eaf707","date":"2026-04-14 20:31:59 +0100","author":"rltakashige","subject":"Add gemma 4 tensor parallelism (#1891)","body":""},{"hash":"8d81811b","date":"2026-04-14 16:37:49 +0100","author":"rltakashige","subject":"Try harder to clean up processes nicely (#1889)","body":"## Motivation\n\nModel loading is actually quite reliable now. No need to kill if you\nhave a slow SSD or it's a massive model; the user can shut the instance\ndown if necessary.\n\nThis was a major cause of signal=9 issues although not the only one (can\nhappen during inference too?).\nThe reason signal=9 is so bad is that RDMA will no longer work until\nrestart if this ever happens.\n\n## Changes\n\n- no more model load timeout\n- no more crazy sigkills\n- try harder to clean up processes on model shutdown\n\n## Test Plan\n\n### Manual Testing\nTested with some RDMA instances"},{"hash":"f2709dcd","date":"2026-04-14 11:12:58 +0100","author":"rltakashige","subject":"Add prefix cache flag to exo bench (#1888)","body":"## Motivation\nFor using Exo-Bench extensively, there are many cases that we could use\nprefix caching to speed up the benchmarks, especially when the focus is\non the token generation.\n\nAt the same time, it's very clear that prefix caching decode tokens is\nnot very useful in most current scenarios. Surprisingly, even for\nnon-thinking models, the chat template means that a continued\nconversation will be formatted such that the existing cache is not\neffective.\n\nWe already (slightly accidentally) do this for the batch generator - we\nshould do it for the sequential generator too.\n\n## Changes\n\nWe can now speed up exo bench by having a use prefix caching flag. Of\ncourse, for most accurate pp results, it is better to not have it, but\nthis speeds up tg and large benchmarking significantly.\nUpdated methodology to match\n\n## Test Plan\n\n### Manual Testing\nTested on many configurations that the difference in results is\nnegligible, even with multiple --pp options."},{"hash":"77ffe039","date":"2026-04-13 18:38:33 +0100","author":"ciaranbor","subject":"Complete responses api usage response field (#1885)","body":"## Motivation\n\nThe Responses API usage response was missing `input_tokens_details` and\n`output_tokens_details`. The chat completions API already reports these.\n\n## Changes\n\n- Added `InputTokensDetails` (`cached_tokens`) and `OutputTokensDetails`\n(`reasoning_tokens`) to `ResponseUsage`\n- Extracted shared `_build_response_usage()` helper for both streaming\nand non-streaming paths\n\n## Test Plan\n\n### Manual Testing\n\n4-node cluster, `Qwen3-30B-A3B-4bit` — verified both detail objects\npresent with correct values in streaming and non-streaming responses.\n\n### Automated Testing\n\n13 tests in `test_openai_responses_api.py`."},{"hash":"3f0df404","date":"2026-04-13 18:32:17 +0100","author":"rltakashige","subject":"Reduce memory consumption by adding Flash Attention to Qwen3.5 and Gemma 4, and fix RotatingKVCache prefix cache memory leak (#1886)","body":"## Motivation\n\nPart 1 of many memory improvements.\n\n## Changes\nAs written in the title\n\n## Test Plan\n\n### Manual Testing\nGemma 4 26B cache reduced from 54GB -> 10GB per 100k tokens, Qwen3.5 35B\nA3B cache reduced from 21GB every 100000 tokens to 7GB."},{"hash":"9b381f7b","date":"2026-04-13 16:45:17 +0100","author":"Evan Quiney","subject":"bump and simplify flake (#1866)","body":"seems like stablepkgs swiftfmt works now! also bump macmon to 0.7"},{"hash":"d2f67b5d","date":"2026-04-13 15:08:15 +0100","author":"Alex Cheema","subject":"dashboard: group Gemma under Google with proper logo (#1883)","body":"## Motivation\n\nIn the dashboard model picker sidebar, the Gemma 4 models were showing\nup under a \"Gemma\" family with the generic fallback tick/checkmark icon\n(the default case in `FamilyLogos.svelte`), since no dedicated logo\nbranch existed for `family === \"gemma\"`. Every other vendor (Meta,\nNVIDIA, OpenAI, DeepSeek, Qwen, …) has its own brand mark.\n\nGemma is Google's model family, so it should live under a **Google**\nbucket that future Google-authored models can join, and it should render\nwith a proper Google logo in the same style as its neighbors.\n\n## Changes\n\n- `dashboard/src/lib/components/FamilyLogos.svelte`: added a `family ===\n\"google\"` branch rendering a monochrome Google \"G\" as a single `<path>`\ninside the shared `24×24` viewBox with `fill=\"currentColor\"`, matching\nthe other vendor logos.\n- `dashboard/src/lib/components/FamilySidebar.svelte`: added `google:\n\"Google\"` to the `familyNames` display map.\n- `dashboard/src/lib/components/ModelPickerModal.svelte`: inserted\n`\"google\"` into the `familyOrder` array (next to `\"llama\"`) so the\nvendor has a deterministic sort position.\n- `resources/inference_model_cards/mlx-community--gemma-4-*.toml` (16\nfiles): changed `family = \"gemma\"` → `family = \"google\"`. `base_model =\n\"Gemma 4 …\"` is unchanged, so the model titles still read \"Gemma\".\n\n## Why It Works\n\nThe sidebar builds its family list from whatever values appear in\n`model.family` across the loaded model cards (`ModelPickerModal.svelte`\n`uniqueFamilies`). Renaming the family string on the 16 Gemma cards from\n`\"gemma\"` to `\"google\"` collapses them into a single \"Google\" bucket,\nand the new logo branch + display-name map entry gives that bucket a\nreal brand mark and label. All other logos share the same `w-6 h-6 /\nviewBox=\"0 0 24 24\" / fill=\"currentColor\"` shape, so inheriting\n`text-exo-yellow` / `text-white/50` just works.\n\n## Test Plan\n\n### Manual Testing\n<!-- Hardware: MacBook Pro M3 Max -->\n- `cd dashboard && npm install && npm run build` — dashboard builds\ncleanly.\n- `uv run exo`, opened `http://localhost:52415`, clicked **SELECT\nMODEL**:\n- sidebar shows a **Google** entry with a monochrome Google \"G\" logo in\nthe same style as Meta / NVIDIA / etc.\n  - old \"Gemma\" entry with the generic tick is gone.\n- clicking **Google** filters to the Gemma 4 variants (e2b / e4b / 26B\nA4B / 31B).\n- hover/selected color states switch between `text-white/50` and\n`text-exo-yellow` correctly.\n\n### Automated Testing\n- No new tests — this is a cosmetic grouping/logo change. Existing\ndashboard build verifies the Svelte + TS compiles."},{"hash":"89735033","date":"2026-04-14 00:01:42 +1000","author":"chaoliang yan","subject":"fix: use configured api_port for IP connectivity probes (#1877)","body":"## Motivation\n\nFixes #1861\n\nWhen `--api-port` is set to a non-default value (e.g., `--api-port\n55555`), the IP connectivity discovery system still probes peers on the\nhardcoded default port 52415. Since the API is not listening on 52415,\nall reachability checks fail, the topology reports zero reachable nodes,\nand the dashboard shows \"No valid configurations for current settings.\"\n\n## Changes\n\nThread the configured `api_port` from `Args` through `Worker` into the\nreachability probe functions:\n\n- `net_profile.py`: `check_reachability()` and `check_reachable()`\naccept an `api_port` parameter (default 52415 for backward\ncompatibility)\n- `worker/main.py`: `Worker` stores `api_port` and passes it to\n`check_reachable()`, uses it in `Multiaddr` construction and the mDNS\nconnection filter\n- `main.py`: passes `args.api_port` to the `Worker` constructor\n\n## Why It Works\n\nThe `/node_id` endpoint used by reachability probes is served by the\nFastAPI app, which binds to `args.api_port`. The probes must use the\nsame port the API is actually listening on. Before this fix, the port\nwas hardcoded in three places in `net_profile.py` and `worker/main.py`;\nnow it uses the value from the CLI flag.\n\n## Test Plan\n\n### Manual Testing\n<!-- Hardware: not available for multi-node testing -->\n- Verified ruff passes on all changed files\n- Code inspection: traced `api_port` flow from `Args.parse()` →\n`Node.create()` → `Worker.__init__()` → `_poll_connection_updates()` →\n`check_reachable()` → `check_reachability()` → HTTP probe URL\n\n### Automated Testing\n- No existing automated tests cover the reachability probe code path\n- The new `api_port` parameter defaults to `52415`, so all existing\nbehavior is preserved when `--api-port` is not specified\n\n---------\n\nCo-authored-by: lawrence3699 <lawrence3699@users.noreply.github.com>\nCo-authored-by: Evan <evanev7@gmail.com>"},{"hash":"eb922861","date":"2026-04-13 14:47:09 +0100","author":"Alex Cheema","subject":"models: add MiniMax M2.7 cards (#1884)","body":"## Motivation\n\nThe mlx-community [MiniMax-M2.7\ncollection](https://huggingface.co/collections/mlx-community/minimax-m27)\nlanded but exo didn't have model cards for any of the variants yet, so\nthey weren't selectable from the dashboard model picker. Adding cards\nalso makes them discoverable under the existing MiniMax family entry.\n\n## Changes\n\nAdded 6 new model cards in `resources/inference_model_cards/`, one per\nquant of MiniMax M2.7:\n\n- `mlx-community--MiniMax-M2.7.toml` (bf16, full precision — 457 GB)\n- `mlx-community--MiniMax-M2.7-4bit.toml` (128 GB)\n- `mlx-community--MiniMax-M2.7-4bit-mxfp4.toml` (121 GB)\n- `mlx-community--MiniMax-M2.7-5bit.toml` (157 GB)\n- `mlx-community--MiniMax-M2.7-6bit.toml` (185 GB)\n- `mlx-community--MiniMax-M2.7-8bit.toml` (243 GB)\n\nAll six use `family = \"minimax\"` and share `base_model = \"MiniMax M2.7\"`\nso they collapse into a single group in the picker with the existing\nMiniMax logo. Architecture fields (`n_layers = 62`, `hidden_size =\n3072`, `num_key_value_heads = 8`, `context_length = 196608`) were read\nfrom each repo's `config.json`; `storage_size.in_bytes` was summed from\nthe HF tree API per repo.\n\n`capabilities = [\"text\", \"thinking\"]` follows the existing MiniMax M2.5\ncards — the chat template always emits `<think>` tags (no toggle),\nmatching M2.5 behavior.\n\n## Why It Works\n\nModel cards in `resources/inference_model_cards/` are auto-loaded by\n`src/exo/shared/models/model_cards.py::get_model_cards`. The dashboard\npicker groups by `base_model` and filters by `family`, so sharing both\nacross all six variants gives a single \"MiniMax M2.7\" group under the\nMiniMax sidebar entry, with the quant variants exposed as selectable\nsub-options.\n\n## Test Plan\n\n### Manual Testing\n<!-- Hardware: MacBook Pro M3 Max -->\n- Ran `uv run python -c \"…await get_model_cards()…\"` and confirmed all 6\nnew cards load with `family=minimax`, `base_model=\"MiniMax M2.7\"`, and\ncorrect quant + byte sizes.\n- `cd dashboard && npm run build` then `uv run exo`, opened the model\npicker → **MiniMax** family → **MiniMax M2.7** group shows all six quant\nvariants.\n\n### Automated Testing\n- No new automated tests — these are data files validated by the\nexisting Pydantic `ModelCard` schema at load time.\n\n---------\n\nCo-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>"},{"hash":"4b13735e","date":"2026-04-11 14:46:03 +0300","author":"MikkoParkkola","subject":"build: remove pyinstaller temp artifacts (#1868)","body":"removes PyInstaller `build/` leftovers after `just package`"},{"hash":"196543ce","date":"2026-04-11 12:29:33 +0100","author":"rltakashige","subject":"Add Gemma 4 + VLM fixes + thinking parsing updates (#1851)","body":"## Motivation\nAdd support for Gemma 4, including VLM!\n\n## Changes\n\n- Add auto parallel strategies and model cards for Gemma 4\n- Normalise Gemma 4's special Vision Transformer handling to be in line\nwith the rest of our vision processors.\n- Also adds reprs to messages and b64 hashes to prevent log spam.\n\n## Test Plan\n\n### Manual Testing\nTested manually on 4bit E2B and 8bit 26B\n\n### Automated Testing\nModel onboarding shows small logit diffs.\n\n---------\n\nCo-authored-by: Evan <evanev7@gmail.com>"},{"hash":"6172617b","date":"2026-04-10 17:39:55 +0100","author":"ciaranbor","subject":"add env override to macos app (#1869)","body":"## Motivation\n\n- Let users pass/override arbitrary exo env vars from the macOS app\nwithout a code change.\n\n## Changes\n\n- `ExoProcessController.swift`: `CustomEnvironmentVariable` struct +\n`@Published` list persisted to `UserDefaults`, injected into the child\nprocess after built-ins.\n- `SettingsView.swift`: new **Environment** tab with add/remove rows,\ntrim + dedup on save, and a POSIX name validator with a warning badge.\n\n## Why It Works\n\n- Custom vars applied last in `makeEnvironment`, so overriding a\nbuilt-in works with no special-casing.\n\n## Test Plan\n\n### Manual Testing\n\n- Set `EXO_LIBP2P_NAMESPACE` via the new UI; confirmed override in\n`~/.exo/exo_log/exo.log`."},{"hash":"93a980a6","date":"2026-04-10 17:25:30 +0100","author":"ciaranbor","subject":"`just package` first builds dashboard (#1867)","body":"## Motivation\n\n<!-- Why is this change needed? What problem does it solve? -->\n<!-- If it fixes an open issue, please link to the issue here -->\n\n## Changes\n\n<!-- Describe what you changed in detail -->\n\n## Why It Works\n\n<!-- Explain why your approach solves the problem -->\n\n## Test Plan\n\n### Manual Testing\n<!-- Hardware: (e.g., MacBook Pro M1 Max 32GB, Mac Mini M2 16GB,\nconnected via Thunderbolt 4) -->\n<!-- What you did: -->\n<!-- - -->\n\n### Automated Testing\n<!-- Describe changes to automated tests, or how existing tests cover\nthis change -->\n<!-- - -->"},{"hash":"2962ebee","date":"2026-04-10 15:59:34 +0100","author":"ciaranbor","subject":"Fix pdf inputs on Safari (#1865)","body":"## Motivation\n\nPDF attachments weren't working on Safari\n\n## Changes\n\nCreate async readable stream if none exists\n\n## Why It Works\npdfjs-dist requires an async readable stream internally\n\n## Test Plan\n\n### Manual Testing\npdf attachments now work on Safari, still work on Firefox"},{"hash":"abd75ae0","date":"2026-04-10 15:53:48 +0100","author":"rltakashige","subject":"Truncate long logs with repr (#1854)","body":"## Motivation\n\n<!-- Why is this change needed? What problem does it solve? -->\n<!-- If it fixes an open issue, please link to the issue here -->\n\n## Changes\n\n<!-- Describe what you changed in detail -->\n\n## Why It Works\n\n<!-- Explain why your approach solves the problem -->\n\n## Test Plan\n\n### Manual Testing\n<!-- Hardware: (e.g., MacBook Pro M1 Max 32GB, Mac Mini M2 16GB,\nconnected via Thunderbolt 4) -->\n<!-- What you did: -->\n<!-- - -->\n\n### Automated Testing\n<!-- Describe changes to automated tests, or how existing tests cover\nthis change -->\n<!-- - -->\n\n---------\n\nCo-authored-by: Evan <evanev7@gmail.com>"},{"hash":"ee2e505b","date":"2026-04-10 17:43:42 +0600","author":"kaiisfree","subject":"fix: handle BrokenResourceError in download progress callback (#1846)","body":"## Summary\n- Wraps progress callback `send()` in try/except to gracefully handle\n`BrokenResourceError` when the memory stream is closed\n- Prevents unhandled `ExceptionGroup` from crashing the process when the\ndownload consumer disconnects during transfer\n\n## Root Cause\nThe download progress callback sends updates through an anyio memory\nobject stream. When the receiving end closes (e.g., client disconnect,\ntimeout, or task cancellation), `send()` raises `BrokenResourceError`.\nInside an anyio `TaskGroup`, this unhandled exception becomes an\n`ExceptionGroup` that propagates up and crashes the coordinator.\n\n## Fix\nCatch `BrokenResourceError` (and `ClosedResourceError` for completeness)\nin the progress callback and handle gracefully — the download continues\nbut progress updates are silently dropped for disconnected consumers.\n\nFixes #1844\n\n🤖 Generated with [Claude Code](https://claude.com/claude-code)"},{"hash":"f2e6b1ef","date":"2026-04-09 12:34:35 +0100","author":"Evan Quiney","subject":"prevent some crash loops (#1827)","body":"extension to #1763 that prevents crash looping in some common scenarios."},{"hash":"e2e17eaf","date":"2026-04-09 12:29:35 +0100","author":"ciaranbor","subject":"Fix reasoning_tokens counting for multi-token thinking tag models (#1848)","body":"## Motivation\n\n`reasoning_tokens` is always 0 in usage stats, even when thinking\ncontent streams correctly via `reasoning_content` SSE deltas. The MLX\ngenerators had their own thinking detection comparing individual\ndetokenized tokens against think tags — this never fires for models\nwhere tags span multiple tokens (e.g. gpt-oss-120b) or are already in\nthe prompt.\n\n## Changes\n\n- Removed broken per-token thinking detection from `batch_generate.py`\nand `generate.py`\n- Added `_count_reasoning_tokens` wrapper in `model_output_parsers.py`\nthat counts `is_thinking=True` responses and patches the total into\nUsage on the final response\n- Wired it as the outermost stage of `apply_all_parsers`, so it works\nregardless of which parser sets `is_thinking`\n- Added 3 tests covering `parse_thinking_models` and `parse_gpt_oss`\npaths\n\n## Why It Works\n\nThe parser pipeline already correctly sets `is_thinking` on each\nresponse. Counting at the output of `apply_all_parsers` means one\ncounting point that works for all model types, replacing the duplicate\nbroken logic in two generators.\n\n## Test Plan\n\n### Manual Testing\n\n- 4-node cluster, `mlx-community/gpt-oss-120b-MXFP4-Q8`\n- Main branch: `reasoning_tokens: 0` — fix branch: `reasoning_tokens:\n25`\n\n### Automated Testing\n\n- 3 new tests: explicit think tags, `starts_in_thinking=True`, and\ngpt-oss Harmony analysis channel"},{"hash":"b12cd1b1","date":"2026-04-08 16:14:28 +0100","author":"ciaranbor","subject":"Cancel SSE keep-alive when instance is deleted (#1828)","body":"## Motivation\n\nWhen a model instance is deleted (e.g. node disconnect, manual\nteardown), any in-flight SSE streaming connections for that instance\nhang indefinitely. The API never closes the response stream, so clients\nblock forever waiting for more chunks.\n\n## Changes\n\n- Listen for `InstanceDeleted` events in the API event loop\n- Add `_close_streams_for_instance()` to find and close any active\ntext/image generation queues tied to tasks on the deleted instance\n- Add unit tests covering text gen, image gen, and\nunrelated-instance-not-closed scenarios\n\n## Why It Works\n\nWhen an instance is deleted, we iterate `state.tasks` to find commands\nrunning on that instance, then close and remove their send-side queue\nhandles. This causes the SSE generator to terminate, unblocking the\nclient.\n\n## Test Plan\n\n### Manual Testing\n- This was causing issues for me on another branch (integration tests).\nIncluding this fix solved the issue\n\n### Automated Testing\n- `test_instance_deleted_stream_cleanup.py`: 3 tests covering text gen\ncleanup, image gen cleanup, and ensuring unrelated streams are not\naffected"},{"hash":"62570227","date":"2026-04-08 11:50:04 +0300","author":"mlpy0","subject":"Catch ClosedResourceError when forwarding chunks to client queues (#1856)","body":"While stress-testing inference with rapid client cancels mid-stream, I\nhit a reproducible crash where the entire exo process exits.\n\nWhen a client cancels a streaming chat completion partway through, its\nreceive stream gets closed cleanly via its context manager. The producer\nin `API._apply_state` then calls `queue.send(event.chunk)`, which raises\n`anyio.ClosedResourceError` rather than `BrokenResourceError`. The\nexisting handler only catches `BrokenResourceError`, so the exception\npropagates through the API task group, kills the Node task group, and\nthe process exits with `EXO Shutdown complete`.\n\nTrace from one of the crashes:\n\n```\nFile \"exo/api/main.py\", line 1818, in _apply_state\n    await queue.send(event.chunk)\nFile \"anyio/streams/memory.py\", line 212, in send_nowait\n    raise ClosedResourceError\nanyio.ClosedResourceError\n```\n\nThe fix is to catch `ClosedResourceError` alongside\n`BrokenResourceError` in both queue handlers (text and image), so the\ndead queue gets dropped and `_apply_state` keeps running for other\nin-flight requests."},{"hash":"645bc209","date":"2026-04-08 02:04:42 +0100","author":"Alex Cheema","subject":"Add Fast Synch Enabled toggle to macOS app settings (#1852)","body":"## Motivation\n\nThe exo backend already supports `--fast-synch` / `--no-fast-synch` CLI\nflags and the `EXO_FAST_SYNCH` environment variable, but there was no\nway to toggle this from the macOS app UI. Users who want fast CPU-to-GPU\nsynchronization for RDMA with Tensor Parallelism had to use CLI flags.\n\n## Changes\n\n- **ExoProcessController.swift**: Added `fastSynchEnabled`\nUserDefaults-backed property and pass `EXO_FAST_SYNCH=on` to the exo\nprocess environment when enabled.\n- **SettingsView.swift**: Added a \"Performance\" section to the Advanced\ntab with a \"Fast Synch Enabled\" toggle, an info icon (ⓘ) tooltip\nexplaining the feature and trade-offs, and a \"Save & Restart\" button.\n\n## Why It Works\n\nFollows the exact same pattern as the existing `offlineMode` and\n`enableImageModels` settings — UserDefaults persistence, `@Published`\nproperty with `didSet`, environment variable passthrough in\n`makeEnvironment()`, and pending state with Save & Restart in the\nsettings UI. The `EXO_FAST_SYNCH=on` value matches what the Python\nbackend already reads in `main.py`.\n\n## Test Plan\n\n### Manual Testing\n<!-- Hardware: macOS app -->\n- Open Settings → Advanced tab → verify \"Performance\" section with \"Fast\nSynch Enabled\" toggle appears\n- Hover the ⓘ icon → verify tooltip explains the feature and GPU lock\ntrade-off\n- Toggle on → click \"Save & Restart\" → verify process restarts with\n`EXO_FAST_SYNCH=on` in env\n- Close and reopen Settings → verify the toggle state persists\n- Verify \"Save & Restart\" button is disabled when no changes are pending\n\n### Automated Testing\n- Existing settings patterns are well-established; no new automated\ntests needed for this UI toggle\n\n---------\n\nCo-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>"},{"hash":"5757c27d","date":"2026-04-08 01:58:39 +0100","author":"rltakashige","subject":"Add download utility script (#1855)","body":"## Motivation\n\n<!-- Why is this change needed? What problem does it solve? -->\n<!-- If it fixes an open issue, please link to the issue here -->\n\n## Changes\n\n<!-- Describe what you changed in detail -->\n\n## Why It Works\n\n<!-- Explain why your approach solves the problem -->\n\n## Test Plan\n\n### Manual Testing\n<!-- Hardware: (e.g., MacBook Pro M1 Max 32GB, Mac Mini M2 16GB,\nconnected via Thunderbolt 4) -->\n<!-- What you did: -->\n<!-- - -->\n\n### Automated Testing\n<!-- Describe changes to automated tests, or how existing tests cover\nthis change -->\n<!-- - -->"},{"hash":"fd5b2328","date":"2026-04-07 18:26:29 +0100","author":"Andrei Cravtov","subject":"Workspace tweaks (#1849)","body":"## Changes\n\nMostly chore changes around vscode and jetbrains workspace settings, and\nsome basedpyright settings tweaks, to allow direnv to work and nixd\nautocomplete with flake parts to work"},{"hash":"43b3df45","date":"2026-04-07 12:50:12 +0100","author":"rltakashige","subject":"Fix BatchGenerator in line with upstream refactor (and prevent Qwen3.5 memory leak) (#1835)","body":"## Motivation\n\nMLX LM has had a massive refactor to their BatchGenerator recently.\nSince we'd like new features from MLX LM such as Gemma 4, we need to\nupdate the code to handle this.\n\nAdditionally this fixes a significant memory leak in GatedDeltaNet (the\ndifference is quite substantial, up to 1GB every 1000 tokens, explaining\nseveral memory issues users were facing with Qwen3.5 models)\n\n## Testing\nBefore\n<img width=\"3146\" height=\"884\" alt=\"image\"\nsrc=\"https://github.com/user-attachments/assets/5af0f55a-393c-4a32-9eed-ae43f1611af4\"\n/>\n\n\nAfter (no memory leak, as one of the changes upstream)\n<img width=\"3190\" height=\"892\" alt=\"image\"\nsrc=\"https://github.com/user-attachments/assets/f0bd128d-fd48-40d4-9bbd-50a564beab14\"\n/>"}]}