Exo
Search the build
-
27b85c8f2026-07-02 local: darwin mlx from stock PyPI wheel 0.31.2 (no Metal toolchain needed; LAN cluster, TB-RDMA fork not applicable) -
6abccd1b2026-07-02 local: darwin mlx from PyPI wheel (no Metal toolchain; LAN cluster, no TB-RDMA) -
b5375f8c2026-06-22 Add Kimi K2.7-Code model card (official INT4 weights + vision) (#2167) -
cdf1add82026-06-22 fix: upgrade devalue to 5.6.2 (CVE-2026-22774) (#2150) -
09f9ea312026-06-03 libp2p -> zenoh (#2132) -
81d7cb0f2026-06-03 docs: add Homebrew cask install instructions (#2140) -
629c55d62026-05-31 Rename exo_pyo3_bindings to exo_rs (#2131) -
f9f8cbb32026-05-29 fix: make app builds work again (#2127) -
051a64e32026-05-28 Capture energy in prefill and ageneration separately (#2124) -
a8602ea62026-05-26 fix(bug): no longer repeated _trigger_notify_user_to_download_model (#2114) -
a1a22b5f2026-05-25 feat: added background/daemon support (#2106) -
74e9fe152026-05-22 fix(bug): EventRouter lifetime-handling fixed, no more process crashes (#2102) -
90f24bef2026-05-15 fix model cards not validating properly after #2071 (#2096) -
5097b2662026-05-15 Tweaked workspace settings (#2095) -
bc6661e62026-05-15 Add node backends to model cards (#2071) -
14aab3562026-05-15 Runner error handling (#2093) -
88d46d462026-05-14 fix: omit null delta fields in streaming chat completions (issue #2082) (#2092) -
e8ec8d502026-05-14 fix ollama API compatibility for VS Code Copilot (#2091) -
1fd15d592026-05-14 create directory on startup (#2089) -
4466cd532026-05-13 use custom mlx sources for linux (#2087) -
ed2d10bd2026-05-12 Redirect runner stdout/stderr to file logs (#2084) -
87c72fc12026-05-11 Fixes issue #2068 (#2083) -
b76bc3012026-05-10 bump rust versions (#2081) -
08ffa5f62026-05-10 Map GLM 4.7 stop tokens to GLM 4 IDs (#2061) -
45df74ba2026-05-09 Andrei/mp capture stdio (#2056) -
ce37bdce2026-05-09 fix: Create directory for PID file if it doesn't exist (#2075) -
e5a1e5da2026-05-08 Create PID file locking for EXO (#2072) -
fa5713132026-05-08 Integration tests infra (#1995) -
414132ae2026-05-07 Use time-weighted power sampling (#2038) -
edef80042026-05-07 Store custom model cards in State (#2024) -
a0c00f9d2026-05-07 fix(placement): gate RDMA on nodeRdmaCtl.enabled at both endpoints (#2014) -
89d20c182026-05-06 fix(inference): prevent TP collective deadlock via agree_on_tasks order (#2048) -
dbcceaa52026-05-05 Initialise _cancelled_tasks in ImageEngine (#2051) -
9c6ff4ce2026-05-01 feat: update rdma_ctl instructions (#1977) -
b26268df2026-05-01 fix(macos-app): disable URL response caching for cluster-state polling (#2005) -
8dae3ecb2026-04-30 A few targeted tweaks to address HF rate limits (#2009) -
fb12b4032026-04-30 fix(app): tighten Share Bug Report prompt layout (#2008) -
1606e6382026-04-30 feat(app): open Share Bug Report in a dedicated window (#2003) -
667a3bb02026-04-28 feat: keep-models option when uninstalling EXO (#1997) -
c80b10c02026-04-28 implement engine abstraction for mlx and mflux (#2000) -
18ffe1df2026-04-28 fix: uninstall-exo.sh removes both current and legacy bridge scripts (#1998) -
f0d1371d2026-04-28 MLX P/D (#1993) -
5d10188d2026-04-27 fix: route by in-flight tasks only — completed tasks were skewing load balance (#1989) -
f2a0db4e2026-04-27 Extend bench/eval tooling (#1905) -
37f6f4f62026-04-27 Add DeepSeek V4 Flash/Pro (#1978) -
48a922fd2026-04-27 fix: map presence_penalty and frequency_penalty from ChatCompletionRequest (#1991) -
fd707de32026-04-23 Add more model cards (#1970) -
45248c5c2026-04-23 chore(app): hardcode bug report presigned-URL endpoint (#1971) -
290e3fd92026-04-23 Keep image cache fresh (temporary fix) (#1961) -
3894cf132026-04-23 Fix Gemma 4 E2B TP + DeepSeek V32 thinking parsing (#1967) -
8993ccaf2026-04-22 feat(app): add friendly context message to bug report prompt (#1959) -
4939fbe92026-04-22 feat(dashboard): add Pi integration tab (#1925) -
73782ecc2026-04-22 Fix event mutation causing indexed vs event mismatch (#1964) -
f6e418ed2026-04-22 Cleanup on #1952 (#1960) -
7a312a172026-04-22 Misc fixes: upstream JACCL all_sum, API, etc. + Add Kimi K2.6 (#1952) -
0a549f882026-04-22 remove layer loading callback (#1890) -
df3320352026-04-22 swap camelcasemodels for frozenmodels globally (#1957) -
af6738452026-04-22 Ignore HF remote repo changes (temporary fix) (#1958) -
49670c862026-04-21 Handle missing total_size in safetensors index files (#1956) -
fcc3718e2026-04-21 Add sampling defaults (#1947) -
8ccfd7fc2026-04-21 Fix some misc build issues (#1948) -
7b4161552026-04-20 bump uv lock for linux builds (#1942) -
93a247482026-04-20 mlx cuda 13 (dgx spark) support (#1874) -
e32829e52026-04-20 chore: bump versions in line with release (#1941) -
09e894dd2026-04-19 Fix vision models on M5 Pro/Max MacBooks (#1927) -
bf8aacfd2026-04-17 Improve build CI (#1920) -
af9e847e2026-04-17 fix: force gc + clear_cache after KV prefix cache eviction (#1832) -
015989602026-04-17 Add model card for Qwen3.6-35B-A3B-8bit (#1917) -
63b8e6472026-04-16 Add model cards for Qwen3.6-35B-A3B variants (#1907) -
28c797842026-04-16 Update mlx and mlx lm to latest (#1906) -
058bb0822026-04-15 Allow copying on dashboard even on HTTP (#1902) -
3eead8022026-04-15 Better environment variables in MacOS app (#1901) -
87329c802026-04-15 Add usage stats to tool calls and handle multiple tool calls correctly (#1899) -
8cdc83382026-04-15 Drain tokens silently skipped in thinking parsing (#1898) -
2cd66ae42026-04-15 Fix out of order event idx causing fatal crashes (#1894) -
2ecefa0c2026-04-14 Fix Qwen3-VL and autodetect vision config (#1893) -
b8eaf7072026-04-14 Add gemma 4 tensor parallelism (#1891) -
8d81811b2026-04-14 Try harder to clean up processes nicely (#1889) -
f2709dcd2026-04-14 Add prefix cache flag to exo bench (#1888) -
77ffe0392026-04-13 Complete responses api usage response field (#1885) -
3f0df4042026-04-13 Reduce memory consumption by adding Flash Attention to Qwen3.5 and Gemma 4, and fix RotatingKVCache prefix cache memory leak (#1886) -
9b381f7b2026-04-13 bump and simplify flake (#1866) -
d2f67b5d2026-04-13 dashboard: group Gemma under Google with proper logo (#1883) -
897350332026-04-14 fix: use configured api_port for IP connectivity probes (#1877) -
eb9228612026-04-13 models: add MiniMax M2.7 cards (#1884) -
4b13735e2026-04-11 build: remove pyinstaller temp artifacts (#1868) -
196543ce2026-04-11 Add Gemma 4 + VLM fixes + thinking parsing updates (#1851) -
6172617b2026-04-10 add env override to macos app (#1869) -
93a980a62026-04-10 `just package` first builds dashboard (#1867) -
2962ebee2026-04-10 Fix pdf inputs on Safari (#1865) -
abd75ae02026-04-10 Truncate long logs with repr (#1854) -
ee2e505b2026-04-10 fix: handle BrokenResourceError in download progress callback (#1846) -
f2e6b1ef2026-04-09 prevent some crash loops (#1827) -
e2e17eaf2026-04-09 Fix reasoning_tokens counting for multi-token thinking tag models (#1848) -
b12cd1b12026-04-08 Cancel SSE keep-alive when instance is deleted (#1828) -
625702272026-04-08 Catch ClosedResourceError when forwarding chunks to client queues (#1856) -
645bc2092026-04-08 Add Fast Synch Enabled toggle to macOS app settings (#1852) -
5757c27d2026-04-08 Add download utility script (#1855) -
fd5b23282026-04-07 Workspace tweaks (#1849) -
43b3df452026-04-07 Fix BatchGenerator in line with upstream refactor (and prevent Qwen3.5 memory leak) (#1835) -
24420eb12026-04-05 Fix reasoning_tokens always reported as 0 for thinking models (#1836) -
59669c112026-04-05 Tighten EXO bench concurrency numbers and explain methodology (#1811) -
1d2ce4642026-04-02 Allow pausing and deleting active downloads (#1829) -
eb6ae9fd2026-04-01 Prevent failed instance retries (#1763) -
4688adb52026-03-31 Support PDFs in dashboard (#1822) -
d9ed94302026-03-30 Fix Nemotron cache leak upstream (#1819) -
c6815bfd2026-03-30 Only update KV prefix cache on a good cache hit (#1817) -
39c39e812026-03-30 Integrations helpers (#1810) -
e5cb7b802026-03-30 Add SSE-keepalive to not time out on long prefill on clients (#1803) -
635801d52026-03-30 Add multimodality! (#1802) -
2efbb8ab2026-03-30 Improve exo harness with path state (#1815) -
c6c5a3e72026-03-30 feat: /state/paths (#1796) -
10ef7ec92026-03-30 feat: add Firefox AI sidebar (?q=) support to dashboard (#1814) -
1e51dc892026-03-27 chore: bump exo-version with release version (#1807) -
5327bdde2026-03-26 Fix custom model add requiring two attempts + enlarge sidebar buttons (#1805) -
15f1b61f2026-03-26 Rework model storage directory management (for external storage) (#1765) -
903430012026-03-26 [Fix] Node hang on reelection (#1801) -
1d1dfaa12026-03-26 Don't download original/ and metal/ folders from HF (#1800) -
7625213d2026-03-26 fix: enable macmon if preflight fails (#1799) -
f318f9ea2026-03-26 Fix macOS build bundling wrong macmon binary (#1797) -
30fd5aa12026-03-25 Prefer higher % downloaded nodes for API placement previews (#1795) -
6de14cfe2026-03-25 Support image generation cancellation (#1774) -
fc1ae9012026-03-25 fix: DeepSeek V3.2 warmup crash and tool calling + add catalog cards (#1769) -
565ed41c2026-03-25 Fix occasional warmup bugs by using mlx_generate (#1794) -
2da740c32026-03-25 Feat/static peer discovery (#1690) -
7117d7482026-03-25 Update dependencies including mlx 0.31.2 (#1789) -
178c617b2026-03-24 Rename Nemotron to NVIDIA in model picker with logo (#1790) -
7277c9032026-03-24 Fix enable thinking (#1786) -
7ee88c1f2026-03-24 override macmon in flake (#1747) -
509533d42026-03-24 send error finish reason on failing to parse a tool call (#1785) -
b6240a972026-03-24 Prevent Qwen3.5 looping by using mlx lm fork (#1784) -
6cdfbb7e2026-03-24 Add HF_ENDPOINT in the app settings (#1783) -
fac6832e2026-03-24 fix warmup consistency for slow machines (#1748) -
7df3774c2026-03-24 Improve batch performance and stats reporting (#1777) -
248919c22026-03-24 Fix first start in offline mode crash (#1782) -
49951e1b2026-03-24 Sync custom model cards across nodes (#1768) -
e06e70a82026-03-24 Prefer higher model download % for placement (#1767) -
e9fdd8d42026-03-23 improve logging: add dates to verbose stderr, match file log level to verbosity (#1772) -
07598a3a2026-03-19 teeny refactor (#1753) -
63f57fc12026-03-19 Ciaran/dashboard download bug (#1755) -
a6519ba02026-03-17 Update mflux to 0.16.9 (#1751) -
b713889f2026-03-17 Fix exo bench again again (#1750) -
6ee673142026-03-17 Fix exo bench prefill and decode tps (#1749) -
ff4d20ee2026-03-17 Fix image models through dashboard (#1746) -
7ed463952026-03-13 use structured concurrency in download coordinator (#1722) -
29d4165f2026-03-14 Add step logo condition to FamilyLogos component (#1676) -
12af7c952026-03-13 fix: shield macmon cleanup from cancellation to prevent orphaned process (#1714) -
f28b2fd02026-03-13 Extract mlx revision from uv lock (#1715) -
ea18a6252026-03-12 fix: guard against ZeroDivisionError in mlx_lm stats (#1707) -
0782d90e2026-03-12 fix: show partial download progress on initial dashboard load (#1706) -
f221a6c82026-03-11 Normalise Responses API tool call format (#1704) -
2994b4102026-03-11 fix: validate num_key_value_heads in tensor sharding placement (#1669) -
38f0c0912026-03-10 fix: use StreamingDetokenizer in batch generator to fix emoji/UTF-8 corruption (#1691) -
f36fd56c2026-03-10 Include power usage in bench responses (#1692) -
82c54dd62026-03-10 Add support for Nemotron sharding (#1693) -
3536161f2026-03-10 Fix stale state.runners (#1684) -
a6aa07ed2026-03-10 feat(dashboard): add mobile drawer support for sidebars (#1677) -
131ad0ff2026-03-09 Implement continuous batching (#1642) -
d01636102026-03-06 Ciaran/gpt oss tool call finish reason (#1673) -
79c8dbea2026-03-06 Ciaran/minor download bugs (#1674) -
009615972026-03-06 up gossipsub limit (#1671) -
7a36d3962026-03-06 docs: Update documentation for v1.0.68 release (#1667) -
eee343272026-03-06 feat(mlx): add repetition_penalty and repetition_context_size to chat completions (#1665) -
e8c3a8732026-03-06 feat(dashboard): improve model picker feedback and HuggingFace search (#1661) -
afab30952026-03-06 fix: `KVPrefixCache` Regression (#1668) -
b9d40e8e2026-03-05 Ciaran/re download bug (#1658) -
3a4d635d2026-03-05 Fix copy code button not working in dashboard (#1659) -
848580502026-03-04 fix(worker): emit error chunks when a runner dies mid-command (#1645) -
4de8f8012026-03-04 #Add reasoning parms to chat completion and responses APIs (#1654) -
5777bf3c2026-03-04 fix: coerce tool-call argument types from tool schema (#1651) -
886192f12026-03-03 ignore closed resource errors when trying to cancel a task (#1652) -
d914acd62026-03-03 check if we have a task before we delete it (#1634) -
37296c822026-03-03 Refactor runner for implementing batching (#1632) -
28817d3e2026-03-03 Add support for Qwen3.5 (#1644) -
0e1b95012026-03-03 fix: mini topology sidebar navigates home on click (#1616) -
f0d4ccbe2026-03-03 feat: add POST /v1/cancel/{command_id} endpoint (#1579) -
858dc8082026-03-02 fix: replace Master event_sender after EventRouter recreation (#1630) (#1637) -
635118ef2026-02-27 Support trace deletion in dashboard (#1628) -
dc0bb5e12026-02-27 fmt: add taplo TOML formatter to treefmt configuration -
152a27ea2026-02-26 Fix pipeline mismatched send after 1587 (#1629) -
db36bd5a2026-02-26 Add custom prefill for pipeline (#1587) -
639243aa2026-02-26 event router (#1572) -
db73c4fd2026-02-26 move messaging into rust (#1549) -
eaed92952026-02-26 Use tmpdir for coordination file (#1624) -
ba611f9c2026-02-25 Revert "report macmon failures more aggressively (#1618)" (#1625) -
eab3e0b42026-02-25 report macmon failures more aggressively (#1618) -
c4e874e92026-02-25 skip nan logprobs on tokens (#1622) -
e23c3a302026-02-25 Address Mac Mini pipeline GPU timeouts (#1620) -
190e63e52026-02-25 fix: log exceptions causing silent node shutdown (#1621) -
766089352026-02-24 fix: replace flaky internet checks with explicit offline mode (#1615) -
9a2d2a4a2026-02-24 bump (#1608) -
bea64fe82026-02-24 fix: prevent stale loading state and conversation loss when switching chats (#1613) -
14526d282026-02-24 update mlx 2 (#1611) -
73e50df82026-02-24 fix glm5 tool calling (#1612) -
2b417f282026-02-24 fix: sync model selectors between sidebar and chat input (#1610) -
b65982dd2026-02-24 fix: improve text contrast on HOME and DOWNLOADS nav links (#1609) -
2fe689312026-02-24 download .model files in exo bench (#1607) -
644c55732026-02-24 fix: improve text contrast on downloads page (#1601) -
12c3015f2026-02-23 fix qwen moe tensor sharding (#1604) -
365dd68d2026-02-23 Final fixes for release (#1603) -
d3d129582026-02-23 test: verify instance deletion cancels ongoing tasks (#1508) -
c90a0cec2026-02-23 fix: suppress closure errors in runnersupervisor and force spawn start method (#1547) -
e8c133712026-02-23 fix: add download/resume buttons to pending downloads (#1581) -
7024ddcf2026-02-23 fix: detect completed downloads by checking final file exists (#1582) -
dc89ba662026-02-23 feat: add info button to model picker variant rows (#1589) -
5cd96b502026-02-23 feat: seamless chat UX with auto model selection and smart recommendations (#1590) -
05986f772026-02-23 add exo bench protobuf dependency (#1596) -
dab7ed482026-02-23 fix: handle gossipsub MessageTooLarge error to prevent silent crash (#1583) -
226101472026-02-23 runner process checks (#1592) -
61d2a2b62026-02-23 add lazy task group (#1569) -
0ff99a2c2026-02-23 fix isinstance for qwen3Moe (#1595) -
fbb80e1c2026-02-23 Address ring slowdown by turning on FAST SYNCH (#1594) -
8d94eab62026-02-21 bench: fix KeyError on DownloadCompleted total field -
f370452d2026-02-23 Better onboarding UX (#1533) -
a4c2aa2b2026-02-23 fix: raise error when MlxJaccl requested without RDMA cycles (#1585) -
7312c5352026-02-22 feat: add user context prompt and GitHub issue option to macOS bug report (#1544) -
187170232026-02-22 chore: remove deprecated MlxIbv dashboard references (#1584) -
1780e4ad2026-02-20 fix: change RDMA AVAILABLE to RDMA NOT ENABLED warning (#1580) -
ab9273e72026-02-20 downloads: add read_only flag to DownloadCompleted for EXO_MODELS_PATH -
71e48c0f2026-02-20 model-cards: add missing metadata for Qwen3 Coder Next variants (#1576) -
42da58c22026-02-20 worker: add EXO_MODELS_PATH for pre-downloaded model directories -
6b5a70592026-02-20 fix: immediate cancel check after prefill completes (#1575) -
6b54a2702026-02-20 fix: add downloaded_bytes to DownloadPending event (#1564) -
e01f50a52026-02-20 Update mlx fork (#1565) -
109308022026-02-20 cancel active downloads on coordinator shutdown (#1567) -
1a2b8b042026-02-20 Refactor runner into separate runners (#1570) -
dc8d42b42026-02-20 add system ids (#1536) -
d484b0622026-02-20 bench: add download timing to bench output (#1566) -
e32b649d2026-02-20 fix: enable psutil fallback for memory monitoring when macmon is missing (#1478) -
bddad7e72026-02-20 feat: show ETA on prefill progress bar (#1557) -
addf73a12026-02-20 Add support for Ollama API (#1560) -
a16ff2c02026-02-20 fix: correct misleading docstring in seed_models (#1561) -
3006c8ea2026-02-20 Ensure coordinator is rank 0 (#1559) -
f662c1292026-02-19 Prioritise tb for ring instances (#1556) -
c45ff9ad2026-02-19 memory tidy (#1558) -
7031901a2026-02-19 Prevent common fatal crashes (#1555) -
cf648a532026-02-19 Add thinking in thinking blocks, and fix DeepSeek interleaved tool calls (#1548) -
94b2ce692026-02-19 feat: Mac Studio en2 RDMA port warning v2 (#1551) -
423ed0f02026-02-19 Strip Claude headers to improve prefix cache hit rates (#1552) -
ed001f242026-02-19 remove prefillprogress event (#1550) -
4c4c6ce92026-02-19 simplify rust ident module -
42e1e7322026-02-19 bench: restore --danger-delete-downloads planning phase (#1542) -
aa3f106f2026-02-19 fix: import ResponsesStreamEvent and DRY up SSE formatting (#1499) -
2e2960512026-02-19 fix: finalize cancel tasks (#1498) -
cacb456c2026-02-19 remove nightly (#1538) -
51021f6f2026-02-19 Add cancellation button and the ability to cancel during prefill (#1540) -
025ed9fd2026-02-18 feat: add prefill progress bar for long prompts (#1181) -
19bc09552026-02-18 Add status=downloaded filter for model endpoint (#1539) -
7cadca4f2026-02-18 Try multiple endpoints for internet connectivity check (#1516) -
24e99ce12026-02-18 Cleanup mistakes (#1537) -
315992542026-02-18 fix: unblock MpReceiver.close() to prevent shutdown hang (#1511) -
ce5a65d32026-02-18 Add MiniMax M2.5 model cards (#1514) -
c2f2111b2026-02-18 Fix tool calling (#1529) -
6c322ebb2026-02-18 feat: only show thinking toggle for models that support it (#1497) -
2ebe62162026-02-18 feat: add explicit --offline mode for air-gapped clusters (#1525) -
f54c80b12026-02-18 Ciaran/image edit api (#1500) -
48b8f8632026-02-18 Add support for GLM 5 (#1526) -
5cbd63772026-02-18 prioritize official model cards over custom model cards -
8f01523d2026-02-18 remove dead code (#1496) -
3addeade2026-02-18 Update mlx-lm to 0.30.7 (#1520) -
f2be92922026-02-17 Leo/address rdma gpu locks 2 (#1515) -
83af8c632026-02-17 Revert "Use custom fork that resolves GPU locks" (#1502) -
eccc62982026-02-17 Revert "Add MetaInstance declarative layer (#1447)" -
c89972172026-02-17 Revert "feat: better onboarding UX for new users (#1479)" -
490d2e462026-02-17 feat: better onboarding UX for new users (#1479) -
facf2d4d2026-02-17 Use custom fork that resolves GPU locks (#1489) -
a962a28a2026-02-17 Add MetaInstance declarative layer (#1447) -
db79c3502026-02-17 Fix graceful process shutdown in macOS app (#1372) -
d6301ed52026-02-17 dashboard: redesign downloads page as model×node table (#1465) -
6d1ca6682026-02-17 don't time out node identities (#1493) -
c01b6fff2026-02-17 eprint banner -
8392e78a2026-02-17 bench: add spec for automatic canary benchmarks (#1483) -
86735ece2026-02-16 begins -
2759e9232026-02-16 api cancellation (#1276) -
131fb1412026-02-11 bench: add --danger-delete-downloads flag with planning phase -
2d8bfc2e2026-02-16 fix: PlaceInstanceParams broken field validator -
042999f72026-02-16 Ciaran/message deletion (#1409) -
b61dc2eb2026-02-16 Prevent image editing without image input (#1410) -
36a7115b2026-02-16 Pass usage and generation stats through all adapters correctly (#1461) -
0b7d88b42026-02-13 python: add hermetic basedpyright typecheck to nix flake check -
1c3cc6992026-02-13 fix: add missing getModelFitStatus prop to Recent tab (#1470) -
5a2864272026-02-13 Add support for Step 3.5 flash! (#1460) -
6950f9412026-02-12 dashboard: show macOS version in debug mode (#1454) -
d0c442732026-02-12 feat: add enable_thinking toggle for thinking-capable models (#1457) -
cc3321382026-02-12 bench: add --settle-timeout for cluster startup retry (#1449) -
62e8110e2026-02-11 fix: prevent DownloadModel TaskCreated event flood (#1452) -
987734372026-02-11 Make info gatherer monitors resilient with retry loops and timeouts (#1448) -
a8acb3ca2026-02-11 dashboard: show available disk space on downloads page -
a0721dbe2026-02-11 feat: warn when cluster nodes have mismatched macOS versions (#1436) -
50e2bcf92026-02-11 fix: RDMA debug labels, TB5 info box, and rdma_ctl status detection (#1437) -
7bed91c92026-02-11 feat: add Recent tab to model picker (#1440) -
48caea4a2026-02-10 feat: add intermediate model-fit state for cluster-capacity-only models (#1441) -
eead50b42026-02-11 Fix setrlimit crash when hard file descriptor limit < 65535 (#1430) -
199df64c2026-02-10 util: remove VecExt trait, inline at call site (#1446) -
dc7ade802026-02-10 set the mlx hash -
dc7814972026-02-10 update mlx to 0.30.6 -
c37eb2432026-02-10 util: remove dead code (#1445) -
8af2af632026-02-10 nix: override apple-sdk to 26.2 and enable MLX_BUILD_CPU (#1443) -
43728b202026-02-10 Send all exo logs (#1439) -
1699fcfb2026-02-10 standardise logs (#1442) -
009b43c62026-02-10 add log rotation for .exo/exo.log (#1438) -
1f242e8e2026-02-10 gossipsub: stop silent message dropping and warn (#1434) -
64179c6f2026-02-10 Dont save to app directory (#1435) -
305a3c8b2026-02-10 event_log: move event log from unbounded in-memory list to disk (#1432) -
ead19bea2026-02-10 Always load image model cards into cache (#1421) -
5a83e5912026-02-10 dashboard: allow typing in chat input while response is generating (#1433) -
5b5577be2026-02-10 build-app: upload DMG to S3 for non-tagged builds (#1428) -
8314a2aa2026-02-10 cleaning up the todos (#1406) -
163cf1832026-02-10 Add error handling to info gatherer monitor loops (#1422) -
2204f6512026-02-10 Yield from reachability checks (#1427) -
4abdaaf72026-02-10 Address GPU timeouts (#1429) -
2fbdb27b2026-02-07 Handle config.json not found (image models) (#1408) -
3f57416d2026-02-07 Add image lightbox (#1414) -
8f3681cf2026-02-07 Synchronize before warmup (#1419) -
9dc4f7862026-02-07 Ciaran/image model listing (#1417) -
dcb4cabc2026-02-06 Update the nix hash for mlx 0.30.5 (#1416) -
d79b3a0e2026-02-06 bench: make exo-bench available via nix run on all platforms (#1415) -
a2f1d4872026-02-06 slow down catchup (#1407) -
3b2f553a2026-02-06 Fix kimi tool calling id (#1413) -
5455a97a2026-02-06 Fix GLM4Moe Tensor Sharding (#1411) -
6f0cb9992026-02-06 Ciaran/flux1 kontext (#1394) -
c8d3154f2026-02-06 More image dimensions (#1395) -
63e9cc4f2026-02-06 Ciaran/num sync steps (#1396) -
9b5cae3d2026-02-06 auto bench (#1405) -
cf7201f92026-02-06 pyproject: set minimum uv version -
b315035a2026-02-06 Add minimax and fix qwen sharding strategies (#1318) -
c8dbbee22026-02-06 skip tensor ring on bench (#1403) -
f0107e962026-02-06 Fix offline no cache (#1402) -
9f5027932026-02-06 fix: retry downloads on transient errors instead of breaking (#1398) -
c83713492026-02-06 add scripts (#1401) -
6b9073982026-02-05 cancel downloads for deleted instances (#1393) -
572e64792026-02-05 better cancellation (#1388) -
e59ebd982026-02-05 set exo as the nix default package (#1391) -
5c2f29f32026-02-05 feat: show download availability in model picker (#1377) -
ffe6396c2026-02-05 Add Qwen3-Coder-Next model cards (#1367) -
3a9baeb92026-02-03 EXO: add CLI flags for root install/uninstall -
01b86a9e2026-02-05 feat: add uncertainty visualization with token-level logprobs (#1180) -
221640a62026-02-05 Acknowledge task after runner status is updated (#1381) -
6177550c2026-02-04 Ciaran/parallel cfg (#1361) -
7b6cad942026-02-04 add resources dir to nix (#1376) -
41ed7afb2026-02-04 feat: add model picker modal with grouped models and HF Hub search (#1369) -
206327892026-02-04 feat: add custom HuggingFace model support (#1368) -
a0f4f3632026-02-03 Reduce reliance on internet (#1363) -
acb971272026-02-03 Normalize TextGenerationTaskParams.input to list[InputMessage] (#1360) -
d90605f12026-02-03 migrate model cards to .toml files (#1354) -
f400b4d72026-02-02 fix InstanceViewModel.swift (#1359) -
d97bca882026-02-02 improve distributed testing (#1300) -
dfce188d2026-02-02 fix: handle unclosed tool calls and GLM arg parsing edge cases (#1344) -
54b198792026-02-02 create config home when checking for config file (#1353) -
19965c7b2026-02-02 Ciaran/profiling (#1345) -
3e27ead72026-02-02 remove mdns discovered peers from appearing in state (#1312) -
d826d3092026-02-02 chore: gitignore hosts_*.json files (#1343) -
c35379802026-02-02 feat: add Claude Messages API and OpenAI Responses API support (#1167) -
21d477f12026-02-02 Update exo bench (#1357) -
b2579c782026-02-02 nix: add macmon to PATH in wrapper scripts on Darwin -
cd9467422026-01-30 fix skipping logic in worker plan (#1342) -
a5bc38ad2026-01-30 Check all nodes to evict (#1341) -
2a4e0d462026-01-30 make node-ids unique per-session (#1338) -
46a141532026-01-30 switch to ModelCard.load outside of download log (#1339) -
9ba61f372026-01-30 improve log message in shard downloader -
d9eca7582026-01-30 Add usage stats (#1333) -
9dabde7e2026-01-29 Fix bench after recent updates (#1331) -
a31942ce2026-01-29 Ciaran/image non streaming (#1328) -
7cc313b22026-01-29 Treat Swift/Xcode build warnings as errors (#1322) -
2837225d2026-01-29 Load pipeline layers sequentially (#1329) -
e4c6a7db2026-01-15 nix: add Python packaging with uv2nix -
b1e88a3d2026-01-29 shfmt -
ebeddfb32026-01-29 mlx: build with Nix (#1285) -
911157592026-01-29 Add startup delay and update network setup message (#1309) -
ffacabe72026-01-29 Fix uninstall button error (#1306) -
9e58a5752026-01-28 Add RDMA caveats to README.md (#1316) -
748a02602026-01-28 fix configdata validation for kimi-k2 (#1314) -
f1a2d0542026-01-28 Update tagline to "Run frontier AI locally" (#1313) -
b3c8f85f2026-01-28 Update MLX to 0.30.4 (#1311) -
a562114b2026-01-28 Add Kimi K2.5 support (#1302) -
991d27812026-01-27 replace nix fmt with treefmt in just lint (#1301) -
c55cbf672026-01-27 Add mlx lm style tensor sharding for Minimax (#1299) -
bd4f0bf02026-01-26 Fix download speed/ETA display for re-downloads (#1294) -
cd8c01b72026-01-26 Fix kv prefix cache (#1262) -
59e991ce2026-01-26 Only ignore message if actually empty (#1292) -
ffba340e2026-01-26 Ciaran/image quantization (#1272) -
9968abe82026-01-26 Leo/fix basic model shard (#1291) -
0e30b0832026-01-26 Fix download system for upstream file changes (#1290) -
44453c4c2026-01-26 Remove change-detection checks from info gatherer monitors (#1283) -
1290e8ed2026-01-26 dashboard: fix prettier-svelte rebuilding on every file change -
d93db3d62026-01-24 re enable the evil network script (#1277) -
ff4a20222026-01-23 Revert state compaction (#1259) (#1275) -
cee48f6f2026-01-23 Parse GPT OSS tool calling (#1271) -
2b67e84a2026-01-23 state compaction (#1259) -
7204fdeb2026-01-23 Restore Thunderbolt Bridge LaunchDaemon (#1270) -
ec345a432026-01-23 fix: deprioritise uncertain ethernet devices (#1267) -
9967dfa72026-01-23 Prevent conversation collision (#1266) -
23fd37fe2026-01-23 Add FLUX.1-Krea-dev model (#1269) -
d229df382026-01-23 Fix placement filter to use subset matching instead of exact match (#1265) -
8a595fee2026-01-23 Fix Thunderbolt bridge cycle detection to include 2-node cycles (#1261) -
c8571a172026-01-23 Fix guidance (#1264) -
771a86332026-01-23 fix instance port assignment (#1268) -
6dbbe7792026-01-19 downloads: add download and delete buttons to downloads UI -
9357503c2026-01-19 downloads: refactor to run at node level -
ba1994082026-01-23 Fix regenerate for image models (#1263) -
f255345a2026-01-23 dashboard: decouple prettier-svelte from dashboard source -
a1939c892026-01-23 Enable UI settings for image editing (#1258) -
cb9c9ee52026-01-23 Enable generating multiple images. Optionally stream partial images (#1251) -
df240f832026-01-23 Fix GLM and Kimi tool calling crashes (#1255) -
cd125b3b2026-01-22 Use icon for image editing models (#1252) -
b783a2132026-01-22 dashboard: add placement filter by clicking topology nodes (#1248) -
43f12f5d2026-01-22 Replace LaunchDaemon with dynamic Thunderbolt Bridge loop detection (#1222) -
8027d7932026-01-22 Ciaran/hf token (#1250) -
ac6efa742026-01-21 add kimi tool parseing -
2e3c33db2026-01-21 implement mlx-lm tool calling -
fc8e6ad02026-01-22 Reduce download log spam (#1249) -
023108a12026-01-21 Disable image model cards temporarily (#1247) -
c9818c302026-01-21 dashboard: show model total size on downloads page for pending downloads -
8f6726d62026-01-21 Fix config.json download errors for image models (#1245) -
ede779212026-01-21 Reduce log spam (#1241) -
a7e205e42026-01-19 treefmt: add Svelte file formatting -
a354aaa32026-01-21 Fix tests broken in recent commits (#1239) -
307f454b2026-01-21 feat: initial image generation support (#1095) -
a31b6ee02026-01-21 Import download utils once all modules are loaded (#1238) -
6a9251b92026-01-21 Add mflux type stubs (#1234) -
758464702026-01-20 Fix GPT OSS tensor sharding with upstream MLX LM (#1223) -
9e2179c82026-01-20 Register original layer in CustomMlxLayer (#1229) -
22b5d8362026-01-20 swap all instances of model_id: str for model_id: ModelId (#1221) -
ea9c6d6b2026-01-20 Remove dead local paths code from download_shard (#1227) -
4ea66d422026-01-20 Reduce download log spam (#1225) -
8b709e682026-01-20 Mark slow tests as slow (#1220) -
4da6eeb12026-01-20 fix a test broken by #1204 (#1219) -
3d2eee482026-01-20 quiet localhost log -
116558832026-01-20 don't clear mdns discovered connections -
d4f551c62026-01-20 Simplify model cards (#1204) -
176ab5ba2026-01-20 Add GLM-4.7-Flash model cards (4bit, 5bit, 6bit, 8bit) (#1214) -
f5e6aa822026-01-20 Load layers individually (#1211) -
39f0ed602026-01-19 Prepend <think> tag to stream for thinking models like GLM-4.7 (#1186) -
ee43b5982026-01-19 Split NodePerformanceProfile into granular state mappings (#1209) -
5fd555942026-01-19 Wrap pipeline models for explicit mx.depends between cache and logits (#1206) -
5ab1f8b32026-01-19 NetworkSetupHelper: detect stale startup script content -
2202685c2026-01-19 refactor all information sources (including ipless rdma discovery) (#928) -
ce3ad3912026-01-19 Update README.md with some changes from release 1.0.61 (#1157) -
fb0151632026-01-15 shard_downloader: make on_progress callback async -
346b13e22026-01-19 Enhance LaTeX rendering in dashboard markdown (#1197) -
ea0588422026-01-19 Custom mlx layer composition (#1201) -
73b3f87e2026-01-19 Set swa_idx and ga_idx for single layer (#1202) -
746589ba2026-01-19 tidy: remove context manager from api (#1199) -
f82f862f2026-01-19 Fix several issues with placement (#1200) -
7ff937d82026-01-19 Add dashboard screenshots to README (#1185) -
d19bf0242026-01-19 re-raise exceptions in the runner (#1198) -
618cee522026-01-18 Resolve test event ordering flakiness (#1194) -
9c29eb7d2026-01-18 Add proxy and custom SSL certificate support for corporate networks (#1189) -
c5158bee2026-01-17 Add pre-commit checks documentation to AGENTS.md (#1184) -
5c8a23792026-01-16 Handle model timeouts (#1177) -
745343c72026-01-16 Return error responses for Chat Completions (#1173) -
5e28664c2026-01-16 Fix draft release detection (attempt 3) (#1176) -
ae0a804c2026-01-16 Fix draft release detection query (#1175) -
07cf2c1a2026-01-16 Add GitHub releases with Sparkle release notes integration (#1172) -
83c5285a2026-01-16 reduce logs -
39ee2bf72026-01-16 switch from synchronous threaded pinging to an async implementation (#1170) -
991adfbd2026-01-16 fix local network warning (#1136) -
4b3de6b92026-01-16 Fix exo bench for transformers 5.x (#1168) -
c8de3b902026-01-16 quiet rust logs -
6e6567a82026-01-16 resolve issue #1070 (#1076) -
a735dad62026-01-15 Parse GPT OSS in runner (#1160) -
aaf4e36b2026-01-15 FIX GPT OSS (#1165) -
3e623ccf2026-01-15 up http timeout to 3 seconds and retry on BadStatusLine (#1164) -
c22dad8a2026-01-15 dashboard: add peer: true to package lock (#1162) -
4bc4d5062026-01-15 rust: remove dead code -
e0aab46f2026-01-15 model_cards.py: clean up commented out code -
82ba42ba2026-01-14 add glm-47, minimax-m21 (#1147) -
3671528f2026-01-13 nix: add dashboard build with dream2nix -
e6434ec42026-01-12 nix: add Rust builds with crane and fenix -
bdb43e1d2026-01-13 nix: drop noisy echos from devshell -
e4a01e2b2026-01-13 chore(deps): nix lock file maintenance -
1200a7db2026-01-13 Add tensor sharding for GPT-OSS (#1144) -
47ceb54b2026-01-13 up the rlimit (#1148) -
f8112fdf2026-01-13 nix: convert to flake-parts -
e388f5942026-01-13 docs: add AGENTS.md for AI coding agents guidance (#1132) -
e5e74e1e2026-01-13 Upgrade mlx-lm to 0.30.2 with transformers 5.x compatibility (#1125) -
b968d6f02026-01-13 ci: remove old commented out job -
3bfffd9b2026-01-12 ci: build all Nix outputs on all platforms and push to cachix -
007eb8002026-01-12 nix: enable cachix -
8d7b67892026-01-09 dashboard: show disk usage for completed models -
3c5b7ea62026-01-12 ci: add workflow_dispatch trigger to build-app -
b74a61052026-01-11 Add a basic documentation to the api interface (#1122) -
18c4e49f2026-01-09 nix: put treefmt in devshell -
d85b5d372026-01-09 feat: uninstall button (#1077) -
caafc4862026-01-09 Forward tools to the models chat template properly (#1106) -
cca8c9982026-01-09 cleanup unused dependencies -
d1e88def2026-01-09 scrollbars fixed (#1113) -
59e7594e2026-01-09 UNKNOWN to PREPARING (#1112) -
c65320ac2026-01-08 Fix mlx seed (#1094) -
b9a78f6f2026-01-08 ci: compute CURRENT_PROJECT_VERSION from semver -
8f7f0e892026-01-08 ci: avoid uploading alpha appcasts -
4759b09d2026-01-08 Use presigned URLs for bug report uploads (#1109) -
ca6801852026-01-08 Display RDMA debug info in macOS app. (#1072) -
383309e22026-01-08 fmt: add typescript formatting -
55463a982026-01-08 fmt: add swift formatting -
56af61fa2026-01-08 add a server for distributed testing in /tests until we work out a stable solution. (#1098) -
f76d543d2026-01-08 We shouldn't fail on an HTTPException in the tier-2 discovery system. (#1104) -
ea841aca2026-01-08 local network check (#1103) -
077b1bc72026-01-06 exo-bench (Benchmark model pp & tg speed) (#1099) -
4963c3312026-01-06 Fix Discord link in README.md. Fixes #1096 (#1097) -
4f6fcd9e2025-12-24 feat(macos-app): add custom namespace UI for cluster isolation -
839b67f32026-01-05 [feat] Add an option to disable the worker (#1091) -
47b8e0ce2026-01-05 feat: remember last launch settings (model, sharding, instance type) (#1028) -
17f9b5832026-01-03 Task Deduplication (#1062) -
844bcc7c2026-01-01 fix: prevent form submission during IME composition (#1069) -
c1be51842025-12-31 Fix tests broken by 283c (#1063) -
1ec550df2025-12-31 Emit download progress on start, and change downloads to be keyed by model_id (#1044) -
283c0e392025-12-31 Placement filters for tensor parallel supports_tensor, tensor dimension and pipeline parallel deepseek v3.1 (#1058) -
35be4c552025-12-30 prioritise mlx jaccl coordinator ip (en0 -> en1 -> non-TB5 -> other) -
31d4cd842025-12-30 set KV_CACHE_BITS to None to disable quantized kv cache -
8a6da5842025-12-30 remove mx.set_cache_limit -
16e2bfd32025-12-30 log EXO_LIBP2P_NAMESPACE on start -
ade3ee7e2025-12-30 fix warmup order. should be rank!=0 then rank=0 -
fea424732025-12-28 Place local node at the top of the dashboard. (#1033) -
ca7adcc22025-12-28 Update README.md with instructions to enable RDMA. (#1031) -
9d9e24f92025-12-28 some dashboard updates (#1017) -
b5d424b62025-12-23 placement: generate per-node host lists for MLX ring backend -
b46513402025-12-28 Fix Kimi K2 Thinking download by adding tiktoken.model to download patterns (#1024) -
eabdcab92025-12-27 Fix linux docs (#1022) -
8e9332d62025-12-27 Separate out the Runner's behaviour into a "connect" phase and a "load" phase (#1006) -
4b65d5f82025-12-27 Fix race condition in mlx_distributed_init with concurrent instances (#1012) -
1c1792f52025-12-23 mlx: update to 0.30.1 and align coordinator naming with MLX conventions -
9afc10432025-12-23 exo: handle -c flag for multiprocessing helpers in frozen apps -
70c423f52025-12-23 feat: conform to XDG Base Directory Specification on Linux (#988) -
a24bdf762025-12-22 exo: enable multiprocessing support in PyInstaller bundles -
e88559592025-12-22 build-app: add branch trigger from named branch -
0a7fe5d92025-12-22 ci: migrate build-app to github hosted runners -
51a5191f2025-12-22 format readme (#978) -
1efbd2632025-12-22 add architecture.md, move images to docs/imgs (#968) -
02c915a82025-12-22 pyproject: drop pathlib dependency -
fc41bfa12025-12-22 Add all prerequisites to README (#975) -
dd0638b72025-12-22 pyproject: add pyinstaller to dev-dependencies -
e06830ce2025-12-22 fix: update macOS app to use correct API port (52415) -
1df5079b2025-12-22 ci: avoid pushing alpha build as latest -
1e75aeb22025-12-22 Add Prerequisites to Readme (#936) -
c582bdd62025-12-20 bugfix: Handle MacMon errors gracefully -
1bae8ebb2025-12-18 ci: add build-app workflow -
abaeb0322025-12-21 Update README.md. (#956) -
7d15fbda2025-12-21 readme tweaks5 (#954) -
4a6e0fe12025-12-21 Update README.md. (#949) -
f4792dce2025-12-21 fix(downloads): use certifi for robust SSL certificate verification (#941) -
a1b14a272025-12-20 Extend eos_token_id fix for other models (#938) -
f8483cfc2025-12-19 Update README.md. (#932) -
8bafd6fe2025-12-19 Update README.md (#925) -
f16afd722025-12-19 nix: get rust build working on linux -
4da004322025-12-18 Update README.md (#917) -
9e2bdeef2025-12-18 LICENSE: Fix company name/year -
379744fe2025-12-18 exo: open source mac app and build process -
74bae3ba2025-12-18 Update README.md -
9815283a2025-12-18 8000 -> 52415 (#915) -
658cf5cc2025-12-18 remove tb_only from master -
170d2dcb2025-12-18 Add Windows as a potential planned platform -
274e35f92025-12-18 update readme -
3fe7bd252025-12-18 update error message -
004fea692025-12-18 clarify platform support -
5c2d254f2025-12-18 add platform support information -
19ca48c42025-12-18 more readme fixups -
57d381362025-12-18 re-add LICENSE -
7cd1527c2025-12-18 update CONTRIBUTING -
ebf0e18c2025-12-18 re-add logos -
28a6151b2025-12-18 remove discord link from README -
2c16e00b2025-12-18 github docs -
0fcee7082025-12-17 prep repo for v1 -
09593c5e2025-12-17 backport the dashboard to staging -
880a18d22025-12-15 fix disconnects -
70298ce02025-12-09 Negative index nack request -
ac3a0a6b2025-12-09 ci: enable `ruff check` in CI through nix -
859233a22025-12-09 Reduce RequestEventLog spam -
c9e2062f2025-12-05 switch from uvicorn to hypercorn -
e8566a3f2025-12-05 placement: pass different ibv_coordinator per node -
39d76aa02025-12-05 nix: move formatting checks to nix and enable in ci -
562998382025-12-05 fmt: format all python/rust/nix files -
7312a7e02025-12-05 plan fix -
9e0a1c232025-12-05 rename ibv to jaccl inline with mlx -
f5783d642025-12-05 proper collection of rdma ports in placement -
e702313b2025-12-05 pingers -
a3f8ecba2025-12-05 prioritise LL4 -
5ef1df1e2025-12-05 rust: move Cargo.toml to the root -
40a0d47d2025-12-03 jaccl -
2b243bd82025-12-03 Consolidate!!! Fixes -
10c905c82025-12-02 worker no longer gets stuck after shutdown -
93f699b62025-11-28 add aarch64-linux for the spark -
b43d30562025-11-27 todo for layer-independent parameters in get_allow_patterns -
20d73e902025-11-26 fix dashboard case sensitive model id -
e56daa7c2025-11-26 render download progress properly -
63c85e122025-11-25 get rid of spammy Finished tokenizing log -
7088988a2025-11-25 bump pyo3 stub-gen -
7b3e3fd62025-11-21 Worker tests 2 -
de5081132025-11-21 Worker tests on staging 1 -
b45cbdee2025-11-21 Consolidate cleanup -
28a917872025-11-20 Demo -
d793f5f92025-11-13 fix kimi eos token ids -
b62f68472025-11-11 improved master error handling -
631cb8102025-11-11 kimi k2 thinking -
364087b92025-11-11 five billion percent better shutdown handling -
aa519b8c2025-11-10 Worker refactor -
9058b1172025-11-07 pipeline parallel fix -
612f58c72025-11-06 Revert dumb merge mistake -
6bcac37d2025-11-06 stop benching on all pushes -
ff00b1652025-11-06 MLX LM type stubs -
19e905722025-11-06 set max_transmit_size on gossipsub to 1MB. Fixes large message erorr -
e60681962025-11-06 show ips on dashboard -
0bb621b62025-11-06 Add mlx nn stubs -
699fd9592025-11-05 fix exo scripts -
6bbb63442025-11-05 mlx.distributed.Group type stubs -
16f724e22025-11-04 Update staging 14 -
3b4096472025-10-31 Squash merge merging_clusters into tensor_parallel94 -
d46c7e6a2025-10-31 fix race condition with downloads where it cancels the download before renaming -
91c635ca2025-10-31 Update mlx and mlx-lm packages -
5f18faec2025-10-30 Update. -
a346af342025-10-22 download fixes -
56f783b32025-10-21 Update. -
363c98a82025-10-15 leaf placement -
f25689d92025-10-15 fix a race condition -
1c6b5ce92025-10-10 new tagged union -
76ed8a512025-10-10 typecheck on ubuntu with install-nix-action -
e8a6efe22025-10-07 add kimi k2 -
a4e833522025-10-07 add just clean -
84dfc8a72025-10-07 Fast memory profiling -
e01f9cf72025-10-07 Disable build macos app -
35ab6b372025-10-07 fix: master tests -
962e5ef42025-10-07 version bump for brew consistency -
b1721e942025-10-01 nix cleanup -
22f0ca2a2025-09-30 FIX: OpenWebUI compat -
57486a432025-09-30 kill go -
38ff949b2025-09-30 big refactor -
7040c9502025-09-17 Multiprocessing Runner -
35c431152025-08-29 Dashboard Status & Bugfixes -
a33787f52025-08-29 Prompt length -
1b8b456c2025-08-26 full mlx caching implementation -
84c90a6d2025-08-26 feat: mlx memory cache for faster ttft -
5efe55622025-08-26 feat: single entrypoint and logging rework -
ef5c5b962025-08-25 changes include: ipc, general utilities, flakes stuff w/ just, autopull script -
5bfc99b42025-08-25 add EXO logo to dashboard -
11f8b4ef2025-08-21 tidy: fix justfile, run.sh, run formatter -
be6f5ae72025-08-21 feat: build system and homebrew compatibility -
40efed442025-08-20 unvendored macmon -
ea9e57342025-08-18 Refactor runner supervisor -
345fafd82025-08-18 Forwarder versioning -
ea3eeea82025-08-15 improved go caching with nix -
a2a37c0e2025-08-15 discovery fixed -
57073f352025-08-15 collection of fixes for Shanghai demo -
7e19804a2025-08-13 Integrate flake parts -
dbcd09aa2025-08-12 No 70b -
c1d5b3812025-08-07 70B model unit test only runs if its downloaded -
473512dd2025-08-04 r1 size -
817c59932025-08-04 fix dem model cards yo -
75ecda552025-08-04 fix gitignore -
c560c55c2025-08-04 build and release on staging -
f51f8f722025-08-04 app launches python modules -
407796d12025-08-04 Minor dashboard fixes -
6daf7f312025-08-04 clean model cards -
f352ddfc2025-08-04 run configure_mlx.sh in run.sh -
6855a7722025-08-03 set a 15 sec timeout for getting initial download progress -
1fe4ed342025-08-02 Worker Exception & Timeout Refactor -
92c9688b2025-08-02 Remove rust -
a46f8c3c2025-08-02 app -
71bafabc2025-08-01 Dashboard with instances -
0e32599e2025-07-31 fix libp2p + other prs that were wrongly overwritten before (111,112,117,118,1119 + misc commits from Alex) -
2031d9482025-07-30 fix api get_state -
b350eded2025-07-30 Test Supervisor Errors. -
ff3d11c72025-07-29 just run -
25fa46c62025-07-29 Update CODEOWNERS -
3f192f202025-07-28 Reinstate dashboard -
a2b4093d2025-07-28 add metrics: gpu_usage, temp, sys_power, pcpu_usage, ecpu_usage, ane_… -
125668652025-07-28 better profiling -
b88abf1c2025-07-28 fix topology disconnects and add heartbeat -
dbd0bdc32025-07-28 fix ci linter -
20241e322025-07-28 some finishing touches to get this working e2e -
176d077c2025-07-28 Fix IPv4 serialisation for topology -
c3c8ddbc2025-07-28 fix forwarder supervisor tests -
36a5d75e2025-07-28 Fix download tests -
e9b803602025-07-28 Add Multiaddr type and refactor Hosts type for creating shard placement -
b285a9f02025-07-28 fix placement tests -
57ca487f2025-07-28 Fixes for running this end to end -
b687dec62025-07-27 Discovery integration master -
98f204d12025-07-26 Fix placement single node -
93330f022025-07-26 Inference Integration Test -
2e4635a82025-07-26 add node started event -
261e57522025-07-25 Serialize topology -
a97fb27c2025-07-25 Glue TWO -
9be08ec72025-07-25 add resource monitor -
a241c92d2025-07-25 Glue -
6f8e34192025-07-24 Placement strategy -
4c0e4ef82025-07-24 Go build -
f41531d92025-07-24 Worker Loop -
67c70b222025-07-24 Best master -
373016042025-07-24 Fix the node-ID test -
df1fe3af2025-07-24 Topology apply -
5097493a2025-07-24 Fix tests -
a6b3ab632025-07-24 Worker plan -
56d356572025-07-24 Add apply functions -
3ab560922025-07-23 wrote race-condition-free persistent NodeID-getting function -
7a452c332025-07-23 Fix tests -
7ac23ce92025-07-23 Refactor tasks / commands / api -
81060b702025-07-23 Made basedpyright work with Jetbrains environment -
8d2536d92025-07-23 Implemented basic discovery library in Rust + python bindings -
76f903502025-07-22 fix -
cd9a1a912025-07-22 Topology update -
14b3c4a62025-07-22 New API! -
596d9fc92025-07-22 add forwarder service -
53c652c32025-07-22 Fix tests! -
5adad08e2025-07-22 New events -
108128b62025-07-21 fix sqlite connector -
449fdac22025-07-21 Downloads -
cb101e3d2025-07-21 Refactor model types -
54efd01d2025-07-21 add forwarder supervisor -
bae58dd32025-07-21 Refactor worker + master state into single state -
d19aa4f92025-07-21 Simplify `Task` type + merge control & data plane types into single type -
2f64e30d2025-07-21 Add sqlite connector -
bb7f1ae92025-07-18 New worker -
cc45c7e92025-07-17 Fixed events issue. -
038cc4cd2025-07-16 fix: Normalize Naming -
e2a793502025-07-16 fix: Fix incorrect logic -
6a6719082025-07-16 fix: FrozenSet Related Bits -
520b11222025-07-16 fix: Many Fixes -
7fa7de8e2025-07-15 more incomplete trash -
9f96b6792025-07-15 fix: Some, still broken -
9b3c105b2025-07-15 fix: Save Andrei's sanity -
806012012025-07-14 tweak -
df6626fa2025-07-14 fix: Event definitions, state definitions -
70f0f09c2025-07-14 Tweaked, Still Broken tho -
8799c2882025-07-14 BROKEN: work thus far -
4e4dbf522025-07-14 fix: Use Nix-compatible LSP set-up -
21acd3792025-07-10 New Runner! -
b0bd95102025-07-09 Merge Basic Interfaces -
74d56e522025-07-07 fix: Improve naming -
fe17aaf92025-07-07 fix: Make master hold a queue of task data -
e1894bc12025-07-07 refactor: A Lot -
81cf6bce2025-07-07 refactor: Simplify networking -
6c8b8b302025-07-07 added rust to flake -
0425422f2025-07-07 Simple fix -
03a1cf592025-07-07 Matt's interfaces -
367e76c82025-07-04 fix: Fix validation over Task types -
cda3de2a2025-07-04 fix: Use state for tasks -
10224d092025-07-03 refactor: Distinguish the topology of the control plane from that of the data plane -
c45693432025-07-03 refactor: Remove timestamp from Wrapped Events -
0b6aadf52025-07-03 refactor: Add safe state mutation method .apply() -
f8039e202025-07-03 feature: Add pretty_name to ModelMetadata -
4bb3a9952025-07-02 feature: Interfaces for graph interfaces -
7dd8a9792025-07-02 feature: Simplest utilities for logging -
40793f1d2025-07-02 refactor: Refactor most things -
8596d5c52025-07-02 refactor: Fix UUID implementation -
6de1f2882025-07-01 feat: Update Interfaces -
73ac89692025-07-01 feat: Add ResourceGraph, runner types, etc. -
df824e2e2025-07-01 fix: Ensure MasterState inherits from SharedState -
d5033e652025-07-01 refactor: Replace Literal with Enum in sources.py -
c0df8e542025-07-01 feat: Implement Many Interfaces -
899d88202025-06-30 Merge Seth's Control Plane API Work into Alex's Events Branch -
53d5d2382025-06-30 refactor: Use enums -
b758df832025-06-30 Chore: Tweak CI -
133ab70d2025-06-30 chore: Run formatter -
aae3e4a82025-06-30 refactor: Put type defs on one line -
596b069f2025-06-30 chore: Fail pipeline if working tree changes instead of committing them in CI -
c0b8bb9c2025-06-29 chore: Rename conditional-commit.yml to action.yml -
0c46adc22025-06-29 refactor: Use official OpenAI types -
4b3e60f82025-06-29 refactor: Add types for model downloading -
784f0ec42025-06-29 chore: Skip protobuf generation if no .proto files exist -
38dcf6982025-06-29 chore: Fix typecheck job in GitHub workflow -
c9d44a162025-06-29 chore: Fix typecheck job in GitHub workflow -
bbdfdac72025-06-29 refactor: Remove redundant comment -
5ba230ed2025-06-29 refactor: Add all event types with Event implementations -
5abf03e32025-06-29 Scaffold Event Sourcing -
d84593582025-06-28 Refactor CI -
c977ce942025-06-28 Ensure `exo-shared` is a Dependency of `exo-master` and `exo-worker` -
74adbc422025-06-28 Remove PoeThePoet -
587a52a92025-06-28 Remove Bad UUID Implementation -
885c7d5c2025-06-28 Add RULES.md and .cursorrules -
e4c4b3e92025-06-28 Overhaul CI Design -
f7f779da2025-06-28 Fix Type Checker; Improve Protobuf Generation -
38bc8ea72025-06-28 Keep Protobuf Directories -
b53c1ba92025-06-28 Use Hatch Build System -
423efe102025-06-28 Add Protobuf Support -
61b8b1cb2025-06-28 Add Protobuf Support -
7f0f71b92025-06-28 Add .gitignore -
da50da2b2025-06-27 Add Simple env.py -
3564d77e2025-06-27 Add Sync to Runner -
77546b952025-06-17 Update pyproject.toml -
c15e402f2025-06-17 Add Simple Groundwork -
c57ed32f2025-06-17 Add Initial Contribution Rules -
41085eef2025-06-17 Prepare Environment Parser -
685c8eff2025-06-17 Configure Runner Tasks to Cover "engines/" -
13b6043c2025-06-17 Add Linter -
180748ee2025-06-17 Update Workspace Configuration, Configure Build Backend -
043253a52025-06-17 Add ML Engines (Backend) -
090265a32025-06-17 Add Formatter To CI -
e2508f342025-06-17 Add Type Checker In CI -
ac2dfa652025-06-17 Initial Structure -
db1a52522025-06-14 Add CODEOWNERS. -
ad3bc6ce2025-03-21 downgrade grpcio, grpcio-tools to 1.70.0 -
50b6800a2025-03-11 m3 ultra flops estimates based on some quick profiling -
2857975b2025-03-11 upgrade grpcio and grpcio-tools to 1.71.0 -
f98d9bac2025-03-05 Changes required to detect AMD GPUs -
013d25732025-03-02 remove dead links in README -
30c3f58a2025-02-28 downgrade grpc to 1.67.0. waiting for fix https://github.com/grpc/grpc/commit/bd8f8a86e06df9d5cf6bc36ba77981d621b80b44 -
4081305e2025-02-28 adjust grpc settings, ensure connected before sending any grpc commands -
971f52402025-02-28 build fix -
36a6389a2025-02-27 bump grpcio and grpcio-tools to 1.70.0 -
ee0957662025-02-25 handle -gzip suffix in etag for integrity check fixes #633 -
f9a1e5342025-02-18 update notice in README -
cb4bee262025-02-17 add notice to README -
9078d0942025-02-16 adding current model name to input container information -
477e3a5e2025-02-14 make max_parallel_downloads configurable, increase download chunk size to 8MB -
b4e6f8ac2025-02-13 always log download errors. some people eg cant access huggingface which causes confusion -
928214d42025-02-08 apt-get debian noninteractive in circleci -
d8c3aed02025-02-08 update discovery / peer networking modules -
2c982d922025-02-08 update README to better reflect support for other devices like NVIDIA and Pi's -
5fe241ec2025-02-06 code-breaking typo -
05ff20fa2025-02-06 workaround f16 cast ambiguity -
5157d80a2025-02-03 remove tenacity dependency, implement simple retry logic instead -
d084dbe52025-02-01 Add toggle to show only models downloaded locally -
72329ba92025-02-01 patch for manual discovery, set known_peers -
51b5c2ca2025-02-01 add model downloading section to README -
2c0d17c32025-02-01 beautiful download -
7034ee0f2025-02-01 resumable downloads with integrity checks -
0bebf8df2025-01-30 fix indent -
55c4385d2025-01-30 cleanup tmp files on failed download -
788c49782025-01-30 retry fetch_file_list also -
6b1c86352025-01-30 ensure exo dir on start, retry with exp backoff on file downloads -
e6b4f2992025-01-29 fix prompt output spacing in tui -
a25e02c92025-01-29 Add 4-bit to the end of DeepSeek V3/R1 model descriptions -
3675804f2025-01-29 throttle repo progress events and only send them out if something changed -
96f1aecb2025-01-29 only in_progress if any given file is in_progress -
23a503062025-01-29 even if part of a file is downloaded it may not be in_progress -
31b56e862025-01-29 make a singleton thread pool executor for tinygrad since we always want it to run on the same thread -
9f6c688d2025-01-29 update tinygrad -
4887be512025-01-29 parallelise model loading -
141de0d02025-01-29 increase chatgpt api response timeout to 900 seconds -
9cf6818f2025-01-28 Fix AMD device capabilities fields -
9c1bea972025-01-28 fix embed_tokens for last layer in qwen models -
af171f062025-01-28 propagate prompts to other nodes so they can display them, cleaner prompt/output output -
4a5b80a92025-01-28 make sure mlx stuff is on separate thread non blocking -
6662d5662025-01-28 load mlx model shard on mlx thread so it doesnt block -
7c6490852025-01-27 fix eta/speed for resuming an existing download, using the session downloaded bytes -
90e0e2762025-01-27 ignore not_started progress updates -
265586f72025-01-27 set timeout on get too -
4748bb7d2025-01-27 increase file download timeout to 30min -
ae770db42025-01-27 increase download chunks to 1MB -
82f75d0c2025-01-27 increase hf download http timeout 15 mins for large downloads -
295f41c52025-01-27 increase bench job timeout to give enough time to download -
19a27c5b2025-01-27 HF_HOME -> EXO_HOME -
d7ca9b772025-01-27 show each node id in the tinychat topology viz -
b349e48b2025-01-27 fix visual bug where frontend would show the full hf repo size, but in some cases that includes redundant files so we should use the model index in those cases too -
215860632025-01-27 use llama-3.2-1b in tinygrad test -
277d63d82025-01-27 special case when a model doesnt have a model index file, then use wildcard for allow_patterns -
74379ef62025-01-27 log download logs with DEBUG>=6 very verbose -
3c7bd48a2025-01-27 get rid of some more hf bloat -
1df023022025-01-27 remove a lot of hf bloat -
b89495f42025-01-27 rewrite ShardDownloader, simplify significantly -
a3766f532025-01-26 add exception for mlx-community/DeepSeek-R1-3bit and mlx-community/DeepSeek-V3-3bit in tokenizers test -
82ef08602025-01-26 add deepseek-v3-3bit and deepseek-r1-3bit -
55ea36692025-01-26 fix post_init deepseek v3 -
fb841a1f2025-01-26 Adjust truncate size in history list for text without any spaces -
451236652025-01-26 Fix bubble behavior when user passes long text without any spaces -
9525c0e72025-01-26 Add adaptive padding for user and assistant messages on width <= 1480px -
fdd05bad2025-01-24 fix tokenizer tests -
59174bdc2025-01-24 we have a lot of models so group them nicely -
cfdaaef82025-01-24 handle thinking outputs nicely, format latex beautifully -
d8ffa59d2025-01-24 add deepseek v1, v3 and all the distills -
4fb01f512025-01-24 chore: update manual_discovery.py -
ad0e0d022025-01-23 fix readme images -
88ac12df2025-01-23 install clang test -
dfd9d3eb2025-01-23 linux install -
200ff4d72025-01-23 linux install -
b2764f172025-01-23 linux install -
e57fa1df2025-01-23 xlarge -
209163c52025-01-23 add linux tinygrad test -
495987b52025-01-23 beef up the instance -
8484eb412025-01-23 fix config -
790c08af2025-01-23 add linux tinygrad test -
a8a9e3ff2025-01-23 explicitly enable TOKENIZERS_PARALLELISM=true -
5c9bcb862025-01-23 set GRPC_VERBOSITY=error; TRANSFORMERS_VERBOSITY=error -
d54e19c22025-01-23 runners back -
cc78738e2025-01-23 remove kern scan intervals -
2391051c2025-01-23 remove kern.timer.scan_interval from bootstrap.sh -
112dea152025-01-23 add back the benchmarks baby -
dc5cdc4d2025-01-22 add back opaque -
f8db4e132025-01-22 fix check for sd2.1 -
bbb685692025-01-22 fix check for sd2.1 -
9ba8bbbc2025-01-22 fix filter to include 169.254.* since thats what mac uses for ethernet -
8ab9977f2025-01-22 fix stable diffusion case for tui, make mlx run on its own thread again and non-blocking -
3a4bae0d2025-01-22 fix issue with eos_token_id -
87d1271d2025-01-22 fix stream: false completion -
55d1846f2025-01-22 clean up DEBUG=2 logs, a few fixes for token -
9954ce8e2025-01-22 fix treating token as a list -
09e12d862025-01-22 temporarily disable github runner benchmarks -
98d6e9862025-01-22 add back .circleci -
d80324fe2025-01-22 disable test-m3-single-node -
97f3bad32025-01-22 fix peer_handle -
27b4577f2025-01-22 directory for images -
a70943f82025-01-22 base images for animation -
5c4ce5392025-01-21 image and text mode fix -
ba5bb3e12025-01-21 fix scripts/build_exo.py: com.exolabs.exo -> net.exolabs.exo -
6b8cd0572025-01-20 fix some issues with results -
b9eccedc2025-01-17 Formatting -
5f06aa272025-01-17 Replace netifaces (unmaintained,outdated) with scapy + add dependencies for previous fixes -
349b53442025-01-16 Minor fix for Shard typing -
df3624d22025-01-14 Add AMD GPU querying + Windows device capabilities -
6737e36e2025-01-14 Fixed MLX import blocking native Windows execution of exo. (Not Final) -
fcc699a52025-01-12 fix -
e7b98f5a2025-01-12 fix unit tests -
ffe78f6d2025-01-12 fix dummy test -
ce5041ee2025-01-12 types -
9b2c01c82025-01-12 ensure dir exists -
2aed3f352025-01-12 handle inference_state properly -
2af5ee022025-01-12 fix exo folder -
40696b212025-01-08 typo in phi test -
2846a9122025-01-08 tok tests -
553ccce72025-01-08 fix prompt and output overflow in tui -
c58759332025-01-08 add phi 3.5, phi 4 -
627bfcae2025-01-06 Fix the /v1/models API to output proper OpenAI compatible endpoint -
29244c632025-01-05 fix args for ensure_shard -
8c1910502025-01-05 download status in parallel, support async ensure shard with using shard_downloader instead -
fe50d4d32025-01-03 Add --system-prompt to exo cli -
178cc4d92024-12-31 add trending badge to README.md -
b13e36832024-12-30 fix inference engine -
9986fb862024-12-30 remove prints and fix download progress for SD -
3475be9e2024-12-30 Remove build -
fff8a1a62024-12-30 fix inference engine for inference state -
b003292b2024-12-28 formatting and fixing tests after rebasing -
1dfd058c2024-11-28 rm unecessary lock -
2eadaa2c2024-11-25 rm redundant cleanup task -
637446ff2024-11-15 rm redundant typing -
a31f9e6c2024-11-15 fix test warnings -
18acb97b2024-11-15 make popping from dict threadsafe -
b066c9442024-11-15 make all I/O ops in manual_discovery.py run inside a ThreadPoolExecutor -
0e34ce212024-11-06 patch after rebasing to main -
90de7ead2024-11-06 changes after rebase -
8d24df2b2024-10-24 fix test runtime warning -
e5eb32592024-10-24 handle when a peer is removed from config, so the known_peers dict gets updated accordingly -
2e8227fc2024-10-24 handle intermediate state for when config is being updated -
98118bab2024-10-24 allow update to manual discovery file -
e08522ee2024-12-27 Revert "Merge pull request #573 from damho1104/feature/add-exaone-3.5-model" -
94a5e9082024-12-24 add exaone-3.5 LLM Model -
185b1e372024-12-24 fix names in dummy tokenizer -
078b80762024-12-24 fix names of qwen models -
188ac4452024-12-24 function calling example with weather tool -
456fbdd22024-12-24 add chatgpt-api-compatible tools for function calling -
c609c05e2024-12-24 add qwen-2.5-1.5b, qwen-2.5-3b, qwen-2.5-32b -
cde912de2024-12-22 - Use `#!/usr/bin/env bash` instead of `#!/bin/bash` for better portability -
154e0f582024-12-21 Implement suggestiond -
6c82365e2024-12-17 Improved clarity, fixed typos, added macOS/Linux examples, and enhanced installation/debugging instructions -
023ddc202024-12-17 support different network interface tests -
2f0b543a2024-12-17 add peer connection info to tinychat -
7ac400432024-12-17 change it back to collecting topology periodically even if peers dont change -
198308b12024-12-17 more robust udp broadcast -
1f108a062024-12-17 remove test sleep -
3a58576f2024-12-17 make sure this is actually doing something -
0a0722302024-12-17 switch to uvloop (faster asyncio event loop) and optimise grpc settings -
58f0a0f52024-12-17 optimise grpc parameters -
5c0cd1832024-12-16 Update strength image to image gen -
e2474c3f2024-12-16 fail if we never get the desired node count -
1b14be602024-12-16 make device_capabilities async running on a thread pool -
036224f82024-12-16 add topology to tinychat ui -
b17faa812024-12-16 dont broadcast every single process_tensor -
8d94b8ae2024-12-16 trigger test -
063964aa2024-12-16 remove redundant sample_logits, put back opaque status for process_prompt so we have a way of preemptively starting downloads -
804ad4702024-12-16 upgrade mlx -
c9ded9ba2024-12-16 optimise networking, remove bloat -
64365d682024-12-15 one two and three m4 pro clusters -
9397464f2024-12-15 add commit to results -
08912d1b2024-12-15 Only collect topology if peers changed -
06c2e2362024-12-14 rip out stats bloat -
cb4615c92024-12-14 fix SendNewToken -
f55a53ae2024-12-14 one token at a time -
470f961f2024-12-15 Only collect topology if peers changed -
a93092102024-12-14 set max-generate-tokens to 250 -
0c6ab3532024-12-14 increase timeout of http request in bench.py up to 10 mins -
b0e079b32024-12-13 fix counts in testmodelhelpers -
e5d54c772024-12-12 add llama-3.3-70b to 3 M4 Pro cluster -
a0bada3b2024-12-12 add llama-3.2-1b-8bit, llama-3.2-3b-8bit, llama-3.2-3b-bf16 -
b6f2385c2024-12-12 run llama-3.1-8b on 3 m4 pro cluster -
9472ab0d2024-12-12 t -
dbb7ad3c2024-12-12 run with three m4 pro -
2abe57be2024-12-12 grasping at straws -
eeecdcb42024-12-12 try a different taskpolicy -
f9f761292024-12-12 better bench system info -
8c6d37d92024-12-12 m4 cluster test -
1194db6e2024-12-12 m3 -
8cb7327d2024-12-12 re-enable m4 cluster run -
bba0aa082024-12-11 single node test 20 -
279354a12024-12-11 single node test 19 -
92e2b7492024-12-11 single node test 18 -
76196b8c2024-12-11 single node test 17 -
8408c8492024-12-11 single node test 16 -
c65d1d912024-12-11 single node test 15 -
0bd44c0f2024-12-11 single node test 14 -
f22bc99f2024-12-11 single node test 13 -
3fda05aa2024-12-11 single node test 12 -
6c322ac02024-12-11 single node test 11 -
c5c27a322024-12-11 single node test 10 -
9f1393dc2024-12-11 single node test 9 -
32ff3ef92024-12-11 single node test 8 -
b23c3fda2024-12-11 single node test 7 -
8b47a9d02024-12-11 single node test 6 -
f89b85b32024-12-11 single node test 5 -
6f097c932024-12-11 single node test 4 -
fb7a0def2024-12-11 single node test 3 -
fe506a532024-12-11 single node test 2 -
3f6ef1c72024-12-11 single node test 1 -
e63c224c2024-12-11 testtt -
20e3065e2024-12-11 les goh -
83892d5b2024-12-11 t -
83470a982024-12-11 t -
92edfa5e2024-12-11 t -
225dcba72024-12-11 t -
6249bee72024-12-11 tes -
741c31832024-12-11 test -
d0b7f1b42024-12-11 t -
906774152024-12-11 t -
6cf2af392024-12-11 t -
5a1a0f5f2024-12-11 t -
dd3fd2792024-12-11 t -
61c096312024-12-11 t -
e698ef6a2024-12-11 t -
26351e712024-12-11 t -
5dee5e552024-12-11 t -
6acfb8182024-12-11 t -
b1142d4f2024-12-11 t -
a932afc02024-12-11 oi -
cdae70262024-12-11 t -
d95f40b62024-12-11 a -
97ffb83e2024-12-11 t -
9a11e27c2024-12-11 ttt -
d6c2146d2024-12-11 t -
63da9fc12024-12-11 a -
7c0c5ef72024-12-11 ttttttt -
739b7d172024-12-11 tttttt -
cacf50cd2024-12-11 tttt -
0904cda32024-12-11 ttt -
6bb389392024-12-11 tt -
1dbe11ca2024-12-11 t -
8d9e3b882024-12-11 t -
9dd33d372024-12-11 t -
a4bb4bb62024-12-11 update bootstrap -
7b99cb4a2024-12-11 t -
9848a45d2024-12-11 TT -
378975812024-12-11 t -
e680e8a12024-12-11 fix name -
7b2282d32024-12-11 run without debug flag -
3b1ea1932024-12-11 use .venv exo -
668766fc2024-12-11 t -
e501eeaf2024-12-11 tweak install -
41902f712024-12-11 tweaks -
b7bab80e2024-12-11 test2 -
6169996c2024-12-11 test -
bbb584602024-12-11 Test on m4 -
cff03fc62024-12-11 perf diag -
f7122d402024-12-11 add system_status check to bench -
c938efb52024-12-11 t -
e2d3a9082024-12-11 runner-token typo -
ba96413a2024-12-11 bootstrap script tweaks -
cb40eb232024-12-11 more robust configure_mlx.sh -
afe71c012024-12-11 check gpu usage -
23158a422024-12-11 add branch name to results -
18e791992024-12-11 test 30 -
0e32a6252024-12-11 test 29 -
04bc163f2024-12-11 test 28 -
949055de2024-12-11 test 27 -
070b163c2024-12-11 test 26 -
fc26ad402024-12-11 test 25 -
5d3be3c62024-12-11 test 24 -
23dd5de32024-12-11 test 23 -
6030b3992024-12-11 test 22 -
4f4ac0fa2024-12-11 test 21 -
16d983902024-12-11 test {i} -
8269b4b12024-12-11 t -
329efb232024-12-11 Model loading and saving for tinygrad -
b1397b492024-12-11 Proper sharding in tinygrad -
7f0c12a92024-12-11 embed fix -
bd3114452024-12-10 Dummied up an abstact save_checkpoint -
cc66a0b72024-12-10 Missed one -
124a03382024-12-10 Slightly simplified waiting for outstanding requests -
a4313da82024-12-08 Removed statefulModel stuff from mlx impl too -
0673d6452024-12-08 Removed ensure_session to clean stuff up. May revisit later -
6aaea8c72024-12-08 Abstract load checkpoint method -
2a3a2e5e2024-12-08 circular include lol -
0c5762d12024-12-11 Node rename -
c2332e242024-12-11 Moved nodes around -
763fbf842024-12-11 Updated node refs -
59af2dd52024-12-08 Do we need casting here? -
b22c21ac2024-12-08 Some session method cleanup -
98edb3932024-12-08 Initialize inference engine session in base class -
bcf87e792024-12-06 Okay let's turn no_grad back on. We'll worry about that when tinygrad training works -
b7bbda332024-12-06 Removed tinygrad StatefulModel class, as it's no longer used -
67f5ae252024-12-06 Fixing tinygrad model -
bfa3b36b2024-12-06 Fixing tinygrad model -
37a75d6b2024-12-06 Fixing tinygrad model -
0d3abfca2024-11-21 Made models save properly -
9283f6d72024-12-06 Correct loss propagation so we can see the actual loss instead of just the requestor shard's loss -
9eadee312024-11-20 Basic model saving -
38e368f02024-11-20 Fixed up the ops so that batches work -
dd3d99042024-11-26 Working distributed training -
175ebc1c2024-11-21 Coordination biz -
3e8690512024-11-19 Okay we should probably await the update -
75c8650f2024-12-06 Naive network-propagated loss implementation on MLX -
836856822024-12-06 WIP: Training works on mlx -
a6fd7a342024-11-26 Generalizing some of the dataset biz while also creating uniform batches -
f5efbe1b2024-12-06 Initial distributed evaluation implementation -
1e869a0f2024-12-10 trigger test -
5a4d128d2024-12-09 trigger test -
8a5d212c2024-12-08 test 20 -
53edb8502024-12-08 test 19 -
29d9df042024-12-08 test 18 -
4d6af6e62024-12-08 test 17 -
8c7c156f2024-12-08 test 16 -
310843482024-12-08 test 15 -
a4b221d02024-12-08 test 14 -
286db8752024-12-08 test 13 -
d714e40f2024-12-08 test 12 -
e78ef7552024-12-08 test 11 -
38eaecf02024-12-08 test 10 -
3cf28f842024-12-08 test 9 -
9ba8bbdd2024-12-08 test 8 -
af6048e32024-12-08 test 7 -
d93b8e892024-12-08 test 6 -
b69cb49a2024-12-08 test 5 -
cc74b1f92024-12-08 test 4 -
e78a52de2024-12-08 test 3 -
f6c2c37c2024-12-08 test 2 -
314a5d972024-12-08 test 1 -
b4e885bb2024-12-08 test range -
bd9d11862024-12-08 sleep before bench -
571b26c52024-12-08 allowed interface types -
b21681932024-12-08 remove -
f584e86d2024-12-08 get rid of lfs stuff -
fd05bca12024-12-08 lfs -
cbac4d6a2024-12-08 git version -
b0977f972024-12-08 t -
1716f6372024-12-08 test -
903a5aab2024-12-08 fix -
b4f864962024-12-08 bootstrap -
8e57f3382024-12-08 trigger test -
3ccbdf192024-12-08 add DEBUG_DISCOVERY -
3687ba182024-12-08 bench logs -
6bb7c11b2024-12-08 enable debug -
c8f937212024-12-08 model matrix -
fb8d87022024-12-08 t -
87865f0c2024-12-08 list exo processes before test, warmup req in bench -
755dd4772024-12-08 jobname -
fb44eb082024-12-08 simplify bench -
be8cbc0f2024-12-08 trigger test -
fe8074922024-12-08 fix -
c3c80c612024-12-08 name -
c138de082024-12-08 job_name -
38bd00392024-12-08 fix -
732ba9152024-12-08 new_conf -
785710352024-12-07 aws -
320892dc2024-12-07 maxtok -
6dae3a472024-12-07 conf -
7b77ef002024-12-06 flush -
6c08b3232024-12-06 nodebug -
4dd617ad2024-12-06 shorter -
acdee16a2024-12-06 debug -
9fc335872024-12-06 path -
f087c0ac2024-12-06 fix -
16b126d82024-12-06 fix -
faf0aaed2024-12-06 jq -
4cac1bb12024-12-06 quotes -
cb3c14772024-12-06 fix -
19a7d5a52024-12-06 fix -
f7e0348f2024-12-06 activate -
c3dfac602024-12-06 debug -
64954aac2024-12-06 fixed -
ccc5415c2024-12-06 try -
1dcc731b2024-12-06 fix -
3662ec402024-12-06 fix -
0739dc952024-12-06 fix -
d16280dd2024-12-06 debug -
f9c236172024-12-06 fix3 -
ce2ccddc2024-12-06 fix2 -
1af28cb52024-12-06 fix -
6b61fc662024-12-06 tweak python install -
bdf417f22024-12-06 tweak -
d154d37a2024-12-06 add exo run -
90fd5c132024-12-06 matrix -
7d223a002024-12-06 matrix -
cb3d89eb2024-12-06 test runner -
8302fd0a2024-12-06 test runner -
deb80d252024-12-06 clang for tinygrad -
976e5f2f2024-12-06 disable mlx test for now..plan to run this on a self-hosted runner -
9dc76ef02024-12-06 tooonygrad -
32cd1f1d2024-12-06 give this a goh -
6b5418812024-12-06 cond -
58bcf5b42024-12-06 check discovery on integration tests too -
3c0297c32024-12-06 more robust discovery log check -
8d433e652024-12-06 run tinygrad and discovery integratrion tests on linux -
676125bf2024-12-06 job -
902e0d352024-12-06 github env vars -
972aea442024-12-06 macos 15 -
0d0338f82024-12-06 migrate from circleci to github actions -
c59343482024-12-07 fix encode endpoint -
9f86737a2024-12-07 fix token encode to use the right model -
24130da42024-12-07 prio mac check for interface -
31d7bc2d2024-12-07 subprocess fork fix -
56842a272024-12-07 ignore topology merges from the non-owner -
69c18d9a2024-12-07 ignore topology merges from the non-owner -
50e4a9662024-12-07 topo fix only take your own as source of truth -
6d09b4ae2024-12-07 add special case for USB adapter over ethernet -
89815b162024-12-06 Applied patch idea from https://github.com/exo-explore/exo/issues/458 -
1d5a7c632024-12-06 pretty name for llama 3.3 70b -
22e4bc232024-12-06 fx -
5f6625a72024-12-06 viz positioning of inteface description -
02ee8b7d2024-12-06 macos interface type -
f4619d462024-12-06 use psutil for mac detection -
f257c21e2024-12-06 mac os network interface name -
9d1f14a22024-12-06 add llama-3.3-70b -
67e05c872024-12-06 topology endpoint to get current topology -
e8ece1152024-12-06 tweak sed, make compile_grpc.sh executable -
97373d532024-12-06 fix noopsharddownloader -
db7c388a2024-12-05 remove origin_node_id -
553442412024-12-05 use double for flops protobuf -
a0e083a12024-12-05 test -
0b9ee8ab2024-12-05 consistnet self.topology -
657520ed2024-12-05 pass origin_node_id to merge -
816322472024-12-05 coll -
f8cc54b92024-12-05 previously was checking all nodes for download status which lead to issues when running on multiple nodes. changed it to just check local node when checking model download status -
272b1e2a2024-12-05 remove unused funcs -
68d70be92024-12-05 always show desc1/desc2 in tui -
12bb315d2024-12-05 adding diable to chatbox and send if download is in progress -
f8d195ee2024-12-05 only collect topology when peers changed -
dba720442024-12-05 handle mutable visited properly -
99b5bf012024-12-05 fix topology merging -
d5b4039f2024-12-05 show interface type and name, fix propagation of topology -
b8c4b46f2024-12-05 prioritise network interfaces and display in the tui -
0f1024492024-12-04 Merge latest -
ca0caad02024-12-04 Image to image generation -
f94c90672024-12-04 trigger test -
438310b42024-12-03 adding option to download model from sidebar -
af7834112024-12-03 trigger start download without waiting for it to finish -
75d45dd92024-12-03 add endpoint for download -
f0bb515d2024-12-02 trigger test -
71db641f2024-12-02 trigger test -
4b8c4a792024-12-01 Images stored in system -
f339f74f2024-12-01 trigger test -
7dc0a7462024-12-01 trigger test -
60a40ecc2024-11-30 typing animation -
2f8f4cdc2024-11-29 new image4 -
9b4d03062024-11-29 adding loading icon to sidebar for model that is being downloaded -
f1a2f1a12024-11-29 opencv-python dep -
153eef682024-11-29 add docs png lfs -
cb948ef12024-11-29 generation anim mp4 -
b045f7612024-11-29 add png to git lfs -
482400c22024-11-28 trigger ci -
d34e67a22024-11-28 add --node-id-filter command line arg to filter by node id -
e0c871132024-11-28 better pagination to avoid rate limits on dashboard -
832e60522024-11-28 add redundant temp and top_p to dummy inference engine -
ac3217052024-11-27 removing console log in initial models -
c26477642024-11-27 add --default-temp option to change sample temperature -
3c81845a2024-11-27 undo diff -
1d1fa8c62024-11-27 move to tinygrad_helpers -
2fdda5172024-11-27 remove unused layers -
be82ac7d2024-11-26 update readme to reflect that mlx and tinygrad are interoperable -
1ab1762e2024-11-26 prio python 3.12 -
6659a18e2024-11-26 add missing top_p_sampling import -
b1b08e682024-11-26 test add line of code -
e6ee942a2024-11-26 fix diff checking on dashboard -
23b5e60b2024-11-26 remove redundant imports -
89b95e442024-11-26 dont run test sound effects on startR -
3fbeeba72024-11-26 use pygame for notification sounds -
97eeac112024-11-26 add mp3 sounds -
39a953472024-11-26 configure git lfs for mp3 files -
e99a73942024-11-25 adding a fetch to get initail model object to show models before going through and checking all the download percentages -
ded80b0f2024-11-25 ensuring requests do not stack up by moving polling to while loop with 5 second delay after -
3f6ea1732024-11-25 remove redundant imports -
2502ed202024-11-25 fix dummy generate so it doesnt have any randomness -
df52c4cb2024-11-25 simpleaudio requirement for dashboard -
3f506ad72024-11-25 cosmetic dashboard improvements -
ca8d59a22024-11-25 add notification sounds -
062f5e3e2024-11-25 dashboard fix limit -
6f8582d82024-11-25 nice dashboard, add benchmark results tokens per second -
ab333ea52024-11-25 circle tail logs -
c35deb6d2024-11-25 fix buffering issue with 2 processes outputting too fast in dummy test -
6b28b3412024-11-25 less strict match on response content -
f601a8302024-11-25 fix dummy inference -
216e7bff2024-11-25 remove redundant jq installation -
5d3ac40f2024-11-25 use jq to check response content circleci -
1331ed762024-11-25 fix dummy tokenizer -
311b4c212024-11-25 skip inference engine selection if running dummy -
e2f71b622024-11-25 remove redundant test_macos_m1 job -
99bf691e2024-11-25 check for response in quotes -
a5addd682024-11-25 fix dummy model id -
370564032024-11-25 require exact match on response from llms in integration tests -
28de4fd72024-11-25 dashboard fix int64 -
e3dc3b202024-11-25 include total line count in dashboard -
ea9b04362024-11-25 fix indent -
c0646be22024-11-25 count lines of code -
0ab63e4e2024-11-24 dashboard reqs -
b000d23b2024-11-24 dashboard async, clickable datapoints that go to commit url -
8c6a6fab2024-11-24 a simple dashboard to track pip package size over time, commit-by-commit -
e16170cf2024-11-23 backend endpoint now uses SSE to send each model as its loaded. also shows loading indicator until first model shows up -
2c5d05532024-11-24 make readme clearer with linux nvidia -
f5fdacd22024-11-24 16.0.0 for all circleci jobs -
84ff7cb32024-11-24 circleci job to get pipsize and store as artifact -
2dfd322f2024-11-24 thanks dnewman -
ab3e76a42024-11-23 change chatgpt-api-response-timeout default back to 90 -
357e33802024-11-22 removed debug -
3384fc722024-11-22 update tinygrad version -
4fdc61722024-11-22 add check for path exist -
1cb1e9e42024-11-22 Undo extra line -
61e9d4bd2024-11-22 Undo extra lines -
729669c92024-11-22 restore the cursor to the terminal on exit from CLI -
39139c142024-11-22 fixiing required engines definition -
fe0f1cdb2024-11-22 fix shutdown -
d3505e032024-11-22 run tests with --disable-tui for faster tests -
8a7414852024-11-22 fix test_inference_engine unittest reshape token output tensor -
e3ec9eaa2024-11-22 Fixed GRPC issues -
bc905cd62024-11-21 formatting deleteModel -
a9838a8f2024-11-21 formatting handle_delete_model -
619df1d42024-11-21 adding functionality to delete the models if there is part of the model downloaded -
fb3baf502024-11-21 adding amount that has been downloaded if model is not fully downloaded -
31ce70f42024-11-21 working with side bar to choose model, show download percentage, select chat, go back to chats, intitiate download of undownloaded models -
7e6c69fd2024-11-21 remvoing console log -
72c3fdab2024-11-21 fix end of request behaviour and add back broadcasting tokens to other nodes -
c77355232024-11-21 target_shard to next_shard -
f5afa4db2024-11-21 compile error fix -
46a8e8fc2024-11-21 typo -
c78973662024-11-21 new tinychat ui -
1ca11ead2024-11-21 defining optional -
9a6af7432024-11-21 fixes -
5269629d2024-11-21 removed unused code -
be6e9ae62024-11-21 pr fixes -
913fdcc62024-11-21 pr fixes -
400e428a2024-11-21 fix -
3210912a2024-11-21 removed unused import -
0f784fff2024-11-21 pr suggestion fixes -
d600bd422024-11-20 transformers version -
e0be8dd52024-11-20 cleaing comments -
5396f0802024-11-20 test clean ups -
4874295b2024-11-20 Image streaming while generation -
907659222024-11-20 added one file -
fece3f0c2024-11-20 gitignore tinychat pngs -
38ee81512024-11-20 static images dir -
416974312024-11-19 error fix -
441182522024-11-19 build error fix -
6b28ef032024-11-19 Stable stable diffusion mlx -
f337e7812024-11-19 pr fixes -
62acc1af2024-11-19 missing lib -
ee6f5dad2024-11-19 clean path -
3a1871c82024-11-19 typo fix -
97ed990a2024-11-19 macos sign -
8bc823222024-11-19 missing lib -
ce9231ad2024-11-19 move model fix -
8f78c7812024-11-15 Refactors to simplify messaging and properly batch inputs -
e15192462024-11-19 error fix -
1fa42f302024-11-19 typo -
6fc0b0442024-11-19 error fix -
520d9d112024-11-19 error fix -
9489b99c2024-11-19 typo -
aae23cec2024-11-19 build error fix -
c82d16482024-11-19 Bump aiohttp from 3.10.2 to 3.10.11 -
1b7e67832024-11-18 fix modelpool, add tests in test/test_model_helpers.py -
559f12e72024-11-18 check if user has read/write access to HF_HOME and warn them if not -
3022aab92024-11-18 remove redundant dummy import -
0ab302a32024-11-18 add --default-model command line arg -
312602fa2024-11-19 fix shard_specific_patterns -
4ece73422024-11-19 always run tinygrad stuff on same thread. tricky because of lazy evaluation -
74b98fdd2024-11-19 update package versions to work on python >= 3.9 -
e79fc3112024-11-19 Bump aiohttp from 3.10.2 to 3.10.11 -
3491b7462024-11-19 moved func -
ef372ab32024-11-19 merge conflict resolve -
c422cea62024-11-19 typo fix -
cb53e7172024-11-19 added file exist check -
65817ab72024-11-19 changes to args -
8ad70b202024-11-19 pr suggestion fixes: -
06c3f5242024-11-19 removed response return -
bcd885dc2024-11-19 cleaned code -
ea3347262024-11-19 code clean -
8ce0fe2b2024-11-19 pr suggestion -
867f348e2024-11-19 moving models -
00d4bda52024-11-18 fix build script -
e991438e2024-11-18 pr suggestions fix -
0ac1b87f2024-11-18 removed unused import -
0d50167d2024-11-18 yapf in download_file -
8ee6cc3b2024-11-18 yapf formatting -
91276ccd2024-11-18 fixing formatting -
8135437c2024-11-18 fixing formatting -
695ab3442024-11-18 removing import get_hf_home -
b77362b42024-11-18 moving os import -
6a7de04d2024-11-18 removing path update -
db610f592024-11-18 removing traceback -
325605132024-11-18 comment -
3ac868722024-11-18 adding redirect for all requests -
4c6fda7c2024-11-18 modifying helper fucntion checking size to follow redirect for .safetensor files to properly check the size with GET request -
dec79ac72024-11-18 modify get_shard_download_status to use helper function -
c61f40c62024-11-18 adding helper funciton to check file download. also modifying download_file to use that helper -
c923ef632024-11-18 modifying how its being displayed becuase now calculating overall percentage in hf_shard_download -
5916defb2024-11-18 update setup -
01e6b9312024-11-18 fix modelpool, add tests in test/test_model_helpers.py -
fea1c0fc2024-11-18 clean branch -
fa8825fa2024-11-18 check if user has read/write access to HF_HOME and warn them if not -
fd84201b2024-11-18 remove redundant dummy import -
a39ca1a42024-11-18 add --default-model command line arg -
a3e7bc002024-11-17 increasing height -
379ee4532024-11-17 adding padding and min height -
649157d42024-11-16 creating HFShardDownloader with quick_check true so it doesnt start downloading models -
3d0e2f1d2024-11-16 fix preemptive downloads with ensure_shard -
f2d5beee2024-11-16 change chatgpt api port from 8000 to 52415 -
dd38924e2024-11-14 removing checking of percentage for models that are not found locally -
972074e92024-11-14 reducing redundent checks -
dfcf513d2024-11-14 removing is_model_downloaded method and changing how downloaded variable is set -
d9aabd782024-11-14 working versions -
f1eec9fa2024-11-14 qwen-2.5-0.5b -
fd8672562024-11-14 healthcheck -
cbeb1b332024-11-13 fix safari issue -
3eb726ce2024-11-13 removing sorting of models by name -
95ce66572024-11-13 removing unneccesary css -
25d67f502024-11-13 cleaning up logging in index.js -
84ce07682024-11-14 Edit configure_mlx.sh for calculate dinamically the value for iogpu.wired_limit_mb and iogpu.wired_lwm_mb. The script limit wired_limit_mb to 80% and wired_lwm_mb to 70%, but this threshold are variables. -
59f5b6d82024-11-13 adding back in set error message -
fb32a8512024-11-13 removing error separtation so I can put in different PR -
7d7bdd832024-11-13 removing uneccesary console logs and fixing order of variables in index.js -
de09e2a82024-11-13 reusing helper function to get cached directory -
c7dd31262024-11-13 adding logic to check which models are downloaded -
b0d7c34e2024-11-13 Edit configure_mlx.sh for calculate dinamically the value for iogpu.wired_limit_mb and iogpu.wired_lwm_mb. The script limit wired_limit_mb to 80% and wired_lwm_mb to 60%, but this threshold are variables. -
b69452242024-11-13 disable configure_mlx.sh for now -
34f3c4a12024-11-13 fix tokenizers test with restructured models -
9712d6962024-11-12 Added a small script to compile grpc -
b787c6762024-11-12 Updated unit tests -
6d12deab2024-11-12 add better error handling: -
c43ad15c2024-11-12 Daniel changes -
b0dc94472024-11-12 First pass at a dynamic model menu in tinychat -
d69a9c4d2024-11-12 Enabled inference engine intercompatibility -
96aaab0b2024-11-12 removing commented underline in css -
42172b2c2024-11-12 Updated unit tests -
8b71d57d2024-11-12 Removed inference state entirely -
f67e18a72024-11-12 adding option to expand error to see stack trace and clearing timeout if message is expanded -
ead8e2892024-11-12 adding timeout back but making it 30 seconds -
325edddd2024-11-12 modifying error handling to include name and stack trace if available. css support multiple lines -
cbe551d12024-11-12 removing timeout for error and adding close button -
4c98108d2024-11-12 increase grpc msg limit -
194637372024-11-12 fix debug log -
f02e62c92024-11-12 Neglected to backpropagate this debug output fix from my training branch -
03924cf92024-11-12 Need tokens. Also, for some reason this gets mad if we have non-integral tokens but this isn't a problem elsewhere? -
e463cd812024-11-12 Ok not sure we're using this but just in case -
7e3ad9ab2024-11-12 Missed a spot -
1cd3efbe2024-11-12 Fixed unit tests -
65fdc99c2024-11-12 Call no longer needs request_id -
90518a3b2024-11-12 Hoisted caching to a wrapper class -
bf33ffde2024-11-11 This doesn't need to be a tuple really -
10e9f44a2024-11-11 one-line output buffering -
52ef6ee42024-11-11 Made temperature and top_p available to the inference engine sample interfaces -
8205a5ae2024-11-11 Implemented per-request caching in tinygrad -
13572e6a2024-11-11 Some stability improvements for tinygrad inference -
aefc0d7c2024-11-11 I think this is more faithful to how it was originally done -
c06b5f3b2024-11-11 Corrected type annotations -
9b66758b2024-11-11 Make sure they're np arrays -
b9d0fb682024-11-11 Since infer_prompt is a thin wrapper that works the same for all inference engines, we can de-abstract it -
527c7a6e2024-11-10 Applied new interface to tinygrad and dummy inference engines -
52b91de82024-11-10 Changed model classname due to the sharding being done elsewhere -
34019e462024-11-10 Forgot an abstractmethod -
82cce4402024-11-10 Some initial inference engine refactors for enabling training -
e9ba815c2024-11-12 add qwen2.5 coder 3b,14b,32b -
5435671c2024-11-11 Add 32b Qwen 2.5 -
167e756b2024-11-11 add documentation of HF_HOME model storage location in README. fixes #427 -
9e4366f32024-11-11 tinygrad ci -
6cd78b942024-11-11 run tinygrad test with CLANG=1 -
49c4394d2024-11-11 enable tinygrad test -
77d789352024-11-11 remove redundant expected_content -
8cc3f51e2024-11-11 test for tinygrad e2e -
472359142024-11-10 ignore 8bit llama 405b from tokenizers test -
989484412024-11-10 add llama 3.1 405b 8bit at mlx-community/Meta-Llama-3.1-405B-Instruct-8bit -
0d8a1ee42024-11-09 Added clear all history button -
f1c747322024-11-09 m4 device capabilities -
fcaebd3b2024-11-09 add Gemma2 9b and Gemma2 27bg -
83d4d2b32024-11-09 update mlx to 0.20.0, mlx-lm to 0.19.3 -
fbec1d2b2024-11-08 formatted changes -
af01b23a2024-11-08 added rope_scaling and tie_word_embeddings to llama transformer -
029dc5f82024-11-08 added new model info for 1B and 3B model sizes -
e80ed76c2024-11-06 ignore in test-tokenizers -
36c1f68c2024-11-06 update llama-3.1-405b-8bit model id to IntuitIntel/Meta-Llama-3.1-405B-Instruct-8bit -
c8438b6d2024-11-06 add llama-3.1-405b-8bit -
6ae6ebeb2024-11-03 revert back to CORRECTED -
e8e05e152024-11-03 remove CORRECTED llama 70b -
436709e52024-11-03 revert back to CORRECTED -
8d524bfe2024-11-03 change mlx-community/Meta-Llama-3.1-70B-Instruct-bf16-CORRECTED to mlx-community/Meta-Llama-3.1-70B-Instruct-bf16 since it works now -
ab91f2022024-11-03 Revert "add max-caches option" -
9e3ae2b42024-11-03 add max-caches option -
bc7acfd32024-11-03 update mlx to 0.19.3, mlx-lm to 0.19.2 -
661035f02024-11-03 use KVCache instead of RotatingKVCache -
d0b7154f2024-10-31 set fallback llama model to use smallest one also -
9987ce842024-10-31 use the smallest possible model as the default_model -
f52371342024-10-31 add a default_model in ChatGPTAPI -
8b23f1242024-10-30 update js -
ac49b2a22024-10-29 revert main.py -
f82410f82024-10-29 clear instructions on formatting with yapf, remove linting -
478db26b2024-10-29 get rid of all the different linters. we just use yapf now -
98ea71ed2024-10-29 run format.py on ./exo -
7dfe856f2024-10-29 remove buggy select logic -
29ca6ad12024-10-29 cehckpoint -
9dc93fd52024-10-28 add traceback.print_exc on topology collection errors from peers -
eb8a444e2024-10-28 fix flops parsing -
dbf40d782024-10-27 Update main.py: Default timeout 90->900 -
d4e26fcb2024-10-27 Update device_capabilities.py -
bc1d88d82024-10-25 ignore dummy -
7f7260482024-10-25 fix prompt ci -
9b8d58c42024-10-25 fix dummy setup -
6e87d7c02024-10-23 simplify manual discovery -
4a7503552024-10-23 remove cbrt which doesnt exist on python 3.9 -
922705952024-10-22 removed logging -
896ea4d92024-10-22 removed logging -
57c62c2c2024-10-22 removed logger -
cd4d324a2024-10-22 fix to creating engines -
da5235732024-10-22 fixed errors -
593d810d2024-10-22 fix to broadcast -
f4a5562c2024-10-22 added logger -
a03f3a2a2024-10-22 added error -
ae47fe112024-10-22 Moving conditonal apple silicon logic to setup.py -
6dd2f7ab2024-10-22 changes to inference engine -
3908b97a2024-10-22 changes to broadcast func -
e1ae6d5a2024-10-22 Getting the setup script to work on Intel based Mac machines -
5e4741352024-10-22 changes to manual_discovery config based on PR comments -
6b48a9362024-10-22 add pydantic dependency -
ad3893632024-10-21 changes to exo/main.py for manual config flags -
1970b9c82024-10-21 tests for manual networking -
f092b08b2024-10-21 initial setup of manual networking config -
fa6d63ba2024-10-21 update gitignore for aider -
82a708f92024-10-20 rm ministral-8b -
16447ba12024-10-20 added yml changes -
0fdd5c7c2024-10-20 feedback 1- changes requested, done -
7d6104752024-10-19 DummyInferenceEngine commit 1 -
6bd0f07a2024-10-17 add ministral-8b -
1e4524b52024-10-16 add nemotron-70b and nemotron-70b-bf16 to tinychat -
61ee67c92024-10-16 add nemotron-70b and nemotron-70b-bf16 -
3e33bca72024-10-15 TFLOPS on GTX 1660 - cabelo@opensuse.org -
d931e4d12024-10-15 Fix null download progress bug -
08ca7adc2024-10-14 remove redundant screenshot image in README -
29a9156e2024-10-14 try screenshot without html -
d43176de2024-10-14 resize screenshot to 80% -
b38715032024-10-14 Move screenshot up in the README to device equality section -
28c291902024-10-14 Rename 376385401-3b6e22d0-ca6a-466c-b1b8-221556fa4163.png to exo-screenshot.png -
76b3f6b12024-10-14 Add screenshot of exo running on 5 nodes -
d554313d2024-10-14 Update README.md -
60bb60f32024-10-14 TFLOPS on GTX 1050 - cabelo@opensuse.org -
03cbcca22024-10-14 Fix gpu capabilities display issue. Also update the capabilities with RTX 2080 TI -
0dd3b2f82024-10-13 Fix TFLOPS on 4060 Ti - cabelo@opensuse.org -
8e3e43bb2024-10-13 Add download progress bar in tinychat -
1dc28e872024-10-12 Fix TFLOPS on 4060 Ti -
83459b772024-10-12 Modify download progress section in tinychat index.css -
53ea57672024-10-11 chore: Support multiple nodes download progress section in tinychat -
073053232024-10-11 fix: tokenize -
17065d872024-10-10 dynamic halfway partition point in unit test -
ae74d2da2024-10-10 run unit test on llama 3.2 1b for faster test -
ad09b4b32024-10-10 also initialize embed_tokens if last layer and tie_word_embeddings true -
fbc407c62024-10-10 make llama-3.2-1b the default for tests so they run faster -
8950d95e2024-10-10 updgrade all mac ci jobs to xcode=16.0.0, resource_class=m2pro.large -
ade9db4d2024-10-10 feat(device_capabilities.py): add support for NVIDIA RTX 4000 ADA generation device capabilities -
a0ad18c62024-10-08 Fix GPU names for RTX Ampere cards -
8a69a7a22024-10-07 one line print -
b7996b9a2024-10-07 race condition in on_listen_message for udp discovery fixes #308 -
e80ee6072024-10-07 fix the race condition in cleanup peers and run the peer checks concurrently. fixes #308 -
aa2056262024-10-07 shield process_prompt so downloads dont get cancelled when chatgpt api request times out -
e8a870232024-10-06 replace tailscale.devices with good old http, removing the need for tailscale dependency -
82c7ce692024-10-05 Point `llama-3.1-70b-bf16` model to the actually bf16 version -
9ffd81162024-10-05 Use official nvidia-ml-py instead of pynvml -
9b9f40d42024-10-03 only stream results for the same request id. this allows multiple concurrent requests on the same LLM without overlapping interference in the streamed outputs -
9223993e2024-10-03 await node process_prompt with timeoout -
b611d0a52024-10-03 fix print -
ac6f1bed2024-10-03 add a priority to broadcast messages where the broadcaster can indicate how to prioritise that particular interface. for now all priorities are set to 1 but in the future this will be based on network latency, bandwidth, jitter, etc.. e.g. Thunderbolt prioritised over WiFi -
c3864f5e2024-10-03 more robust handling of timeouts -
4746ffdd2024-10-03 clean up download progress -
5521dcbf2024-10-02 add script to calculate pipsize -
3ed8a52a2024-10-02 Fixed the retry -
552c04fb2024-10-02 add back Jinja2 -
b655f3552024-10-02 resolved conflicts by git pull -
0079e7352024-10-02 remove unused imports -
b33159262024-10-02 remove blobfile, tiktoken, tokenizers -
413ecb1b2024-10-02 remove hf-transfer, huggingface-hub, Jinja2 unused dependencies -
4443d3ce2024-10-02 update README with docs on exo run command -
8ae59b702024-10-02 add exo run command. usage: exo run <model-name> e.g. exo run llama-3.1-8b -
90a88f312024-10-02 update readme with editable pip install -
fc65765b2024-10-02 always install interactively -
be9f2e792024-10-02 fix ci to use exo command instead of python3 main.py -
4923eb7e2024-10-02 also use tempdir for .exo_node_id to keep the dir clean -
1ccfdc3c2024-10-02 give examples of device configurations in readme -
71e93a102024-10-02 simplify hardware requirements in readme -
bef585d12024-10-01 download progress optimization in tinychat -
39a8a8142024-10-01 Download Progression in tiny chat -
a7a9124d2024-10-01 Hardware Requirements -
09e721622024-10-01 Hardware Requirements -
c498930f2024-10-01 Hardware Requirement notes -
8d0e8a8a2024-10-01 Hardware Requirementes -
c5b38f452024-09-10 move tinychat inside exo package -
31e4454b2024-09-08 add missing __init__.py files -
fa67ee9b2024-09-08 add entry_point exo to run main.py -
6abf48772024-09-08 move main.py script to package dir -
0120891c2024-10-01 update readme with PyTorch inference engine and llama.cpp link to issue -
67f789b62024-10-01 clearer example docs and add 405b example -
abca3bfa2024-09-30 add support for qwen2.5 coder 1.5b and 7b -
3e13e5ed2024-09-28 upgrade mlx to 0.18.0 -
073b3ffc2024-09-28 move udp and tailscale into their own modules -
2ebcf5f42024-09-28 fix llama 3.2 issue with apply_chat_template assuming messages is a list if its not a dict fixes #239 -
4a2992182024-09-26 cleanup -
7f9810c62024-09-26 Add error toast popup -
777102c92024-09-25 add support for llama 3.2 -
9831c26d2024-09-25 update README with Install Certificates SSL troubleshooting -
04df31b12024-09-25 bump up tinygrad to 232edcfd4f8b388807c64fb1817a7668ce27cbad -
41053c552024-09-24 change flags for unit test ci -
10812c452024-09-24 test ci -
c770f19a2024-09-24 update resources classes -
115f0eac2024-09-24 change the message we search for in ci -
428bb6062024-09-24 health check udp discovered peers before adding them -
da06fb3c2024-09-24 check before removing -
8aab93042024-09-24 if any peers changed from last time, we should always update the topology -
b1cf10852024-09-24 discovery should not include unhealthy peers -
7fa9f2cf2024-09-24 increase default max-generate-tokens to 10,000 -
4c4d8d2e2024-09-24 connect with 5 sec timeout -
4db674e82024-09-23 simplify health check -
5207edbd2024-09-23 ensure connected when health checking -
4f0c91ef2024-09-23 fix -
db3c603d2024-09-23 more robust health checks -
26dc19892024-09-23 implement grpc health check -
de2f6d2e2024-09-23 implement a health check for peers and discovery should only return healthy peers -
6e3239492024-09-23 comparing timestamps across distributed systems of consumer devices is probably a bad idea. -
417fe82b2024-09-23 formatting -
b480e7ed2024-09-23 delete expired peers -
69f1fe182024-09-23 faster discovery_interval, separate update_interval for tailscale -
15a2165d2024-09-23 periodically update exo_updated_at attribute for tailscale -
2e74db8f2024-09-23 if it doesnt have exo node attributes then skip -
e6808a5c2024-09-23 more robust sanitization for chip and model -
1798fc072024-09-23 support node_id, node_port and device_capabilities with tailscale attributes -
2244ff4a2024-09-23 clarify which models exo supports -
cb575f5d2024-09-23 ndim check in llama -
7dd7fe492024-09-22 fix allow patterns -
f7a4eaf12024-09-22 update logs for tailscale discovery -
09ed39182024-09-22 Update README.md -
27bf50692024-09-21 Fix issue where offline node cannot detect online node over thunderbolt due to the online node not broadcasting over the Thunderbolt bridge when non-bridge alternatives exist (ie. wifi or eth) Extract broadcast logic into a DatagramProtocol to follow similar pattern as the listener -
e7d40faf2024-09-21 prevent duplication -
932247992024-09-21 implement tailscale discovery module -
d7fff0d62024-09-21 fix allow patterns -
cb6338682024-09-21 update readme with tips for performance for apple silicon macs -
e138aa6a2024-09-20 trust_remote_code=True when loading tokenizer -
9ee8a0902024-09-20 trigger ci -
b6d239af2024-09-20 ignore deepseek v2.5 from tokenizers test as it requires remote code -
b1ec5ae22024-09-20 tweak ci for unit tests -
2caccf892024-09-20 update gpu rich/poor calc -
6ce8fd872024-09-20 script to configure mlx -
835e20972024-09-20 add deepseek-coder-v2.5 -
744b95762024-09-20 bump mlx to 0.17.3, bump mlx-lm to 0.18.2 -
311c81972024-09-19 update twitter handle exolabs_ -> exolabs -
68028cc92024-09-19 ignore Qwen models in tokenizers test until bos issue is fixed -
dee83e482024-09-18 add more qwen2.5 models: mlx-community/Qwen2.5-7B-Instruct-4bit mlx-community/Qwen2.5-Math-7B-Instruct-4bit mlx-community/Qwen2.5-72B-Instruct-4bit mlx-community/Qwen2.5-Math-72B-Instruct-4bit -
3597fba32024-09-18 add support for qwen2.5, initially adding mlx-community/Qwen2.5-14B-Instruct-4bit -
b39a251d2024-09-18 fix: remove extraneous '/' -
a0024fd42024-09-14 feat: support HF_ENDPOINT base url ENV VAR -
db9f44d12024-09-13 website link -
6c875dcc2024-09-13 update hiring link -
074228e32024-09-13 update README with hiring -
198cd6fb2024-09-13 trigger ci -
20522e062024-09-11 update docs to make tinygrad usage clearer -
4b0094012024-09-08 move `.exo_used_ports` to `/tmp` -
ca6445622024-09-08 fix broken links in README -
874886ab2024-09-05 simplify mlx non blocking -
e616d4e82024-09-05 run realize on the result in tinygrad -
9345684b2024-09-05 closely match prev impl mlx non blocking -
d6e661fd2024-09-05 match previous impl with np.array in mlx -
caf9b57a2024-09-05 trigger ci -
841871132024-09-05 add a test for hf get_weight_map -
4ec613d42024-09-05 simplify tinygrad non blocking -
a1a0ffac2024-09-05 add tinychat option for llama-3.1-70b-bf16 -
2948a8342024-09-05 add llama-3.1-70b-bf16 model option -
11dd952d2024-09-05 use set for shard specific patterns -
ea3322de2024-09-05 remove comment -
e0fda94d2024-09-05 use sets for shard specific patterns -
8f65e1e62024-09-05 fix weight_map resolution. previously we were always defaulting to allow pattern *.safetensors -
6881722b2024-09-05 simplify non-blocking mlx inference -
9db16f8d2024-09-05 use a queue for non-blocking mlx inference -
0ca5c2602024-09-05 run mlx inference engine on a single thread too -
58f535d02024-09-05 formatting -
2950373d2024-09-05 experiment with tinygrad on its own thread, so it doesnt block event loop -
41f0a22e2024-09-04 DEBUG>=8 for SendOpaqueStatus logs -
01cc6a4c2024-09-04 fix Mistral-Large special case when we pass in a path -
41dd700f2024-09-04 less aggressive logs for opaque status / download progress. too much spam -
4537d6142024-09-04 circleci use tee to output logs in realtime as well as capture them -
56c1bf9a2024-09-04 consistent remove _secs / -secs suffix -
f342cdca2024-09-04 get rid of -secs suffix -
a0d9c90e2024-09-04 shorten cli name --chatgpt-api-response-timeout -
8cb678e72024-09-04 better logs around peer connecting / disconnecting -
c97da5482024-09-04 add id to set -
80c48b9e2024-09-04 update visited with self.id, timeout on collecting topology from a peer 5s -
355c57992024-09-04 more robust discovery / peer handling. now we track if the same node id changes address, then we immediately conenct to it -
8114a79e2024-09-04 add back listen and cleanup tasks -
dcb3ac762024-09-04 test kill pids -
3dd81a1e2024-09-04 fix UDPDiscovery params, create a new transport every time we broadcast -
15b5043d2024-09-04 test for reconnect -
baf6efd32024-09-04 cleaner discovery -
572150412024-09-01 todo for speculative model -
dc3b2bde2024-08-30 use NousResearch/Meta-Llama-3.1-70B-Instruct as tinygrad llama-3.1-70b model, previously using non-instruct model -
12609cb62024-08-30 integration test for udp discovery with grpc server -
f93f811d2024-08-29 generalise UDPDiscovery to any kind of PeerHandle that accepts an address. test it -
d4a932e42024-08-29 fix merge -
5a9f4ba52024-08-28 update examples: remove old llama3_distributed, add chatgpt_api -
581856892024-08-27 clean up unused, formatting -
62e372622024-08-26 add RTX 20 series to device capabilities -
ebff636a2024-08-26 script ot start openwebui -
394935712024-08-26 add all chat endpoints without v1 prefix to support ollama / openwebui. related: #175 -
70172d7c2024-08-26 add /v1/models endpoint and change Content-Type of stremed response to text/event-stream. fixes #175 -
d917778e2024-08-25 update mlx to 0.17.1 (not sure where 0.17.0 went on PyPi disappeared)g -
2667c8af2024-08-25 cleaner download_progress -
f46d077b2024-08-24 fix font dependencies for tinychat. related: #172 -
8a4928f82024-08-24 fix gitignore to not ignore tinychat static files -
d515d9ef2024-08-24 explicitly use absolute paths for tinychat deps -
a386c35f2024-08-24 script to update tinychat deps -
3791e6692024-08-24 download tinychat dependencies all to local dir so we dont need internet -
85bab25a2024-08-24 fix local check if dir does not exist -
59c4393d2024-08-24 first try loading tokenizer from local path instead of always going to the internet first. significant speed ups -
784e6bae2024-08-24 print traceback on topology collection error -
8cad0e182024-08-24 only use_fast tokenizer for Mistral Large until this inconsistency bug is fixed #171 -
852790072024-08-23 hotfix edge case where we try to render before tokenizer is set -
09a846832024-08-23 upgrade mlx to 0.17.0 -
1f9d16ec2024-08-23 run tokenizers test in ci, run all models available -
6243846e2024-08-22 ci logs -
cfe980bd2024-08-22 simplify ci -
9513c4fd2024-08-22 ci tail log files -
7a02acdc2024-08-22 fix ci output streaming -
ad6956962024-08-22 run on every commit on main, reuqire approval on other branches -
710e5a312024-08-22 TODO for why use_fast=False is giving inconsistent behaviour (no spaces decoding invididual tokens) for Mistral-Large-Instruct-2407-4bit -
e17e5f9a2024-08-22 tests for tokenizers. unfortunately use_fast=False and use_fast=True give different behaviour -
0d218e242024-08-22 use fast AutoProcessor fixes #164 tokenizer issues with mistral-large. -
23ae5e922024-08-22 hold circleci tests for approval on non-main branches -
d54944f42024-08-22 stream outputs from chatgpt api integration test -
f53056de2024-08-22 more compact operator formatting -
14f2846a2024-08-22 yapf set blank_line_before_nested_class_or_def to false -
ea70c9fb2024-08-22 reformat with yapf format.py -
2e2707662024-08-22 simplify formatting with yapf -
417114fa2024-08-22 fix mistral nemo -
5101f0332024-08-22 keep 4 in RotatingKVCache -
6db73fab2024-08-21 laptop gpu device capabilites -
647ffb942024-08-21 increase cli generation timeout -
dd24e7db2024-08-21 only ignore CancelledError inside stop -
2e1233352024-08-21 ignore CancelledError when stopping the server -
ae35ada12024-08-21 fix headless mode with --disable-tui -
b95916e02024-08-21 show prompts and outputs in tui -
e84304312024-08-21 add a cli that can be triggered with --run-model <model> --prompt <prompt> -
65e0488e2024-08-21 logs for file filtering, grpc_discovery -> udp_discovery -
cea9b48d2024-08-20 update mlx-lm to 0.17.0, use lru caches for kv_cache with RotatingKVCache to optimise memory fixes #158 -
430d4c0c2024-08-20 astra clarify readme, it's an example app -
e87e72602024-08-20 astra: live camera with overlay debug info / ui -
1a419f1f2024-08-19 astra better ui with camera vlm -
850354382024-08-18 fix streaming, change default model to llava -
c94ffa0e2024-08-18 fix audio buffering -
23c713c02024-08-18 better readme for astra -
2fe3a52d2024-08-18 readme for astra example -
ff71ccc62024-08-18 send api request in astra example -
b85d19562024-08-18 open source astra example -
e2e98c302024-08-17 Update README.md -
c4b261da2024-08-16 Update README.md -
0e2ae28d2024-08-15 trigger test -
92dbb3202024-08-15 update mlx to 0.16.3 -
9b8e1bcd2024-08-14 trigger test -
1819df362024-08-14 Add common RTX A series cards to device_capabilities.py -
a930be4f2024-08-13 t -
53ec180d2024-08-13 fix test import -
611085b32024-08-13 trigger test -
2a214db72024-08-13 rm tokenizer from test -
803dffd12024-08-13 always call convert_from_huggingface with tinygrad models. this was broken by shard layer filtering which made the check sometimes fail. fixes #144 -
a8b877bc2024-08-12 Update README.md -
7ddb80e22024-08-11 f-string expression part cannot include a backslash fixes #142 -
75681a972024-08-10 use async for all file ops, cache fetch_file_list, cache commit hash, quickly check file sizes on disk before making requests -
6c1bf1272024-08-10 add --max-parallel-downloads flag that limits the number of downloads at a time with asyncio.semaphore -
440fd35e2024-08-10 upgrade aiohttp -
8e6414b22024-08-10 spacing -
e8267e732024-08-10 LAPTOP GPU and Laptop GPU prefixes -
31641d102024-08-10 tinygrad select model size -
e6902b2f2024-08-09 add --download-quick-check flag to bypass the hf api calls / remote file checks -
84afdbcb2024-08-09 add --download-quick-check flag to bypass the hf api calls / remote file checks -
71591d2e2024-08-09 display all interfaces web chat and chatgpt api are available on fixes #134 -
3bd5a1162024-08-09 ignore files that dont match allow patterns -
3db3e8292024-08-09 make download panel slightly larger -
5112f53a2024-08-08 trigger ci -
047ef48c2024-08-08 use separate hf cache dirs for chatgpt api integration test. its an unusual setup where we're running 2 exo instances on the same device which share a disk and hf cache -
2be446542024-08-08 refactor tinygrad, only load necessary layers for each shard fixes #128, enable JIT (much faster), prefill all layers not just the first shard fixes #12, use new ShardDownloader for more robust, parallel downloads -
357331c52024-08-08 remove some logs, make get_allow_patterns out of class -
b1eb05ed2024-08-08 debug level 7 for tests -
09a9abc02024-08-08 fix inference engine test -
dd41026c2024-08-07 cache completed download paths -
706488732024-08-07 disable prefix matching on prompts. causes subsequent requests to fail with cannot be broadcast. hotfix for #130 -
35b7042e2024-08-07 upgrade mlx to 0.16.1 -
b181f8aa2024-08-07 handle writing responses errors -
7ec660bb2024-08-07 fix shard download -
29f154592024-08-07 init active_downloadsa -
6bddb2a92024-08-07 download edge cases -
f29963f42024-08-07 preemptively start downloads when any node starts processing a prompt. this fixes #104 -
7a65a96e2024-08-07 download progress styling -
c59ceab82024-08-07 viz spacing -
0a588d042024-08-07 viz styles -
d9f232b32024-08-07 cleaner download progress ui -
476a714b2024-08-07 make a separate ShardDownloader abstract class w HFShardDownloader. this opens up plugging in different methods of downloading model shards e.g. #79 / #16 -
d22ed12e2024-08-06 bring tinygrad to parity with mlx on llama models, show progress of each download file -
45142dab2024-08-06 tests -
545a486e2024-08-05 separate hf_helpers, make extra dir with download_hf script, unify downloading so tinygrad uses the same method as mlx and interoperable model formats -
9014efae2024-08-05 minimal script to download from hf async with progress -
55bcad982024-08-04 standardise tinygrad models/tokenizers so it can handle mlx hf -
6b1960ba2024-08-04 fix nvidia capabilities -
4a5c6cc52024-08-02 t -
f93ae2b52024-08-02 disable tinygrad test for now. need a larger runner or smalelr model -
201996af2024-08-02 macos -
0eb5c0c62024-08-02 mac runners -
e10aa5812024-08-02 cuda test -
23d2432c2024-08-02 test cuda -
fcdf57b82024-08-02 t -
f9c427462024-08-02 t -
cf32ec9a2024-08-02 t -
c5646ae02024-08-02 t -
6b3001122024-08-02 t -
4095dea82024-08-02 t -
911712542024-08-02 t -
ec98b9cf2024-08-02 t -
fecd08102024-08-02 t -
49664c462024-08-02 t -
2564d7c22024-08-02 docker runner -
aa99c8ae2024-08-02 trigger circleci -
331315102024-08-02 test cuda -
7599df7f2024-08-02 trigger circleci -
0c3638f42024-08-02 trigger circleci -
a3b0650f2024-08-02 use gpu.nvidia.medium -
3a0c9c852024-08-02 linux-cuda image -
af2f98ba2024-08-02 run tinygrad tests on gpu.nvidia.small.gen2 (NVIDIA A10G 24GB) -
0fd6bd912024-08-02 run tinygrad with llama-3-8b -
eafed8e12024-08-02 fix legacy model loading -
08e8cacf2024-08-02 run integration test for each inference engine -
32bb44b32024-08-02 request to both nodes in integration test, dont preload the model - exo should be robust against that -
c06124c62024-08-02 fix circleci badge -
f4b0f1ca2024-08-02 replace tests badge with circleci -
96969e352024-08-02 prefix matching logic -
7c0923272024-08-02 trigger circleci -
2a8f1ae42024-08-02 get rid of caches. circeci downloads fast enough -
54ea5dbb2024-08-02 only remove the matching prefix from the prompt if its length is less than the prompt -
67ad3f572024-08-02 use llama 3.1 in tests -
92b66e272024-08-02 fix cache -
cf3dae952024-08-02 circleci: separate hf, tinygrad caches -
faadfa292024-08-02 circleci chatgpt integration test -
9a03991f2024-08-02 cache key -
2065c7752024-08-02 orb -
cf4cddcc2024-08-02 port github workflow to circleci -
21b0bf872024-08-02 Update config.yml -
997fcaff2024-08-02 CircleCI Commit -
f5755ea12024-08-01 give explicit node ids when running on the same instance in tests, otherewise they use the same one because of sticky node ids -
4faa6c062024-08-01 add support for selective model downloading. related: #16 -
be3a09c22024-07-31 rm logs -
0bfb8e3b2024-07-31 sticky node ids #16 -
980d5d2c2024-07-31 bring back TINYGRAD_DEBUG. not sure why it was removed -
767662532024-07-31 fix regression introduced by image_str for tinygrad -
1d54f1052024-07-31 pass on tinygrad set_on_download_progress -
d6a7e4632024-07-31 async model downloading with download progress. fixes #102. related: #16 #104 -
5c67e24c2024-07-31 smart prompt longest prefix matching to avoid sending the same text through the NN again. speeds up prefill significantly -
94ac94632024-07-31 fix model id for llama 3.1 405b now its finally on the hub -
178fb75c2024-07-30 fix image api prompt encoding -
2d2000092024-07-30 use AutoProcessor with use_fast=False since there's a bug with use_fast=True where whitespace is removed on single token decodes -
af1c7ce32024-07-30 add support for image upload to tinychat for vision models -
0d45a8552024-07-30 increase max request size to send raw images, make image download from url async, use chatgpt-compatible convention for images -
e68d06f42024-07-30 move model-selector styles to index.css -
78db451d2024-07-30 add pillow to main dependencies -
142682642024-07-30 bump up tinygrad version -
8d3d3df12024-07-28 update readme -
acc94b502024-07-28 chatgpt api integration -
33cbacf52024-07-27 fix llava sanitize -
2fb961fc2024-07-27 stick to same convention as new llama -
b44b91712024-07-27 add pillow as testing dependency -
2aa1e24e2024-07-27 remove unused torch import -
833e7f332024-07-27 rename sharded_llava -> llava to match new convention -
63e51a822024-07-27 formatting -
6695b0192024-07-27 format format.py -
1dc08fec2024-07-27 increase max line length to 200 -
444137772024-07-27 formatting -
a6bb8ddf2024-07-28 update deepseek sanitize to shard layers first before handle switch -
cb217b7b2024-07-27 format format.py -
4cb36a7f2024-07-27 increase max line length to 200 -
d94e3f9c2024-07-27 formatting -
666b1c832024-07-28 refactor(mlx): model sharding and add deepseek v2 support -
931ced7c2024-07-27 fix a few more linter errors -
57b2f2a42024-07-27 fix ruff lint errors -
ce7610382024-07-27 formatting / linting -
f051ebe62024-07-27 remove accidentally added files -
5eafd5a32024-07-27 try/except for decode, #75 -
2849128d2024-07-28 processor load -
549939952024-07-28 conflicts -
9d2616b92024-07-28 shareded inference -
faa131942024-07-26 disable chatgpt api integration test, github changed something in their mac runners? perhaps time to switch over to circleci like mlx -
67a1aaa82024-07-26 check processes in github workflow -
628d86792024-07-26 force mlx inference engine in github workflow, where it defaults to tinygrad because it's running on 'model': 'Apple Virtual Machine 1', 'chip': 'Apple M1 (Virtual)' -
e856d7f72024-07-26 log chatgpt integration test output from each process on github workflow failure -
d2fa7b242024-07-26 Showing the message only if successfully decoded, #75 -
4f5ab78d2024-07-26 Addressing issue #75 to avoid decoding binary packets -
7cbf6a352024-07-26 working test -
5a2337602024-07-25 add log_request middleware if DEBUG>=2 to chatgpt api to debug api issues, default always to llama-3.1-8b -
803a44212024-07-26 init -
208478442024-07-25 per-request kv cache, remove all explicit reset functionality as it wasnt used. fixes #67 -
dd8c5d632024-07-25 add support for mistral nemo and mistral large -
03fe7a052024-07-25 more robust message parsing fixes #81 -
0770c59d2024-07-25 Update main.py -
e1792e292024-07-25 chore: Update argparse action for --disable-tui flag -
2c71a4b12024-07-25 Update device_capabilities.py -
942012572024-07-24 styling for tinychat model selector -
5ac6b6a72024-07-24 clearer documentation on accessing web UI and chatgpt-api -
9a373c2b2024-07-23 make configurable discovery timeout -
63a05d5b2024-07-23 make configurable discovery timeout -
8d2bb8192024-07-23 add llama-3.1 notice to README -
7a2fbf222024-07-23 add model selection to tinychat -
bbfd5adc2024-07-23 add support for llama3.1 (8b, 70b, 405b). bump mlx up to 0.16.0 and mlx-lm up to 0.16.1. fixes #66 -
5496cd852024-07-22 Revert "smart model downloading for mlx #16" -
3a230f3b2024-07-22 smart model downloading for mlx #16 -
b0e7dd9d2024-07-22 add max-generate-tokens flag fixes #54 -
f2f61cce2024-07-22 inference engine selection improvements -
4e4623232024-07-22 add simple prometheus metrics collection, with a prometheus / grafana instance for live dashboard. related: #22 -
e93466412024-07-20 implement dynamic inference engine selection -
1fcbe18b2024-07-20 fix m2 ultra flops -
9d9d257e2024-07-20 reduce chatgpt api response timeout in test -
8850187b2024-07-20 tell the mofo in the workflow to keep responses concise -
052ee1c72024-07-20 cache isolation per workflow job -
ce41e6532024-07-20 check cached files in workflow -
3d82338c2024-07-20 debug cached files in workflow -
aec58b3b2024-07-20 remove redaudant discovery check in automated test -
9785e2502024-07-20 formatting if -
08b2f3752024-07-20 test output spacing -
db583a862024-07-20 disable tui flag -
821f114b2024-07-20 add tests badge -
71b8c6602024-07-20 test workflow -
6c8715622024-07-20 fix huggingface cache -
cf98cc502024-07-20 trigger workflow -
719e149a2024-07-20 test trigger workflow -
9d939b372024-07-20 disable tinygrad test again, we need a smaller model or a machine with more memory otherwise we get Metal OOM -
774e62092024-07-20 add space between outputs in github workflow integration test -
a2a7ca1f2024-07-20 cleaner node info = -
04f2aa2a2024-07-20 try with METAL_XCODE=1 for tinygrad metal -
d2ed4c2a2024-07-20 disable tinygrad infernece engine test waiting Waiting on https://github.com/tinygrad/tinygrad/issues/5549 -
115aab0d2024-07-20 cache tinygrad models in github workflow -
a4cc66772024-07-20 async model downloading fixes #30 -
e49924e12024-07-20 add chatgpt-api-response-timeout-secs flag, set this to 20 mins in test -
7dd7ccab2024-07-20 do one request to load the model then another to check the response -
144af1062024-07-20 separate discovery and chatgpt api integration test -
93df43d02024-07-20 redundant sh -
bf7aa51b2024-07-20 rename to discovery integration test as thats all it checks -
b9a2c0f72024-07-20 fix tests -
d9516d2e2024-07-20 insstall in workflow -
8efd65632024-07-19 set different api ports so they dont conlict -
8dd17fe02024-07-19 integration test with discovery -
4d962ffc2024-07-19 fix hardcoded path in debug_inference_engine -
30ab126c2024-07-19 fix test_inference_engine -
56e5e34e2024-07-19 fix invalid escape sequence exo_text -
62a240732024-07-19 github workflow: use python3 consistently -
ba1916a32024-07-19 github workflow for tests -
10a043772024-07-19 check for the last file that downloads in case it fails part way through -
1475c7352024-07-19 fix inference_state serialization. related: #40 #44 #45 -
e18549e92024-07-19 rm print -
0c5a927f2024-07-19 spacing in viz -
9fa0cb1a2024-07-19 add gpu poor/rich bar in panel. fixes #33 -
5b8f127b2024-07-19 fix opaque broadcast -
a342e1ab2024-07-19 add web url and chatgpt api endpoint to panel (fixes #43), fix a rounding error in the partition to shard mapping implementation -
8939f8882024-07-19 remove spammy log -
d94849062024-07-18 remove the spammy logs -
dd09c5972024-07-18 fix issues with chatgpt api where it would generate too long output. avoid nonlocal -
4b592f9d2024-07-18 exo topology visualisation that shows the topology of the network, device capabilities and the currently active node using opaque statuses. fixes #36. ready for #33 -
351776902024-07-18 by default find an ephemeral node port fixes #35, more robust topology updates. both fix #15 and #14 -
54c986072024-07-18 more robust grpc discovery with asyncio and proper error handling, add flops to device capabilities. fixes #23 and progress on #33 -
fa9d41692024-07-18 rm unused imports -
0af164f02024-07-18 remove old PartitioningStrategy -
1b194b432024-07-18 reference the code for each feature listed in README -
945f90f62024-07-18 allow overriding inference_engine and separate flag for TINYGRAD_DEBUG -
47163d222024-07-18 broadcast results concurrently fixes #31 -
621a5f5d2024-07-18 Add license badge -
46d618ab2024-07-18 tiny fixes -
d4f550022024-07-18 sort topology by memory descending (works well for now to workaround #12 -
071b1caa2024-07-18 drop exo to 0.0.1 (still experimental) -
e7dcdac22024-07-18 fix exo text -
17ecf2662024-07-17 typo bullet point -
9958ac392024-07-17 Make known issues more prominent -
72fe29372024-07-17 exo text on start and stop -
fbbb45c32024-07-17 install script -
3778301b2024-07-17 add alternative installation through install.sh -
4f4696e02024-07-17 remove calls to updateTotalTokens in tiny, not sure why its there -
d4e0a7d12024-07-17 add endpoint to get number of encoded tokens -
127b8e012024-07-17 dont explicitly specify show_index -
a94bdbb92024-07-17 serve tinychat static -
e82dab1d2024-07-17 remove tinygrad hidden files -
c2fcee432024-07-17 only retail examples/tinychat from tinygrad subtree -
0870e6bf2024-07-17 Squashed 'tinychat/' content from commit fa7e734b4 -
d8c40bb42024-07-17 print a warning if stream task ever times out -
1e1e11cd2024-07-17 check if inference_engine has tokenizer before printing with it -
8df2f4d82024-07-17 support for streaming from non-tail nodes on chatgpt api, addresses #20 -
8762effa2024-07-17 chatgpt api repsonse streaming solves #20 -
5de2ea512024-07-17 default to llama-3-8b and temperature=0 if not provided -
5c3f0e3a2024-07-17 faster initial node discovery -
8a35fd832024-07-17 support chatgpt api endpoint fron any node #24 -
ba7abb982024-07-17 fix ring topology img -
c432871e2024-07-17 replace the ring topology image as it was not rendering sometimes -
bcab97cb2024-07-17 instructions for how to force a node to be the tail -
eb92da2c2024-07-17 cleaner chatgpt api impl with async callbacks -
7c97ef522024-07-17 DEBUG should be imported from exo -
5055e3782024-07-17 separate prerequisities seciton / troubleshooting section in installation of readme -
998d48432024-07-17 match psutil platform detection might catch some edge cases -
99d40b1d2024-07-17 download tinygrad model log -
442c7d8c2024-07-17 update readme -
879b96902024-07-17 Removed requirements.txt -
4ea7db852024-07-17 more complete python gitignore template -
bfaeccc72024-07-17 added setup py -
12fcbc0d2024-07-17 switch over to psutil, more robust system detection -
7545e0602024-07-17 fix: syntax error in requirements.txt -
6b3727f02024-07-16 fetch model if doesnt exist on tinygrad -
b1f3204e2024-07-16 add Jinja requirement for linux -
365114ec2024-07-16 add stars to readme -
71e007452024-07-16 fix tokenizer inconsistencies -
c673b4c32024-07-16 clarify readme -
7cb1ba552024-07-16 Clarify readme iOS -
c819f6752024-07-16 fix linux/amd gpu memory, convert to MB -
e93a319c2024-07-16 typo -
ce46f0002024-07-16 linux device capabilities -
dbbc7be52024-07-16 remove hard dependency on MLX fixes #8 -
5e8bfc8a2024-07-16 remove deprecated main_static -
e049f7012024-07-16 debugging instructions in README -
bde1e53f2024-07-16 add license -
dd8d18122024-07-16 add an opaque inference_state that inference engines can use to pass around small state to other devices -
03ba31c02024-07-16 wip state -
ed7672e32024-07-16 Update README.md -
9324fc152024-07-16 Fix broken links -
b897fa442024-07-16 Typo ring memory weighted partitioning strategy -
50d5e9482024-07-16 Update README.md -
b6d919722024-07-16 add notice of python>=3.12.0 -
403abcfa2024-07-16 smaller ring topology img -
231cde5f2024-07-16 ring topology image -
d78d5b202024-07-16 explain device equality in README -
94b6a2492024-07-16 print debug only -
bf565f942024-07-16 fix #7 no module named aiohttp -
bdf105a62024-07-16 clarify example in readme -
9759408a2024-07-16 trim off the eos_token_id from chatgpt api response -
f2895cbc2024-07-16 revive the chatgpt api endpoint on :8000 -
1d5c28ae2024-07-15 (partially) restore exo node equality by forwarding prompts to the dynamically selected head -
1ec92b732024-07-15 README notice about api endpoint -
108a904a2024-07-15 Readme whitespace -
e17905e22024-07-15 add global reset -
de9b89ea2024-07-15 readme typo -
199eeb032024-07-15 known issues section in readme -
d2184f582024-07-15 keep track of already visited peers in global operations: collect_topology -
4502da5b2024-07-15 readme bug notice -
f9a201dd2024-07-15 docs dir -
4d43cb912024-07-15 logo -
4e1f01ee2024-07-15 update links -
1dea8b9c2024-07-15 update discord link -
544c229e2024-07-15 discord link -
98b30e052024-07-15 readme tweak -
963f8eb62024-07-14 better logs for DEBUG>=1 -
a009f7d62024-07-14 move examples to examples dir -
b6595bac2024-07-14 add llama-3-70b to the examples -
54e8cad22024-07-14 remove uneeded prints -
c69120552024-07-14 empty space -
bcd589382024-07-14 clean debug logs -
b9c323bb2024-07-14 memory-efficient shard loading -
53a5b3fc2024-07-14 add uuid requirement -
05b9fa492024-07-14 initialize node id to uuid4 if not set -
ff597d952024-07-14 fix discovery -
a04974162024-07-14 fix model import path -
b8a2a0fb2024-07-14 update readme run instruction -
a933352a2024-07-14 add DEBUG flag for controlling debug logs -
dd882fe62024-07-14 experimental notice -
c8753ba52024-07-14 reshuffle readme -
ee5204fb2024-07-14 readme installation instructions -
78da11e12024-07-14 slightly nicer readme -
2fc472c82024-07-14 slightly nicer readme -
8ff3e2632024-07-14 slightly nicer readme -
32f2e36f2024-07-14 main rename -
5bbde22a2024-07-14 move everything under exo module -
c851644a2024-07-14 update requirements, specify exact versions -
329720332024-07-14 update readme -
5ef07d412024-07-14 readme -
490fa1022024-07-14 tinygrad inference engine -
e6f387a62024-07-13 handle is_finished -
b01f69bb2024-07-13 add support for multiple concurrent requests with request ids -
7077652c2024-07-13 graceful node shutdown -
ca6095c02024-07-13 a generic test for every inference engine -
850b72d32024-07-13 make StatefulShardedModel callable, add some tests for mlx sharded inference -
6ee0547e2024-07-13 fix layer calculation for sharded llama -
445eda152024-06-25 dynamically assign shards to nodes deterministically weighted by memory -
36b845672024-06-25 collect global topology with local peer visibility, ring memory weighted partitioning strategy -
3a66a0a42024-06-24 add requirements.txt -
ee96c6b02024-06-24 add another test for device capabiities on MacBook Air -
6c8c9ee72024-06-24 topology with partitioning strategy -
563dcb562024-06-24 mlx sharded implementation with example of distributed inference -
a21f59ff2024-06-23 scaffolding for networking, inference and orchestration
Authors
- Alex Cheema1091
- rltakashige139
- Evan Quiney119
- Jake Hillion73
- Nel Nibcord62
- Arbion Halili60
- ciaranbor55
- josh54
- cadenmackenzie48
- Glen43
- Evan25
- Matt Beton21
- Gelu Vrabie21
- Andrei Cravtov20
- Ian Paul19
- Sami Khan16
- Seth Howes12
- Pranav Veldurthi10
- Mustafa Alp Yılmaz9
- Varshith8
- Gaetan Lepage7
- DevEmilio967
- Rory Clear6
- Daniel Newman6
- Sandesh Bharadwaj5
- Sean-fn5
- vskiwi4
- thenatlog4
- Alessandro de Oliveira Faria (A.K.A. CABELO)4
- Alex4
- Heidar3
- Drifter42423
- Adam Durham3
- mlpy03
- DeftDawg3
- Caden MacKenzie3
- sigseg53
- Ogden Wells3
- Mukund Mauji3
- rahat21343
- Mark Kockerbeck3
- Cloud15903
- Steve2
- ecohash-co2
- ArvidSU2
- Michael Harrigan2
- Andrei Onel2
- wysie2
- Hunter Bown2
- Ryuichi Leo Takashige2
- Heath Dutton🕴️2
- divinity762
- Piyush Acharya2
- dependabot[bot]2
- Giuseppe Gambino2
- FFAMax2
- willschneider152
- Mark Van Aken2
- drew2
- Baye Dieng2
- Anchen2
- aidiffuser1
- OrbisAI Security1
- Sakutaro1
- team-wcv1
- Kerollos Magdy1
- Sam Bradbury1
- Nadeem Hilal Wani1
- chaoliang yan1
- MikkoParkkola1
- kaiisfree1
- DeepZima1
- Mazin Sharaf1
- Miguel Cruz1
- Miguel Miranda Dias1
- Owleksiy1
- Daiz1
- Antonio Lujano Luna1
- PG1
- Chris A1
- madanlalit1
- RickyChen / 陳昭儒1
- Matiwos Kebede1
- majiayu0001
- Nightguarder1
- Olimbek Nizomov1
- mags0ft1
- Rodionov Pavel1
- Nirav Patel1
- Ikko Eltociear Ashimine1
- Carsen Klock1
- pepebruari1
- damho.lee1
- Will Bickford1
- Austin1
- Alessandro de Oliveira Faria (A.K.A.CABELO)1
- LIPERE Benjamin1
- prajeshElEvEn1
- Yazan Maarouf1
- James Alexander Shield1
- James Shield1
- Sam1
- jinke1
- JakobDylanC1
- itsknk1
- Alec Potluri1
- Andrey Varfolomeev1
- Matt Royer1
Agents used
Skills used
- /exo130
- /github104
- /user-attachments62
- /assets62
- /claude51
- /claude-code50
- /models34
- /exo-explore31
- /include28
- /shared25
- /master22
- /issues19
- /main19
- /download19
- /localhost16
- /state15
- /worker14
- /chat11
- /completions11
- /bench11
- /types10
- /data10
- /huggingface9
- /json9
- /stderr9
- /except9
- /gpt-oss-120b-9
- /components8
- /users8
- /jake8
Creative ideas + design notes
b5375f8c · 2026-06-22 · Add Kimi K2.7-Code model card (official INT4 weights + vision) (#2167)
Adds a model card for [moonshotai/Kimi-K2.7-Code](https://huggingface.co/moonshotai/Kimi-K2.7-Code), released 2026-06-12. Same architecture as Kimi K2.6 (`kimi_k25`, 61 layers, official INT4), so the card mirrors the existing `moonshotai--Kimi-K2.6.toml`. Sampling defaults per the model card (temperature 1.0 / top_p 0.95 for thinking mode). **Vision:** the official repo ships MoonViT weights inline, so I extracted the 335 `vision_tower.*` / `mm_projector.*` tensors (unmodified bf16) into [aidiffuser/Kimi-K2.7-Code-vision](https://huggingface.co/aidiffuser/Kimi-K2.7-Code-vision), following the `exolabs/Kimi-K2.6-vision` format. The vision config is byte-identical to K2.6's; the extraction script is included in the repo for verification. Happy to have this re-hosted under the exolabs org if you prefer — it's a one-line change to the card. **Tested:** distributed serving on 2× Mac Studio M3 Ultra (512 GB), tensor parallelism, text + thinking + image understanding all confirmed working. Co-authored-by: aidiffuser <your-noreply-email@users.noreply.github.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
cdf1add8 · 2026-06-22 · fix: upgrade devalue to 5.6.2 (CVE-2026-22774) (#2150)
## Summary Upgrade devalue from 5.5.0 to 5.6.2 to fix CVE-2026-22774. ## Vulnerability | Field | Value | |-------|-------| | **ID** | CVE-2026-22774 | | **Severity** | HIGH | | **Scanner** | trivy | | **Rule** | `CVE-2026-22774` | | **File** | `dashboard/package-lock.json` | | **Assessment** | Likely exploitable | **Description**: devalue: devalue: Denial of Service due to excessive resource consumption from untrusted input ## Evidence **Scanner confirmation**: trivy rule `CVE-2026-22774` flagged this pattern. **Production code**: This file is in the production codebase, not test-only code. ## Threat Model Context This is a web service - vulnerabilities in request handlers are directly exploitable by remote attackers. ## Changes - `dashboard/package.json` - `dashboard/package-lock.json` ## Verification - [x] Build passes - [x] Scanner re-scan confirms fix - [x] LLM code review passed --- *This change addresses a pattern flagged by static analysis. The code path handles user-influenced input and the fix reduces the attack surface against both manual and automated exploitation.* --- *Automated security fix by [OrbisAI Security](https://orbisappsec.com)*
81d7cb0f · 2026-06-03 · docs: add Homebrew cask install instructions (#2140)
## Motivation exo is now available as a Homebrew cask, so the README should show the simplest macOS installation path alongside the existing DMG download. Fixes https://github.com/exo-explore/exo/issues/2105 https://github.com/exo-explore/exo/issues/176 ## Changes - Added `brew install --cask exo` to the macOS App section of `README.md` - Kept the existing DMG download link as the first installation option ## Why It Works Adding the Homebrew cask command gives macOS users a package-manager-managed installation path while preserving the existing DMG download option. ## Test Plan ### Manual Testing - Reviewed the rendered Markdown structure in `README.md` ### Automated Testing - Not run. Documentation-only change. ## Related - https://github.com/Homebrew/homebrew-cask/pull/265956
629c55d6 · 2026-05-31 · Rename exo_pyo3_bindings to exo_rs (#2131)
## Motivation (I think it) Makes Evan's massive PR easier to merge later on ## Changes - Renamed exo_pyo3_bindings to exo_rs - Upgraded versions of pyo3-based dependencies - Renamed PyFromSwarm to just FromSwarm, and PyNetworkingHandle to just NetworkingHandle
051a64e3 · 2026-05-28 · Capture energy in prefill and ageneration separately (#2124)
## Motivation Energy was reported as a single aggregate. Split into prefill vs. generation so each phase can be analysed independently. ## Changes - `PowerSampler`: `mark_prefill_done()` + `trapezoidal_energy_range()` helper; `result()` now emits per-phase splits. - `PowerUsage` / `NodePowerStats`: optional `prefill_*` / `generation_*` fields (back-compat: `None` if unmarked). - API marks the boundary on the first non-`PrefillProgressChunk`. - `bench/exo_bench.py` surfaces the split in the log line and persists `power_usage` to JSON. - METHODOLOGY: one sentence + one bullet. ## Why It Works First non-prefill chunk *is* the boundary. Anchoring a sample there and interpolating power at the boundary makes phase energies sum exactly to the unsplit total. ## Test Plan ### Manual Testing `eco`-reserved nodes: - M3 Ultra, Qwen3-VL-4B, pp=8192/tg=1024: server 1940 J vs client 1931 J (+0.5 %) - M4 Pro, Qwen3.6-27B, pp=16384/tg=2048: server 20,292 J vs client 20,221 J (+0.35 %) ### Automated Testing 5 new tests in `test_power_sampler.py` (range integrator, splits-sum-to-total, `None`-when-unmarked, idempotency). 14/14 pass.
a8602ea6 · 2026-05-26 · fix(bug): no longer repeated _trigger_notify_user_to_download_model (#2114)
## Motivation Partially fixes [this](https://github.com/exo-explore/exo/issues/2098) issue. Removed erroneous logic for telling user to download when they already downloaded. Could not figure out about the "spontaneous crashes" in that issue, author should consolidate more logs and open a new issue dedicated to that. I believe [this](https://github.com/exo-explore/exo/commit/74e9fe15e62fe189dc7e019db86e75c83eca2721) commit solved some EventRouter-related crashes, which was mentioned in [this](https://github.com/exo-explore/exo/issues/2098) issue, so it may have already been solved. If not, should be re-submitted as a new issue. ## Changes - Consolidated _resolve_and_validate_text_model and _validate_image_model into one function: _validate_model_has_instance; - + They already had virtually identical logic, it being different seems to be an artifact of history - + Added logic to ensure that _trigger_notify_user_to_download_model is only called when no such model is downloaded, not just if there is no instance of it - Added a new `/instance/await` SSE streaming endpoint to wait for when a model has an instance available. Complements instance-placement API, so we can wait till that is done without client-side polling. - Updated docs and a /tmp script to reflect some of the changes - Updated dashboard `getModelForRequest` to only return model ID if an instance exists for it, and updated bits to use `handleChatSend` instead of `sendMessage` because that checks for if a model instance exists first. ## Why It Works The problem was that there was erroneous logging for model not downloaded. I fixed that logic. The rest is extra.
a1a22b5f · 2026-05-25 · feat: added background/daemon support (#2106)
## Motivation Addresses [this](https://github.com/exo-explore/exo/issues/1931) issue. ## Changes You can now launch Exo as a legacy SysV-style daemin (in the background) with `--legacy-daemon` flag. NOTE: don't use it if you're managing Exo with systemd or launchd SIDE FIX: the macmon process not found trace is no longer displayed on process shutdown via ctrl+c, that error is supressed. ## Why It Works Because I used a daemonization library and tweaked it not to break multiprocessing. ## Test Plan I ran it in daemon mode, non daemon mode, etc., and pid locking + inference + everything else works just fine. Also ran it `ssh user@host -t 'cd exo && nohup nix run .#exo -- --legacy-daemon'` on a 4-node TB mac-mini cluster and the mDNS didn't die
74e9fe15 · 2026-05-22 · fix(bug): EventRouter lifetime-handling fixed, no more process crashes (#2102)
## Motivation Trying to (partially) fix [this](https://github.com/exo-explore/exo/issues/2101) issue. ## Changes Changed channels (in channels.py) to support exception overriding. Made EventRouter channels throw a subclass of the resource closed/broken errors. The current lifetime logic of EventRouter in event loop no longer blows up because components that use channels from EventRouter now catch the subclass exceptions in the run method: Worker, Master, DownloadCoordinator, RunnerSupervisor. Added logic to throw when API server exits without being asked to shut down - this kill the sleep-forever in the task-group.
14aab356 · 2026-05-15 · Runner error handling (#2093)
# Runner error handling ## Motivation Runner failures were mostly surfaced as plain shutdown messages, which made root cause hard to spot from API errors or runner status. This adds a MVP path for preserving runner crash context and attaching known stderr diagnostics to failure reports. ## Changes - Added `RunnerTerminationError` for Python exceptions raised inside runner bootstrap - Changed runner bootstrap to send `Event | RunnerTerminationError` over the private runner channel - Moved public `RunnerFailed` emission back into supervisor - Added stderr-only `RunnerDiagnosticCollector` - + Added known diagnostics for Metal GPU timeout, ring socket receive errno, and ring transport abort - Added diagnostics to `RunnerFailed` and `ErrorChunk` - Tweaked async process termination to join briefly before terminate/kill - Updated tests/fixtures for new failure payload shape - Added Ruff VS Code formatter settings ## Why It Works Runner child now reports raw-ish failure context to supervisor instead of publishing failed status directly. Supervisor still owns process lifecycle, exit code/signal handling, in-flight task error chunks, and final runner status. Stderr diagnostics stay best effort and only known root-cause variants are surfaced. ## Test Plan ### Manual Testing Hardware: remote runner logs from e16/e11/e4/e2 What you did: - inspected live runner stderr logs - used observed Metal GPU timeout and ring socket errors as initial diagnostic targets ### Automated Testing - `nix flake check` - supervisor test covers error chunk + failed status emission - plan lifecycle test updated for failed runner diagnostics - type/lint checks cover new runner channel union --------- Co-authored-by: Evan Quiney <evanev7@gmail.com>
88d46d46 · 2026-05-14 · fix: omit null delta fields in streaming chat completions (issue #2082) (#2092)
## Motivation
Streaming /v1/chat/completions responses emitted null for tool_calls,
function_call, name, and tool_call_id in every delta chunk. The OpenAI
streaming spec marks these fields as non-nullable — they must either
carry a
real value or be absent entirely. Spec-correct clients doing
delta.get("tool_calls", []) receive None and crash with 'NoneType'
object is
not iterable.
Root cause: the streaming serialisation path called model_dump_json()
without
exclude_none=True, while the request-parsing path already used it
correctly.
Three call sites in chat_completions.py and two in responses.py were
affected.
## Testing
Before — every delta carries explicit nulls:
$ curl -sN -X POST http://localhost:52415/v1/chat/completions \
-H 'Content-Type: application/json' \
-d
'{"model":"mlx-community/Qwen3.5-2B-MLX-8bit","messages":[{"role":"user","
content":"hi"}],"max_tokens":3,"stream":true}' \
| grep "^data: "
data:
{"id":"7c4dae10-...","choices":[{"index":0,"delta":{"role":"assistant","c
ontent":null,"reasoning_content":"Okay","name":null,"tool_calls":null,"tool_cal
l_id":null,"function_call":null},"logprobs":null,"finish_reason":null,"usage":n
ull}],"usage":null,"service_tier":null}
data:
{"id":"7c4dae10-...","choices":[{"index":0,"delta":{"role":"assistant","c
ontent":null,"reasoning_content":",","name":null,"tool_calls":null,"tool_call_i
d":null,"function_call":null},"logprobs":null,"finish_reason":null,"usage":null
}],"usage":null,"service_tier":null}
data:
{"id":"7c4dae10-...","choices":[{"index":0,"delta":{"role":"assistant","c
ontent":"
the","reasoning_content":null,"name":null,"tool_calls":null,"tool_cal
l_id":null,"function_call":null},"logprobs":null,"finish_reason":"length","usag
e":{"prompt_tokens":11,...}}],"usage":null,"service_tier":null}
data: [DONE]
After — only populated fields are emitted:
data:
{"id":"demo","object":"chat.completion","created":...,"model":"mlx-commun
ity/Qwen3.5-2B-MLX-8bit","choices":[{"index":0,"delta":{"role":"assistant","rea
soning_content":"Okay"}}]}
data:
{"id":"demo","object":"chat.completion","created":...,"model":"mlx-commun
ity/Qwen3.5-2B-MLX-8bit","choices":[{"index":0,"delta":{"role":"assistant","rea
soning_content":","}}]}
data:
{"id":"demo","object":"chat.completion","created":...,"model":"mlx-commun
ity/Qwen3.5-2B-MLX-8bit","choices":[{"index":0,"delta":{"role":"assistant","con
tent":"
the"},"finish_reason":"length"}],"usage":{"prompt_tokens":11,"completio
n_tokens":3,"total_tokens":14,...}}
data: [DONE]
e8ec8d50 · 2026-05-14 · fix ollama API compatibility for VS Code Copilot (#2091)
Ollama adapter fixes for VS Code Copilot (#2042): - /api/version: bare semver "1.0.0" - Copilot parseInts each segment. - /api/show: populate model_info + capabilities - Copilot crashes on null model_info and filters by `tools`. - Add POST /ollama/v1/chat/completions - ollama serves the OpenAI-compat route here, BYOK clients 405 without it. Before: <img width="1380" height="144" alt="image" src="https://github.com/user-attachments/assets/99d5464f-187d-4432-9a31-8229c55aa209" /> After: <img width="1362" height="181" alt="image" src="https://github.com/user-attachments/assets/361dc006-d8df-435f-8d8b-4fa4f44a8c23" /> <img width="279" height="909" alt="image" src="https://github.com/user-attachments/assets/4621aba7-bd57-4762-8568-34a3383a6025" /> Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
1fd15d59 · 2026-05-14 · create directory on startup (#2089)
## Motivation <!-- Why is this change needed? What problem does it solve? --> <!-- If it fixes an open issue, please link to the issue here --> When you first run `uv run exo` you get an error like : `FileNotFoundError: [Errno 2] No such file or directory: '/Users/heidar/.exo/models'` Manually tested on Macbook Pro M1 32GB Fixes issue - https://github.com/exo-explore/exo/issues/2090
ed2d10bd · 2026-05-12 · Redirect runner stdout/stderr to file logs (#2084)
## Motivation We want to use log mining tools like [Drain3](https://github.com/logpai/Drain3) to get standardized error formats, but for that we should record runner stdout/stderr in a massive append-only log to gather training data for such tools. Also useful for future opt-in telemetry. ## Changes The stdout/stderr from runner now splits into 3 tasks: 1) raw write to dedicated runner logs 2) sanitized line-by-line logging with log-guru 3) stub for further error-processing (i.e. turning lines into errors) ### Manual Testing Works on 4x mac mini clusted connected as TB4 ring.
87c72fc1 · 2026-05-11 · Fixes issue #2068 (#2083)
## Motivation To fix https://github.com/exo-explore/exo/issues/2068 ## Changes Adds queue shutdown logic & hard-timeouts for closing server. ## Why It Works Prevents API from hanging more than 5 seconds.
08ffa5f6 · 2026-05-10 · Map GLM 4.7 stop tokens to GLM 4 IDs (#2061)
## Motivation GLM 4.7 reuses the GLM 4 chat-template tokenizer, but the model card and EOS-detection path didn't have an explicit mapping for it, so OpenAI-compatible clients didn't see a clean stop and the runner emitted follow-on role turns (e.g. \`<|user|>\` continuations after \`<|assistant|>\`'s output). ## Changes \`src/exo/worker/engines/mlx/utils_mlx.py\` — add the GLM 4 stop-token IDs as the EOS set when the loaded model's tokenizer matches GLM 4 / 4.7 chat templates. ## Why It Works The GLM 4 tokenizer's \`<|user|>\`, \`<|observation|>\`, and \`<|endoftext|>\` IDs are stable across the GLM 4 / 4.7 line; treating any of them as EOS lets the runner stop at the assistant turn boundary the same way it stops at \`</s>\` for Llama-style models. No prompt-template changes — only the stop set widens. ## Test Plan ### Automated Testing New unit test \`src/exo/worker/tests/unittests/test_mlx/test_eos_token_ids.py\` covering: GLM 4 / 4.7 path returns the expected stop ID set; non-GLM path returns the standard EOS only. \`\`\` src/exo/worker/tests/unittests/test_mlx/test_eos_token_ids.py .. === 2 passed in 0.01s === \`\`\` \`uv run basedpyright\` and \`uv run ruff check\` both clean. ### Manual Testing Hardware: 4-node Apple Silicon cluster, M5 Max master. - Loaded \`mlx-community/GLM-4.7-Air-mlx-4bit\`, ran chat completion via \`/v1/chat/completions\`. Before this fix the assistant turn ran on into a synthetic \`<|user|>\` continuation; after the fix the response stops cleanly at the assistant boundary. --------- Co-authored-by: jw-wcv <101585096+jw-wcv@users.noreply.github.com> Co-authored-by: Evan Quiney <evanev7@gmail.com>
45df74ba · 2026-05-09 · Andrei/mp capture stdio (#2056)
## Motivation Process-isolated runner crashes and C-extension failures can write directly to fd-level stdout/stderr, bypassing Python/loguru. We need to capture that output per runner process without polluting the main process or other workers, and without breaking operation when the parent stdio is detached. ## Changes - Added `AsyncProcess`, a spawn-only multiprocessing wrapper that redirects child stdout/stderr to pipes and exposes them as in-memory `Receiver[bytes]`s - Replaced runner-supervisor's raw `multiprocessing.Process` usage with `AsyncProcess` - Added `--no-stdio`, redirecting stdin/stdout/stderr to `/dev/null` after logging is configured - Disabled verbose MLX - Added tests covering stdio capture, child crashes, repeated bad children, SIGTERM/SIGKILL shutdown escalation, stdio detachment, and spawning captured children from a stdio-detached parent ## Why It Works The parent can redirect its own stdio fds to `/dev/null`, while `AsyncProcess` installs fresh pipe fds over fd 1 and 2 inside each spawned child. That keeps stdio-detached parents quiet while preserving per-runner stdout/stderr capture. Runner shutdown is still bounded: SIGTERM grace first, then SIGKILL escalation if needed. Next direction: the runner supervisor currently drains captured output and logs it as stdout/debug and stderr/warning. This should be split into more useful process-isolated error reporting instead of just log forwarding (regex match on errors to obtain "reason" string, best effort). ## Test Plan ### Manual Testing Ran on 4 Mac Minis in a Thunderbolt 4 ring, can see that runner's stdout/stderr contents are being captured. ### Automated Testing - Added async-process tests for fd-level stdout/stderr capture, Python traceback capture, bounded-buffer output, child `exit`/abort, parent stdio preservation, fd leak checks, spawn-context mp channels, and SIGTERM/SIGKILL shutdown behavior - Added stdio-detach tests proving stdio detaches to `/dev/null`, a stdio-detached parent can still spawn and capture a child, and the same stdio-detached parent can spawn/capture multiple children sequentially - Updated runner-supervisor tests for the new `AsyncProcess.exitcode` path
ce37bdce · 2026-05-09 · fix: Create directory for PID file if it doesn't exist (#2075)
Ensure the directory for the PID file exists before creating it. ## Motivation Fixes https://github.com/exo-explore/exo/issues/2074 ## Changes <!-- Describe what you changed in detail --> ## Why It Works <!-- Explain why your approach solves the problem --> ## Test Plan ### Manual Testing <!-- Hardware: (e.g., MacBook Pro M1 Max 32GB, Mac Mini M2 16GB, connected via Thunderbolt 4) --> <!-- What you did: --> <!-- - --> ### Automated Testing <!-- Describe changes to automated tests, or how existing tests cover this change --> <!-- - -->
e5a1e5da · 2026-05-08 · Create PID file locking for EXO (#2072)
## Motivation EXO should be PID file locked, to prevent duplicate processes from clobbering the log, right now this isn't the case. ## Changes I added a wrapper around a Rust PID file lock library, and used it to implement PID locking for EXO, with the PID file being in exo cache directory. ## Test Plan ### Manual Testing Tested on e11, trying to spawn duplicate EXO processes prevented.
fa571313 · 2026-05-08 · Integration tests infra (#1995)
## Motivation No automated integration tests exist for exo. Manual testing against real hardware clusters is slow and error-prone. We need a pytest framework that deploys clusters via `eco`, runs inference scenarios, and tears down cleanly. ## Changes - **`tools/src/exo_tools/`** — New workspace member shared by bench, eval, and tests: - `client.py` — `ExoClient` HTTP client (extracted from `bench/harness.py`) - `harness.py` — instance lifecycle helpers (placement, wait-for-ready, etc.) - `cluster.py` — `EcoSession` for eco cluster lifecycle (deploy/stop/start/release/logs/exec) with unique `USER=<prefix>-<uuid>` per session and atexit/signal cleanup - **`tests/integration/`** — 17 pytest tests across 5 files: - `test_1node.py` — place, chat, multi-turn, delete, state/models endpoints, cluster snapshot, download-from-scratch - `test_2node.py` — parametrized tensor/jaccl + pipeline/ring inference and multi-turn - `test_4node.py` — parametrized 4-node pipeline/ring inference, cluster state - `test_resilience.py` — full disconnect/reconnect cycle (2-node → disconnect → 1-node → reconnect → 2-node) - `test_dashboard.py` — Playwright: dashboard loads, shows node info, chat flow - `helpers.py` — placement/inference helpers, re-exports from `exo_tools` - `conftest.py` — session-scoped cluster fixtures with constraint-based eco reservations; `--hosts` override; `EXO_REF` env var for CI deployments from a GitHub branch - **`bench/`** — Updated imports from `exo_tools.client` / `exo_tools.harness` - **`pyproject.toml`** — Added `tools` workspace member, `playwright` dev dep, `--ignore=tests/integration` ## Why It Works Tests use `eco` for cluster lifecycle and `ExoClient` for API interactions — same tools humans use. Session-scoped fixtures deploy once per file. Unique eco users prevent test runs from interfering with each other or manual usage. ## Test Plan ### Automated Testing - `uv run pytest tests/integration/ -v -s` — full suite (~4-5 min, 17/17 passing) - `uv run pytest tests/integration/ -v -s --hosts s4,s9,s10,s22` — pin specific hosts - `EXO_REF=main uv run pytest tests/integration/ -v` — deploy from a GitHub branch (CI) - `uv run pytest` — confirms integration tests are excluded from default runs
414132ae · 2026-05-07 · Use time-weighted power sampling (#2038)
## Why The power sampler currently averages sampled wattage values arithmetically. That can be materially wrong when sample intervals are uneven: a short high-power spike gets the same weight as a long steady interval. Energy should be computed by integrating power over time, and average power should be derived from energy / elapsed time. ## How - Store each power sample with its relative timestamp. - Anchor the first sample at `t=0` and take a final sample at `elapsed` when producing results. - Integrate per-node power using the trapezoidal rule. - Sum node energy for total cluster energy, then derive total average system power from total energy / elapsed. - Add focused unit tests for uneven sample intervals and the single-sample fallback. ## Tests - `uv run pytest src/exo/utils/tests/test_power_sampler.py` - `uv run basedpyright` - `uv run ruff check src/exo/utils/power_sampler.py src/exo/utils/tests/test_power_sampler.py` - `nix fmt`
edef8004 · 2026-05-07 · Store custom model cards in State (#2024)
## Why Workers currently update their custom model-card cache by reacting to `CustomModelCardAdded` / `CustomModelCardDeleted` events directly. That is another snapshot footgun: a worker restored from State may never see the historical add/delete event, so the durable State must include the desired custom-card set. ## How - Add `State.custom_model_cards`, keyed by `ModelId`. - Reduce `CustomModelCardAdded` into State. - Reduce `CustomModelCardDeleted` into State. - Add focused reducer tests for add and delete. This PR only makes custom cards durable in State. A follow-up PR will make workers reconcile their on-disk custom-card cache from this state instead of relying on those events directly. ## Tests - `uv run pytest src/exo/shared/tests/test_apply/test_apply_custom_model_cards.py src/exo/shared/tests/test_state_serialization.py` - `uv run pytest` - `uv run ruff check src/exo/shared/types/state.py src/exo/shared/apply.py src/exo/shared/tests/test_apply/test_apply_custom_model_cards.py` - `uv run basedpyright` - `nix fmt`
a0c00f9d · 2026-05-07 · fix(placement): gate RDMA on nodeRdmaCtl.enabled at both endpoints (#2014)
## Summary
- Fixes a bug where `POST /place_instance` (and the dashboard UI) would
accept an MlxJaccl/RDMA instance spanning nodes whose
`nodeRdmaCtl.enabled` was `false`, because topology + placement
consulted Thunderbolt-derived RDMA edges without checking the per-node
`rdma_ctl` status.
- Three-layer fix: topology only emits `RDMAConnection` edges when both
endpoints have `nodeRdmaCtl.enabled = true`; flipping a node to disabled
immediately purges every RDMA edge touching it; `place_instance`
additionally rejects RDMA cycles containing any disabled or unobserved
node as a defense-in-depth check on the API/master path.
## Details
- `src/exo/shared/apply.py`
- `MacThunderboltConnections` case now filters out RDMA connections
whose source or sink lacks observed-and-enabled `rdma_ctl` status
(missing entry → treated as disabled).
- `RdmaCtlStatus` case now calls
`topology.remove_all_rdma_connections_touching(node_id)` when the node
reports disabled, so consumers don't have to wait for the next TB poll.
- `src/exo/shared/topology.py`
- New `Topology.remove_all_rdma_connections_touching(node_id)` removes
every RDMA edge incident to the node (incoming and outgoing) while
leaving socket edges intact.
- `src/exo/master/placement.py`
- `place_instance` accepts `node_rdma_ctl: Mapping[NodeId,
NodeRdmaCtlStatus] | None`. The `is_rdma_cycle` filter now also requires
`nodeRdmaCtl.enabled` for every node in the cycle. MlxJaccl placement
raises the existing "no RDMA-connected cycles available" error if no
qualifying cycle remains.
- `src/exo/api/main.py`, `src/exo/master/main.py`
- Both placement entrypoints now pass `state.node_rdma_ctl` through.
## Tests
- `src/exo/shared/tests/test_apply/test_apply_rdma_gating.py` (new): six
unit tests covering enabled/disabled/missing combinations on apply, the
immediate-purge transition, and that purging RDMA edges leaves socket
edges untouched.
- `src/exo/master/tests/test_placement.py`: existing
`test_tensor_rdma_backend_connectivity_matrix` updated to pass
`node_rdma_ctl`. Two new tests assert MlxJaccl placement is rejected
when any cycle node is `enabled=false` or has no `rdma_ctl` entry.
## Test plan
- [x] `uv run basedpyright` — 0 errors
- [x] `uv run ruff check` — clean
- [x] `nix fmt`
- [x] `uv run pytest` — 429 passed, 1 skipped
- [ ] On a real mixed cluster (s15/s16 disabled, s17/s18 enabled),
confirm:
- [ ] `POST /place_instance` for an RDMA instance including s15 or s16
returns an error
- [ ] An RDMA instance can still be placed across {s17, s18}
- [ ] `GET /state` shows no `sourceRdmaIface`/`sinkRdmaIface` on s15↔s16
connections
- [ ] Dashboard previews don't surface RDMA-spanning options that
include s15/s16
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
89d20c18 · 2026-05-06 · fix(inference): prevent TP collective deadlock via agree_on_tasks order (#2048)
If you have two machines and make two requests at the same time, it can crash. This is because the tasks can sometimes end up in different orders on different machines. We need to sort the tasks and mx_all_gather_tasks already sorts the tasks but the code ignores that ordering. The fix is to make sure the sort order is preserved. The rest is written by Sonnet (reviewed by me): Tensor-parallel inference requires that every rank enqueues tasks in the same order before running agree_on_tasks collectives. The old implementation filtered from _maybe_queue: self._queue.extend(task for task in self._maybe_queue if task in agreed) self._maybe_queue = [task for task in self._maybe_queue if task in different] Because _maybe_queue is independently ordered per-rank (tasks arrive via gRPC in whatever order the API server sends them), two concurrent requests could produce different _maybe_queue orderings on rank 0 vs rank 1. The filter then preserved those different orders into _queue, so each rank started processing tasks in a different sequence. The next mlx collective (all_reduce, all_gather, etc.) on rank 0 corresponded to a different task than on rank 1 → permanent deadlock. Fix: extend from agreed directly. mx_all_gather_tasks returns agreed as a list sorted by task_id on all ranks, so every rank appends the same sequence regardless of local arrival order. Applies to both SequentialGenerator and BatchGenerator. ## Motivation `agree_on_tasks` is called on every rank after accumulating new requests in `_maybe_queue`. Its job is to run an `all_gather` collective so all ranks agree on which tasks to promote to `_queue` before the next inference step. The old implementation re-imposed **local arrival order** when extending `_queue`: ```python self._queue.extend(task for task in self._maybe_queue if task in agreed) ``` `mx_all_gather_tasks` already returns `agreed` sorted by `task_id` — the same deterministic order on every rank. But iterating `self._maybe_queue` instead of `agreed` discarded that sort and substituted the local gRPC arrival order, which differs per rank under concurrent load. Two concurrent requests arriving in `[A, B]` order on rank 0 and `[B, A]` on rank 1 caused the first MLX collective in the next step to hang permanently: each rank was executing a different task's collective and would never match. ## Changes `SequentialGenerator.agree_on_tasks` and `BatchGenerator.agree_on_tasks`: ```python # Before self._queue.extend(task for task in self._maybe_queue if task in agreed) self._maybe_queue = [task for task in self._maybe_queue if task in different] # After self._queue.extend(agreed) # preserves mx_all_gather_tasks sort order self._maybe_queue = list(different) # already in local order; filter was redundant ``` ## Why It Works `mx_all_gather_tasks` (in `utils_mlx.py`) computes the agreed set then sorts by `task_id`: ```python agreed = [local_tasks[tid] for tid in sorted(agreed_ids)] ``` Because `task_id` is a UUID and the sort is lexicographic, every rank produces the same `agreed` list regardless of local arrival order. Using `agreed` directly preserves this guarantee. The `different` list (tasks not yet seen on all ranks) is built by iterating `tasks` in local order, which is already correct. ## Test Plan ### Manual Testing **Hardware:** 2× Mac Studio M3 Ultra 512 GB, Thunderbolt 5 direct bridge, `MlxJaccl` RDMA tensor-parallel (`moonshotai/Kimi-K2.6`, 595 GB INT4, 61 layers). - Sent concurrent streaming requests; confirmed all complete without deadlock. - This hardware configuration (sub-millisecond inter-node latency) is the most likely to trigger the race, as requests from separate HTTP connections can reach rank 0 and rank 1 in opposite order before `agree_on_tasks` runs. ### Automated Testing All existing tests pass: `pytest src -m "not slow" --import-mode=importlib` — 422/422 passed. The existing `test_event_ordering.py` covers the `agree_on_tasks` call path with a mock that returns tasks in consistent order; the race requires real distributed hardware to reproduce deterministically.
dbcceaa5 · 2026-05-05 · Initialise _cancelled_tasks in ImageEngine (#2051)
we yielded nonsense chunks from engines; we didn't initialize the image engine correctly. mostly rewrite of #2049 --------- Co-authored-by: ciaranbor <ciaranborourke-dev@proton.me>
9c6ff4ce · 2026-05-01 · feat: update rdma_ctl instructions (#1977)
## Motivation The RDMA setup instructions were missing a step: after booting to Recovery mode, users need to open Terminal from the Utilities menu before they can run the `rdma_ctl` command. Without this step, users following the instructions wouldn't know how to access a terminal in Recovery mode. This step was already in the README just not in the UI notifications. ## Changes Added a missing instruction step — "Open Terminal from the Utilities menu" — to three instances of the RDMA setup flow in `dashboard/src/routes/+page.svelte`. ## Why It Works N/A copy change only. ## Test Plan ### Manual Testing Hardware: MacBook Pro M4 Max 48GB ### Automated Testing No automated tests affected; this is a UI copy change only. Co-authored-by: Sam Bradbury <sam@consultbradbury.com>
b26268df · 2026-05-01 · fix(macos-app): disable URL response caching for cluster-state polling (#2005)
Fixes #2004.
`ClusterStateService` polls `/state` at 2 Hz via `URLSession.shared`,
which keeps an on-disk `URLCache` attached by default. Every polled
response body gets persisted under `~/Library/Caches/exolabs.EXO/`,
sustaining ~500–620 KB/sec of file-backed memory dirtied — far above
macOS's ~25 KB/sec per-process daily-average baseline. Six
microstackshot reports observed on a single Mac Studio M3 Ultra over
eight days, with one 15-hour run accumulating 34.36 GB of cache writes.
Heaviest stack on every diagnostic report (96–98% of samples):
```
_dispatch_workloop_worker_thread → _dispatch_block_async_invoke2 →
__CFURLCache::CreateAndStoreCacheNode → write
```
Full diagnostic data and analysis in #2004.
## What changed
`ClusterStateService` now defaults to an ephemeral, non-caching
`URLSession` instead of `URLSession.shared`. Cluster-state responses are
time-sensitive and small; nothing benefits from being cached on disk.
```swift
private static func makeNonCachingSession() -> URLSession {
let config = URLSessionConfiguration.ephemeral
config.urlCache = nil
config.requestCachePolicy = .reloadIgnoringLocalCacheData
return URLSession(configuration: config)
}
```
The existing per-request `request.cachePolicy =
.reloadIgnoringLocalCacheData` calls are kept as defense in depth — they
only affect read behavior, but harmless to leave alongside the
session-level config.
## Scope
- **Behavioral**: none. Polled requests still go out at the same
cadence; responses still parse the same; no semantic change to any API
surface.
- **Test injection**: the `session:` parameter remains in `init`, so
tests can still inject a custom mock session unchanged.
- **`BugReportService` and other `URLSession.shared` callers**:
untouched. If maintainers prefer an app-wide URLCache disable instead,
happy to switch the approach (issue body has the alternative spelled
out).
## Verification
Verified locally that compiling EXO with this change produces a working
menubar app and `ClusterStateService` continues to fetch state
correctly. After ~30 min of idle polling, no new entries in
`/Library/Logs/DiagnosticReports/EXO_*.diag` and no growth in
`~/Library/Caches/exolabs.EXO/`.
## Test plan
- [ ] Build EXO from this branch on macOS 26.4
- [ ] Launch, let cluster state polling run for 30+ min
- [ ] Confirm no new microstackshot diagnostic reports
- [ ] Confirm `~/Library/Caches/exolabs.EXO/Cache.db*` does not grow
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-authored-by: Jordan Miller <jordan.d.miller@gmail.com>
8dae3ecb · 2026-04-30 · A few targeted tweaks to address HF rate limits (#2009)
## Motivation - exo bursts ~200 HF Hub-API requests on every cold start, blowing past the anonymous 500-req/5-min budget. - The existing retry loop catches 429 generically and gives up in ~3s — well before HF's reset window. - `file_meta` and `_download_file` had no 429 handling at all (became `AssertionError`). - Disk file-list cache was bypassed on every process restart. ## Changes All in `src/exo/download/download_utils.py` + tests. - Parse `t=` from HF's `RateLimit` header on 429; sleep `min(t, 300s) + jitter`. - Handle 429 at all three call sites (`_fetch_file_list`, `file_meta`, `_download_file`). - `n_attempts`: 3 → 5. - Disk cache now primary across restarts (24h mtime TTL). - `?recursive=true` instead of N+1 subdir walks. ## Why It Works `t=<seconds>` is HF's "wait this long and you'll be unblocked" — sleeping that long lets the window reset. Disk-cache-as-primary plus recursive listing cuts cold-start Hub-API traffic by ~10×. ## Test Plan ### Manual Testing MacBook Pro M1 Max. Tripped the real HF 429. Pre-fix: failed in 3.4s. Post-fix: slept (HF returned `t=158`) and recovered. ### Automated Testing - New `test_rate_limit_handling.py` (19 tests) — header parsing, retry-loop behaviour, plus HTTP-level coverage that mocks aiohttp to return a 429 and asserts each call site raises `HuggingFaceRateLimitError(retry_after=52.0)`. - New `TestFileListCacheTTL` in `test_offline_mode.py` — fresh cache hits, stale cache refetches. - 421 tests pass; basedpyright / ruff / nix fmt clean.
fb12b403 · 2026-04-30 · fix(app): tighten Share Bug Report prompt layout (#2008)
## Summary Follow-ups to #2003 based on feedback that the Share Bug Report window felt visually weighty: too much padding above and below, and a description editor that invited an essay rather than a one-liner. ## Changes (one file) `app/EXO/EXO/Views/BugReportWindowController.swift`: - **Auto-size the window to its content.** Switched from `NSHostingView` + fixed `contentRect: 480x380` + SwiftUI `frame(minHeight: 320)` to `NSHostingController` with `sizingOptions = [.preferredContentSize, .minSize]`. The fixed-min combo was centering the form in dead vertical space. - **Smaller, lower-pressure editor.** Field is now labeled `Description (optional)` with a placeholder hint (`What were you doing when it broke?`) inside the editor. Editor height fixed at 72pt (was 120pt min). Replaced the long lead-in paragraph and headline with a single one-line caption between field and buttons: `Diagnostic logs will be uploaded with your report.` - **Tighter spacing.** Outer padding 20 -> 16, root spacing 16 -> 12, prompting-section spacing 12 -> 8. - **Remove em dash from copy.** `BugReportService` and the menu wiring are unchanged. ## Test plan - [ ] Click `Share Bug Report...` from the menu bar. - [ ] The window opens centered and sized to its content (no big empty bands top/bottom). - [ ] Description editor is visibly compact, with the placeholder hint showing when empty. - [ ] The optional-ness is conveyed by the field label (no separate help paragraph). - [ ] Caption `Diagnostic logs will be uploaded with your report.` appears in `.caption` style under the editor, above the buttons. - [ ] Resize the window: persists across re-opens (frame autosave still works). - [ ] Send/Cancel/Try Again/Done flows behave the same as before. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
1606e638 · 2026-04-30 · feat(app): open Share Bug Report in a dedicated window (#2003)
## Summary
- Adds a top-level **Share Bug Report…** menu item to the macOS popover
(between *Check for Updates* and *Quit*) with SF Symbol `ladybug`.
- Clicking it opens a dedicated resizable `NSWindow` ("Send a Bug
Report") that hosts the prompting / sending / success / failure flow.
- Removes the description-less duplicate from Settings → Debug Info, and
the dead `debugSection` it nominally lived behind.
## Why
PR #1959 added a user-description prompt to the bug-report flow, but its
trigger lived inside `ContentView.debugSection` — a view that's defined
but never rendered in the body. The path users actually hit was
`SettingsView.sendBugReportButton`, which called
`BugReportService.sendReport(isManual: true)` without ever passing
`userDescription`. So the description prompt was unreachable in the
built app.
## Approach
Per Apple HIG, an action that requires further input before completing
should open a dialog, not transform the menu inline. So:
- Add a top-level menu entry that ends in `…` (HIG: ellipsis indicates
"further input required").
- Move the prompting/sending/success/failure state machine into a
standalone `BugReportWindowController` modeled after the existing
`SettingsWindowController`.
- Single-instance window with frame-autosave name, sensible
`contentMinSize`, resizable, native button layout (`.cancelAction` /
`.defaultAction` keyboard shortcuts), light/dark-mode-correct
`.textBackgroundColor` and `.separatorColor`.
- Auto-focus the description field on open. `Try Again` from failure,
`Open GitHub Issue` + `Done` from success.
## Files
- `app/EXO/EXO/Views/BugReportWindowController.swift` (new) — controller
+ view.
- `app/EXO/EXO/EXOApp.swift` — wire `BugReportWindowController` as a
`@StateObject` and inject as environment object.
- `app/EXO/EXO/ContentView.swift` — replace inline state machine with
menu item that calls `bugReportWindowController.open()`. Remove
now-unused state, helpers, and dead `debugSection`.
- `app/EXO/EXO/Views/SettingsView.swift` — remove duplicate
`sendBugReportButton`, `sendBugReport()`, and related `@State`. Section
"Debug Info" keeps Thunderbolt / interface / RDMA info.
`BugReportService` is unchanged.
## Test plan
- [ ] Open the menu-bar popover → confirm **Share Bug Report…** appears
between *Check for Updates* and *Quit*, with a ladybug icon.
- [ ] Click it → a window titled "Send a Bug Report" appears, centered,
with the description editor focused.
- [ ] Resize the window → size persists across re-opens (frame
autosave).
- [ ] Type a description, press Return → upload succeeds, success card
with **Open GitHub Issue** + **Done** appears.
- [ ] Click **Open GitHub Issue** → browser opens with the description
pre-filled into the issue template.
- [ ] Send with empty description → upload still succeeds.
- [ ] Press Esc from the prompting state → window closes.
- [ ] On failure (e.g., offline) → error card with **Try Again** +
**Close** appears; Try Again returns to the editor with the description
preserved.
- [ ] Open the Settings window → Debug Info section is unchanged except
the Send Bug Report button is gone.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
---------
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
667a3bb0 · 2026-04-28 · feat: keep-models option when uninstalling EXO (#1997)
## Summary - Adds a **Keep downloaded models (~/.exo/models)** checkbox to the macOS uninstall confirmation dialog (Settings → Advanced → Danger Zone). The full `~/.exo` directory is now removed on uninstall by default; if the checkbox is checked, `~/.exo/models` is preserved. - The standalone `app/EXO/uninstall-exo.sh` gains a matching `--keep-models` flag and the same `~/.exo` cleanup so GUI and CLI flows stay in sync. Resolves the user home via `$SUDO_USER` since the script runs under `sudo`. Previously, "Uninstall EXO" only cleaned up system-level components (LaunchDaemon, network location, logs, app bundle) and left the entire `~/.exo` directory behind. Now uninstalling actually removes EXO's user data, with a one-click opt-out for the (potentially many GB) of downloaded models.  > Note: the rendered icon in the screenshot above is the generic system folder icon because it was captured from a small standalone Swift binary (no app bundle / icon resource). When triggered from the actual EXO.app, the EXO app icon is shown. ## Test plan - [ ] Build EXO.app locally; open Settings → Advanced → Danger Zone → Uninstall EXO; confirm the new "Keep downloaded models (~/.exo/models)" checkbox is present and unchecked by default. - [ ] Uninstall with the checkbox **checked** → `~/.exo/models/` survives, all other entries under `~/.exo` are gone, system components removed, app moved to Trash. - [ ] Uninstall with the checkbox **unchecked** → `~/.exo` is fully removed. - [ ] `sudo app/EXO/uninstall-exo.sh --keep-models` → `~/.exo/models/` is preserved, the rest of `~/.exo` is removed. - [ ] `sudo app/EXO/uninstall-exo.sh` (no flag) → `~/.exo` is fully removed. - [ ] `app/EXO/uninstall-exo.sh --help` prints usage and exits 0; unknown args exit 2 with a usage hint. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> Co-authored-by: Evan <evanev7@gmail.com>
18ffe1df · 2026-04-28 · fix: uninstall-exo.sh removes both current and legacy bridge scripts (#1998)
## Summary The standalone `app/EXO/uninstall-exo.sh` only knew about the legacy filename `disable_bridge_enable_dhcp.sh`. On machines installed with newer EXO versions, the current `/Library/Application Support/EXO/disable_bridge.sh` was left behind, and the script then reported `EXO support directory not empty, leaving in place`. This PR makes the script try both filenames, removing whichever ones exist. Tolerates **either**, **both**, or **neither** being present without erroring. The Swift `NetworkSetupHelper.makeUninstallScript()` already handles both paths correctly, so the GUI uninstall flow is unaffected — this is a script-only fix. Caught while running an end-to-end uninstall on a real machine for #1997. ## Test plan Verified the new block in isolation against all four states: - [x] both `disable_bridge.sh` and `disable_bridge_enable_dhcp.sh` present → both removed - [x] only `disable_bridge.sh` present → removed cleanly - [x] only `disable_bridge_enable_dhcp.sh` present → removed cleanly (legacy install) - [x] neither present → prints the existing "already removed?" warning, exits 0 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
5d10188d · 2026-04-27 · fix: route by in-flight tasks only — completed tasks were skewing load balance (#1989)
The load balancer counted ALL tasks (Complete, Cancelled, TimedOut, Failed) instead of only Pending/Running ones. With 138 accumulated tasks and only 7 active, routing decisions were based on historical distribution, causing one node to appear permanently 'busier' and starving the other of work. Co-authored-by: Adam Durham <adam@example.com> Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
f2a0db4e · 2026-04-27 · Extend bench/eval tooling (#1905)
## Motivation Extend bench/eval tooling with robustness features, streaming support, and align model configs with vllm eval for reproducible comparisons. ## Changes - **exo_eval**: Checkpoint/resume (JSONL), instance health monitoring + early abort, `top_k`/`min_p`/`enable_thinking` params, LCB `--release-version`/`--offset` - **exo_bench**: Streaming SSE (`--stream`), Kimi tokenizer fix for transformers 5.x - **Both tools**: Auto-detect running instances instead of requiring `--skip-instance-setup`; `--fresh-instance` to override - **harness**: SSE streaming client, `find_existing_instance()` shared helper, removed download timeout, settle-timeout default 0→7200s - **models.toml**: Added `enable_thinking`, aligned `max_tokens`/temps with vllm, added new models - **API**: Streaming SSE for `/bench/chat/completions` ## Why It Works - Checkpoint/resume uses append-only JSONL + skip-on-load so interrupted evals resume without re-running completed questions - Health monitoring races an `asyncio.Event` against API calls for fast abort when the instance dies - Auto-detection queries `/state` for existing instances matching the model ID before attempting placement - Streaming reuses the existing `generate_chat_stream` infrastructure from the regular chat endpoint
48a922fd · 2026-04-27 · fix: map presence_penalty and frequency_penalty from ChatCompletionRequest (#1991)
Upstream PR #1947 added `presence_penalty` and `frequency_penalty` to `TextGenerationTaskParams` and the mlx-lm generator call sites, but missed wiring them up in the API adapter so they were silently dropped from incoming requests. This fixes the API mapping. Co-authored-by: Adam Durham <adam@example.com>
45248c5c · 2026-04-23 · chore(app): hardcode bug report presigned-URL endpoint (#1971)
## Motivation
The bug-report presigned-URL endpoint
(`https://reports.exolabs.net/presigned-urls`) was injected at build
time from the `EXO_BUG_REPORT_PRESIGNED_URL_ENDPOINT` GitHub Actions
secret into `Info.plist`, then read at runtime by `BugReportService`. It
isn't actually a secret — the POST body is just `{"keys":[...]}` with no
credential (see `app/EXO/EXO/Services/BugReportService.swift:136-142`),
abuse prevention lives server-side on the lambda, and the URL is already
visible in every publicly-distributed DMG's `Info.plist`. Treating it as
a repo secret added plumbing with no security benefit and broke local
dev builds — hitting **Send Bug Report** on an uncustomised `just
build-app` raised "Bug report endpoint is invalid".
## Changes
- `app/EXO/EXO/Info.plist`: replace
`$(EXO_BUG_REPORT_PRESIGNED_URL_ENDPOINT)` with the literal URL.
- `.github/workflows/build-app.yml`: drop the
`EXO_BUG_REPORT_PRESIGNED_URL_ENDPOINT` job-level env var and the
xcodebuild build-setting passthrough. No other workflow changes.
Swift code is unchanged — `BugReportService` still reads from
`Info.plist`, which leaves an escape hatch if anyone ever needs to
override via `xcodebuild EXOBugReportPresignedUrlEndpoint=...` without
recompiling.
Follow-up: the `EXO_BUG_REPORT_PRESIGNED_URL_ENDPOINT` repo secret can
now be deleted in the GitHub Actions settings UI.
## Why It Works
`Info.plist` variable substitution turns `$(FOO)` into whatever build
setting `FOO` resolves to. CI was setting `FOO` via xcodebuild; local
dev wasn't, so the key resolved to an empty string, which
`BugReportService.fetchPresignedUploadUrls` rejects via the
`!trimmedEndpointString.isEmpty` guard at `BugReportService.swift:131`.
Hardcoding the literal string removes the substitution entirely, so
every build — local or CI — gets the right value.
## Test Plan
### Manual Testing
<!-- Hardware: MacBook Pro (macOS app build via Xcode) -->
- `just build-app` with no extra env vars (reproduces the failure path
on `main`).
- `/usr/libexec/PlistBuddy -c "Print :EXOBugReportPresignedUrlEndpoint"
app/EXO/build/Build/Products/Debug/EXO.app/Contents/Info.plist` →
returns `https://reports.exolabs.net/presigned-urls` (was empty before
this change).
- `open app/EXO/build/Build/Products/Debug/EXO.app` → menubar → **Debug
Info** → **Send Bug Report** → type a description → **Send** → upload
succeeds and the **Create GitHub Issue** button appears (was failing
with "Bug report endpoint is invalid" before).
- Cross-check on the Slack side that the uploaded `report.json` lands
under `reports/YYYY/MM/DD/<ts>/` as before.
### Automated Testing
<!-- Describe changes to automated tests, or how existing tests cover
this change -->
- No new tests. This is a single-string change to `Info.plist` plus a
workflow cleanup. `nix flake check` in CI verifies formatting/lint for
the rest of the tree.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
290e3fd9 · 2026-04-23 · Keep image cache fresh (temporary fix) (#1961)
## Motivation When a new node joins, it might not have the cache. Caveat: This is potentially fallible if a new node joins and updates real topology, but the API topology hasn't caught up with this fact and the user queues up a new text generation. In practice, there is only a split second where this is the case, and this is only for users of the dashboard interface. We should fix this properly after the release.
8993ccaf · 2026-04-22 · feat(app): add friendly context message to bug report prompt (#1959)
## Motivation When a user clicks **Send Bug Report** in the macOS app, we already give them the option to add more context via an optional text field. But the current prompt is just a terse label — `"What's the issue? (optional)"` — which doesn't tell the user why bothering to fill it in matters. A friendly one-line explanation increases the chance they'll describe what went wrong, which is the single most useful signal when we triage the resulting diagnostic bundle. ## Changes - `app/EXO/EXO/ContentView.swift`: In the `.prompting` phase of `sendBugReportButton`, replace the single label with a two-line hierarchy: - Primary: `Tell us what went wrong (optional)` - Helper: `A quick description of what you were doing and what happened helps us track down the bug for you.` - The helper uses `.caption2` + `.secondary` + `.opacity(0.8)` + `.fixedSize(horizontal: false, vertical: true)` so it stays visually subordinate and wraps cleanly inside the 340pt popover. No changes to `BugReportService`, the `user_description` payload, or any other flow. ## Why It Works The optional description is already plumbed end-to-end (text editor → `bugReportUserDescription` state → `BugReportService.sendReport(..., userDescription:)` → `report.json`'s `user_description` field → GitHub issue pre-fill). The only gap was user-facing motivation, so this is purely a copy/layout tweak inside the existing `.prompting` case — no new state, bindings, or service changes. ## Test Plan ### Manual Testing <!-- Hardware: MacBook Pro (macOS app build via Xcode) --> - Build the macOS app in Xcode (`app/EXO/EXO.xcodeproj`) and launch it. - Open the menubar popover → expand **Debug Info** → click **Send Bug Report**. - Verify the new primary label and helper sentence both appear above the text editor and wrap cleanly within the popover width. - Leave the field empty → click **Send** → upload should succeed (no `user_description` in payload, same as before). - Fill in a description → click **Send** → upload succeeds and the success card with **Create GitHub Issue** appears; clicking it opens GitHub with the description pre-filled. - Click **Cancel** from the prompting state → returns to idle. ### Automated Testing <!-- Describe changes to automated tests, or how existing tests cover this change --> - No new automated tests. This is a SwiftUI copy/layout change; existing `EXOTests` are smoke-level and don't cover `ContentView` view bodies, and UI snapshot tests aren't worth adding for a two-line copy tweak. - `nix fmt` reports 0 files changed after the edit; `nix flake check` in CI will verify formatting/lint for the rest of the tree. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
4939fbe9 · 2026-04-22 · feat(dashboard): add Pi integration tab (#1925)
## Summary
Adds a new **Pi** tab to the Integrations page (`/#/integrations`)
alongside the existing Claude Code, OpenCode, Codex, OpenClaw, Open
WebUI, n8n, and Firefox tabs.
[pi](https://pi.dev) (`@mariozechner/pi-coding-agent`) is a terminal
coding agent that supports custom OpenAI-compatible providers via
`~/.pi/agent/models.json`.
This tab gives users a copy-pasteable config to wire pi up to their exo
cluster.
## What's in the tab
- **Model selector** (shown when multiple models are running) — picks
the default model for the generated shell command.
- **Models Config card** — generates `~/.pi/agent/models.json`
registering `exo` as a custom provider:
- `baseUrl` → `<apiUrl>/v1`
- `api` → `openai-completions`
- `apiKey` → `"exo"` (placeholder; exo ignores it)
- `compat.supportsDeveloperRole: false` and
`compat.supportsReasoningEffort: false`, per pi docs recommendation for
local OpenAI-compatible servers
- Auto-populates every running model with `id`, `contextWindow` (from
`/v1/models`), and `input: ["text", "image"]` for vision-capable models
- **Shell Command card** — `pi --provider exo --model <model>` for quick
launch.
The tab gracefully falls back to `your-model-id` when no models are
running, matching the behavior of the other tabs.
## Usage
1. `npm install -g @mariozechner/pi-coding-agent`
2. Paste the generated config into `~/.pi/agent/models.json`
3. Run `pi` and pick an exo model via `/model` — or run the shell
command directly
## Changes
- `dashboard/src/routes/integrations/+page.svelte` — adds `"Pi"` to the
`tabs` tuple, `piModel` state, `piModelsJson` + `piShellCommand`
derivations, and the tab content block.
Single-file, scoped change — no backend or type changes.
## Testing
- `cd dashboard && npm run build` — ✅ builds cleanly
- `svelte-check` on the edited file — no new errors
- Manually verified the tab renders, the model selector updates the
generated JSON, and the config reflects `/v1/models` capabilities
(vision → `input: ["text","image"]`,
`context_length` → `contextWindow`).
## Screenshots
<img width="1545" height="1236" alt="pi-tab"
src="https://github.com/user-attachments/assets/38aa179f-4ed9-4a1e-9783-d3baa7738263"
/>
7a312a17 · 2026-04-22 · Misc fixes: upstream JACCL all_sum, API, etc. + Add Kimi K2.6 (#1952)
## Motivation This fixes a bunch of observed model quality issues introduced upstream in JACCL, as well as API issues and prefix cache calculation. ## Test Plan ### Manual Testing Tested a bunch ### Automated Testing Added a test, automated eval tool calls on Kimi K2.6, Minimax M2.7, GPT OSS and Qwen3.6 models. --------- Co-authored-by: Evan <evanev7@gmail.com>
af673845 · 2026-04-22 · Ignore HF remote repo changes (temporary fix) (#1958)
## Motivation Fixes #1918. Downloaded model status reverts from "completed" to "pending" during each download scan. Reproduced with `zai-org/GLM-5.1`. ## Changes - `coordinator.py`: In the periodic rescan, don't downgrade already-completed models; fall back to `resolve_existing_model()` (safetensors weight check) when per-file size check reports incomplete - New `test_download_status_not_lost.py`: 3 regression tests ## Why It Works The rescan compares local file sizes against HF's `main` revision. When HF updates text files (README, jinja, etc.), remote sizes change but local files still match the old revision — causing a false "incomplete". The fix uses the safetensors weight check as ground truth instead. Long-term: pin the downloaded revision SHA rather than always checking against `main`. ## Test Plan ### Manual Testing - Mac Studio M3 Ultra with GLM-5.1 downloaded (natural reproduction of the issue) - Confirmed GLM-5.1 stays `DownloadCompleted` through multiple rescan cycles ### Automated Testing - 3 new tests: completed-not-downgraded, fallback-to-resolve, genuinely-incomplete-stays-pending
49670c86 · 2026-04-21 · Handle missing total_size in safetensors index files (#1956)
## Motivation Image models fail to load after a mid-download instance deletion and recreation. The system skips the download and crashes with `FileNotFoundError: No safetensors files found in .../vae`. ## Changes - Make `ModelSafetensorsIndexMetadata.total_size` optional (`PositiveInt | None = None`) - Add null guard in `fetch_safetensors_size` - Add regression test ## Why It Works Exolabs quantized image models have safetensors index files with mflux metadata (`quantization_level`, `mflux_version`) but no `total_size`. The required `PositiveInt` field caused Pydantic validation to fail, which was silently swallowed by `except Exception: continue` in `_scan_model_directory`. This skipped all weight map checks, making incomplete models appear complete. ## Test Plan ### Manual Testing - Hardware: Mac Studio - Before: `CreateRunner → LoadModel` (crash). After: `CreateRunner → DownloadModel` (correct). ### Automated Testing - `test_safetensors_index.py`: 3 cases covering missing, valid, and null metadata
8ccfd7fc · 2026-04-21 · Fix some misc build issues (#1948)
## Motivation <!-- Why is this change needed? What problem does it solve? --> <!-- If it fixes an open issue, please link to the issue here --> ## Changes <!-- Describe what you changed in detail --> ## Why It Works <!-- Explain why your approach solves the problem --> ## Test Plan ### Manual Testing <!-- Hardware: (e.g., MacBook Pro M1 Max 32GB, Mac Mini M2 16GB, connected via Thunderbolt 4) --> <!-- What you did: --> <!-- - --> ### Automated Testing <!-- Describe changes to automated tests, or how existing tests cover this change --> <!-- - -->
09e894dd · 2026-04-19 · Fix vision models on M5 Pro/Max MacBooks (#1927)
## Motivation Vision models don't understand images on M5 series MacBooks. The upstream NAX addmm fix (https://github.com/ml-explore/mlx/pull/3422) fixes this. ## Why It Works Same conclusion I came to when I was debugging the issue on an M5 Max. It works after this fix. ## Test Plan ### Manual Testing Works for Qwen3.5 27B
bf8aacfd · 2026-04-17 · Improve build CI (#1920)
## Motivation <!-- Why is this change needed? What problem does it solve? --> <!-- If it fixes an open issue, please link to the issue here --> ## Changes <!-- Describe what you changed in detail --> ## Why It Works <!-- Explain why your approach solves the problem --> ## Test Plan ### Manual Testing <!-- Hardware: (e.g., MacBook Pro M1 Max 32GB, Mac Mini M2 16GB, connected via Thunderbolt 4) --> <!-- What you did: --> <!-- - --> ### Automated Testing <!-- Describe changes to automated tests, or how existing tests cover this change --> <!-- - -->
af9e847e · 2026-04-17 · fix: force gc + clear_cache after KV prefix cache eviction (#1832)
## Summary - After `KVPrefixCache` evicts LRU entries, the MLX Metal buffers stay allocated until Python's GC runs - This leaks ~3-4 GB between long-context requests, reducing the effective context ceiling for back-to-back requests - Adding `gc.collect()` + `mx.clear_cache()` after eviction frees Metal buffers promptly ## Test plan - [x] Measured on 2-node PP cluster with Qwen3.5-397B-A17B-4bit at 63K context - [x] Before: 108.88 GB retained after eviction (3.78 GB above baseline) - [x] After: 105.48 GB retained after eviction (0.38 GB above baseline — draft model KV + minor overhead) - [x] `gc.collect()` adds ~2-3ms latency, runs once per eviction cycle (not per token) - [ ] Verify with `uv run pytest` 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Adam Durham <adam@example.com> Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> Co-authored-by: rltakashige <rl.takashige@gmail.com>
01598960 · 2026-04-17 · Add model card for Qwen3.6-35B-A3B-8bit (#1917)
Adds the 8bit variant missing from #1907 — the safetensors index is now live on HF. - `mlx-community/Qwen3.6-35B-A3B-8bit` (~35 GB) Architectural fields match the existing 4bit/5bit/bf16 cards. `storage_size.in_bytes` is taken from `metadata.total_size` of the upstream `model.safetensors.index.json`.
63b8e647 · 2026-04-16 · Add model cards for Qwen3.6-35B-A3B variants (#1907)
## Motivation
`mlx-community` has just published the new **Qwen3.6-35B-A3B**
multimodal MoE family on HuggingFace. Without static model cards exo
doesn't surface these models in the dashboard picker or match its
placement / prefill logic, so users can't one-click launch them. This PR
adds cards for the three quants whose safetensors indexes are already
live on HF (4bit / 5bit / bf16).
## Changes
Three new TOML files in `resources/inference_model_cards/`:
- `mlx-community--Qwen3.6-35B-A3B-4bit.toml` (~19 GB)
- `mlx-community--Qwen3.6-35B-A3B-5bit.toml` (~23 GB)
- `mlx-community--Qwen3.6-35B-A3B-bf16.toml` (~65 GB)
All three share the same architectural fields (`n_layers = 40`,
`hidden_size = 2048`, `num_key_value_heads = 2`, `context_length =
262144`, capabilities `text, thinking, thinking_toggle, vision`,
`base_model = "Qwen3.6 35B A3B"`) — only `model_id`, `quantization`, and
`storage_size.in_bytes` differ between variants.
## Why It Works
- Qwen3.6-35B-A3B reuses the `qwen3_5_moe` architecture
(`Qwen3_5MoeForConditionalGeneration`) — the same one already wired into
exo's MLX runner at `src/exo/worker/engines/mlx/auto_parallel.py:47` via
`Qwen3_5MoeModel`. The architectural fields are taken verbatim from the
HF `config.json.text_config` and match the existing `Qwen3.5-35B-A3B-*`
cards.
- Storage sizes are the exact `metadata.total_size` read from each
variant's `model.safetensors.index.json` on HF, so download progress and
cluster-memory-fit checks are accurate.
- Vision support is flagged in `capabilities`; the `[vision]` block is
auto-detected by `ModelCard._autodetect_vision` from the upstream
`config.json`, so no hand-written vision config is required.
- The card loader (`_refresh_card_cache` in
`src/exo/shared/models/model_cards.py`) globs every `.toml` in
`resources/inference_model_cards/` on startup, so nothing else needs to
change — the `/models` endpoint and the dashboard picker pick them up
automatically.
The `mxfp4` / `mxfp8` / `nvfp4` variants are still uploading upstream
(index JSONs currently 404) and can be added in a follow-up PR once HF
completes.
## Test Plan
### Manual Testing
Hardware: MacBook Pro M4 Max, 48 GB unified memory.
- Built the dashboard, ran `uv run exo`, waited for the API to come up
on `http://localhost:52415`.
- `curl -s http://localhost:52415/models` returns the three new model
ids (`mlx-community/Qwen3.6-35B-A3B-{4bit,5bit,bf16}`) alongside
existing models.
- Opened the dashboard, clicked SELECT MODEL, typed "Qwen3.6" into the
search box. A single **"Qwen3.6 35B A3B"** group appears showing `3
variants (19GB-65GB)`. Expanding it lists the `4bit` / `5bit` / `bf16`
quants with sizes `19GB` / `23GB` / `65GB`, exactly as expected:

- Programmatically loaded each TOML via `ModelCard.load_from_path(...)`
and confirmed the parsed fields (layers / hidden / KV heads / context /
quant / base_model / caps / bytes) match what's written in the files.
### Automated Testing
No code paths were touched — these are pure TOML data files that plug
into the existing model-card loader. The existing pytest suite covers
TOML parsing and card serving; adding new TOMLs doesn't require new test
scaffolding. `uv run ruff check` and `nix fmt` are clean.
---------
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: Ryuichi Leo Takashige <rl.takashige@gmail.com>
058bb082 · 2026-04-15 · Allow copying on dashboard even on HTTP (#1902)
## Motivation <!-- Why is this change needed? What problem does it solve? --> <!-- If it fixes an open issue, please link to the issue here --> ## Changes <!-- Describe what you changed in detail --> ## Why It Works <!-- Explain why your approach solves the problem --> ## Test Plan ### Manual Testing <!-- Hardware: (e.g., MacBook Pro M1 Max 32GB, Mac Mini M2 16GB, connected via Thunderbolt 4) --> <!-- What you did: --> <!-- - --> ### Automated Testing <!-- Describe changes to automated tests, or how existing tests cover this change --> <!-- - -->
3eead802 · 2026-04-15 · Better environment variables in MacOS app (#1901)
## Motivation Closes #1858 ## Changes <!-- Describe what you changed in detail --> ## Why It Works <!-- Explain why your approach solves the problem --> ## Test Plan ### Manual Testing <!-- Hardware: (e.g., MacBook Pro M1 Max 32GB, Mac Mini M2 16GB, connected via Thunderbolt 4) --> <!-- What you did: --> <!-- - --> ### Automated Testing <!-- Describe changes to automated tests, or how existing tests cover this change --> <!-- - -->
2cd66ae4 · 2026-04-15 · Fix out of order event idx causing fatal crashes (#1894)
## Motivation <img width="828" height="373" alt="Screenshot 2026-04-14 at 22 56 52" src="https://github.com/user-attachments/assets/f8f48c1d-68c5-4acc-a6de-9d180672da9d" /> if is_new_master=True, _elect_loop creates a new EventRouter before the worker has receivers. Then, event router runs _run_ext_in and buf.drain_indexed() will pick off events, even though self.internal_outbound is not populated fully. Finally, when the worker does try requesting events, the next event it receives is not the first event, meaning the worker crashes. ## Changes Start the event router after all the receivers are registered ## Why It Works self.internal_outbound is populated before the loop begins. ## Test Plan ### Manual Testing No more crashes observed in testing (it's actually quite easy to reproduce the issue if you have one node with this fix but the other node on main). I'm convinced this is a fix, at least.
File tree
- .clauderules
- .cursorrules
- .envrc
- .githooks/post-checkout
- .githooks/post-commit
- .githooks/post-merge
- .githooks/pre-push
- .github/CODEOWNERS
- .github/ISSUE_TEMPLATE/bug_report.md
- .github/ISSUE_TEMPLATE/feature_request.md
- .github/actions/conditional-commit/action.yml
- .github/actions/format/action.yml
- .github/actions/lint-check/action.yml
- .github/actions/lint/action.yml
- .github/actions/regenerate-protobufs/action.yml
- .github/actions/setup-python-uv/action.yml
- .github/actions/unit-test/action.yml
- .github/actions/verify-clean/action.yml
- .github/pull_request_template.md
- .github/workflows/build-app.yml
- .github/workflows/pipeline.yml
- .gitignore
- .idea/.gitignore
- .idea/LanguageServersSettings.xml
- .idea/externalDependencies.xml
- .idea/misc.xml
- .idea/modules.xml
- .idea/pyright-overrides.xml
- .idea/pyright.xml
- .idea/vcs.xml
- .python-version
- .swift-format
- .typings/.gitkeep
- .typings/mflux/__init__.pyi
- .typings/mflux/callbacks/__init__.pyi
- .typings/mflux/callbacks/callback.pyi
- .typings/mflux/callbacks/callback_registry.pyi
- .typings/mflux/callbacks/generation_context.pyi
- .typings/mflux/cli/__init__.pyi
- .typings/mflux/cli/defaults/defaults.pyi
- .typings/mflux/models/__init__.pyi
- .typings/mflux/models/common/__init__.pyi
- .typings/mflux/models/common/cli/__init__.pyi
- .typings/mflux/models/common/config/__init__.pyi
- .typings/mflux/models/common/config/config.pyi
- .typings/mflux/models/common/config/model_config.pyi
- .typings/mflux/models/common/latent_creator/__init__.pyi
- .typings/mflux/models/common/latent_creator/latent_creator.pyi
- .typings/mflux/models/common/lora/__init__.pyi
- .typings/mflux/models/common/lora/layer/fused_linear_lora_layer.pyi
- .typings/mflux/models/common/lora/layer/linear_lora_layer.pyi
- .typings/mflux/models/common/lora/mapping/lora_loader.pyi
- .typings/mflux/models/common/lora/mapping/lora_mapping.pyi
- .typings/mflux/models/common/lora/mapping/lora_saver.pyi
- .typings/mflux/models/common/lora/mapping/lora_transforms.pyi
- .typings/mflux/models/common/resolution/__init__.pyi
- .typings/mflux/models/common/resolution/actions.pyi
- .typings/mflux/models/common/resolution/config_resolution.pyi
- .typings/mflux/models/common/resolution/lora_resolution.pyi
- .typings/mflux/models/common/resolution/path_resolution.pyi
- .typings/mflux/models/common/resolution/quantization_resolution.pyi
- .typings/mflux/models/common/schedulers/__init__.pyi
- .typings/mflux/models/common/schedulers/base_scheduler.pyi
- .typings/mflux/models/common/schedulers/flow_match_euler_discrete_scheduler.pyi
- .typings/mflux/models/common/schedulers/linear_scheduler.pyi
- .typings/mflux/models/common/schedulers/seedvr2_euler_scheduler.pyi
- .typings/mflux/models/common/tokenizer/__init__.pyi
- .typings/mflux/models/common/tokenizer/tokenizer.pyi
- .typings/mflux/models/common/tokenizer/tokenizer_loader.pyi
- .typings/mflux/models/common/tokenizer/tokenizer_output.pyi
- .typings/mflux/models/common/vae/__init__.pyi
- .typings/mflux/models/common/vae/tiling_config.pyi
- .typings/mflux/models/common/vae/vae_tiler.pyi
- .typings/mflux/models/common/vae/vae_util.pyi
- .typings/mflux/models/common/weights/__init__.pyi
- .typings/mflux/models/common/weights/loading/loaded_weights.pyi
- .typings/mflux/models/common/weights/loading/weight_applier.pyi
- .typings/mflux/models/common/weights/loading/weight_definition.pyi
- .typings/mflux/models/common/weights/loading/weight_loader.pyi
- .typings/mflux/models/common/weights/mapping/weight_mapper.pyi
- .typings/mflux/models/common/weights/mapping/weight_mapping.pyi
- .typings/mflux/models/common/weights/mapping/weight_transforms.pyi
- .typings/mflux/models/common/weights/saving/model_saver.pyi
- .typings/mflux/models/depth_pro/depth_pro_initializer.pyi
- .typings/mflux/models/depth_pro/model/decoder/feature_fusion_block_2d.pyi
- .typings/mflux/models/depth_pro/model/decoder/multires_conv_decoder.pyi
- .typings/mflux/models/depth_pro/model/decoder/residual_block.pyi
- .typings/mflux/models/depth_pro/model/depth_pro.pyi
- .typings/mflux/models/depth_pro/model/depth_pro_model.pyi
- .typings/mflux/models/depth_pro/model/depth_pro_util.pyi
- .typings/mflux/models/depth_pro/model/dino_v2/attention.pyi
- .typings/mflux/models/depth_pro/model/dino_v2/dino_vision_transformer.pyi
- .typings/mflux/models/depth_pro/model/dino_v2/layer_scale.pyi
- .typings/mflux/models/depth_pro/model/dino_v2/mlp.pyi
- .typings/mflux/models/depth_pro/model/dino_v2/patch_embed.pyi
- .typings/mflux/models/depth_pro/model/dino_v2/transformer_block.pyi
- .typings/mflux/models/depth_pro/model/encoder/depth_pro_encoder.pyi
- .typings/mflux/models/depth_pro/model/encoder/upsample_block.pyi
- .typings/mflux/models/depth_pro/model/head/fov_head.pyi
- .typings/mflux/models/depth_pro/weights/depth_pro_weight_definition.pyi
- .typings/mflux/models/depth_pro/weights/depth_pro_weight_mapping.pyi
- .typings/mflux/models/fibo/latent_creator/fibo_latent_creator.pyi
- .typings/mflux/models/fibo/weights/fibo_weight_definition.pyi
- .typings/mflux/models/fibo/weights/fibo_weight_mapping.pyi
- .typings/mflux/models/fibo_vlm/tokenizer/qwen2vl_image_processor.pyi
- .typings/mflux/models/fibo_vlm/tokenizer/qwen2vl_processor.pyi
- .typings/mflux/models/fibo_vlm/weights/fibo_vlm_weight_definition.pyi
- .typings/mflux/models/fibo_vlm/weights/fibo_vlm_weight_mapping.pyi
- .typings/mflux/models/flux/__init__.pyi
- .typings/mflux/models/flux/cli/__init__.pyi
- .typings/mflux/models/flux/flux_initializer.pyi
- .typings/mflux/models/flux/latent_creator/__init__.pyi
- .typings/mflux/models/flux/latent_creator/flux_latent_creator.pyi
- .typings/mflux/models/flux/model/__init__.pyi
- .typings/mflux/models/flux/model/flux_text_encoder/clip_encoder/clip_embeddings.pyi
- .typings/mflux/models/flux/model/flux_text_encoder/clip_encoder/clip_encoder.pyi
- .typings/mflux/models/flux/model/flux_text_encoder/clip_encoder/clip_encoder_layer.pyi
- .typings/mflux/models/flux/model/flux_text_encoder/clip_encoder/clip_mlp.pyi
- .typings/mflux/models/flux/model/flux_text_encoder/clip_encoder/clip_sdpa_attention.pyi
- .typings/mflux/models/flux/model/flux_text_encoder/clip_encoder/clip_text_model.pyi
- .typings/mflux/models/flux/model/flux_text_encoder/clip_encoder/encoder_clip.pyi
- .typings/mflux/models/flux/model/flux_text_encoder/prompt_encoder.pyi
- .typings/mflux/models/flux/model/flux_text_encoder/t5_encoder/t5_attention.pyi
- .typings/mflux/models/flux/model/flux_text_encoder/t5_encoder/t5_block.pyi
- .typings/mflux/models/flux/model/flux_text_encoder/t5_encoder/t5_dense_relu_dense.pyi
- .typings/mflux/models/flux/model/flux_text_encoder/t5_encoder/t5_encoder.pyi
- .typings/mflux/models/flux/model/flux_text_encoder/t5_encoder/t5_feed_forward.pyi
- .typings/mflux/models/flux/model/flux_text_encoder/t5_encoder/t5_layer_norm.pyi
- .typings/mflux/models/flux/model/flux_text_encoder/t5_encoder/t5_self_attention.pyi
- .typings/mflux/models/flux/model/flux_transformer/ada_layer_norm_continuous.pyi
- .typings/mflux/models/flux/model/flux_transformer/ada_layer_norm_zero.pyi
- .typings/mflux/models/flux/model/flux_transformer/ada_layer_norm_zero_single.pyi
- .typings/mflux/models/flux/model/flux_transformer/common/attention_utils.pyi
- .typings/mflux/models/flux/model/flux_transformer/embed_nd.pyi
- .typings/mflux/models/flux/model/flux_transformer/feed_forward.pyi
- .typings/mflux/models/flux/model/flux_transformer/guidance_embedder.pyi
- .typings/mflux/models/flux/model/flux_transformer/joint_attention.pyi
- .typings/mflux/models/flux/model/flux_transformer/joint_transformer_block.pyi
- .typings/mflux/models/flux/model/flux_transformer/single_block_attention.pyi
- .typings/mflux/models/flux/model/flux_transformer/single_transformer_block.pyi
- .typings/mflux/models/flux/model/flux_transformer/text_embedder.pyi
- .typings/mflux/models/flux/model/flux_transformer/time_text_embed.pyi
- .typings/mflux/models/flux/model/flux_transformer/timestep_embedder.pyi
- .typings/mflux/models/flux/model/flux_transformer/transformer.pyi
- .typings/mflux/models/flux/model/flux_vae/common/attention.pyi
- .typings/mflux/models/flux/model/flux_vae/common/resnet_block_2d.pyi
- .typings/mflux/models/flux/model/flux_vae/common/unet_mid_block.pyi
- .typings/mflux/models/flux/model/flux_vae/decoder/conv_in.pyi
- .typings/mflux/models/flux/model/flux_vae/decoder/conv_norm_out.pyi
- .typings/mflux/models/flux/model/flux_vae/decoder/conv_out.pyi
- .typings/mflux/models/flux/model/flux_vae/decoder/decoder.pyi
- .typings/mflux/models/flux/model/flux_vae/decoder/up_block_1_or_2.pyi
- .typings/mflux/models/flux/model/flux_vae/decoder/up_block_3.pyi
- .typings/mflux/models/flux/model/flux_vae/decoder/up_block_4.pyi
- .typings/mflux/models/flux/model/flux_vae/decoder/up_sampler.pyi
- .typings/mflux/models/flux/model/flux_vae/encoder/conv_in.pyi
- .typings/mflux/models/flux/model/flux_vae/encoder/conv_norm_out.pyi
- .typings/mflux/models/flux/model/flux_vae/encoder/conv_out.pyi
- .typings/mflux/models/flux/model/flux_vae/encoder/down_block_1.pyi
- .typings/mflux/models/flux/model/flux_vae/encoder/down_block_2.pyi
- .typings/mflux/models/flux/model/flux_vae/encoder/down_block_3.pyi
- .typings/mflux/models/flux/model/flux_vae/encoder/down_block_4.pyi
- .typings/mflux/models/flux/model/flux_vae/encoder/down_sampler.pyi
- .typings/mflux/models/flux/model/flux_vae/encoder/encoder.pyi
- .typings/mflux/models/flux/model/flux_vae/vae.pyi
- .typings/mflux/models/flux/model/redux_encoder/redux_encoder.pyi
- .typings/mflux/models/flux/model/siglip_vision_transformer/siglip_encoder.pyi
- .typings/mflux/models/flux/model/siglip_vision_transformer/siglip_encoder_layer.pyi
- .typings/mflux/models/flux/model/siglip_vision_transformer/siglip_mlp.pyi
- .typings/mflux/models/flux/model/siglip_vision_transformer/siglip_multi_head_attention_pooling_head.pyi
- .typings/mflux/models/flux/model/siglip_vision_transformer/siglip_sdpa_attention.pyi
- .typings/mflux/models/flux/model/siglip_vision_transformer/siglip_vision_embeddings.pyi
- .typings/mflux/models/flux/model/siglip_vision_transformer/siglip_vision_transformer.pyi
- .typings/mflux/models/flux/variants/__init__.pyi
- .typings/mflux/models/flux/variants/concept_attention/attention_data.pyi
- .typings/mflux/models/flux/variants/concept_attention/joint_attention_concept.pyi
- .typings/mflux/models/flux/variants/concept_attention/joint_transformer_block_concept.pyi
- .typings/mflux/models/flux/variants/concept_attention/transformer_concept.pyi
- .typings/mflux/models/flux/variants/controlnet/transformer_controlnet.pyi
- .typings/mflux/models/flux/variants/kontext/__init__.pyi
- .typings/mflux/models/flux/variants/kontext/flux_kontext.pyi
- .typings/mflux/models/flux/variants/kontext/kontext_util.pyi
- .typings/mflux/models/flux/variants/txt2img/flux.pyi
- .typings/mflux/models/flux/weights/__init__.pyi
- .typings/mflux/models/flux/weights/flux_lora_mapping.pyi
- .typings/mflux/models/flux/weights/flux_weight_definition.pyi
- .typings/mflux/models/flux/weights/flux_weight_mapping.pyi
- .typings/mflux/models/qwen/__init__.pyi
- .typings/mflux/models/qwen/cli/__init__.pyi
- .typings/mflux/models/qwen/latent_creator/__init__.pyi
- .typings/mflux/models/qwen/latent_creator/qwen_latent_creator.pyi
- .typings/mflux/models/qwen/model/__init__.pyi
- .typings/mflux/models/qwen/model/qwen_text_encoder/qwen_attention.pyi
- .typings/mflux/models/qwen/model/qwen_text_encoder/qwen_encoder.pyi
- .typings/mflux/models/qwen/model/qwen_text_encoder/qwen_encoder_layer.pyi
- .typings/mflux/models/qwen/model/qwen_text_encoder/qwen_mlp.pyi
- .typings/mflux/models/qwen/model/qwen_text_encoder/qwen_patch_merger.pyi
- .typings/mflux/models/qwen/model/qwen_text_encoder/qwen_prompt_encoder.pyi
- .typings/mflux/models/qwen/model/qwen_text_encoder/qwen_rms_norm.pyi
- .typings/mflux/models/qwen/model/qwen_text_encoder/qwen_rope.pyi
- .typings/mflux/models/qwen/model/qwen_text_encoder/qwen_text_encoder.pyi
- .typings/mflux/models/qwen/model/qwen_text_encoder/qwen_vision_attention.pyi
- .typings/mflux/models/qwen/model/qwen_text_encoder/qwen_vision_block.pyi
- .typings/mflux/models/qwen/model/qwen_text_encoder/qwen_vision_language_encoder.pyi
- .typings/mflux/models/qwen/model/qwen_text_encoder/qwen_vision_mlp.pyi
- .typings/mflux/models/qwen/model/qwen_text_encoder/qwen_vision_patch_embed.pyi
- .typings/mflux/models/qwen/model/qwen_text_encoder/qwen_vision_rotary_embedding.pyi
- .typings/mflux/models/qwen/model/qwen_text_encoder/qwen_vision_transformer.pyi
- .typings/mflux/models/qwen/model/qwen_transformer/qwen_attention.pyi
- .typings/mflux/models/qwen/model/qwen_transformer/qwen_feed_forward.pyi
- .typings/mflux/models/qwen/model/qwen_transformer/qwen_rope.pyi
- .typings/mflux/models/qwen/model/qwen_transformer/qwen_time_text_embed.pyi
- .typings/mflux/models/qwen/model/qwen_transformer/qwen_timestep_embedding.pyi
- .typings/mflux/models/qwen/model/qwen_transformer/qwen_timesteps.pyi
- .typings/mflux/models/qwen/model/qwen_transformer/qwen_transformer.pyi
- .typings/mflux/models/qwen/model/qwen_transformer/qwen_transformer_block.pyi
- .typings/mflux/models/qwen/model/qwen_transformer/qwen_transformer_rms_norm.pyi
- .typings/mflux/models/qwen/model/qwen_vae/qwen_image_attention_block_3d.pyi
- .typings/mflux/models/qwen/model/qwen_vae/qwen_image_causal_conv_3d.pyi
- .typings/mflux/models/qwen/model/qwen_vae/qwen_image_decoder_3d.pyi
- .typings/mflux/models/qwen/model/qwen_vae/qwen_image_down_block_3d.pyi
- .typings/mflux/models/qwen/model/qwen_vae/qwen_image_encoder_3d.pyi
- .typings/mflux/models/qwen/model/qwen_vae/qwen_image_mid_block_3d.pyi
- .typings/mflux/models/qwen/model/qwen_vae/qwen_image_res_block_3d.pyi
- .typings/mflux/models/qwen/model/qwen_vae/qwen_image_resample_3d.pyi
- .typings/mflux/models/qwen/model/qwen_vae/qwen_image_rms_norm.pyi
- .typings/mflux/models/qwen/model/qwen_vae/qwen_image_up_block_3d.pyi
- .typings/mflux/models/qwen/model/qwen_vae/qwen_vae.pyi
- .typings/mflux/models/qwen/qwen_initializer.pyi
- .typings/mflux/models/qwen/tokenizer/__init__.pyi
- .typings/mflux/models/qwen/tokenizer/qwen_image_processor.pyi
- .typings/mflux/models/qwen/tokenizer/qwen_vision_language_processor.pyi
- .typings/mflux/models/qwen/tokenizer/qwen_vision_language_tokenizer.pyi
- .typings/mflux/models/qwen/variants/__init__.pyi
- .typings/mflux/models/qwen/variants/edit/qwen_edit_util.pyi
- .typings/mflux/models/qwen/variants/edit/qwen_image_edit.pyi
- .typings/mflux/models/qwen/variants/txt2img/qwen_image.pyi
- .typings/mflux/models/qwen/weights/__init__.pyi
- .typings/mflux/models/qwen/weights/qwen_lora_mapping.pyi
- .typings/mflux/models/qwen/weights/qwen_weight_definition.pyi
- .typings/mflux/models/qwen/weights/qwen_weight_mapping.pyi
- .typings/mflux/models/seedvr2/weights/seedvr2_weight_definition.pyi
- .typings/mflux/models/seedvr2/weights/seedvr2_weight_mapping.pyi
- .typings/mflux/models/z_image/latent_creator/z_image_latent_creator.pyi
- .typings/mflux/models/z_image/weights/z_image_weight_definition.pyi
- .typings/mflux/models/z_image/weights/z_image_weight_mapping.pyi
- .typings/mflux/release/__init__.pyi
- .typings/mflux/utils/__init__.pyi
- .typings/mflux/utils/box_values.pyi
- .typings/mflux/utils/exceptions.pyi
- .typings/mflux/utils/generated_image.pyi
- .typings/mflux/utils/image_util.pyi
- .typings/mflux/utils/metadata_builder.pyi
- .typings/mflux/utils/version_util.pyi
- .typings/mlx/core/__init__.pyi
- .typings/mlx/core/cuda/__init__.pyi
- .typings/mlx/core/distributed/__init__.pyi
- .typings/mlx/core/metal/__init__.pyi
- .typings/mlx/core/random/__init__.pyi
- .typings/mlx/nn/__init__.pyi
- .typings/mlx/nn/init.pyi
- .typings/mlx/nn/layers/__init__.pyi
- .typings/mlx/nn/layers/activations.pyi
- .typings/mlx/nn/layers/base.pyi
- .typings/mlx/nn/layers/containers.pyi
- .typings/mlx/nn/layers/convolution.pyi
- .typings/mlx/nn/layers/convolution_transpose.pyi
- .typings/mlx/nn/layers/distributed.pyi
- .typings/mlx/nn/layers/dropout.pyi
- .typings/mlx/nn/layers/embedding.pyi
- .typings/mlx/nn/layers/linear.pyi
- .typings/mlx/nn/layers/normalization.pyi
- .typings/mlx/nn/layers/pooling.pyi
- .typings/mlx/nn/layers/positional_encoding.pyi
- .typings/mlx/nn/layers/quantized.pyi
- .typings/mlx/nn/layers/recurrent.pyi
- .typings/mlx/nn/layers/transformer.pyi
- .typings/mlx/nn/layers/upsample.pyi
- .typings/mlx/nn/losses.pyi
- .typings/mlx/nn/utils.pyi
- .typings/mlx/utils.pyi
- .typings/mlx_lm/__init__.pyi
- .typings/mlx_lm/_version.pyi
- .typings/mlx_lm/convert.pyi
- .typings/mlx_lm/generate.pyi
- .typings/mlx_lm/models/__init__.pyi
- .typings/mlx_lm/models/activations.pyi
- .typings/mlx_lm/models/base.pyi
- .typings/mlx_lm/models/bitlinear_layers.pyi
- .typings/mlx_lm/models/cache.pyi
- .typings/mlx_lm/models/deepseek_v3.pyi
- .typings/mlx_lm/models/deepseek_v4.pyi
- .typings/mlx_lm/models/gated_delta.pyi
- .typings/mlx_lm/models/gemma4.pyi
- .typings/mlx_lm/models/gemma4_text.pyi
- .typings/mlx_lm/models/glm4_moe.pyi
- .typings/mlx_lm/models/glm_moe_dsa.pyi
- .typings/mlx_lm/models/gpt_oss.pyi
- .typings/mlx_lm/models/minimax.pyi
- .typings/mlx_lm/models/nemotron_h.pyi
- .typings/mlx_lm/models/qwen3_5.pyi
- .typings/mlx_lm/models/qwen3_5_moe.pyi
- .typings/mlx_lm/models/qwen3_next.pyi
- .typings/mlx_lm/models/rope_utils.pyi
- .typings/mlx_lm/models/step3p5.pyi
- .typings/mlx_lm/models/switch_layers.pyi
- .typings/mlx_lm/sample_utils.pyi
- .typings/mlx_lm/tokenizer_utils.pyi
- .typings/mlx_lm/tuner/dora.pyi
- .typings/mlx_lm/tuner/lora.pyi
- .typings/mlx_lm/tuner/utils.pyi
- .typings/mlx_lm/utils.pyi
- .typings/mlx_vlm/__init__.pyi
- .typings/mlx_vlm/prompt_utils.pyi
- .typings/mlx_vlm/utils.pyi
- .typings/pynvml/__init__.pyi
- .typings/safetensors/__init__.pyi
- .vscode/extensions.json
- .vscode/settings.json
- .zed/settings.json
- AGENTS.md
- CLAUDE.md
- CONTRIBUTING.md
- Cargo.lock
- Cargo.toml
- LICENSE
- PLATFORMS.md
- README.md
- RULES.md
- TODO.md
- app/EXO/EXO.xcodeproj/project.pbxproj
- app/EXO/EXO.xcodeproj/project.xcworkspace/contents.xcworkspacedata
- app/EXO/EXO.xcodeproj/project.xcworkspace/xcshareddata/swiftpm/Package.resolved
- app/EXO/EXO.xcodeproj/xcshareddata/xcschemes/EXO.xcscheme
- app/EXO/EXO.xcodeproj/xcuserdata/samikhan.xcuserdatad/xcschemes/xcschememanagement.plist
- app/EXO/EXO/Assets.xcassets/AccentColor.colorset/Contents.json
- app/EXO/EXO/Assets.xcassets/AppIcon.appiconset/1024-mac.png
- app/EXO/EXO/Assets.xcassets/AppIcon.appiconset/128-mac.png
- app/EXO/EXO/Assets.xcassets/AppIcon.appiconset/16-mac.png
- app/EXO/EXO/Assets.xcassets/AppIcon.appiconset/256-mac 1.png
- app/EXO/EXO/Assets.xcassets/AppIcon.appiconset/256-mac.png
- app/EXO/EXO/Assets.xcassets/AppIcon.appiconset/32-mac 1.png
- app/EXO/EXO/Assets.xcassets/AppIcon.appiconset/32-mac.png
- app/EXO/EXO/Assets.xcassets/AppIcon.appiconset/512-mac 1.png
- app/EXO/EXO/Assets.xcassets/AppIcon.appiconset/512-mac.png
- app/EXO/EXO/Assets.xcassets/AppIcon.appiconset/64-mac.png
- app/EXO/EXO/Assets.xcassets/AppIcon.appiconset/Contents.json
- app/EXO/EXO/Assets.xcassets/Contents.json
- app/EXO/EXO/Assets.xcassets/menubar-icon.imageset/Contents.json
- app/EXO/EXO/Assets.xcassets/menubar-icon.imageset/exo-logo-hq-square-transparent-bg.png
- app/EXO/EXO/ContentView.swift
- app/EXO/EXO/EXO.entitlements
- app/EXO/EXO/EXOApp.swift
- app/EXO/EXO/ExoProcessController.swift
- app/EXO/EXO/Info.plist
- app/EXO/EXO/Models/ClusterState.swift
- app/EXO/EXO/Preview Content/Preview Assets.xcassets/Contents.json
- app/EXO/EXO/Services/BugReportService.swift
- app/EXO/EXO/Services/ClusterStateService.swift
- app/EXO/EXO/Services/LocalNetworkChecker.swift
- app/EXO/EXO/Services/NetworkSetupHelper.swift
- app/EXO/EXO/Services/NetworkStatusService.swift
- app/EXO/EXO/Services/ThunderboltBridgeDetector.swift
- app/EXO/EXO/Services/ThunderboltBridgeService.swift
- app/EXO/EXO/ViewModels/InstanceViewModel.swift
- app/EXO/EXO/ViewModels/NodeViewModel.swift
- app/EXO/EXO/Views/BugReportWindowController.swift
- app/EXO/EXO/Views/FirstLaunchPopout.swift
- app/EXO/EXO/Views/InstanceRowView.swift
- app/EXO/EXO/Views/NodeDetailView.swift
- app/EXO/EXO/Views/NodeRowView.swift
- app/EXO/EXO/Views/SettingsView.swift
- app/EXO/EXO/Views/SettingsWindowController.swift
- app/EXO/EXO/Views/TopologyMiniView.swift
- app/EXO/EXO/main.swift
- app/EXO/EXOTests/EXOTests.swift
- app/EXO/EXOUITests/EXOUITests.swift
- app/EXO/EXOUITests/EXOUITestsLaunchTests.swift
- app/EXO/uninstall-exo.sh
- bench/METHODOLOGY.md
- bench/bench.toml
- bench/eval_configs/models.toml
- bench/eval_tool_calls.py
- bench/exo_bench.py
- bench/exo_eval.py
- bench/parallel_requests.py
- bench/prefill-decode.toml
- bench/prefill_decode_bench.py
- bench/pyproject.toml
- bench/scenarios.toml
- bench/single-m3-ultra.toml
- bench/vendor/__init__.py
- bench/vendor/lcb_testing_util.py
- dashboard/dashboard.nix
- dashboard/exo-logo-hq-square-black-bg.jpg
- dashboard/exo-logo-hq-square-black-bg.png
- dashboard/exo-logo-hq-square-black-bg.webp
- dashboard/exo-logo.png
- dashboard/favicon.ico
- dashboard/package-lock.json
- dashboard/package.json
- dashboard/parts.nix
- dashboard/src/app.css
- dashboard/src/app.d.ts
- dashboard/src/app.html
- dashboard/src/lib/components/ChatAttachments.svelte
- dashboard/src/lib/components/ChatForm.svelte
- dashboard/src/lib/components/ChatMessages.svelte
- dashboard/src/lib/components/ChatModelSelector.svelte
- dashboard/src/lib/components/ChatSidebar.svelte
- dashboard/src/lib/components/ConnectionBanner.svelte
- dashboard/src/lib/components/DeviceIcon.svelte
- dashboard/src/lib/components/FamilyLogos.svelte
- dashboard/src/lib/components/FamilySidebar.svelte
- dashboard/src/lib/components/HeaderNav.svelte
- dashboard/src/lib/components/HuggingFaceResultItem.svelte
- dashboard/src/lib/components/ImageLightbox.svelte
- dashboard/src/lib/components/ImageParamsPanel.svelte
- dashboard/src/lib/components/IntegrationCard.svelte
- dashboard/src/lib/components/MarkdownContent.svelte
- dashboard/src/lib/components/ModelCard.svelte
- dashboard/src/lib/components/ModelFilterPopover.svelte
- dashboard/src/lib/components/ModelPickerGroup.svelte
- dashboard/src/lib/components/ModelPickerModal.svelte
- dashboard/src/lib/components/PrefillDecodeDisaggregation.svelte
- dashboard/src/lib/components/PrefillProgressBar.svelte
- dashboard/src/lib/components/ToastContainer.svelte
- dashboard/src/lib/components/TokenHeatmap.svelte
- dashboard/src/lib/components/TopologyGraph.svelte
- dashboard/src/lib/components/index.ts
- dashboard/src/lib/stores/app.svelte.ts
- dashboard/src/lib/stores/favorites.svelte.ts
- dashboard/src/lib/stores/recents.svelte.ts
- dashboard/src/lib/stores/toast.svelte.ts
- dashboard/src/lib/types/files.ts
- dashboard/src/lib/utils/clipboard.ts
- dashboard/src/lib/utils/downloads.ts
- dashboard/src/lib/utils/model_family.ts
- dashboard/src/routes/+layout.svelte
- dashboard/src/routes/+page.svelte
- dashboard/src/routes/advanced/+page.svelte
- dashboard/src/routes/downloads/+page.svelte
- dashboard/src/routes/integrations/+page.svelte
- dashboard/src/routes/traces/+page.svelte
- dashboard/src/routes/traces/[taskId]/+page.svelte
- dashboard/static/exo-logo.png
- dashboard/static/favicon.ico
- dashboard/svelte.config.js
- dashboard/tsconfig.json
- dashboard/vite.config.ts
- docs/api.md
- docs/benchmarks/jeffgeerling/mac-studio-cluster-ai-full-1-qwen3-235b.jpeg
- docs/benchmarks/jeffgeerling/mac-studio-cluster-ai-full-2-deepseek-3.1-671b.jpeg
- docs/benchmarks/jeffgeerling/mac-studio-cluster-ai-full-3-kimi-k2-thinking.jpeg
- docs/imgs/dashboard-cluster-view.png
- docs/imgs/exo-logo-black-bg.jpg
- docs/imgs/exo-logo-transparent-black-text.png
- docs/imgs/exo-logo-transparent.png
- docs/imgs/exo-rounded.png
- docs/imgs/exo-screenshot.jpg
- docs/imgs/four-mac-studio-topology.png
- docs/imgs/macos-app-one-macbook.png
- flake.lock
- flake.nix
- justfile
- nix/apple-sdk-overlay.nix
- nix/apple-sdk/metadata/versions.json
- nix/darwin-build-fixes.patch
- nix/metal-toolchain.nix
- package-lock.json
- packaging/dmg/background.png
- packaging/dmg/create-dmg.sh
- packaging/dmg/generate-background.py
- packaging/pyinstaller/exo.spec
- pyproject.toml
- python/parts.nix
- resources/image_model_cards/exolabs--FLUX.1-Kontext-dev-4bit.toml
- resources/image_model_cards/exolabs--FLUX.1-Kontext-dev-8bit.toml
- resources/image_model_cards/exolabs--FLUX.1-Kontext-dev.toml
- resources/image_model_cards/exolabs--FLUX.1-Krea-dev-4bit.toml
- resources/image_model_cards/exolabs--FLUX.1-Krea-dev-8bit.toml
- resources/image_model_cards/exolabs--FLUX.1-Krea-dev.toml
- resources/image_model_cards/exolabs--FLUX.1-dev-4bit.toml
- resources/image_model_cards/exolabs--FLUX.1-dev-8bit.toml
- resources/image_model_cards/exolabs--FLUX.1-dev.toml
- resources/image_model_cards/exolabs--FLUX.1-schnell-4bit.toml
- resources/image_model_cards/exolabs--FLUX.1-schnell-8bit.toml
- resources/image_model_cards/exolabs--FLUX.1-schnell.toml
- resources/image_model_cards/exolabs--Qwen-Image-4bit.toml
- resources/image_model_cards/exolabs--Qwen-Image-8bit.toml
- resources/image_model_cards/exolabs--Qwen-Image-Edit-2509-4bit.toml
- resources/image_model_cards/exolabs--Qwen-Image-Edit-2509-8bit.toml
- resources/image_model_cards/exolabs--Qwen-Image-Edit-2509.toml
- resources/image_model_cards/exolabs--Qwen-Image.toml
- resources/inference_model_cards/mlx-community--DeepSeek-V3.1-4bit.toml
- resources/inference_model_cards/mlx-community--DeepSeek-V3.1-8bit.toml
- resources/inference_model_cards/mlx-community--DeepSeek-V3.2-4bit.toml
- resources/inference_model_cards/mlx-community--DeepSeek-V3.2-8bit.toml
- resources/inference_model_cards/mlx-community--DeepSeek-V4-Flash.toml
- resources/inference_model_cards/mlx-community--DeepSeek-V4-Pro.toml