← back to Exo
feat: add model picker modal with grouped models and HF Hub search (#1369)
41ed7afb3bd6263c70cde9d29e5f440ce98782e8 · 2026-02-04 05:56:23 -0800 · Alex Cheema
## Motivation
Reimplements the model picker modal from #1191 on top of the custom
model support branch. Replaces the inline model dropdown with a
full-featured modal that groups models by base model, supports
filtering, favorites, and HuggingFace Hub search.
## Changes
**Backend:**
- Add `family`, `quantization`, `base_model`, `capabilities` metadata
fields to `ModelCard` and all 40 TOML model cards
- Pass new fields through `ModelListModel` and `get_models()` API
response
- Add `GET /models/search` endpoint using
`huggingface_hub.list_models()`
**Dashboard (7 new files):**
- `ModelPickerModal.svelte` — Main modal with search, family filtering,
HuggingFace Hub tab
- `ModelPickerGroup.svelte` — Expandable model group row with
quantization variants
- `FamilySidebar.svelte` — Vertical sidebar with family icons (All,
Favorites, Hub, model families)
- `FamilyLogos.svelte` — SVG icons for each model family
- `ModelFilterPopover.svelte` — Capability and size range filters
- `HuggingFaceResultItem.svelte` — HF search result item with
download/like counts
- `favorites.svelte.ts` — localStorage-backed favorites store
**Integration:**
- Replace inline dropdown in `+page.svelte` with button that opens
`ModelPickerModal`
- Custom models shown in Hub tab with delete support
**Polish:**
- Real brand logos (Meta, Qwen, DeepSeek, OpenAI, GLM, MiniMax, Kimi,
HuggingFace) from Simple Icons / LobeHub
- Clean SVG stroke icons for capabilities (thinking, code, vision, image
gen)
- Consistent `border-exo-yellow/10` borders, descriptive tooltips
throughout
- Cluster memory (used/total) shown in modal header
- Selected model highlight with checkmark for both single and
multi-variant groups
- Cursor pointer on all interactive elements, fix filter popover
click-outside bug
- Custom models now appear in All tab alongside built-in models
## Bug Fix: Gemma 3 EOS tokens
Also included in this branch: fix for Gemma 3 models generating infinite
`<end_of_turn>` tokens. The tokenizer's `eos_token_ids` was missing
token ID 106 (`<end_of_turn>`), so generation never stopped. The fix
appends this token to the EOS list after loading the tokenizer. Also
handles `eos_token_ids` being a `set` (not just a `list`).
## Why It Works
Model metadata (family, capabilities, etc.) is stored directly in TOML
cards rather than derived from heuristics, ensuring accuracy. The modal
groups models by `base_model` field so quantization variants appear
together. Custom models are separated into the Hub tab since they lack
grouping metadata.
## Test Plan
### Manual Testing
- Open dashboard, click model selector to open modal
- Browse models by family sidebar, search, and filters
- Expand model groups to see quantization variants
- Star favorites and verify persistence across page reloads
- Navigate to Hub tab, search and add models
- Verify error messages shown for invalid model IDs
- Run a Gemma 3 model and verify generation stops at `<end_of_turn>`
### Automated Testing
- `uv run basedpyright` — 0 errors
- `uv run ruff check` — passes
- `nix fmt` — clean
- `uv run pytest src/` — 173 passed
- `cd dashboard && npm run build` — builds successfully
---------
Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com>
Files touched
M .gitignoreM .mlx_typings/mlx_lm/tokenizer_utils.pyiA dashboard/src/lib/components/FamilyLogos.svelteA dashboard/src/lib/components/FamilySidebar.svelteA dashboard/src/lib/components/HuggingFaceResultItem.svelteA dashboard/src/lib/components/ModelFilterPopover.svelteA dashboard/src/lib/components/ModelPickerGroup.svelteA dashboard/src/lib/components/ModelPickerModal.svelteM dashboard/src/lib/components/index.tsA dashboard/src/lib/stores/favorites.svelte.tsM dashboard/src/routes/+page.svelteM resources/inference_model_cards/mlx-community--DeepSeek-V3.1-4bit.tomlM resources/inference_model_cards/mlx-community--DeepSeek-V3.1-8bit.tomlM resources/inference_model_cards/mlx-community--GLM-4.5-Air-8bit.tomlM resources/inference_model_cards/mlx-community--GLM-4.5-Air-bf16.tomlM resources/inference_model_cards/mlx-community--GLM-4.7-4bit.tomlM resources/inference_model_cards/mlx-community--GLM-4.7-6bit.tomlM resources/inference_model_cards/mlx-community--GLM-4.7-8bit-gs32.tomlM resources/inference_model_cards/mlx-community--GLM-4.7-Flash-4bit.tomlM resources/inference_model_cards/mlx-community--GLM-4.7-Flash-5bit.tomlM resources/inference_model_cards/mlx-community--GLM-4.7-Flash-6bit.tomlM resources/inference_model_cards/mlx-community--GLM-4.7-Flash-8bit.tomlM resources/inference_model_cards/mlx-community--Kimi-K2-Instruct-4bit.tomlM resources/inference_model_cards/mlx-community--Kimi-K2-Thinking.tomlM resources/inference_model_cards/mlx-community--Kimi-K2.5.tomlM resources/inference_model_cards/mlx-community--Llama-3.2-1B-Instruct-4bit.tomlM resources/inference_model_cards/mlx-community--Llama-3.2-3B-Instruct-4bit.tomlM resources/inference_model_cards/mlx-community--Llama-3.2-3B-Instruct-8bit.tomlM resources/inference_model_cards/mlx-community--Llama-3.3-70B-Instruct-4bit.tomlM resources/inference_model_cards/mlx-community--Llama-3.3-70B-Instruct-8bit.tomlM resources/inference_model_cards/mlx-community--Meta-Llama-3.1-70B-Instruct-4bit.tomlM resources/inference_model_cards/mlx-community--Meta-Llama-3.1-8B-Instruct-4bit.tomlM resources/inference_model_cards/mlx-community--Meta-Llama-3.1-8B-Instruct-8bit.tomlM resources/inference_model_cards/mlx-community--Meta-Llama-3.1-8B-Instruct-bf16.tomlM resources/inference_model_cards/mlx-community--MiniMax-M2.1-3bit.tomlM resources/inference_model_cards/mlx-community--MiniMax-M2.1-8bit.tomlM resources/inference_model_cards/mlx-community--Qwen3-0.6B-4bit.tomlM resources/inference_model_cards/mlx-community--Qwen3-0.6B-8bit.tomlM resources/inference_model_cards/mlx-community--Qwen3-235B-A22B-Instruct-2507-4bit.tomlM resources/inference_model_cards/mlx-community--Qwen3-235B-A22B-Instruct-2507-8bit.tomlM resources/inference_model_cards/mlx-community--Qwen3-30B-A3B-4bit.tomlM resources/inference_model_cards/mlx-community--Qwen3-30B-A3B-8bit.tomlM resources/inference_model_cards/mlx-community--Qwen3-Coder-480B-A35B-Instruct-4bit.tomlM resources/inference_model_cards/mlx-community--Qwen3-Coder-480B-A35B-Instruct-8bit.tomlM resources/inference_model_cards/mlx-community--Qwen3-Next-80B-A3B-Instruct-4bit.tomlM resources/inference_model_cards/mlx-community--Qwen3-Next-80B-A3B-Instruct-8bit.tomlM resources/inference_model_cards/mlx-community--Qwen3-Next-80B-A3B-Thinking-4bit.tomlM resources/inference_model_cards/mlx-community--Qwen3-Next-80B-A3B-Thinking-8bit.tomlM resources/inference_model_cards/mlx-community--gpt-oss-120b-MXFP4-Q8.tomlM resources/inference_model_cards/mlx-community--gpt-oss-20b-MXFP4-Q8.tomlM resources/inference_model_cards/mlx-community--llama-3.3-70b-instruct-fp16.tomlM src/exo/master/api.pyM src/exo/shared/models/model_cards.pyM src/exo/shared/types/api.pyM src/exo/worker/engines/mlx/auto_parallel.pyM src/exo/worker/engines/mlx/utils_mlx.py
Diff
commit 41ed7afb3bd6263c70cde9d29e5f440ce98782e8
Author: Alex Cheema <41707476+AlexCheema@users.noreply.github.com>
Date: Wed Feb 4 05:56:23 2026 -0800
feat: add model picker modal with grouped models and HF Hub search (#1369)
## Motivation
Reimplements the model picker modal from #1191 on top of the custom
model support branch. Replaces the inline model dropdown with a
full-featured modal that groups models by base model, supports
filtering, favorites, and HuggingFace Hub search.
## Changes
**Backend:**
- Add `family`, `quantization`, `base_model`, `capabilities` metadata
fields to `ModelCard` and all 40 TOML model cards
- Pass new fields through `ModelListModel` and `get_models()` API
response
- Add `GET /models/search` endpoint using
`huggingface_hub.list_models()`
**Dashboard (7 new files):**
- `ModelPickerModal.svelte` — Main modal with search, family filtering,
HuggingFace Hub tab
- `ModelPickerGroup.svelte` — Expandable model group row with
quantization variants
- `FamilySidebar.svelte` — Vertical sidebar with family icons (All,
Favorites, Hub, model families)
- `FamilyLogos.svelte` — SVG icons for each model family
- `ModelFilterPopover.svelte` — Capability and size range filters
- `HuggingFaceResultItem.svelte` — HF search result item with
download/like counts
- `favorites.svelte.ts` — localStorage-backed favorites store
**Integration:**
- Replace inline dropdown in `+page.svelte` with button that opens
`ModelPickerModal`
- Custom models shown in Hub tab with delete support
**Polish:**
- Real brand logos (Meta, Qwen, DeepSeek, OpenAI, GLM, MiniMax, Kimi,
HuggingFace) from Simple Icons / LobeHub
- Clean SVG stroke icons for capabilities (thinking, code, vision, image
gen)
- Consistent `border-exo-yellow/10` borders, descriptive tooltips
throughout
- Cluster memory (used/total) shown in modal header
- Selected model highlight with checkmark for both single and
multi-variant groups
- Cursor pointer on all interactive elements, fix filter popover
click-outside bug
- Custom models now appear in All tab alongside built-in models
## Bug Fix: Gemma 3 EOS tokens
Also included in this branch: fix for Gemma 3 models generating infinite
`<end_of_turn>` tokens. The tokenizer's `eos_token_ids` was missing
token ID 106 (`<end_of_turn>`), so generation never stopped. The fix
appends this token to the EOS list after loading the tokenizer. Also
handles `eos_token_ids` being a `set` (not just a `list`).
## Why It Works
Model metadata (family, capabilities, etc.) is stored directly in TOML
cards rather than derived from heuristics, ensuring accuracy. The modal
groups models by `base_model` field so quantization variants appear
together. Custom models are separated into the Hub tab since they lack
grouping metadata.
## Test Plan
### Manual Testing
- Open dashboard, click model selector to open modal
- Browse models by family sidebar, search, and filters
- Expand model groups to see quantization variants
- Star favorites and verify persistence across page reloads
- Navigate to Hub tab, search and add models
- Verify error messages shown for invalid model IDs
- Run a Gemma 3 model and verify generation stops at `<end_of_turn>`
### Automated Testing
- `uv run basedpyright` — 0 errors
- `uv run ruff check` — passes
- `nix fmt` — clean
- `uv run pytest src/` — 173 passed
- `cd dashboard && npm run build` — builds successfully
---------
Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com>
---
.gitignore | 1 +
.mlx_typings/mlx_lm/tokenizer_utils.pyi | 3 +-
dashboard/src/lib/components/FamilyLogos.svelte | 73 ++
dashboard/src/lib/components/FamilySidebar.svelte | 142 ++++
.../lib/components/HuggingFaceResultItem.svelte | 127 ++++
.../src/lib/components/ModelFilterPopover.svelte | 182 +++++
.../src/lib/components/ModelPickerGroup.svelte | 324 +++++++++
.../src/lib/components/ModelPickerModal.svelte | 748 +++++++++++++++++++++
dashboard/src/lib/components/index.ts | 6 +
dashboard/src/lib/stores/favorites.svelte.ts | 97 +++
dashboard/src/routes/+page.svelte | 341 ++--------
.../mlx-community--DeepSeek-V3.1-4bit.toml | 4 +
.../mlx-community--DeepSeek-V3.1-8bit.toml | 4 +
.../mlx-community--GLM-4.5-Air-8bit.toml | 4 +
.../mlx-community--GLM-4.5-Air-bf16.toml | 4 +
.../mlx-community--GLM-4.7-4bit.toml | 4 +
.../mlx-community--GLM-4.7-6bit.toml | 4 +
.../mlx-community--GLM-4.7-8bit-gs32.toml | 4 +
.../mlx-community--GLM-4.7-Flash-4bit.toml | 4 +
.../mlx-community--GLM-4.7-Flash-5bit.toml | 4 +
.../mlx-community--GLM-4.7-Flash-6bit.toml | 4 +
.../mlx-community--GLM-4.7-Flash-8bit.toml | 4 +
.../mlx-community--Kimi-K2-Instruct-4bit.toml | 4 +
.../mlx-community--Kimi-K2-Thinking.toml | 4 +
.../mlx-community--Kimi-K2.5.toml | 4 +
.../mlx-community--Llama-3.2-1B-Instruct-4bit.toml | 4 +
.../mlx-community--Llama-3.2-3B-Instruct-4bit.toml | 4 +
.../mlx-community--Llama-3.2-3B-Instruct-8bit.toml | 4 +
...mlx-community--Llama-3.3-70B-Instruct-4bit.toml | 4 +
...mlx-community--Llama-3.3-70B-Instruct-8bit.toml | 4 +
...ommunity--Meta-Llama-3.1-70B-Instruct-4bit.toml | 4 +
...community--Meta-Llama-3.1-8B-Instruct-4bit.toml | 4 +
...community--Meta-Llama-3.1-8B-Instruct-8bit.toml | 4 +
...community--Meta-Llama-3.1-8B-Instruct-bf16.toml | 4 +
.../mlx-community--MiniMax-M2.1-3bit.toml | 4 +
.../mlx-community--MiniMax-M2.1-8bit.toml | 4 +
.../mlx-community--Qwen3-0.6B-4bit.toml | 4 +
.../mlx-community--Qwen3-0.6B-8bit.toml | 4 +
...munity--Qwen3-235B-A22B-Instruct-2507-4bit.toml | 4 +
...munity--Qwen3-235B-A22B-Instruct-2507-8bit.toml | 4 +
.../mlx-community--Qwen3-30B-A3B-4bit.toml | 4 +
.../mlx-community--Qwen3-30B-A3B-8bit.toml | 4 +
...unity--Qwen3-Coder-480B-A35B-Instruct-4bit.toml | 4 +
...unity--Qwen3-Coder-480B-A35B-Instruct-8bit.toml | 4 +
...ommunity--Qwen3-Next-80B-A3B-Instruct-4bit.toml | 4 +
...ommunity--Qwen3-Next-80B-A3B-Instruct-8bit.toml | 4 +
...ommunity--Qwen3-Next-80B-A3B-Thinking-4bit.toml | 4 +
...ommunity--Qwen3-Next-80B-A3B-Thinking-8bit.toml | 4 +
.../mlx-community--gpt-oss-120b-MXFP4-Q8.toml | 4 +
.../mlx-community--gpt-oss-20b-MXFP4-Q8.toml | 4 +
...mlx-community--llama-3.3-70b-instruct-fp16.toml | 4 +
src/exo/master/api.py | 30 +
src/exo/shared/models/model_cards.py | 4 +
src/exo/shared/types/api.py | 13 +
src/exo/worker/engines/mlx/auto_parallel.py | 6 +
src/exo/worker/engines/mlx/utils_mlx.py | 11 +
56 files changed, 1998 insertions(+), 270 deletions(-)
diff --git a/.gitignore b/.gitignore
index d0b8299e..139fc326 100644
--- a/.gitignore
+++ b/.gitignore
@@ -31,3 +31,4 @@ dashboard/.svelte-kit/
# host config snapshots
hosts_*.json
+.swp
diff --git a/.mlx_typings/mlx_lm/tokenizer_utils.pyi b/.mlx_typings/mlx_lm/tokenizer_utils.pyi
index 251e3d28..83eb4e33 100644
--- a/.mlx_typings/mlx_lm/tokenizer_utils.pyi
+++ b/.mlx_typings/mlx_lm/tokenizer_utils.pyi
@@ -108,6 +108,7 @@ class TokenizerWrapper:
_tokenizer: PreTrainedTokenizerFast
eos_token_id: int | None
eos_token: str | None
+ eos_token_ids: list[int] | set[int] | None
bos_token_id: int | None
bos_token: str | None
vocab_size: int
@@ -117,7 +118,7 @@ class TokenizerWrapper:
self,
tokenizer: Any,
detokenizer_class: Any = ...,
- eos_token_ids: list[int] | None = ...,
+ eos_token_ids: list[int] | set[int] | None = ...,
chat_template: Any = ...,
tool_parser: Any = ...,
tool_call_start: str | None = ...,
diff --git a/dashboard/src/lib/components/FamilyLogos.svelte b/dashboard/src/lib/components/FamilyLogos.svelte
new file mode 100644
index 00000000..8e2919fe
--- /dev/null
+++ b/dashboard/src/lib/components/FamilyLogos.svelte
@@ -0,0 +1,73 @@
+<script lang="ts">
+ type FamilyLogoProps = {
+ family: string;
+ class?: string;
+ };
+
+ let { family, class: className = "" }: FamilyLogoProps = $props();
+</script>
+
+{#if family === "favorites"}
+ <svg class="w-6 h-6 {className}" viewBox="0 0 24 24" fill="currentColor">
+ <path
+ d="M12 2l3.09 6.26L22 9.27l-5 4.87 1.18 6.88L12 17.77l-6.18 3.25L7 14.14 2 9.27l6.91-1.01L12 2z"
+ />
+ </svg>
+{:else if family === "llama" || family === "meta"}
+ <svg class="w-6 h-6 {className}" viewBox="0 0 24 24" fill="currentColor">
+ <path
+ d="M6.915 4.03c-1.968 0-3.683 1.28-4.871 3.113C.704 9.208 0 11.883 0 14.449c0 .706.07 1.369.21 1.973a6.624 6.624 0 0 0 .265.86 5.297 5.297 0 0 0 .371.761c.696 1.159 1.818 1.927 3.593 1.927 1.497 0 2.633-.671 3.965-2.444.76-1.012 1.144-1.626 2.663-4.32l.756-1.339.186-.325c.061.1.121.196.183.3l2.152 3.595c.724 1.21 1.665 2.556 2.47 3.314 1.046.987 1.992 1.22 3.06 1.22 1.075 0 1.876-.355 2.455-.843a3.743 3.743 0 0 0 .81-.973c.542-.939.861-2.127.861-3.745 0-2.72-.681-5.357-2.084-7.45-1.282-1.912-2.957-2.93-4.716-2.93-1.047 0-2.088.467-3.053 1.308-.652.57-1.257 1.29-1.82 2.05-.69-.875-1.335-1.547-1.958-2.056-1.182-.966-2.315-1.303-3.454-1.303zm10.16 2.053c1.147 0 2.188.758 2.992 1.999 1.132 1.748 1.647 4.195 1.647 6.4 0 1.548-.368 2.9-1.839 2.9-.58 0-1.027-.23-1.664-1.004-.496-.601-1.343-1.878-2.832-4.358l-.617-1.028a44.908 44.908 0 0 0-1.255-1.98c.07-.109.141-.224.211-.327 1.12-1.667 2.118-2.602 3.358-2.602zm-10.201.553c1.265 0 2.058.791 2.675 1.446.307.327.737.871 1.234 1.579l-1.02 1.566c-.757 1.163-1.882 3.017-2.837 4.338-1.191 1.649-1.81 1.817-2.486 1.817-.524 0-1.038-.237-1.383-.794-.263-.426-.464-1.13-.464-2.046 0-2.221.63-4.535 1.66-6.088.454-.687.964-1.226 1.533-1.533a2.264 2.264 0 0 1 1.088-.285z"
+ />
+ </svg>
+{:else if family === "qwen"}
+ <svg class="w-6 h-6 {className}" viewBox="0 0 24 24" fill="currentColor">
+ <path
+ d="M12.604 1.34c.393.69.784 1.382 1.174 2.075a.18.18 0 00.157.091h5.552c.174 0 .322.11.446.327l1.454 2.57c.19.337.24.478.024.837-.26.43-.513.864-.76 1.3l-.367.658c-.106.196-.223.28-.04.512l2.652 4.637c.172.301.111.494-.043.77-.437.785-.882 1.564-1.335 2.34-.159.272-.352.375-.68.37-.777-.016-1.552-.01-2.327.016a.099.099 0 00-.081.05 575.097 575.097 0 01-2.705 4.74c-.169.293-.38.363-.725.364-.997.003-2.002.004-3.017.002a.537.537 0 01-.465-.271l-1.335-2.323a.09.09 0 00-.083-.049H4.982c-.285.03-.553-.001-.805-.092l-1.603-2.77a.543.543 0 01-.002-.54l1.207-2.12a.198.198 0 000-.197 550.951 550.951 0 01-1.875-3.272l-.79-1.395c-.16-.31-.173-.496.095-.965.465-.813.927-1.625 1.387-2.436.132-.234.304-.334.584-.335a338.3 338.3 0 012.589-.001.124.124 0 00.107-.063l2.806-4.895a.488.488 0 01.422-.246c.524-.001 1.053 0 1.583-.006L11.704 1c.341-.003.724.032.9.34zm-3.432.403a.06.06 0 00-.052.03L6.254 6.788a.157.157 0 01-.135.078H3.253c-.056 0-.07.025-.041.074l5.81 10.156c.025.042.013.062-.034.063l-2.795.015a.218.218 0 00-.2.116l-1.32 2.31c-.044.078-.021.118.068.118l5.716.008c.046 0 .08.02.104.061l1.403 2.454c.046.081.092.082.139 0l5.006-8.76.783-1.382a.055.055 0 01.096 0l1.424 2.53a.122.122 0 00.107.062l2.763-.02a.04.04 0 00.035-.02.041.041 0 000-.04l-2.9-5.086a.108.108 0 010-.113l.293-.507 1.12-1.977c.024-.041.012-.062-.035-.062H9.2c-.059 0-.073-.026-.043-.077l1.434-2.505a.107.107 0 000-.114L9.225 1.774a.06.06 0 00-.053-.031zm6.29 8.02c.046 0 .058.02.034.06l-.832 1.465-2.613 4.585a.056.056 0 01-.05.029.058.058 0 01-.05-.029L8.498 9.841c-.02-.034-.01-.052.028-.054l.216-.012 6.722-.012z"
+ />
+ </svg>
+{:else if family === "deepseek"}
+ <svg class="w-6 h-6 {className}" viewBox="0 0 24 24" fill="currentColor">
+ <path
+ d="M23.748 4.482c-.254-.124-.364.113-.512.234-.051.039-.094.09-.137.136-.372.397-.806.657-1.373.626-.829-.046-1.537.214-2.163.848-.133-.782-.575-1.248-1.247-1.548-.352-.156-.708-.311-.955-.65-.172-.241-.219-.51-.305-.774-.055-.16-.11-.323-.293-.35-.2-.031-.278.136-.356.276-.313.572-.434 1.202-.422 1.84.027 1.436.633 2.58 1.838 3.393.137.093.172.187.129.323-.082.28-.18.552-.266.833-.055.179-.137.217-.329.14a5.526 5.526 0 01-1.736-1.18c-.857-.828-1.631-1.742-2.597-2.458a11.365 11.365 0 00-.689-.471c-.985-.957.13-1.743.388-1.836.27-.098.093-.432-.779-.428-.872.004-1.67.295-2.687.684a3.055 3.055 0 01-.465.137 9.597 9.597 0 00-2.883-.102c-1.885.21-3.39 1.102-4.497 2.623C.082 8.606-.231 10.684.152 12.85c.403 2.284 1.569 4.175 3.36 5.653 1.858 1.533 3.997 2.284 6.438 2.14 1.482-.085 3.133-.284 4.994-1.86.47.234.962.327 1.78.397.63.059 1.236-.03 1.705-.128.735-.156.684-.837.419-.961-2.155-1.004-1.682-.595-2.113-.926 1.096-1.296 2.746-2.642 3.392-7.003.05-.347.007-.565 0-.845-.004-.17.035-.237.23-.256a4.173 4.173 0 001.545-.475c1.396-.763 1.96-2.015 2.093-3.517.02-.23-.004-.467-.247-.588zM11.581 18c-2.089-1.642-3.102-2.183-3.52-2.16-.392.024-.321.471-.235.763.09.288.207.486.371.739.114.167.192.416-.113.603-.673.416-1.842-.14-1.897-.167-1.361-.802-2.5-1.86-3.301-3.307-.774-1.393-1.224-2.887-1.298-4.482-.02-.386.093-.522.477-.592a4.696 4.696 0 011.529-.039c2.132.312 3.946 1.265 5.468 2.774.868.86 1.525 1.887 2.202 2.891.72 1.066 1.494 2.082 2.48 2.914.348.292.625.514.891.677-.802.09-2.14.11-3.054-.614zm1-6.44a.306.306 0 01.415-.287.302.302 0 01.2.288.306.306 0 01-.31.307.303.303 0 01-.304-.308zm3.11 1.596c-.2.081-.399.151-.59.16a1.245 1.245 0 01-.798-.254c-.274-.23-.47-.358-.552-.758a1.73 1.73 0 01.016-.588c.07-.327-.008-.537-.239-.727-.187-.156-.426-.199-.688-.199a.559.559 0 01-.254-.078c-.11-.054-.2-.19-.114-.358.028-.054.16-.186.192-.21.356-.202.767-.136 1.146.016.352.144.618.408 1.001.782.391.451.462.576.685.914.176.265.336.537.445.848.067.195-.019.354-.25.452z"
+ />
+ </svg>
+{:else if family === "openai" || family === "gpt-oss"}
+ <svg class="w-6 h-6 {className}" viewBox="0 0 24 24" fill="currentColor">
+ <path
+ d="M22.2819 9.8211a5.9847 5.9847 0 0 0-.5157-4.9108 6.0462 6.0462 0 0 0-6.5098-2.9A6.0651 6.0651 0 0 0 4.9807 4.1818a5.9847 5.9847 0 0 0-3.9977 2.9 6.0462 6.0462 0 0 0 .7427 7.0966 5.98 5.98 0 0 0 .511 4.9107 6.051 6.051 0 0 0 6.5146 2.9001A5.9847 5.9847 0 0 0 13.2599 24a6.0557 6.0557 0 0 0 5.7718-4.2058 5.9894 5.9894 0 0 0 3.9977-2.9001 6.0557 6.0557 0 0 0-.7475-7.0729zm-9.022 12.6081a4.4755 4.4755 0 0 1-2.8764-1.0408l.1419-.0804 4.7783-2.7582a.7948.7948 0 0 0 .3927-.6813v-6.7369l2.02 1.1686a.071.071 0 0 1 .038.052v5.5826a4.504 4.504 0 0 1-4.4945 4.4944zm-9.6607-4.1254a4.4708 4.4708 0 0 1-.5346-3.0137l.142.0852 4.783 2.7582a.7712.7712 0 0 0 .7806 0l5.8428-3.3685v2.3324a.0804.0804 0 0 1-.0332.0615L9.74 19.9502a4.4992 4.4992 0 0 1-6.1408-1.6464zM2.3408 7.8956a4.485 4.485 0 0 1 2.3655-1.9728V11.6a.7664.7664 0 0 0 .3879.6765l5.8144 3.3543-2.0201 1.1685a.0757.0757 0 0 1-.071 0l-4.8303-2.7865A4.504 4.504 0 0 1 2.3408 7.872zm16.5963 3.8558L13.1038 8.364 15.1192 7.2a.0757.0757 0 0 1 .071 0l4.8303 2.7913a4.4944 4.4944 0 0 1-.6765 8.1042v-5.6772a.79.79 0 0 0-.407-.667zm2.0107-3.0231l-.142-.0852-4.7735-2.7818a.7759.7759 0 0 0-.7854 0L9.409 9.2297V6.8974a.0662.0662 0 0 1 .0284-.0615l4.8303-2.7866a4.4992 4.4992 0 0 1 6.6802 4.66zM8.3065 12.863l-2.02-1.1638a.0804.0804 0 0 1-.038-.0567V6.0742a4.4992 4.4992 0 0 1 7.3757-3.4537l-.142.0805L8.704 5.459a.7948.7948 0 0 0-.3927.6813zm1.0976-2.3654l2.602-1.4998 2.6069 1.4998v2.9994l-2.5974 1.4997-2.6067-1.4997Z"
+ />
+ </svg>
+{:else if family === "glm"}
+ <svg class="w-6 h-6 {className}" viewBox="0 0 24 24" fill="currentColor">
+ <path
+ d="M11.991 23.503a.24.24 0 00-.244.248.24.24 0 00.244.249.24.24 0 00.245-.249.24.24 0 00-.22-.247l-.025-.001zM9.671 5.365a1.697 1.697 0 011.099 2.132l-.071.172-.016.04-.018.054c-.07.16-.104.32-.104.498-.035.71.47 1.279 1.186 1.314h.366c1.309.053 2.338 1.173 2.286 2.523-.052 1.332-1.152 2.38-2.478 2.327h-.174c-.715.018-1.274.64-1.239 1.368 0 .124.018.23.053.337.209.373.54.658.96.8.75.23 1.517-.125 1.9-.782l.018-.035c.402-.64 1.17-.96 1.92-.711.854.284 1.378 1.226 1.099 2.167a1.661 1.661 0 01-2.077 1.102 1.711 1.711 0 01-.907-.711l-.017-.035c-.2-.323-.463-.58-.851-.711l-.056-.018a1.646 1.646 0 00-1.954.746 1.66 1.66 0 01-1.065.764 1.677 1.677 0 01-1.989-1.279c-.209-.906.332-1.83 1.257-2.043a1.51 1.51 0 01.296-.035h.018c.68-.071 1.151-.622 1.116-1.333a1.307 1.307 0 00-.227-.693 2.515 2.515 0 01-.366-1.403 2.39 2.39 0 01.366-1.208c.14-.195.21-.444.227-.693.018-.71-.506-1.261-1.186-1.332l-.07-.018a1.43 1.43 0 01-.299-.07l-.05-.019a1.7 1.7 0 01-1.047-2.114 1.68 1.68 0 012.094-1.101zm-5.575 10.11c.26-.264.639-.367.994-.27.355.096.633.379.728.74.095.362-.007.748-.267 1.013-.402.41-1.053.41-1.455 0a1.062 1.062 0 010-1.482zm14.845-.294c.359-.09.738.024.992.297.254.274.344.665.237 1.025-.107.36-.396.634-.756.718-.551.128-1.1-.22-1.23-.781a1.05 1.05 0 01.757-1.26zm-.064-4.39c.314.32.49.753.49 1.206 0 .452-.176.886-.49 1.206-.315.32-.74.5-1.185.5-.444 0-.87-.18-1.184-.5a1.727 1.727 0 010-2.412 1.654 1.654 0 012.369 0zm-11.243.163c.364.484.447 1.128.218 1.691a1.665 1.665 0 01-2.188.923c-.855-.36-1.26-1.358-.907-2.228a1.68 1.68 0 011.33-1.038c.593-.08 1.183.169 1.547.652zm11.545-4.221c.368 0 .708.2.892.524.184.324.184.724 0 1.048a1.026 1.026 0 01-.892.524c-.568 0-1.03-.47-1.03-1.048 0-.579.462-1.048 1.03-1.048zm-14.358 0c.368 0 .707.2.891.524.184.324.184.724 0 1.048a1.026 1.026 0 01-.891.524c-.569 0-1.03-.47-1.03-1.048 0-.579.461-1.048 1.03-1.048zm10.031-1.475c.925 0 1.675.764 1.675 1.706s-.75 1.705-1.675 1.705-1.674-.763-1.674-1.705c0-.942.75-1.706 1.674-1.706zm-2.626-.684c.362-.082.653-.356.761-.718a1.062 1.062 0 00-.238-1.028 1.017 1.017 0 00-.996-.294c-.547.14-.881.7-.752 1.257.13.558.675.907 1.225.783zm0 16.876c.359-.087.644-.36.75-.72a1.062 1.062 0 00-.237-1.019 1.018 1.018 0 00-.985-.301 1.037 1.037 0 00-.762.717c-.108.361-.017.754.239 1.028.245.263.606.377.953.305l.043-.01zM17.19 3.5a.631.631 0 00.628-.64c0-.355-.279-.64-.628-.64a.631.631 0 00-.628.64c0 .355.28.64.628.64zm-10.38 0a.631.631 0 00.628-.64c0-.355-.28-.64-.628-.64a.631.631 0 00-.628.64c0 .355.279.64.628.64zm-5.182 7.852a.631.631 0 00-.628.64c0 .354.28.639.628.639a.63.63 0 00.627-.606l.001-.034a.62.62 0 00-.628-.64zm5.182 9.13a.631.631 0 00-.628.64c0 .355.279.64.628.64a.631.631 0 00.628-.64c0-.355-.28-.64-.628-.64zm10.38.018a.631.631 0 00-.628.64c0 .355.28.64.628.64a.631.631 0 00.628-.64c0-.355-.279-.64-.628-.64zm5.182-9.148a.631.631 0 00-.628.64c0 .354.279.639.628.639a.631.631 0 00.628-.64c0-.355-.28-.64-.628-.64zm-.384-4.992a.24.24 0 00.244-.249.24.24 0 00-.244-.249.24.24 0 00-.244.249c0 .142.122.249.244.249zM11.991.497a.24.24 0 00.245-.248A.24.24 0 0011.99 0a.24.24 0 00-.244.249c0 .133.108.236.223.247l.021.001zM2.011 6.36a.24.24 0 00.245-.249.24.24 0 00-.244-.249.24.24 0 00-.244.249.24.24 0 00.244.249zm0 11.263a.24.24 0 00-.243.248.24.24 0 00.244.249.24.24 0 00.244-.249.252.252 0 00-.244-.248zm19.995-.018a.24.24 0 00-.245.248.24.24 0 00.245.25.24.24 0 00.244-.25.252.252 0 00-.244-.248z"
+ />
+ </svg>
+{:else if family === "minimax"}
+ <svg class="w-6 h-6 {className}" viewBox="0 0 24 24" fill="currentColor">
+ <path
+ d="M16.278 2c1.156 0 2.093.927 2.093 2.07v12.501a.74.74 0 00.744.709.74.74 0 00.743-.709V9.099a2.06 2.06 0 012.071-2.049A2.06 2.06 0 0124 9.1v6.561a.649.649 0 01-.652.645.649.649 0 01-.653-.645V9.1a.762.762 0 00-.766-.758.762.762 0 00-.766.758v7.472a2.037 2.037 0 01-2.048 2.026 2.037 2.037 0 01-2.048-2.026v-12.5a.785.785 0 00-.788-.753.785.785 0 00-.789.752l-.001 15.904A2.037 2.037 0 0113.441 22a2.037 2.037 0 01-2.048-2.026V18.04c0-.356.292-.645.652-.645.36 0 .652.289.652.645v1.934c0 .263.142.506.372.638.23.131.514.131.744 0a.734.734 0 00.372-.638V4.07c0-1.143.937-2.07 2.093-2.07zm-5.674 0c1.156 0 2.093.927 2.093 2.07v11.523a.648.648 0 01-.652.645.648.648 0 01-.652-.645V4.07a.785.785 0 00-.789-.78.785.785 0 00-.789.78v14.013a2.06 2.06 0 01-2.07 2.048 2.06 2.06 0 01-2.071-2.048V9.1a.762.762 0 00-.766-.758.762.762 0 00-.766.758v3.8a2.06 2.06 0 01-2.071 2.049A2.06 2.06 0 010 12.9v-1.378c0-.357.292-.646.652-.646.36 0 .653.29.653.646V12.9c0 .418.343.757.766.757s.766-.339.766-.757V9.099a2.06 2.06 0 012.07-2.048 2.06 2.06 0 012.071 2.048v8.984c0 .419.343.758.767.758.423 0 .766-.339.766-.758V4.07c0-1.143.937-2.07 2.093-2.07z"
+ />
+ </svg>
+{:else if family === "kimi"}
+ <svg class="w-6 h-6 {className}" viewBox="0 0 24 24" fill="currentColor">
+ <path
+ d="M19.738 5.776c.163-.209.306-.4.457-.585.07-.087.064-.153-.004-.244-.655-.861-.717-1.817-.34-2.787.283-.73.909-1.072 1.674-1.145.477-.045.945.004 1.379.236.57.305.902.77 1.01 1.412.086.512.07 1.012-.075 1.508-.257.878-.888 1.333-1.753 1.448-.718.096-1.446.108-2.17.157-.056.004-.113 0-.178 0z"
+ />
+ <path
+ d="M17.962 1.844h-4.326l-3.425 7.81H5.369V1.878H1.5V22h3.87v-8.477h6.824a3.025 3.025 0 002.743-1.75V22h3.87v-8.477a3.87 3.87 0 00-3.588-3.86v-.01h-2.125a3.94 3.94 0 002.323-2.12l2.545-5.689z"
+ />
+ </svg>
+{:else if family === "huggingface"}
+ <svg class="w-6 h-6 {className}" viewBox="0 0 24 24" fill="currentColor">
+ <path
+ d="M12.025 1.13c-5.77 0-10.449 4.647-10.449 10.378 0 1.112.178 2.181.503 3.185.064-.222.203-.444.416-.577a.96.96 0 0 1 .524-.15c.293 0 .584.124.84.284.278.173.48.408.71.694.226.282.458.611.684.951v-.014c.017-.324.106-.622.264-.874s.403-.487.762-.543c.3-.047.596.06.787.203s.31.313.4.467c.15.257.212.468.233.542.01.026.653 1.552 1.657 2.54.616.605 1.01 1.223 1.082 1.912.055.537-.096 1.059-.38 1.572.637.121 1.294.187 1.967.187.657 0 1.298-.063 1.921-.178-.287-.517-.44-1.041-.384-1.581.07-.69.465-1.307 1.081-1.913 1.004-.987 1.647-2.513 1.657-2.539.021-.074.083-.285.233-.542.09-.154.208-.323.4-.467a1.08 1.08 0 0 1 .787-.203c.359.056.604.29.762.543s.247.55.265.874v.015c.225-.34.457-.67.683-.952.23-.286.432-.52.71-.694.257-.16.547-.284.84-.285a.97.97 0 0 1 .524.151c.228.143.373.388.43.625l.006.04a10.3 10.3 0 0 0 .534-3.273c0-5.731-4.678-10.378-10.449-10.378M8.327 6.583a1.5 1.5 0 0 1 .713.174 1.487 1.487 0 0 1 .617 2.013c-.183.343-.762-.214-1.102-.094-.38.134-.532.914-.917.71a1.487 1.487 0 0 1 .69-2.803m7.486 0a1.487 1.487 0 0 1 .689 2.803c-.385.204-.536-.576-.916-.71-.34-.12-.92.437-1.103.094a1.487 1.487 0 0 1 .617-2.013 1.5 1.5 0 0 1 .713-.174m-10.68 1.55a.96.96 0 1 1 0 1.921.96.96 0 0 1 0-1.92m13.838 0a.96.96 0 1 1 0 1.92.96.96 0 0 1 0-1.92M8.489 11.458c.588.01 1.965 1.157 3.572 1.164 1.607-.007 2.984-1.155 3.572-1.164.196-.003.305.12.305.454 0 .886-.424 2.328-1.563 3.202-.22-.756-1.396-1.366-1.63-1.32q-.011.001-.02.006l-.044.026-.01.008-.03.024q-.018.017-.035.036l-.032.04a1 1 0 0 0-.058.09l-.014.025q-.049.088-.11.19a1 1 0 0 1-.083.116 1.2 1.2 0 0 1-.173.18q-.035.029-.075.058a1.3 1.3 0 0 1-.251-.243 1 1 0 0 1-.076-.107c-.124-.193-.177-.363-.337-.444-.034-.016-.104-.008-.2.022q-.094.03-.216.087-.06.028-.125.063l-.13.074q-.067.04-.136.086a3 3 0 0 0-.135.096 3 3 0 0 0-.26.219 2 2 0 0 0-.12.121 2 2 0 0 0-.106.128l-.002.002a2 2 0 0 0-.09.132l-.001.001a1.2 1.2 0 0 0-.105.212q-.013.036-.024.073c-1.139-.875-1.563-2.317-1.563-3.203 0-.334.109-.457.305-.454m.836 10.354c.824-1.19.766-2.082-.365-3.194-1.13-1.112-1.789-2.738-1.789-2.738s-.246-.945-.806-.858-.97 1.499.202 2.362c1.173.864-.233 1.45-.685.64-.45-.812-1.683-2.896-2.322-3.295s-1.089-.175-.938.647 2.822 2.813 2.562 3.244-1.176-.506-1.176-.506-2.866-2.567-3.49-1.898.473 1.23 2.037 2.16c1.564.932 1.686 1.178 1.464 1.53s-3.675-2.511-4-1.297c-.323 1.214 3.524 1.567 3.287 2.405-.238.839-2.71-1.587-3.216-.642-.506.946 3.49 2.056 3.522 2.064 1.29.33 4.568 1.028 5.713-.624m5.349 0c-.824-1.19-.766-2.082.365-3.194 1.13-1.112 1.789-2.738 1.789-2.738s.246-.945.806-.858.97 1.499-.202 2.362c-1.173.864.233 1.45.685.64.451-.812 1.683-2.896 2.322-3.295s1.089-.175.938.647-2.822 2.813-2.562 3.244 1.176-.506 1.176-.506 2.866-2.567 3.49-1.898-.473 1.23-2.037 2.16c-1.564.932-1.686 1.178-1.464 1.53s3.675-2.511 4-1.297c.323 1.214-3.524 1.567-3.287 2.405.238.839 2.71-1.587 3.216-.642.506.946-3.49 2.056-3.522 2.064-1.29.33-4.568 1.028-5.713-.624"
+ />
+ </svg>
+{:else}
+ <svg class="w-6 h-6 {className}" viewBox="0 0 24 24" fill="currentColor">
+ <path
+ d="M12 2C6.48 2 2 6.48 2 12s4.48 10 10 10 10-4.48 10-10S17.52 2 12 2zm-2 15l-5-5 1.41-1.41L10 14.17l7.59-7.59L19 8l-9 9z"
+ />
+ </svg>
+{/if}
diff --git a/dashboard/src/lib/components/FamilySidebar.svelte b/dashboard/src/lib/components/FamilySidebar.svelte
new file mode 100644
index 00000000..886a68d5
--- /dev/null
+++ b/dashboard/src/lib/components/FamilySidebar.svelte
@@ -0,0 +1,142 @@
+<script lang="ts">
+ import FamilyLogos from "./FamilyLogos.svelte";
+
+ type FamilySidebarProps = {
+ families: string[];
+ selectedFamily: string | null;
+ hasFavorites: boolean;
+ onSelect: (family: string | null) => void;
+ };
+
+ let { families, selectedFamily, hasFavorites, onSelect }: FamilySidebarProps =
+ $props();
+
+ // Family display names
+ const familyNames: Record<string, string> = {
+ favorites: "Favorites",
+ huggingface: "Hub",
+ llama: "Meta",
+ qwen: "Qwen",
+ deepseek: "DeepSeek",
+ "gpt-oss": "OpenAI",
+ glm: "GLM",
+ minimax: "MiniMax",
+ kimi: "Kimi",
+ };
+
+ function getFamilyName(family: string): string {
+ return (
+ familyNames[family] || family.charAt(0).toUpperCase() + family.slice(1)
+ );
+ }
+</script>
+
+<div
+ class="flex flex-col gap-1 py-2 px-1 border-r border-exo-yellow/10 bg-exo-medium-gray/30 min-w-[64px]"
+>
+ <!-- All models (no filter) -->
+ <button
+ type="button"
+ onclick={() => onSelect(null)}
+ class="group flex flex-col items-center justify-center p-2 rounded transition-all duration-200 cursor-pointer {selectedFamily ===
+ null
+ ? 'bg-exo-yellow/20 border-l-2 border-exo-yellow'
+ : 'hover:bg-white/5 border-l-2 border-transparent'}"
+ title="All models"
+ >
+ <svg
+ class="w-5 h-5 {selectedFamily === null
+ ? 'text-exo-yellow'
+ : 'text-white/50 group-hover:text-white/70'}"
+ viewBox="0 0 24 24"
+ fill="currentColor"
+ >
+ <path
+ d="M4 8h4V4H4v4zm6 12h4v-4h-4v4zm-6 0h4v-4H4v4zm0-6h4v-4H4v4zm6 0h4v-4h-4v4zm6-10v4h4V4h-4zm-6 4h4V4h-4v4zm6 6h4v-4h-4v4zm0 6h4v-4h-4v4z"
+ />
+ </svg>
+ <span
+ class="text-[9px] font-mono mt-0.5 {selectedFamily === null
+ ? 'text-exo-yellow'
+ : 'text-white/40 group-hover:text-white/60'}">All</span
+ >
+ </button>
+
+ <!-- Favorites (only show if has favorites) -->
+ {#if hasFavorites}
+ <button
+ type="button"
+ onclick={() => onSelect("favorites")}
+ class="group flex flex-col items-center justify-center p-2 rounded transition-all duration-200 cursor-pointer {selectedFamily ===
+ 'favorites'
+ ? 'bg-exo-yellow/20 border-l-2 border-exo-yellow'
+ : 'hover:bg-white/5 border-l-2 border-transparent'}"
+ title="Show favorited models"
+ >
+ <FamilyLogos
+ family="favorites"
+ class={selectedFamily === "favorites"
+ ? "text-amber-400"
+ : "text-white/50 group-hover:text-amber-400/70"}
+ />
+ <span
+ class="text-[9px] font-mono mt-0.5 {selectedFamily === 'favorites'
+ ? 'text-amber-400'
+ : 'text-white/40 group-hover:text-white/60'}">Faves</span
+ >
+ </button>
+ {/if}
+
+ <!-- HuggingFace Hub -->
+ <button
+ type="button"
+ onclick={() => onSelect("huggingface")}
+ class="group flex flex-col items-center justify-center p-2 rounded transition-all duration-200 cursor-pointer {selectedFamily ===
+ 'huggingface'
+ ? 'bg-orange-500/20 border-l-2 border-orange-400'
+ : 'hover:bg-white/5 border-l-2 border-transparent'}"
+ title="Browse and add models from Hugging Face"
+ >
+ <FamilyLogos
+ family="huggingface"
+ class={selectedFamily === "huggingface"
+ ? "text-orange-400"
+ : "text-white/50 group-hover:text-orange-400/70"}
+ />
+ <span
+ class="text-[9px] font-mono mt-0.5 {selectedFamily === 'huggingface'
+ ? 'text-orange-400'
+ : 'text-white/40 group-hover:text-white/60'}">Hub</span
+ >
+ </button>
+
+ <div class="h-px bg-exo-yellow/10 my-1"></div>
+
+ <!-- Model families -->
+ {#each families as family}
+ <button
+ type="button"
+ onclick={() => onSelect(family)}
+ class="group flex flex-col items-center justify-center p-2 rounded transition-all duration-200 cursor-pointer {selectedFamily ===
+ family
+ ? 'bg-exo-yellow/20 border-l-2 border-exo-yellow'
+ : 'hover:bg-white/5 border-l-2 border-transparent'}"
+ title={getFamilyName(family)}
+ >
+ <FamilyLogos
+ {family}
+ class={selectedFamily === family
+ ? "text-exo-yellow"
+ : "text-white/50 group-hover:text-white/70"}
+ />
+ <span
+ class="text-[9px] font-mono mt-0.5 truncate max-w-full {selectedFamily ===
+ family
+ ? 'text-exo-yellow'
+ : 'text-white/40 group-hover:text-white/60'}"
+ >
+ {getFamilyName(family)}
+ </span>
+ </button>
+ {/each}
+</div>
diff --git a/dashboard/src/lib/components/HuggingFaceResultItem.svelte b/dashboard/src/lib/components/HuggingFaceResultItem.svelte
new file mode 100644
index 00000000..566d8e17
--- /dev/null
+++ b/dashboard/src/lib/components/HuggingFaceResultItem.svelte
@@ -0,0 +1,127 @@
+<script lang="ts">
+ interface HuggingFaceModel {
+ id: string;
+ author: string;
+ downloads: number;
+ likes: number;
+ last_modified: string;
+ tags: string[];
+ }
+
+ type HuggingFaceResultItemProps = {
+ model: HuggingFaceModel;
+ isAdded: boolean;
+ isAdding: boolean;
+ onAdd: () => void;
+ onSelect: () => void;
+ };
+
+ let {
+ model,
+ isAdded,
+ isAdding,
+ onAdd,
+ onSelect,
+ }: HuggingFaceResultItemProps = $props();
+
+ function formatNumber(num: number): string {
+ if (num >= 1000000) {
+ return `${(num / 1000000).toFixed(1)}M`;
+ } else if (num >= 1000) {
+ return `${(num / 1000).toFixed(1)}k`;
+ }
+ return num.toString();
+ }
+
+ // Extract model name from full ID (e.g., "mlx-community/Llama-3.2-1B" -> "Llama-3.2-1B")
+ const modelName = $derived(model.id.split("/").pop() || model.id);
+</script>
+
+<div
+ class="flex items-center justify-between gap-3 px-3 py-2.5 hover:bg-white/5 transition-colors border-b border-white/5 last:border-b-0"
+>
+ <div class="flex-1 min-w-0">
+ <div class="flex items-center gap-2">
+ <span class="text-sm font-mono text-white truncate" title={model.id}
+ >{modelName}</span
+ >
+ {#if isAdded}
+ <span
+ class="px-1.5 py-0.5 text-[10px] font-mono bg-green-500/20 text-green-400 rounded"
+ >Added</span
+ >
+ {/if}
+ </div>
+ <div class="flex items-center gap-3 mt-0.5 text-xs text-white/40">
+ <span class="truncate">{model.author}</span>
+ <span
+ class="flex items-center gap-1 shrink-0"
+ title="Downloads in the last 30 days"
+ >
+ <svg
+ class="w-3 h-3"
+ fill="none"
+ stroke="currentColor"
+ viewBox="0 0 24 24"
+ >
+ <path
+ stroke-linecap="round"
+ stroke-linejoin="round"
+ stroke-width="2"
+ d="M4 16v1a3 3 0 003 3h10a3 3 0 003-3v-1m-4-4l-4 4m0 0l-4-4m4 4V4"
+ />
+ </svg>
+ {formatNumber(model.downloads)}
+ </span>
+ <span
+ class="flex items-center gap-1 shrink-0"
+ title="Community likes on Hugging Face"
+ >
+ <svg
+ class="w-3 h-3"
+ fill="none"
+ stroke="currentColor"
+ viewBox="0 0 24 24"
+ >
+ <path
+ stroke-linecap="round"
+ stroke-linejoin="round"
+ stroke-width="2"
+ d="M4.318 6.318a4.5 4.5 0 000 6.364L12 20.364l7.682-7.682a4.5 4.5 0 00-6.364-6.364L12 7.636l-1.318-1.318a4.5 4.5 0 00-6.364 0z"
+ />
+ </svg>
+ {formatNumber(model.likes)}
+ </span>
+ </div>
+ </div>
+
+ <div class="flex items-center gap-2 shrink-0">
+ {#if isAdded}
+ <button
+ type="button"
+ onclick={onSelect}
+ class="px-3 py-1.5 text-xs font-mono tracking-wider uppercase bg-exo-yellow/10 text-exo-yellow border border-exo-yellow/30 hover:bg-exo-yellow/20 transition-colors rounded cursor-pointer"
+ >
+ Select
+ </button>
+ {:else}
+ <button
+ type="button"
+ onclick={onAdd}
+ disabled={isAdding}
+ class="px-3 py-1.5 text-xs font-mono tracking-wider uppercase bg-orange-500/10 text-orange-400 border border-orange-400/30 hover:bg-orange-500/20 transition-colors rounded cursor-pointer disabled:opacity-50 disabled:cursor-not-allowed"
+ >
+ {#if isAdding}
+ <span class="flex items-center gap-1.5">
+ <span
+ class="w-3 h-3 border-2 border-orange-400 border-t-transparent rounded-full animate-spin"
+ ></span>
+ Adding...
+ </span>
+ {:else}
+ + Add
+ {/if}
+ </button>
+ {/if}
+ </div>
+</div>
diff --git a/dashboard/src/lib/components/ModelFilterPopover.svelte b/dashboard/src/lib/components/ModelFilterPopover.svelte
new file mode 100644
index 00000000..5406618a
--- /dev/null
+++ b/dashboard/src/lib/components/ModelFilterPopover.svelte
@@ -0,0 +1,182 @@
+<script lang="ts">
+ import { fly } from "svelte/transition";
+ import { cubicOut } from "svelte/easing";
+
+ interface FilterState {
+ capabilities: string[];
+ sizeRange: { min: number; max: number } | null;
+ }
+
+ type ModelFilterPopoverProps = {
+ filters: FilterState;
+ onChange: (filters: FilterState) => void;
+ onClear: () => void;
+ onClose: () => void;
+ };
+
+ let { filters, onChange, onClear, onClose }: ModelFilterPopoverProps =
+ $props();
+
+ // Available capabilities
+ const availableCapabilities = [
+ { id: "text", label: "Text" },
+ { id: "thinking", label: "Thinking" },
+ { id: "code", label: "Code" },
+ { id: "vision", label: "Vision" },
+ ];
+
+ // Size ranges
+ const sizeRanges = [
+ { label: "< 10GB", min: 0, max: 10 },
+ { label: "10-50GB", min: 10, max: 50 },
+ { label: "50-200GB", min: 50, max: 200 },
+ { label: "> 200GB", min: 200, max: 10000 },
+ ];
+
+ function toggleCapability(cap: string) {
+ const next = filters.capabilities.includes(cap)
+ ? filters.capabilities.filter((c) => c !== cap)
+ : [...filters.capabilities, cap];
+ onChange({ ...filters, capabilities: next });
+ }
+
+ function selectSizeRange(range: { min: number; max: number } | null) {
+ // Toggle off if same range is clicked
+ if (
+ filters.sizeRange &&
+ range &&
+ filters.sizeRange.min === range.min &&
+ filters.sizeRange.max === range.max
+ ) {
+ onChange({ ...filters, sizeRange: null });
+ } else {
+ onChange({ ...filters, sizeRange: range });
+ }
+ }
+
+ function handleClickOutside(e: MouseEvent) {
+ const target = e.target as HTMLElement;
+ if (
+ !target.closest(".filter-popover") &&
+ !target.closest(".filter-toggle")
+ ) {
+ onClose();
+ }
+ }
+</script>
+
+<svelte:window onclick={handleClickOutside} />
+
+<!-- svelte-ignore a11y_no_static_element_interactions -->
+<div
+ class="filter-popover absolute right-0 top-full mt-2 w-64 bg-exo-dark-gray border border-exo-yellow/10 rounded-lg shadow-xl z-10"
+ transition:fly={{ y: -10, duration: 200, easing: cubicOut }}
+ onclick={(e) => e.stopPropagation()}
+ role="dialog"
+ aria-label="Filter options"
+>
+ <div class="p-3 space-y-4">
+ <!-- Capabilities -->
+ <div>
+ <h4 class="text-xs font-mono text-white/50 mb-2">Capabilities</h4>
+ <div class="flex flex-wrap gap-1.5">
+ {#each availableCapabilities as cap}
+ {@const isSelected = filters.capabilities.includes(cap.id)}
+ <button
+ type="button"
+ class="px-2 py-1 text-xs font-mono rounded transition-colors {isSelected
+ ? 'bg-exo-yellow/20 text-exo-yellow border border-exo-yellow/30'
+ : 'bg-white/5 text-white/60 hover:bg-white/10 border border-transparent'}"
+ onclick={() => toggleCapability(cap.id)}
+ >
+ {#if cap.id === "text"}
+ <svg
+ class="w-3.5 h-3.5 inline-block"
+ viewBox="0 0 24 24"
+ fill="none"
+ stroke="currentColor"
+ stroke-width="1.5"
+ ><path
+ d="M21 15a2 2 0 0 1-2 2H7l-4 4V5a2 2 0 0 1 2-2h14a2 2 0 0 1 2 2z"
+ stroke-linecap="round"
+ stroke-linejoin="round"
+ /></svg
+ >
+ {:else if cap.id === "thinking"}
+ <svg
+ class="w-3.5 h-3.5 inline-block"
+ viewBox="0 0 24 24"
+ fill="none"
+ stroke="currentColor"
+ stroke-width="1.5"
+ ><path
+ d="M12 2a7 7 0 0 0-7 7c0 2.38 1.19 4.47 3 5.74V17a1 1 0 0 0 1 1h6a1 1 0 0 0 1-1v-2.26c1.81-1.27 3-3.36 3-5.74a7 7 0 0 0-7-7zM9 20h6M10 22h4"
+ stroke-linecap="round"
+ stroke-linejoin="round"
+ /></svg
+ >
+ {:else if cap.id === "code"}
+ <svg
+ class="w-3.5 h-3.5 inline-block"
+ viewBox="0 0 24 24"
+ fill="none"
+ stroke="currentColor"
+ stroke-width="1.5"
+ ><path
+ d="M16 18l6-6-6-6M8 6l-6 6 6 6"
+ stroke-linecap="round"
+ stroke-linejoin="round"
+ /></svg
+ >
+ {:else if cap.id === "vision"}
+ <svg
+ class="w-3.5 h-3.5 inline-block"
+ viewBox="0 0 24 24"
+ fill="none"
+ stroke="currentColor"
+ stroke-width="1.5"
+ ><path
+ d="M1 12s4-8 11-8 11 8 11 8-4 8-11 8-11-8-11-8z"
+ stroke-linecap="round"
+ stroke-linejoin="round"
+ /><circle cx="12" cy="12" r="3" /></svg
+ >
+ {/if}
+ <span class="ml-1">{cap.label}</span>
+ </button>
+ {/each}
+ </div>
+ </div>
+
+ <!-- Size range -->
+ <div>
+ <h4 class="text-xs font-mono text-white/50 mb-2">Model Size</h4>
+ <div class="flex flex-wrap gap-1.5">
+ {#each sizeRanges as range}
+ {@const isSelected =
+ filters.sizeRange &&
+ filters.sizeRange.min === range.min &&
+ filters.sizeRange.max === range.max}
+ <button
+ type="button"
+ class="px-2 py-1 text-xs font-mono rounded transition-colors {isSelected
+ ? 'bg-exo-yellow/20 text-exo-yellow border border-exo-yellow/30'
+ : 'bg-white/5 text-white/60 hover:bg-white/10 border border-transparent'}"
+ onclick={() => selectSizeRange(range)}
+ >
+ {range.label}
+ </button>
+ {/each}
+ </div>
+ </div>
+
+ <!-- Clear button -->
+ <button
+ type="button"
+ class="w-full py-1.5 text-xs font-mono text-white/50 hover:text-white/70 hover:bg-white/5 rounded transition-colors"
+ onclick={onClear}
+ >
+ Clear all filters
+ </button>
+ </div>
+</div>
diff --git a/dashboard/src/lib/components/ModelPickerGroup.svelte b/dashboard/src/lib/components/ModelPickerGroup.svelte
new file mode 100644
index 00000000..b3ad425b
--- /dev/null
+++ b/dashboard/src/lib/components/ModelPickerGroup.svelte
@@ -0,0 +1,324 @@
+<script lang="ts">
+ interface ModelInfo {
+ id: string;
+ name?: string;
+ storage_size_megabytes?: number;
+ base_model?: string;
+ quantization?: string;
+ supports_tensor?: boolean;
+ capabilities?: string[];
+ family?: string;
+ is_custom?: boolean;
+ }
+
+ interface ModelGroup {
+ id: string;
+ name: string;
+ capabilities: string[];
+ family: string;
+ variants: ModelInfo[];
+ smallestVariant: ModelInfo;
+ hasMultipleVariants: boolean;
+ }
+
+ type ModelPickerGroupProps = {
+ group: ModelGroup;
+ isExpanded: boolean;
+ isFavorite: boolean;
+ selectedModelId: string | null;
+ canModelFit: (id: string) => boolean;
+ onToggleExpand: () => void;
+ onSelectModel: (modelId: string) => void;
+ onToggleFavorite: (baseModelId: string) => void;
+ onShowInfo: (group: ModelGroup) => void;
+ };
+
+ let {
+ group,
+ isExpanded,
+ isFavorite,
+ selectedModelId,
+ canModelFit,
+ onToggleExpand,
+ onSelectModel,
+ onToggleFavorite,
+ onShowInfo,
+ }: ModelPickerGroupProps = $props();
+
+ // Format storage size
+ function formatSize(mb: number | undefined): string {
+ if (!mb) return "";
+ if (mb >= 1024) {
+ return `${(mb / 1024).toFixed(0)}GB`;
+ }
+ return `${mb}MB`;
+ }
+
+ // Check if any variant can fit
+ const anyVariantFits = $derived(
+ group.variants.some((v) => canModelFit(v.id)),
+ );
+
+ // Check if this group's model is currently selected (for single-variant groups)
+ const isMainSelected = $derived(
+ !group.hasMultipleVariants &&
+ group.variants.some((v) => v.id === selectedModelId),
+ );
+</script>
+
+<div
+ class="border-b border-white/5 last:border-b-0 {!anyVariantFits
+ ? 'opacity-50'
+ : ''}"
+>
+ <!-- Main row -->
+ <div
+ class="flex items-center gap-2 px-3 py-2.5 transition-colors {anyVariantFits
+ ? 'hover:bg-white/5 cursor-pointer'
+ : 'cursor-not-allowed'} {isMainSelected
+ ? 'bg-exo-yellow/10 border-l-2 border-exo-yellow'
+ : 'border-l-2 border-transparent'}"
+ onclick={() => {
+ if (group.hasMultipleVariants) {
+ onToggleExpand();
+ } else {
+ const modelId = group.variants[0]?.id;
+ if (modelId && canModelFit(modelId)) {
+ onSelectModel(modelId);
+ }
+ }
+ }}
+ role="button"
+ tabindex="0"
+ onkeydown={(e) => {
+ if (e.key === "Enter" || e.key === " ") {
+ e.preventDefault();
+ if (group.hasMultipleVariants) {
+ onToggleExpand();
+ } else {
+ const modelId = group.variants[0]?.id;
+ if (modelId && canModelFit(modelId)) {
+ onSelectModel(modelId);
+ }
+ }
+ }
+ }}
+ >
+ <!-- Expand/collapse chevron (for groups with variants) -->
+ {#if group.hasMultipleVariants}
+ <svg
+ class="w-4 h-4 text-white/40 transition-transform duration-200 flex-shrink-0 {isExpanded
+ ? 'rotate-90'
+ : ''}"
+ viewBox="0 0 24 24"
+ fill="currentColor"
+ >
+ <path d="M8.59 16.59L13.17 12 8.59 7.41 10 6l6 6-6 6-1.41-1.41z" />
+ </svg>
+ {:else}
+ <div class="w-4 flex-shrink-0"></div>
+ {/if}
+
+ <!-- Model name -->
+ <div class="flex-1 min-w-0">
+ <div class="flex items-center gap-2">
+ <span class="font-mono text-sm text-white truncate">
+ {group.name}
+ </span>
+ <!-- Capability icons -->
+ {#each group.capabilities.filter((c) => c !== "text") as cap}
+ {#if cap === "thinking"}
+ <svg
+ class="w-3.5 h-3.5 text-white/40 flex-shrink-0"
+ viewBox="0 0 24 24"
+ fill="none"
+ stroke="currentColor"
+ stroke-width="1.5"
+ title="Supports Thinking"
+ >
+ <path
+ d="M12 2a7 7 0 0 0-7 7c0 2.38 1.19 4.47 3 5.74V17a1 1 0 0 0 1 1h6a1 1 0 0 0 1-1v-2.26c1.81-1.27 3-3.36 3-5.74a7 7 0 0 0-7-7zM9 20h6M10 22h4"
+ stroke-linecap="round"
+ stroke-linejoin="round"
+ />
+ </svg>
+ {:else if cap === "code"}
+ <svg
+ class="w-3.5 h-3.5 text-white/40 flex-shrink-0"
+ viewBox="0 0 24 24"
+ fill="none"
+ stroke="currentColor"
+ stroke-width="1.5"
+ title="Supports code generation"
+ >
+ <path
+ d="M16 18l6-6-6-6M8 6l-6 6 6 6"
+ stroke-linecap="round"
+ stroke-linejoin="round"
+ />
+ </svg>
+ {:else if cap === "vision"}
+ <svg
+ class="w-3.5 h-3.5 text-white/40 flex-shrink-0"
+ viewBox="0 0 24 24"
+ fill="none"
+ stroke="currentColor"
+ stroke-width="1.5"
+ title="Supports image input"
+ >
+ <path
+ d="M1 12s4-8 11-8 11 8 11 8-4 8-11 8-11-8-11-8z"
+ stroke-linecap="round"
+ stroke-linejoin="round"
+ />
+ <circle cx="12" cy="12" r="3" />
+ </svg>
+ {:else if cap === "image_gen"}
+ <svg
+ class="w-3.5 h-3.5 text-white/40 flex-shrink-0"
+ viewBox="0 0 24 24"
+ fill="none"
+ stroke="currentColor"
+ stroke-width="1.5"
+ title="Supports image generation"
+ >
+ <rect x="3" y="3" width="18" height="18" rx="2" ry="2" />
+ <circle cx="8.5" cy="8.5" r="1.5" />
+ <path d="M21 15l-5-5L5 21" />
+ </svg>
+ {/if}
+ {/each}
+ </div>
+ </div>
+
+ <!-- Size indicator (smallest variant) -->
+ {#if !group.hasMultipleVariants && group.smallestVariant?.storage_size_megabytes}
+ <span class="text-xs font-mono text-white/30 flex-shrink-0">
+ {formatSize(group.smallestVariant.storage_size_megabytes)}
+ </span>
+ {/if}
+
+ <!-- Variant count -->
+ {#if group.hasMultipleVariants}
+ <span class="text-xs font-mono text-white/30 flex-shrink-0">
+ {group.variants.length} variants
+ </span>
+ {/if}
+
+ <!-- Check mark if selected (single-variant) -->
+ {#if isMainSelected}
+ <svg
+ class="w-4 h-4 text-exo-yellow flex-shrink-0"
+ viewBox="0 0 24 24"
+ fill="currentColor"
+ >
+ <path d="M9 16.17L4.83 12l-1.42 1.41L9 19 21 7l-1.41-1.41L9 16.17z" />
+ </svg>
+ {/if}
+
+ <!-- Favorite star -->
+ <button
+ type="button"
+ class="p-1 rounded hover:bg-white/10 transition-colors flex-shrink-0"
+ onclick={(e) => {
+ e.stopPropagation();
+ onToggleFavorite(group.id);
+ }}
+ title={isFavorite ? "Remove from favorites" : "Add to favorites"}
+ >
+ {#if isFavorite}
+ <svg
+ class="w-4 h-4 text-amber-400"
+ viewBox="0 0 24 24"
+ fill="currentColor"
+ >
+ <path
+ d="M12 2l3.09 6.26L22 9.27l-5 4.87 1.18 6.88L12 17.77l-6.18 3.25L7 14.14 2 9.27l6.91-1.01L12 2z"
+ />
+ </svg>
+ {:else}
+ <svg
+ class="w-4 h-4 text-white/30 hover:text-white/50"
+ viewBox="0 0 24 24"
+ fill="none"
+ stroke="currentColor"
+ stroke-width="2"
+ >
+ <path
+ d="M12 2l3.09 6.26L22 9.27l-5 4.87 1.18 6.88L12 17.77l-6.18 3.25L7 14.14 2 9.27l6.91-1.01L12 2z"
+ />
+ </svg>
+ {/if}
+ </button>
+
+ <!-- Info button -->
+ <button
+ type="button"
+ class="p-1 rounded hover:bg-white/10 transition-colors flex-shrink-0"
+ onclick={(e) => {
+ e.stopPropagation();
+ onShowInfo(group);
+ }}
+ title="View model details"
+ >
+ <svg
+ class="w-4 h-4 text-white/30 hover:text-white/50"
+ viewBox="0 0 24 24"
+ fill="currentColor"
+ >
+ <path
+ d="M12 2C6.48 2 2 6.48 2 12s4.48 10 10 10 10-4.48 10-10S17.52 2 12 2zm1 15h-2v-6h2v6zm0-8h-2V7h2v2z"
+ />
+ </svg>
+ </button>
+ </div>
+
+ <!-- Expanded variants -->
+ {#if isExpanded && group.hasMultipleVariants}
+ <div class="bg-black/20 border-t border-white/5">
+ {#each group.variants as variant}
+ {@const modelCanFit = canModelFit(variant.id)}
+ {@const isSelected = selectedModelId === variant.id}
+ <button
+ type="button"
+ class="w-full flex items-center gap-3 px-3 py-2 pl-10 hover:bg-white/5 transition-colors text-left {!modelCanFit
+ ? 'opacity-50 cursor-not-allowed'
+ : 'cursor-pointer'} {isSelected
+ ? 'bg-exo-yellow/10 border-l-2 border-exo-yellow'
+ : 'border-l-2 border-transparent'}"
+ disabled={!modelCanFit}
+ onclick={() => {
+ if (modelCanFit) {
+ onSelectModel(variant.id);
+ }
+ }}
+ >
+ <!-- Quantization badge -->
+ <span
+ class="text-xs font-mono px-1.5 py-0.5 rounded bg-white/10 text-white/70 flex-shrink-0"
+ >
+ {variant.quantization || "default"}
+ </span>
+
+ <!-- Size -->
+ <span class="text-xs font-mono text-white/40 flex-1">
+ {formatSize(variant.storage_size_megabytes)}
+ </span>
+
+ <!-- Check mark if selected -->
+ {#if isSelected}
+ <svg
+ class="w-4 h-4 text-exo-yellow"
+ viewBox="0 0 24 24"
+ fill="currentColor"
+ >
+ <path
+ d="M9 16.17L4.83 12l-1.42 1.41L9 19 21 7l-1.41-1.41L9 16.17z"
+ />
+ </svg>
+ {/if}
+ </button>
+ {/each}
+ </div>
+ {/if}
+</div>
diff --git a/dashboard/src/lib/components/ModelPickerModal.svelte b/dashboard/src/lib/components/ModelPickerModal.svelte
new file mode 100644
index 00000000..40827a0e
--- /dev/null
+++ b/dashboard/src/lib/components/ModelPickerModal.svelte
@@ -0,0 +1,748 @@
+<script lang="ts">
+ import { fade, fly } from "svelte/transition";
+ import { cubicOut } from "svelte/easing";
+ import FamilySidebar from "./FamilySidebar.svelte";
+ import ModelPickerGroup from "./ModelPickerGroup.svelte";
+ import ModelFilterPopover from "./ModelFilterPopover.svelte";
+ import HuggingFaceResultItem from "./HuggingFaceResultItem.svelte";
+
+ interface ModelInfo {
+ id: string;
+ name?: string;
+ storage_size_megabytes?: number;
+ base_model?: string;
+ quantization?: string;
+ supports_tensor?: boolean;
+ capabilities?: string[];
+ family?: string;
+ is_custom?: boolean;
+ tasks?: string[];
+ hugging_face_id?: string;
+ }
+
+ interface ModelGroup {
+ id: string;
+ name: string;
+ capabilities: string[];
+ family: string;
+ variants: ModelInfo[];
+ smallestVariant: ModelInfo;
+ hasMultipleVariants: boolean;
+ }
+
+ interface FilterState {
+ capabilities: string[];
+ sizeRange: { min: number; max: number } | null;
+ }
+
+ interface HuggingFaceModel {
+ id: string;
+ author: string;
+ downloads: number;
+ likes: number;
+ last_modified: string;
+ tags: string[];
+ }
+
+ type ModelPickerModalProps = {
+ isOpen: boolean;
+ models: ModelInfo[];
+ selectedModelId: string | null;
+ favorites: Set<string>;
+ existingModelIds: Set<string>;
+ canModelFit: (modelId: string) => boolean;
+ onSelect: (modelId: string) => void;
+ onClose: () => void;
+ onToggleFavorite: (baseModelId: string) => void;
+ onAddModel: (modelId: string) => Promise<void>;
+ onDeleteModel: (modelId: string) => Promise<void>;
+ totalMemoryGB: number;
+ usedMemoryGB: number;
+ };
+
+ let {
+ isOpen,
+ models,
+ selectedModelId,
+ favorites,
+ existingModelIds,
+ canModelFit,
+ onSelect,
+ onClose,
+ onToggleFavorite,
+ onAddModel,
+ onDeleteModel,
+ totalMemoryGB,
+ usedMemoryGB,
+ }: ModelPickerModalProps = $props();
+
+ // Local state
+ let searchQuery = $state("");
+ let selectedFamily = $state<string | null>(null);
+ let expandedGroups = $state<Set<string>>(new Set());
+ let showFilters = $state(false);
+ let filters = $state<FilterState>({ capabilities: [], sizeRange: null });
+ let infoGroup = $state<ModelGroup | null>(null);
+
+ // HuggingFace Hub state
+ let hfSearchQuery = $state("");
+ let hfSearchResults = $state<HuggingFaceModel[]>([]);
+ let hfTrendingModels = $state<HuggingFaceModel[]>([]);
+ let hfIsSearching = $state(false);
+ let hfIsLoadingTrending = $state(false);
+ let addingModelId = $state<string | null>(null);
+ let hfSearchDebounceTimer: ReturnType<typeof setTimeout> | null = null;
+ let manualModelId = $state("");
+ let addModelError = $state<string | null>(null);
+
+ // Reset state when modal opens
+ $effect(() => {
+ if (isOpen) {
+ searchQuery = "";
+ selectedFamily = null;
+ expandedGroups = new Set();
+ showFilters = false;
+ hfSearchQuery = "";
+ hfSearchResults = [];
+ manualModelId = "";
+ addModelError = null;
+ }
+ });
+
+ // Fetch trending models when HuggingFace is selected
+ $effect(() => {
+ if (
+ selectedFamily === "huggingface" &&
+ hfTrendingModels.length === 0 &&
+ !hfIsLoadingTrending
+ ) {
+ fetchTrendingModels();
+ }
+ });
+
+ async function fetchTrendingModels() {
+ hfIsLoadingTrending = true;
+ try {
+ const response = await fetch("/models/search?query=&limit=20");
+ if (response.ok) {
+ hfTrendingModels = await response.json();
+ }
+ } catch (error) {
+ console.error("Failed to fetch trending models:", error);
+ } finally {
+ hfIsLoadingTrending = false;
+ }
+ }
+
+ async function searchHuggingFace(query: string) {
+ if (query.length < 2) {
+ hfSearchResults = [];
+ return;
+ }
+
+ hfIsSearching = true;
+ try {
+ const response = await fetch(
+ `/models/search?query=${encodeURIComponent(query)}&limit=20`,
+ );
+ if (response.ok) {
+ hfSearchResults = await response.json();
+ } else {
+ hfSearchResults = [];
+ }
+ } catch (error) {
+ console.error("Failed to search models:", error);
+ hfSearchResults = [];
+ } finally {
+ hfIsSearching = false;
+ }
+ }
+
+ function handleHfSearchInput(query: string) {
+ hfSearchQuery = query;
+ addModelError = null;
+
+ if (hfSearchDebounceTimer) {
+ clearTimeout(hfSearchDebounceTimer);
+ }
+
+ if (query.length >= 2) {
+ hfSearchDebounceTimer = setTimeout(() => {
+ searchHuggingFace(query);
+ }, 300);
+ } else {
+ hfSearchResults = [];
+ }
+ }
+
+ async function handleAddModel(modelId: string) {
+ addingModelId = modelId;
+ addModelError = null;
+ try {
+ await onAddModel(modelId);
+ } catch (error) {
+ addModelError =
+ error instanceof Error ? error.message : "Failed to add model";
+ } finally {
+ addingModelId = null;
+ }
+ }
+
+ async function handleAddManualModel() {
+ if (!manualModelId.trim()) return;
+ await handleAddModel(manualModelId.trim());
+ if (!addModelError) {
+ manualModelId = "";
+ }
+ }
+
+ function handleSelectHfModel(modelId: string) {
+ onSelect(modelId);
+ onClose();
+ }
+
+ // Models to display in HuggingFace view
+ const hfDisplayModels = $derived.by((): HuggingFaceModel[] => {
+ if (hfSearchQuery.length >= 2) {
+ return hfSearchResults;
+ }
+ return hfTrendingModels;
+ });
+
+ // Group models by base_model
+ const groupedModels = $derived.by((): ModelGroup[] => {
+ const groups = new Map<string, ModelGroup>();
+
+ for (const model of models) {
+ const groupId = model.base_model || model.id;
+ const groupName = model.base_model || model.name || model.id;
+
+ if (!groups.has(groupId)) {
+ groups.set(groupId, {
+ id: groupId,
+ name: groupName,
+ capabilities: model.capabilities || ["text"],
+ family: model.family || "",
+ variants: [],
+ smallestVariant: model,
+ hasMultipleVariants: false,
+ });
+ }
+
+ const group = groups.get(groupId)!;
+ group.variants.push(model);
+
+ // Track smallest variant
+ if (
+ (model.storage_size_megabytes || 0) <
+ (group.smallestVariant.storage_size_megabytes || Infinity)
+ ) {
+ group.smallestVariant = model;
+ }
+
+ // Update capabilities if not set
+ if (
+ group.capabilities.length <= 1 &&
+ model.capabilities &&
+ model.capabilities.length > 1
+ ) {
+ group.capabilities = model.capabilities;
+ }
+ if (!group.family && model.family) {
+ group.family = model.family;
+ }
+ }
+
+ // Sort variants within each group by size
+ for (const group of groups.values()) {
+ group.variants.sort(
+ (a, b) =>
+ (a.storage_size_megabytes || 0) - (b.storage_size_megabytes || 0),
+ );
+ group.hasMultipleVariants = group.variants.length > 1;
+ }
+
+ // Convert to array and sort by smallest variant size (biggest first)
+ return Array.from(groups.values()).sort((a, b) => {
+ return (
+ (b.smallestVariant.storage_size_megabytes || 0) -
+ (a.smallestVariant.storage_size_megabytes || 0)
+ );
+ });
+ });
+
+ // Get unique families
+ const uniqueFamilies = $derived.by((): string[] => {
+ const families = new Set<string>();
+ for (const group of groupedModels) {
+ if (group.family) {
+ families.add(group.family);
+ }
+ }
+ const familyOrder = [
+ "kimi",
+ "qwen",
+ "glm",
+ "minimax",
+ "deepseek",
+ "gpt-oss",
+ "llama",
+ ];
+ return Array.from(families).sort((a, b) => {
+ const aIdx = familyOrder.indexOf(a);
+ const bIdx = familyOrder.indexOf(b);
+ if (aIdx === -1 && bIdx === -1) return a.localeCompare(b);
+ if (aIdx === -1) return 1;
+ if (bIdx === -1) return -1;
+ return aIdx - bIdx;
+ });
+ });
+
+ // Filter models based on search, family, and filters
+ const filteredGroups = $derived.by((): ModelGroup[] => {
+ let result: ModelGroup[] = [...groupedModels];
+
+ // Filter by family
+ if (selectedFamily === "favorites") {
+ result = result.filter((g) => favorites.has(g.id));
+ } else if (selectedFamily && selectedFamily !== "huggingface") {
+ result = result.filter((g) => g.family === selectedFamily);
+ }
+
+ // Filter by search query
+ if (searchQuery.trim()) {
+ const query = searchQuery.toLowerCase().trim();
+ result = result.filter(
+ (g) =>
+ g.name.toLowerCase().includes(query) ||
+ g.variants.some(
+ (v) =>
+ v.id.toLowerCase().includes(query) ||
+ (v.name || "").toLowerCase().includes(query),
+ ),
+ );
+ }
+
+ // Filter by capabilities
+ if (filters.capabilities.length > 0) {
+ result = result.filter((g) =>
+ filters.capabilities.every((cap) => g.capabilities.includes(cap)),
+ );
+ }
+
+ // Filter by size range
+ if (filters.sizeRange) {
+ const { min, max } = filters.sizeRange;
+ result = result.filter((g) => {
+ const sizeGB = (g.smallestVariant.storage_size_megabytes || 0) / 1024;
+ return sizeGB >= min && sizeGB <= max;
+ });
+ }
+
+ // Sort: models that fit first, then by size (largest first)
+ result.sort((a, b) => {
+ const aFits = a.variants.some((v) => canModelFit(v.id));
+ const bFits = b.variants.some((v) => canModelFit(v.id));
+
+ if (aFits && !bFits) return -1;
+ if (!aFits && bFits) return 1;
+
+ return (
+ (b.smallestVariant.storage_size_megabytes || 0) -
+ (a.smallestVariant.storage_size_megabytes || 0)
+ );
+ });
+
+ return result;
+ });
+
+ // Check if any favorites exist
+ const hasFavorites = $derived(favorites.size > 0);
+
+ function toggleGroupExpanded(groupId: string) {
+ const next = new Set(expandedGroups);
+ if (next.has(groupId)) {
+ next.delete(groupId);
+ } else {
+ next.add(groupId);
+ }
+ expandedGroups = next;
+ }
+
+ function handleSelect(modelId: string) {
+ onSelect(modelId);
+ onClose();
+ }
+
+ function handleKeydown(e: KeyboardEvent) {
+ if (e.key === "Escape") {
+ onClose();
+ }
+ }
+
+ function handleFiltersChange(newFilters: FilterState) {
+ filters = newFilters;
+ }
+
+ function clearFilters() {
+ filters = { capabilities: [], sizeRange: null };
+ }
+
+ const hasActiveFilters = $derived(
+ filters.capabilities.length > 0 || filters.sizeRange !== null,
+ );
+</script>
+
+<svelte:window onkeydown={handleKeydown} />
+
+{#if isOpen}
+ <!-- Backdrop -->
+ <div
+ class="fixed inset-0 z-50 bg-black/80 backdrop-blur-sm"
+ transition:fade={{ duration: 200 }}
+ onclick={onClose}
+ role="presentation"
+ ></div>
+
+ <!-- Modal -->
+ <div
+ class="fixed z-50 top-1/2 left-1/2 -translate-x-1/2 -translate-y-1/2 w-[min(90vw,600px)] h-[min(80vh,700px)] bg-exo-dark-gray border border-exo-yellow/10 rounded-lg shadow-2xl overflow-hidden flex flex-col"
+ transition:fly={{ y: 20, duration: 300, easing: cubicOut }}
+ role="dialog"
+ aria-modal="true"
+ aria-label="Select a model"
+ >
+ <!-- Header with search -->
+ <div
+ class="flex items-center gap-2 p-3 border-b border-exo-yellow/10 bg-exo-medium-gray/30"
+ >
+ {#if selectedFamily === "huggingface"}
+ <!-- HuggingFace search -->
+ <svg
+ class="w-5 h-5 text-orange-400/60 flex-shrink-0"
+ viewBox="0 0 24 24"
+ fill="none"
+ stroke="currentColor"
+ stroke-width="2"
+ >
+ <circle cx="11" cy="11" r="8" />
+ <path d="M21 21l-4.35-4.35" />
+ </svg>
+ <input
+ type="search"
+ class="flex-1 bg-transparent border-none outline-none text-sm font-mono text-white placeholder-white/40"
+ placeholder="Search mlx-community models..."
+ value={hfSearchQuery}
+ oninput={(e) => handleHfSearchInput(e.currentTarget.value)}
+ />
+ {#if hfIsSearching}
+ <div class="flex-shrink-0">
+ <span
+ class="w-4 h-4 border-2 border-orange-400 border-t-transparent rounded-full animate-spin block"
+ ></span>
+ </div>
+ {/if}
+ {:else}
+ <!-- Normal model search -->
+ <svg
+ class="w-5 h-5 text-white/40 flex-shrink-0"
+ viewBox="0 0 24 24"
+ fill="none"
+ stroke="currentColor"
+ stroke-width="2"
+ >
+ <circle cx="11" cy="11" r="8" />
+ <path d="M21 21l-4.35-4.35" />
+ </svg>
+ <input
+ type="search"
+ class="flex-1 bg-transparent border-none outline-none text-sm font-mono text-white placeholder-white/40"
+ placeholder="Search models..."
+ bind:value={searchQuery}
+ />
+ <!-- Cluster memory -->
+ <span
+ class="text-xs font-mono flex-shrink-0"
+ title="Cluster memory usage"
+ ><span class="text-exo-yellow">{Math.round(usedMemoryGB)}GB</span
+ ><span class="text-white/40">/{Math.round(totalMemoryGB)}GB</span
+ ></span
+ >
+ <!-- Filter button -->
+ <div class="relative filter-toggle">
+ <button
+ type="button"
+ class="p-1.5 rounded hover:bg-white/10 transition-colors {hasActiveFilters
+ ? 'text-exo-yellow'
+ : 'text-white/50'}"
+ onclick={() => (showFilters = !showFilters)}
+ title="Filter by capability or size"
+ >
+ <svg class="w-5 h-5" viewBox="0 0 24 24" fill="currentColor">
+ <path d="M10 18h4v-2h-4v2zM3 6v2h18V6H3zm3 7h12v-2H6v2z" />
+ </svg>
+ </button>
+ {#if showFilters}
+ <ModelFilterPopover
+ {filters}
+ onChange={handleFiltersChange}
+ onClear={clearFilters}
+ onClose={() => (showFilters = false)}
+ />
+ {/if}
+ </div>
+ {/if}
+ <!-- Close button -->
+ <button
+ type="button"
+ class="p-1.5 rounded hover:bg-white/10 transition-colors text-white/50 hover:text-white/70"
+ onclick={onClose}
+ title="Close model picker"
+ >
+ <svg class="w-5 h-5" viewBox="0 0 24 24" fill="currentColor">
+ <path
+ d="M19 6.41L17.59 5 12 10.59 6.41 5 5 6.41 10.59 12 5 17.59 6.41 19 12 13.41 17.59 19 19 17.59 13.41 12 19 6.41z"
+ />
+ </svg>
+ </button>
+ </div>
+
+ <!-- Body -->
+ <div class="flex flex-1 overflow-hidden">
+ <!-- Family sidebar -->
+ <FamilySidebar
+ families={uniqueFamilies}
+ {selectedFamily}
+ {hasFavorites}
+ onSelect={(family) => (selectedFamily = family)}
+ />
+
+ <!-- Model list -->
+ <div class="flex-1 overflow-y-auto flex flex-col">
+ {#if selectedFamily === "huggingface"}
+ <!-- HuggingFace Hub view -->
+ <div class="flex-1 flex flex-col min-h-0">
+ <!-- Section header -->
+ <div
+ class="sticky top-0 z-10 px-3 py-2 bg-exo-dark-gray/95 border-b border-exo-yellow/10"
+ >
+ <span class="text-xs font-mono text-white/40">
+ {#if hfSearchQuery.length >= 2}
+ Search results for "{hfSearchQuery}"
+ {:else}
+ Trending on mlx-community
+ {/if}
+ </span>
+ </div>
+
+ <!-- Results list -->
+ <div class="flex-1 overflow-y-auto">
+ {#if hfIsLoadingTrending && hfTrendingModels.length === 0}
+ <div
+ class="flex items-center justify-center py-12 text-white/40"
+ >
+ <span
+ class="w-5 h-5 border-2 border-orange-400 border-t-transparent rounded-full animate-spin mr-2"
+ ></span>
+ <span class="font-mono text-sm"
+ >Loading trending models...</span
+ >
+ </div>
+ {:else if hfDisplayModels.length === 0}
+ <div
+ class="flex flex-col items-center justify-center py-12 text-white/40"
+ >
+ <svg
+ class="w-10 h-10 mb-2"
+ viewBox="0 0 24 24"
+ fill="currentColor"
+ >
+ <path
+ d="M12 2C6.48 2 2 6.48 2 12s4.48 10 10 10 10-4.48 10-10S17.52 2 12 2zm-2 13.5c-.83 0-1.5-.67-1.5-1.5s.67-1.5 1.5-1.5 1.5.67 1.5 1.5-.67 1.5-1.5 1.5zm4 0c-.83 0-1.5-.67-1.5-1.5s.67-1.5 1.5-1.5 1.5.67 1.5 1.5-.67 1.5-1.5 1.5zm2-4.5H8c0-2.21 1.79-4 4-4s4 1.79 4 4z"
+ />
+ </svg>
+ <p class="font-mono text-sm">No models found</p>
+ {#if hfSearchQuery}
+ <p class="font-mono text-xs mt-1">
+ Try a different search term
+ </p>
+ {/if}
+ </div>
+ {:else}
+ {#each hfDisplayModels as model}
+ <HuggingFaceResultItem
+ {model}
+ isAdded={existingModelIds.has(model.id)}
+ isAdding={addingModelId === model.id}
+ onAdd={() => handleAddModel(model.id)}
+ onSelect={() => handleSelectHfModel(model.id)}
+ />
+ {/each}
+ {/if}
+ </div>
+
+ <!-- Manual input footer -->
+ <div
+ class="sticky bottom-0 border-t border-exo-yellow/10 bg-exo-dark-gray p-3"
+ >
+ {#if addModelError}
+ <div
+ class="bg-red-500/10 border border-red-500/30 rounded px-3 py-2 mb-2"
+ >
+ <p class="text-red-400 text-xs font-mono break-words">
+ {addModelError}
+ </p>
+ </div>
+ {/if}
+ <div class="flex gap-2">
+ <input
+ type="text"
+ class="flex-1 bg-exo-black/60 border border-exo-yellow/30 rounded px-3 py-1.5 text-xs font-mono text-white placeholder-white/30 focus:outline-none focus:border-exo-yellow/50"
+ placeholder="Or paste model ID directly..."
+ bind:value={manualModelId}
+ onkeydown={(e) => {
+ if (e.key === "Enter") handleAddManualModel();
+ }}
+ />
+ <button
+ type="button"
+ onclick={handleAddManualModel}
+ disabled={!manualModelId.trim() || addingModelId !== null}
+ class="px-3 py-1.5 text-xs font-mono tracking-wider uppercase bg-orange-500/10 text-orange-400 border border-orange-400/30 hover:bg-orange-500/20 transition-colors rounded disabled:opacity-50 disabled:cursor-not-allowed"
+ >
+ Add
+ </button>
+ </div>
+ </div>
+ </div>
+ {:else if filteredGroups.length === 0}
+ <div
+ class="flex flex-col items-center justify-center h-full text-white/40 p-8"
+ >
+ <svg class="w-12 h-12 mb-3" viewBox="0 0 24 24" fill="currentColor">
+ <path
+ d="M12 2C6.48 2 2 6.48 2 12s4.48 10 10 10 10-4.48 10-10S17.52 2 12 2zm-2 15l-5-5 1.41-1.41L10 14.17l7.59-7.59L19 8l-9 9z"
+ />
+ </svg>
+ <p class="font-mono text-sm">No models found</p>
+ {#if hasActiveFilters || searchQuery}
+ <button
+ type="button"
+ class="mt-2 text-xs text-exo-yellow hover:underline"
+ onclick={() => {
+ searchQuery = "";
+ clearFilters();
+ }}
+ >
+ Clear filters
+ </button>
+ {/if}
+ </div>
+ {:else}
+ {#each filteredGroups as group}
+ <ModelPickerGroup
+ {group}
+ isExpanded={expandedGroups.has(group.id)}
+ isFavorite={favorites.has(group.id)}
+ {selectedModelId}
+ {canModelFit}
+ onToggleExpand={() => toggleGroupExpanded(group.id)}
+ onSelectModel={handleSelect}
+ {onToggleFavorite}
+ onShowInfo={(g) => (infoGroup = g)}
+ />
+ {/each}
+ {/if}
+ </div>
+ </div>
+
+ <!-- Footer with active filters indicator -->
+ {#if hasActiveFilters}
+ <div
+ class="flex items-center gap-2 px-3 py-2 border-t border-exo-yellow/10 bg-exo-medium-gray/20 text-xs font-mono text-white/50"
+ >
+ <span>Filters:</span>
+ {#each filters.capabilities as cap}
+ <span class="px-1.5 py-0.5 bg-exo-yellow/20 text-exo-yellow rounded"
+ >{cap}</span
+ >
+ {/each}
+ {#if filters.sizeRange}
+ <span class="px-1.5 py-0.5 bg-exo-yellow/20 text-exo-yellow rounded">
+ {filters.sizeRange.min}GB - {filters.sizeRange.max}GB
+ </span>
+ {/if}
+ <button
+ type="button"
+ class="ml-auto text-white/40 hover:text-white/60"
+ onclick={clearFilters}
+ >
+ Clear all
+ </button>
+ </div>
+ {/if}
+ </div>
+
+ <!-- Info modal -->
+ {#if infoGroup}
+ <div
+ class="fixed inset-0 z-[60] bg-black/60"
+ transition:fade={{ duration: 150 }}
+ onclick={() => (infoGroup = null)}
+ role="presentation"
+ ></div>
+ <div
+ class="fixed z-[60] top-1/2 left-1/2 -translate-x-1/2 -translate-y-1/2 w-[min(80vw,400px)] bg-exo-dark-gray border border-exo-yellow/10 rounded-lg shadow-2xl p-4"
+ transition:fly={{ y: 10, duration: 200, easing: cubicOut }}
+ role="dialog"
+ aria-modal="true"
+ >
+ <div class="flex items-start justify-between mb-3">
+ <h3 class="font-mono text-lg text-white">{infoGroup.name}</h3>
+ <button
+ type="button"
+ class="p-1 rounded hover:bg-white/10 transition-colors text-white/50"
+ onclick={() => (infoGroup = null)}
+ title="Close model details"
+ aria-label="Close info dialog"
+ >
+ <svg class="w-4 h-4" viewBox="0 0 24 24" fill="currentColor">
+ <path
+ d="M19 6.41L17.59 5 12 10.59 6.41 5 5 6.41 10.59 12 5 17.59 6.41 19 12 13.41 17.59 19 19 17.59 13.41 12 19 6.41z"
+ />
+ </svg>
+ </button>
+ </div>
+ <div class="space-y-2 text-xs font-mono">
+ <div class="flex items-center gap-2">
+ <span class="text-white/40">Family:</span>
+ <span class="text-white/70">{infoGroup.family || "Unknown"}</span>
+ </div>
+ <div class="flex items-center gap-2">
+ <span class="text-white/40">Capabilities:</span>
+ <span class="text-white/70">{infoGroup.capabilities.join(", ")}</span>
+ </div>
+ <div class="flex items-center gap-2">
+ <span class="text-white/40">Variants:</span>
+ <span class="text-white/70">{infoGroup.variants.length}</span>
+ </div>
+ {#if infoGroup.variants.length > 0}
+ <div class="mt-3 pt-3 border-t border-exo-yellow/10">
+ <span class="text-white/40">Available quantizations:</span>
+ <div class="flex flex-wrap gap-1 mt-1">
+ {#each infoGroup.variants as variant}
+ <span
+ class="px-1.5 py-0.5 bg-white/10 text-white/60 rounded text-[10px]"
+ >
+ {variant.quantization || "default"} ({Math.round(
+ (variant.storage_size_megabytes || 0) / 1024,
+ )}GB)
+ </span>
+ {/each}
+ </div>
+ </div>
+ {/if}
+ </div>
+ </div>
+ {/if}
+{/if}
diff --git a/dashboard/src/lib/components/index.ts b/dashboard/src/lib/components/index.ts
index dc8a7d76..b0424bf4 100644
--- a/dashboard/src/lib/components/index.ts
+++ b/dashboard/src/lib/components/index.ts
@@ -6,3 +6,9 @@ export { default as ChatSidebar } from "./ChatSidebar.svelte";
export { default as ModelCard } from "./ModelCard.svelte";
export { default as MarkdownContent } from "./MarkdownContent.svelte";
export { default as ImageParamsPanel } from "./ImageParamsPanel.svelte";
+export { default as FamilyLogos } from "./FamilyLogos.svelte";
+export { default as FamilySidebar } from "./FamilySidebar.svelte";
+export { default as HuggingFaceResultItem } from "./HuggingFaceResultItem.svelte";
+export { default as ModelFilterPopover } from "./ModelFilterPopover.svelte";
+export { default as ModelPickerGroup } from "./ModelPickerGroup.svelte";
+export { default as ModelPickerModal } from "./ModelPickerModal.svelte";
diff --git a/dashboard/src/lib/stores/favorites.svelte.ts b/dashboard/src/lib/stores/favorites.svelte.ts
new file mode 100644
index 00000000..877b059c
--- /dev/null
+++ b/dashboard/src/lib/stores/favorites.svelte.ts
@@ -0,0 +1,97 @@
+/**
+ * FavoritesStore - Manages favorite models with localStorage persistence
+ */
+
+import { browser } from "$app/environment";
+
+const FAVORITES_KEY = "exo-favorite-models";
+
+class FavoritesStore {
+ favorites = $state<Set<string>>(new Set());
+
+ constructor() {
+ if (browser) {
+ this.loadFromStorage();
+ }
+ }
+
+ private loadFromStorage() {
+ try {
+ const stored = localStorage.getItem(FAVORITES_KEY);
+ if (stored) {
+ const parsed = JSON.parse(stored) as string[];
+ this.favorites = new Set(parsed);
+ }
+ } catch (error) {
+ console.error("Failed to load favorites:", error);
+ }
+ }
+
+ private saveToStorage() {
+ try {
+ const array = Array.from(this.favorites);
+ localStorage.setItem(FAVORITES_KEY, JSON.stringify(array));
+ } catch (error) {
+ console.error("Failed to save favorites:", error);
+ }
+ }
+
+ add(baseModelId: string) {
+ const next = new Set(this.favorites);
+ next.add(baseModelId);
+ this.favorites = next;
+ this.saveToStorage();
+ }
+
+ remove(baseModelId: string) {
+ const next = new Set(this.favorites);
+ next.delete(baseModelId);
+ this.favorites = next;
+ this.saveToStorage();
+ }
+
+ toggle(baseModelId: string) {
+ if (this.favorites.has(baseModelId)) {
+ this.remove(baseModelId);
+ } else {
+ this.add(baseModelId);
+ }
+ }
+
+ isFavorite(baseModelId: string): boolean {
+ return this.favorites.has(baseModelId);
+ }
+
+ getAll(): string[] {
+ return Array.from(this.favorites);
+ }
+
+ getSet(): Set<string> {
+ return new Set(this.favorites);
+ }
+
+ hasAny(): boolean {
+ return this.favorites.size > 0;
+ }
+
+ clearAll() {
+ this.favorites = new Set();
+ this.saveToStorage();
+ }
+}
+
+export const favoritesStore = new FavoritesStore();
+
+export const favorites = () => favoritesStore.favorites;
+export const hasFavorites = () => favoritesStore.hasAny();
+export const isFavorite = (baseModelId: string) =>
+ favoritesStore.isFavorite(baseModelId);
+export const toggleFavorite = (baseModelId: string) =>
+ favoritesStore.toggle(baseModelId);
+export const addFavorite = (baseModelId: string) =>
+ favoritesStore.add(baseModelId);
+export const removeFavorite = (baseModelId: string) =>
+ favoritesStore.remove(baseModelId);
+export const getFavorites = () => favoritesStore.getAll();
+export const getFavoritesSet = () => favoritesStore.getSet();
+export const clearFavorites = () => favoritesStore.clearAll();
diff --git a/dashboard/src/routes/+page.svelte b/dashboard/src/routes/+page.svelte
index 6ecce7cc..288c3991 100644
--- a/dashboard/src/routes/+page.svelte
+++ b/dashboard/src/routes/+page.svelte
@@ -5,7 +5,13 @@
ChatMessages,
ChatSidebar,
ModelCard,
+ ModelPickerModal,
} from "$lib/components";
+ import {
+ favorites,
+ toggleFavorite,
+ getFavoritesSet,
+ } from "$lib/stores/favorites.svelte";
import {
hasStartedChat,
isTopologyMinimized,
@@ -101,6 +107,10 @@
tasks?: string[];
hugging_face_id?: string;
is_custom?: boolean;
+ family?: string;
+ quantization?: string;
+ base_model?: string;
+ capabilities?: string[];
}>
>([]);
@@ -212,14 +222,11 @@
let launchingModelId = $state<string | null>(null);
let instanceDownloadExpandedNodes = $state<Set<string>>(new Set());
- // Custom dropdown state
- let isModelDropdownOpen = $state(false);
- let modelDropdownSearch = $state("");
+ // Model picker modal state
+ let isModelPickerOpen = $state(false);
- // Custom model add state
- let customModelInput = $state("");
- let isAddingCustomModel = $state(false);
- let customModelError = $state("");
+ // Favorites state (reactive)
+ const favoritesSet = $derived(getFavoritesSet());
// Slider dragging state
let isDraggingSlider = $state(false);
@@ -536,41 +543,25 @@
}
}
- async function addCustomModel() {
- const modelId = customModelInput.trim();
- if (!modelId) return;
-
- isAddingCustomModel = true;
- customModelError = "";
-
- try {
- const response = await fetch("/models/add", {
- method: "POST",
- headers: { "Content-Type": "application/json" },
- body: JSON.stringify({ model_id: modelId }),
- });
+ async function addModelFromPicker(modelId: string) {
+ const response = await fetch("/models/add", {
+ method: "POST",
+ headers: { "Content-Type": "application/json" },
+ body: JSON.stringify({ model_id: modelId }),
+ });
- if (!response.ok) {
- try {
- const err = await response.json();
- customModelError =
- err.detail || `Failed to add model (${response.status})`;
- } catch {
- customModelError = `Failed to add model (${response.status}: ${response.statusText})`;
- }
- return;
+ if (!response.ok) {
+ let message = `Failed to add model (${response.status}: ${response.statusText})`;
+ try {
+ const err = await response.json();
+ if (err.detail) message = err.detail;
+ } catch {
+ // use default message
}
-
- const added = await response.json();
- customModelInput = "";
- await fetchModels();
- selectPreviewModel(added.id);
- isModelDropdownOpen = false;
- } catch {
- customModelError = "Network error";
- } finally {
- isAddingCustomModel = false;
+ throw new Error(message);
}
+
+ await fetchModels();
}
async function deleteCustomModel(modelId: string) {
@@ -587,6 +578,12 @@
}
}
+ function handleModelPickerSelect(modelId: string) {
+ selectPreviewModel(modelId);
+ saveLaunchDefaults();
+ isModelPickerOpen = false;
+ }
+
async function launchInstance(
modelId: string,
specificPreview?: PlacementPreview | null,
@@ -2417,14 +2414,12 @@
>
</div>
- <!-- Model Dropdown (Custom) -->
- <div class="flex-shrink-0 mb-3 relative">
+ <!-- Model Picker Button -->
+ <div class="flex-shrink-0 mb-3">
<button
type="button"
- onclick={() => (isModelDropdownOpen = !isModelDropdownOpen)}
- class="w-full bg-exo-medium-gray/50 border border-exo-yellow/30 rounded pl-3 pr-8 py-2.5 text-sm font-mono text-left tracking-wide cursor-pointer transition-all duration-200 hover:border-exo-yellow/50 focus:outline-none focus:border-exo-yellow/70 {isModelDropdownOpen
- ? 'border-exo-yellow/70'
- : ''}"
+ onclick={() => (isModelPickerOpen = true)}
+ class="w-full bg-exo-medium-gray/50 border border-exo-yellow/30 rounded pl-3 pr-8 py-2.5 text-sm font-mono text-left tracking-wide cursor-pointer transition-all duration-200 hover:border-exo-yellow/50 focus:outline-none focus:border-exo-yellow/70 relative"
>
{#if selectedModelId}
{@const foundModel = models.find(
@@ -2432,54 +2427,12 @@
)}
{#if foundModel}
{@const sizeGB = getModelSizeGB(foundModel)}
- {@const isImageModel = modelSupportsImageGeneration(
- foundModel.id,
- )}
- {@const isImageEditModel = modelSupportsImageEditing(
- foundModel.id,
- )}
<span
class="flex items-center justify-between gap-2 w-full pr-4"
>
<span
class="flex items-center gap-2 text-exo-light-gray truncate"
>
- {#if isImageModel}
- <svg
- class="w-4 h-4 flex-shrink-0 text-exo-yellow"
- fill="none"
- viewBox="0 0 24 24"
- stroke="currentColor"
- stroke-width="2"
- >
- <rect
- x="3"
- y="3"
- width="18"
- height="18"
- rx="2"
- ry="2"
- />
- <circle cx="8.5" cy="8.5" r="1.5" />
- <polyline points="21 15 16 10 5 21" />
- </svg>
- {/if}
- {#if isImageEditModel}
- <svg
- class="w-4 h-4 flex-shrink-0 text-exo-yellow"
- fill="none"
- viewBox="0 0 24 24"
- stroke="currentColor"
- stroke-width="2"
- >
- <path
- d="M11 4H4a2 2 0 0 0-2 2v14a2 2 0 0 0 2 2h14a2 2 0 0 0 2-2v-7"
- />
- <path
- d="M18.5 2.5a2.121 2.121 0 0 1 3 3L12 15l-4 1 1-4 9.5-9.5z"
- />
- </svg>
- {/if}
<span class="truncate"
>{foundModel.name || foundModel.id}</span
>
@@ -2496,193 +2449,24 @@
{:else}
<span class="text-white/50">— SELECT MODEL —</span>
{/if}
- </button>
- <div
- class="absolute right-3 top-1/2 -translate-y-1/2 pointer-events-none transition-transform duration-200 {isModelDropdownOpen
- ? 'rotate-180'
- : ''}"
- >
- <svg
- class="w-4 h-4 text-exo-yellow/60"
- fill="none"
- viewBox="0 0 24 24"
- stroke="currentColor"
- >
- <path
- stroke-linecap="round"
- stroke-linejoin="round"
- stroke-width="2"
- d="M19 9l-7 7-7-7"
- />
- </svg>
- </div>
-
- {#if isModelDropdownOpen}
- <!-- Backdrop to close dropdown -->
- <button
- type="button"
- class="fixed inset-0 z-40 cursor-default"
- onclick={() => (isModelDropdownOpen = false)}
- aria-label="Close dropdown"
- ></button>
-
- <!-- Dropdown Panel -->
<div
- class="absolute top-full left-0 right-0 mt-1 bg-exo-dark-gray border border-exo-yellow/30 rounded shadow-lg shadow-black/50 z-50 max-h-64 overflow-y-auto"
+ class="absolute right-3 top-1/2 -translate-y-1/2 pointer-events-none"
>
- <!-- Search within dropdown -->
- <div
- class="sticky top-0 bg-exo-dark-gray border-b border-exo-medium-gray/30 p-2 space-y-2"
+ <svg
+ class="w-4 h-4 text-exo-yellow/60"
+ fill="none"
+ viewBox="0 0 24 24"
+ stroke="currentColor"
>
- <input
- type="text"
- placeholder="Search models..."
- bind:value={modelDropdownSearch}
- class="w-full bg-exo-dark-gray/60 border border-exo-medium-gray/30 rounded px-2 py-1.5 text-xs font-mono text-white/80 placeholder:text-white/40 focus:outline-none focus:border-exo-yellow/50"
+ <path
+ stroke-linecap="round"
+ stroke-linejoin="round"
+ stroke-width="2"
+ d="M19 9l-7 7-7-7"
/>
- <!-- Add custom model -->
- <form
- onsubmit={(e) => {
- e.preventDefault();
- addCustomModel();
- }}
- class="flex gap-1.5"
- >
- <input
- type="text"
- placeholder="org/model-name"
- bind:value={customModelInput}
- class="flex-1 bg-exo-dark-gray/60 border border-exo-medium-gray/30 rounded px-2 py-1.5 text-xs font-mono text-white/80 placeholder:text-white/40 focus:outline-none focus:border-exo-yellow/50"
- />
- <button
- type="submit"
- disabled={!customModelInput.trim() ||
- isAddingCustomModel}
- class="px-2 py-1.5 text-xs font-mono bg-exo-yellow/20 text-exo-yellow border border-exo-yellow/30 rounded hover:bg-exo-yellow/30 disabled:opacity-40 disabled:cursor-not-allowed"
- >
- {isAddingCustomModel ? "..." : "+ ADD"}
- </button>
- </form>
- {#if customModelError}
- <div class="text-xs text-red-400 font-mono">
- {customModelError}
- </div>
- {/if}
- </div>
-
- <!-- Options -->
- <div class="py-1">
- {#each sortedModels().filter((m) => !modelDropdownSearch || (m.name || m.id)
- .toLowerCase()
- .includes(modelDropdownSearch.toLowerCase())) as model}
- {@const sizeGB = getModelSizeGB(model)}
- {@const modelCanFit = hasEnoughMemory(model)}
- {@const isImageModel = modelSupportsImageGeneration(
- model.id,
- )}
- {@const isImageEditModel = modelSupportsImageEditing(
- model.id,
- )}
- <button
- type="button"
- onclick={() => {
- if (modelCanFit) {
- selectPreviewModel(model.id);
- saveLaunchDefaults();
- isModelDropdownOpen = false;
- modelDropdownSearch = "";
- }
- }}
- disabled={!modelCanFit}
- class="w-full px-3 py-2 text-left text-sm font-mono tracking-wide transition-colors duration-100 flex items-center justify-between gap-2 {selectedModelId ===
- model.id
- ? 'bg-transparent text-exo-yellow cursor-pointer'
- : modelCanFit
- ? 'text-white/80 hover:text-exo-yellow cursor-pointer'
- : 'text-white/30 cursor-default'}"
- >
- <span class="flex items-center gap-2 truncate flex-1">
- {#if isImageModel}
- <svg
- class="w-4 h-4 flex-shrink-0 text-exo-yellow"
- fill="none"
- viewBox="0 0 24 24"
- stroke="currentColor"
- stroke-width="2"
- aria-label="Image generation model"
- >
- <rect
- x="3"
- y="3"
- width="18"
- height="18"
- rx="2"
- ry="2"
- />
- <circle cx="8.5" cy="8.5" r="1.5" />
- <polyline points="21 15 16 10 5 21" />
- </svg>
- {/if}
- {#if isImageEditModel}
- <svg
- class="w-4 h-4 flex-shrink-0 text-exo-yellow"
- fill="none"
- viewBox="0 0 24 24"
- stroke="currentColor"
- stroke-width="2"
- aria-label="Image editing model"
- >
- <path
- d="M11 4H4a2 2 0 0 0-2 2v14a2 2 0 0 0 2 2h14a2 2 0 0 0 2-2v-7"
- />
- <path
- d="M18.5 2.5a2.121 2.121 0 0 1 3 3L12 15l-4 1 1-4 9.5-9.5z"
- />
- </svg>
- {/if}
- <span class="truncate">{model.name || model.id}</span>
- </span>
- <span class="flex items-center gap-1.5 flex-shrink-0">
- <span
- class="text-xs {modelCanFit
- ? 'text-white/50'
- : 'text-red-400/60'}"
- >
- {sizeGB >= 1
- ? sizeGB.toFixed(0)
- : sizeGB.toFixed(1)}GB
- </span>
- {#if model.is_custom}
- <button
- type="button"
- onclick={(e) => {
- e.stopPropagation();
- deleteCustomModel(model.id);
- }}
- class="text-white/30 hover:text-red-400 transition-colors p-0.5"
- title="Remove custom model"
- >
- <svg
- class="w-3 h-3"
- fill="none"
- viewBox="0 0 24 24"
- stroke="currentColor"
- stroke-width="2"
- >
- <path d="M18 6L6 18M6 6l12 12" />
- </svg>
- </button>
- {/if}
- </span>
- </button>
- {:else}
- <div class="px-3 py-2 text-xs text-white/50 font-mono">
- No models found
- </div>
- {/each}
- </div>
+ </svg>
</div>
- {/if}
+ </button>
</div>
<!-- Configuration Options -->
@@ -3462,3 +3246,22 @@
{/if}
</main>
</div>
+
+<ModelPickerModal
+ isOpen={isModelPickerOpen}
+ {models}
+ {selectedModelId}
+ favorites={favoritesSet}
+ existingModelIds={new Set(models.map((m) => m.id))}
+ canModelFit={(modelId) => {
+ const model = models.find((m) => m.id === modelId);
+ return model ? hasEnoughMemory(model) : false;
+ }}
+ onSelect={handleModelPickerSelect}
+ onClose={() => (isModelPickerOpen = false)}
+ onToggleFavorite={toggleFavorite}
+ onAddModel={addModelFromPicker}
+ onDeleteModel={deleteCustomModel}
+ totalMemoryGB={clusterMemory().total / (1024 * 1024 * 1024)}
+ usedMemoryGB={clusterMemory().used / (1024 * 1024 * 1024)}
+/>
diff --git a/resources/inference_model_cards/mlx-community--DeepSeek-V3.1-4bit.toml b/resources/inference_model_cards/mlx-community--DeepSeek-V3.1-4bit.toml
index 26de8de8..41784cf6 100644
--- a/resources/inference_model_cards/mlx-community--DeepSeek-V3.1-4bit.toml
+++ b/resources/inference_model_cards/mlx-community--DeepSeek-V3.1-4bit.toml
@@ -3,6 +3,10 @@ n_layers = 61
hidden_size = 7168
supports_tensor = true
tasks = ["TextGeneration"]
+family = "deepseek"
+quantization = "4bit"
+base_model = "DeepSeek V3.1"
+capabilities = ["text", "thinking"]
[storage_size]
in_bytes = 405874409472
diff --git a/resources/inference_model_cards/mlx-community--DeepSeek-V3.1-8bit.toml b/resources/inference_model_cards/mlx-community--DeepSeek-V3.1-8bit.toml
index 13cf367b..a5d77bcd 100644
--- a/resources/inference_model_cards/mlx-community--DeepSeek-V3.1-8bit.toml
+++ b/resources/inference_model_cards/mlx-community--DeepSeek-V3.1-8bit.toml
@@ -3,6 +3,10 @@ n_layers = 61
hidden_size = 7168
supports_tensor = true
tasks = ["TextGeneration"]
+family = "deepseek"
+quantization = "8bit"
+base_model = "DeepSeek V3.1"
+capabilities = ["text", "thinking"]
[storage_size]
in_bytes = 765577920512
diff --git a/resources/inference_model_cards/mlx-community--GLM-4.5-Air-8bit.toml b/resources/inference_model_cards/mlx-community--GLM-4.5-Air-8bit.toml
index 288392f6..a7acea44 100644
--- a/resources/inference_model_cards/mlx-community--GLM-4.5-Air-8bit.toml
+++ b/resources/inference_model_cards/mlx-community--GLM-4.5-Air-8bit.toml
@@ -3,6 +3,10 @@ n_layers = 46
hidden_size = 4096
supports_tensor = false
tasks = ["TextGeneration"]
+family = "glm"
+quantization = "8bit"
+base_model = "GLM 4.5 Air"
+capabilities = ["text", "thinking"]
[storage_size]
in_bytes = 122406567936
diff --git a/resources/inference_model_cards/mlx-community--GLM-4.5-Air-bf16.toml b/resources/inference_model_cards/mlx-community--GLM-4.5-Air-bf16.toml
index 00a19df2..4258c225 100644
--- a/resources/inference_model_cards/mlx-community--GLM-4.5-Air-bf16.toml
+++ b/resources/inference_model_cards/mlx-community--GLM-4.5-Air-bf16.toml
@@ -3,6 +3,10 @@ n_layers = 46
hidden_size = 4096
supports_tensor = true
tasks = ["TextGeneration"]
+family = "glm"
+quantization = "bf16"
+base_model = "GLM 4.5 Air"
+capabilities = ["text", "thinking"]
[storage_size]
in_bytes = 229780750336
diff --git a/resources/inference_model_cards/mlx-community--GLM-4.7-4bit.toml b/resources/inference_model_cards/mlx-community--GLM-4.7-4bit.toml
index 816c9657..0672d664 100644
--- a/resources/inference_model_cards/mlx-community--GLM-4.7-4bit.toml
+++ b/resources/inference_model_cards/mlx-community--GLM-4.7-4bit.toml
@@ -3,6 +3,10 @@ n_layers = 91
hidden_size = 5120
supports_tensor = true
tasks = ["TextGeneration"]
+family = "glm"
+quantization = "4bit"
+base_model = "GLM 4.7"
+capabilities = ["text", "thinking"]
[storage_size]
in_bytes = 198556925568
diff --git a/resources/inference_model_cards/mlx-community--GLM-4.7-6bit.toml b/resources/inference_model_cards/mlx-community--GLM-4.7-6bit.toml
index b087164b..bcf1cae4 100644
--- a/resources/inference_model_cards/mlx-community--GLM-4.7-6bit.toml
+++ b/resources/inference_model_cards/mlx-community--GLM-4.7-6bit.toml
@@ -3,6 +3,10 @@ n_layers = 91
hidden_size = 5120
supports_tensor = true
tasks = ["TextGeneration"]
+family = "glm"
+quantization = "6bit"
+base_model = "GLM 4.7"
+capabilities = ["text", "thinking"]
[storage_size]
in_bytes = 286737579648
diff --git a/resources/inference_model_cards/mlx-community--GLM-4.7-8bit-gs32.toml b/resources/inference_model_cards/mlx-community--GLM-4.7-8bit-gs32.toml
index 6f221cef..0f56c2f7 100644
--- a/resources/inference_model_cards/mlx-community--GLM-4.7-8bit-gs32.toml
+++ b/resources/inference_model_cards/mlx-community--GLM-4.7-8bit-gs32.toml
@@ -3,6 +3,10 @@ n_layers = 91
hidden_size = 5120
supports_tensor = true
tasks = ["TextGeneration"]
+family = "glm"
+quantization = "8bit"
+base_model = "GLM 4.7"
+capabilities = ["text", "thinking"]
[storage_size]
in_bytes = 396963397248
diff --git a/resources/inference_model_cards/mlx-community--GLM-4.7-Flash-4bit.toml b/resources/inference_model_cards/mlx-community--GLM-4.7-Flash-4bit.toml
index 43eb0dcd..8637cef0 100644
--- a/resources/inference_model_cards/mlx-community--GLM-4.7-Flash-4bit.toml
+++ b/resources/inference_model_cards/mlx-community--GLM-4.7-Flash-4bit.toml
@@ -3,6 +3,10 @@ n_layers = 47
hidden_size = 2048
supports_tensor = true
tasks = ["TextGeneration"]
+family = "glm"
+quantization = "4bit"
+base_model = "GLM 4.7 Flash"
+capabilities = ["text", "thinking"]
[storage_size]
in_bytes = 19327352832
diff --git a/resources/inference_model_cards/mlx-community--GLM-4.7-Flash-5bit.toml b/resources/inference_model_cards/mlx-community--GLM-4.7-Flash-5bit.toml
index 6a512c0a..b9a9da4d 100644
--- a/resources/inference_model_cards/mlx-community--GLM-4.7-Flash-5bit.toml
+++ b/resources/inference_model_cards/mlx-community--GLM-4.7-Flash-5bit.toml
@@ -3,6 +3,10 @@ n_layers = 47
hidden_size = 2048
supports_tensor = true
tasks = ["TextGeneration"]
+family = "glm"
+quantization = "5bit"
+base_model = "GLM 4.7 Flash"
+capabilities = ["text", "thinking"]
[storage_size]
in_bytes = 22548578304
diff --git a/resources/inference_model_cards/mlx-community--GLM-4.7-Flash-6bit.toml b/resources/inference_model_cards/mlx-community--GLM-4.7-Flash-6bit.toml
index 86c65489..e3cb1fa8 100644
--- a/resources/inference_model_cards/mlx-community--GLM-4.7-Flash-6bit.toml
+++ b/resources/inference_model_cards/mlx-community--GLM-4.7-Flash-6bit.toml
@@ -3,6 +3,10 @@ n_layers = 47
hidden_size = 2048
supports_tensor = true
tasks = ["TextGeneration"]
+family = "glm"
+quantization = "6bit"
+base_model = "GLM 4.7 Flash"
+capabilities = ["text", "thinking"]
[storage_size]
in_bytes = 26843545600
diff --git a/resources/inference_model_cards/mlx-community--GLM-4.7-Flash-8bit.toml b/resources/inference_model_cards/mlx-community--GLM-4.7-Flash-8bit.toml
index eb69183f..bd6df312 100644
--- a/resources/inference_model_cards/mlx-community--GLM-4.7-Flash-8bit.toml
+++ b/resources/inference_model_cards/mlx-community--GLM-4.7-Flash-8bit.toml
@@ -3,6 +3,10 @@ n_layers = 47
hidden_size = 2048
supports_tensor = true
tasks = ["TextGeneration"]
+family = "glm"
+quantization = "8bit"
+base_model = "GLM 4.7 Flash"
+capabilities = ["text", "thinking"]
[storage_size]
in_bytes = 34359738368
diff --git a/resources/inference_model_cards/mlx-community--Kimi-K2-Instruct-4bit.toml b/resources/inference_model_cards/mlx-community--Kimi-K2-Instruct-4bit.toml
index d7acabec..3f21d4c0 100644
--- a/resources/inference_model_cards/mlx-community--Kimi-K2-Instruct-4bit.toml
+++ b/resources/inference_model_cards/mlx-community--Kimi-K2-Instruct-4bit.toml
@@ -3,6 +3,10 @@ n_layers = 61
hidden_size = 7168
supports_tensor = true
tasks = ["TextGeneration"]
+family = "kimi"
+quantization = "4bit"
+base_model = "Kimi K2"
+capabilities = ["text"]
[storage_size]
in_bytes = 620622774272
diff --git a/resources/inference_model_cards/mlx-community--Kimi-K2-Thinking.toml b/resources/inference_model_cards/mlx-community--Kimi-K2-Thinking.toml
index 8d21727b..0a955b04 100644
--- a/resources/inference_model_cards/mlx-community--Kimi-K2-Thinking.toml
+++ b/resources/inference_model_cards/mlx-community--Kimi-K2-Thinking.toml
@@ -3,6 +3,10 @@ n_layers = 61
hidden_size = 7168
supports_tensor = true
tasks = ["TextGeneration"]
+family = "kimi"
+quantization = ""
+base_model = "Kimi K2"
+capabilities = ["text", "thinking"]
[storage_size]
in_bytes = 706522120192
diff --git a/resources/inference_model_cards/mlx-community--Kimi-K2.5.toml b/resources/inference_model_cards/mlx-community--Kimi-K2.5.toml
index c44cf9b1..806c6b30 100644
--- a/resources/inference_model_cards/mlx-community--Kimi-K2.5.toml
+++ b/resources/inference_model_cards/mlx-community--Kimi-K2.5.toml
@@ -3,6 +3,10 @@ n_layers = 61
hidden_size = 7168
supports_tensor = true
tasks = ["TextGeneration"]
+family = "kimi"
+quantization = ""
+base_model = "Kimi K2.5"
+capabilities = ["text", "thinking"]
[storage_size]
in_bytes = 662498705408
diff --git a/resources/inference_model_cards/mlx-community--Llama-3.2-1B-Instruct-4bit.toml b/resources/inference_model_cards/mlx-community--Llama-3.2-1B-Instruct-4bit.toml
index db334221..b38ec20f 100644
--- a/resources/inference_model_cards/mlx-community--Llama-3.2-1B-Instruct-4bit.toml
+++ b/resources/inference_model_cards/mlx-community--Llama-3.2-1B-Instruct-4bit.toml
@@ -3,6 +3,10 @@ n_layers = 16
hidden_size = 2048
supports_tensor = true
tasks = ["TextGeneration"]
+family = "llama"
+quantization = "4bit"
+base_model = "Llama 3.2 1B"
+capabilities = ["text"]
[storage_size]
in_bytes = 729808896
diff --git a/resources/inference_model_cards/mlx-community--Llama-3.2-3B-Instruct-4bit.toml b/resources/inference_model_cards/mlx-community--Llama-3.2-3B-Instruct-4bit.toml
index 001b4a1d..81ce4567 100644
--- a/resources/inference_model_cards/mlx-community--Llama-3.2-3B-Instruct-4bit.toml
+++ b/resources/inference_model_cards/mlx-community--Llama-3.2-3B-Instruct-4bit.toml
@@ -3,6 +3,10 @@ n_layers = 28
hidden_size = 3072
supports_tensor = true
tasks = ["TextGeneration"]
+family = "llama"
+quantization = "4bit"
+base_model = "Llama 3.2 3B"
+capabilities = ["text"]
[storage_size]
in_bytes = 1863319552
diff --git a/resources/inference_model_cards/mlx-community--Llama-3.2-3B-Instruct-8bit.toml b/resources/inference_model_cards/mlx-community--Llama-3.2-3B-Instruct-8bit.toml
index 358db81b..ac9a203b 100644
--- a/resources/inference_model_cards/mlx-community--Llama-3.2-3B-Instruct-8bit.toml
+++ b/resources/inference_model_cards/mlx-community--Llama-3.2-3B-Instruct-8bit.toml
@@ -3,6 +3,10 @@ n_layers = 28
hidden_size = 3072
supports_tensor = true
tasks = ["TextGeneration"]
+family = "llama"
+quantization = "8bit"
+base_model = "Llama 3.2 3B"
+capabilities = ["text"]
[storage_size]
in_bytes = 3501195264
diff --git a/resources/inference_model_cards/mlx-community--Llama-3.3-70B-Instruct-4bit.toml b/resources/inference_model_cards/mlx-community--Llama-3.3-70B-Instruct-4bit.toml
index cf6eece0..24c7cbaa 100644
--- a/resources/inference_model_cards/mlx-community--Llama-3.3-70B-Instruct-4bit.toml
+++ b/resources/inference_model_cards/mlx-community--Llama-3.3-70B-Instruct-4bit.toml
@@ -3,6 +3,10 @@ n_layers = 80
hidden_size = 8192
supports_tensor = true
tasks = ["TextGeneration"]
+family = "llama"
+quantization = "4bit"
+base_model = "Llama 3.3 70B"
+capabilities = ["text"]
[storage_size]
in_bytes = 40652242944
diff --git a/resources/inference_model_cards/mlx-community--Llama-3.3-70B-Instruct-8bit.toml b/resources/inference_model_cards/mlx-community--Llama-3.3-70B-Instruct-8bit.toml
index 15f0c551..3bfc97dc 100644
--- a/resources/inference_model_cards/mlx-community--Llama-3.3-70B-Instruct-8bit.toml
+++ b/resources/inference_model_cards/mlx-community--Llama-3.3-70B-Instruct-8bit.toml
@@ -3,6 +3,10 @@ n_layers = 80
hidden_size = 8192
supports_tensor = true
tasks = ["TextGeneration"]
+family = "llama"
+quantization = "8bit"
+base_model = "Llama 3.3 70B"
+capabilities = ["text"]
[storage_size]
in_bytes = 76799803392
diff --git a/resources/inference_model_cards/mlx-community--Meta-Llama-3.1-70B-Instruct-4bit.toml b/resources/inference_model_cards/mlx-community--Meta-Llama-3.1-70B-Instruct-4bit.toml
index b766164d..27d0b724 100644
--- a/resources/inference_model_cards/mlx-community--Meta-Llama-3.1-70B-Instruct-4bit.toml
+++ b/resources/inference_model_cards/mlx-community--Meta-Llama-3.1-70B-Instruct-4bit.toml
@@ -3,6 +3,10 @@ n_layers = 80
hidden_size = 8192
supports_tensor = true
tasks = ["TextGeneration"]
+family = "llama"
+quantization = "4bit"
+base_model = "Llama 3.1 70B"
+capabilities = ["text"]
[storage_size]
in_bytes = 40652242944
diff --git a/resources/inference_model_cards/mlx-community--Meta-Llama-3.1-8B-Instruct-4bit.toml b/resources/inference_model_cards/mlx-community--Meta-Llama-3.1-8B-Instruct-4bit.toml
index b6d10c40..1fe34ba8 100644
--- a/resources/inference_model_cards/mlx-community--Meta-Llama-3.1-8B-Instruct-4bit.toml
+++ b/resources/inference_model_cards/mlx-community--Meta-Llama-3.1-8B-Instruct-4bit.toml
@@ -3,6 +3,10 @@ n_layers = 32
hidden_size = 4096
supports_tensor = true
tasks = ["TextGeneration"]
+family = "llama"
+quantization = "4bit"
+base_model = "Llama 3.1 8B"
+capabilities = ["text"]
[storage_size]
in_bytes = 4637851648
diff --git a/resources/inference_model_cards/mlx-community--Meta-Llama-3.1-8B-Instruct-8bit.toml b/resources/inference_model_cards/mlx-community--Meta-Llama-3.1-8B-Instruct-8bit.toml
index 4cfe47cf..5310a2a0 100644
--- a/resources/inference_model_cards/mlx-community--Meta-Llama-3.1-8B-Instruct-8bit.toml
+++ b/resources/inference_model_cards/mlx-community--Meta-Llama-3.1-8B-Instruct-8bit.toml
@@ -3,6 +3,10 @@ n_layers = 32
hidden_size = 4096
supports_tensor = true
tasks = ["TextGeneration"]
+family = "llama"
+quantization = "8bit"
+base_model = "Llama 3.1 8B"
+capabilities = ["text"]
[storage_size]
in_bytes = 8954839040
diff --git a/resources/inference_model_cards/mlx-community--Meta-Llama-3.1-8B-Instruct-bf16.toml b/resources/inference_model_cards/mlx-community--Meta-Llama-3.1-8B-Instruct-bf16.toml
index 9e04f27b..eb6405e0 100644
--- a/resources/inference_model_cards/mlx-community--Meta-Llama-3.1-8B-Instruct-bf16.toml
+++ b/resources/inference_model_cards/mlx-community--Meta-Llama-3.1-8B-Instruct-bf16.toml
@@ -3,6 +3,10 @@ n_layers = 32
hidden_size = 4096
supports_tensor = true
tasks = ["TextGeneration"]
+family = "llama"
+quantization = "bf16"
+base_model = "Llama 3.1 8B"
+capabilities = ["text"]
[storage_size]
in_bytes = 16882073600
diff --git a/resources/inference_model_cards/mlx-community--MiniMax-M2.1-3bit.toml b/resources/inference_model_cards/mlx-community--MiniMax-M2.1-3bit.toml
index 4bf81136..92ec6746 100644
--- a/resources/inference_model_cards/mlx-community--MiniMax-M2.1-3bit.toml
+++ b/resources/inference_model_cards/mlx-community--MiniMax-M2.1-3bit.toml
@@ -3,6 +3,10 @@ n_layers = 61
hidden_size = 3072
supports_tensor = true
tasks = ["TextGeneration"]
+family = "minimax"
+quantization = "3bit"
+base_model = "MiniMax M2.1"
+capabilities = ["text", "thinking"]
[storage_size]
in_bytes = 100086644736
diff --git a/resources/inference_model_cards/mlx-community--MiniMax-M2.1-8bit.toml b/resources/inference_model_cards/mlx-community--MiniMax-M2.1-8bit.toml
index 54a49f97..c1388d2f 100644
--- a/resources/inference_model_cards/mlx-community--MiniMax-M2.1-8bit.toml
+++ b/resources/inference_model_cards/mlx-community--MiniMax-M2.1-8bit.toml
@@ -3,6 +3,10 @@ n_layers = 61
hidden_size = 3072
supports_tensor = true
tasks = ["TextGeneration"]
+family = "minimax"
+quantization = "8bit"
+base_model = "MiniMax M2.1"
+capabilities = ["text", "thinking"]
[storage_size]
in_bytes = 242986745856
diff --git a/resources/inference_model_cards/mlx-community--Qwen3-0.6B-4bit.toml b/resources/inference_model_cards/mlx-community--Qwen3-0.6B-4bit.toml
index 212cdef6..7929aaba 100644
--- a/resources/inference_model_cards/mlx-community--Qwen3-0.6B-4bit.toml
+++ b/resources/inference_model_cards/mlx-community--Qwen3-0.6B-4bit.toml
@@ -3,6 +3,10 @@ n_layers = 28
hidden_size = 1024
supports_tensor = false
tasks = ["TextGeneration"]
+family = "qwen"
+quantization = "4bit"
+base_model = "Qwen3 0.6B"
+capabilities = ["text", "thinking"]
[storage_size]
in_bytes = 342884352
diff --git a/resources/inference_model_cards/mlx-community--Qwen3-0.6B-8bit.toml b/resources/inference_model_cards/mlx-community--Qwen3-0.6B-8bit.toml
index ac591d6c..d9fcc368 100644
--- a/resources/inference_model_cards/mlx-community--Qwen3-0.6B-8bit.toml
+++ b/resources/inference_model_cards/mlx-community--Qwen3-0.6B-8bit.toml
@@ -3,6 +3,10 @@ n_layers = 28
hidden_size = 1024
supports_tensor = false
tasks = ["TextGeneration"]
+family = "qwen"
+quantization = "8bit"
+base_model = "Qwen3 0.6B"
+capabilities = ["text", "thinking"]
[storage_size]
in_bytes = 698351616
diff --git a/resources/inference_model_cards/mlx-community--Qwen3-235B-A22B-Instruct-2507-4bit.toml b/resources/inference_model_cards/mlx-community--Qwen3-235B-A22B-Instruct-2507-4bit.toml
index 020c11be..ef835c6a 100644
--- a/resources/inference_model_cards/mlx-community--Qwen3-235B-A22B-Instruct-2507-4bit.toml
+++ b/resources/inference_model_cards/mlx-community--Qwen3-235B-A22B-Instruct-2507-4bit.toml
@@ -3,6 +3,10 @@ n_layers = 94
hidden_size = 4096
supports_tensor = true
tasks = ["TextGeneration"]
+family = "qwen"
+quantization = "4bit"
+base_model = "Qwen3 235B"
+capabilities = ["text", "thinking"]
[storage_size]
in_bytes = 141733920768
diff --git a/resources/inference_model_cards/mlx-community--Qwen3-235B-A22B-Instruct-2507-8bit.toml b/resources/inference_model_cards/mlx-community--Qwen3-235B-A22B-Instruct-2507-8bit.toml
index 64afc366..f6e079ab 100644
--- a/resources/inference_model_cards/mlx-community--Qwen3-235B-A22B-Instruct-2507-8bit.toml
+++ b/resources/inference_model_cards/mlx-community--Qwen3-235B-A22B-Instruct-2507-8bit.toml
@@ -3,6 +3,10 @@ n_layers = 94
hidden_size = 4096
supports_tensor = true
tasks = ["TextGeneration"]
+family = "qwen"
+quantization = "8bit"
+base_model = "Qwen3 235B"
+capabilities = ["text", "thinking"]
[storage_size]
in_bytes = 268435456000
diff --git a/resources/inference_model_cards/mlx-community--Qwen3-30B-A3B-4bit.toml b/resources/inference_model_cards/mlx-community--Qwen3-30B-A3B-4bit.toml
index 1b9f92a6..48a6666f 100644
--- a/resources/inference_model_cards/mlx-community--Qwen3-30B-A3B-4bit.toml
+++ b/resources/inference_model_cards/mlx-community--Qwen3-30B-A3B-4bit.toml
@@ -3,6 +3,10 @@ n_layers = 48
hidden_size = 2048
supports_tensor = true
tasks = ["TextGeneration"]
+family = "qwen"
+quantization = "4bit"
+base_model = "Qwen3 30B"
+capabilities = ["text", "thinking"]
[storage_size]
in_bytes = 17612931072
diff --git a/resources/inference_model_cards/mlx-community--Qwen3-30B-A3B-8bit.toml b/resources/inference_model_cards/mlx-community--Qwen3-30B-A3B-8bit.toml
index f8e59ac1..c283396f 100644
--- a/resources/inference_model_cards/mlx-community--Qwen3-30B-A3B-8bit.toml
+++ b/resources/inference_model_cards/mlx-community--Qwen3-30B-A3B-8bit.toml
@@ -3,6 +3,10 @@ n_layers = 48
hidden_size = 2048
supports_tensor = true
tasks = ["TextGeneration"]
+family = "qwen"
+quantization = "8bit"
+base_model = "Qwen3 30B"
+capabilities = ["text", "thinking"]
[storage_size]
in_bytes = 33279705088
diff --git a/resources/inference_model_cards/mlx-community--Qwen3-Coder-480B-A35B-Instruct-4bit.toml b/resources/inference_model_cards/mlx-community--Qwen3-Coder-480B-A35B-Instruct-4bit.toml
index 25a0cf2b..b390bd21 100644
--- a/resources/inference_model_cards/mlx-community--Qwen3-Coder-480B-A35B-Instruct-4bit.toml
+++ b/resources/inference_model_cards/mlx-community--Qwen3-Coder-480B-A35B-Instruct-4bit.toml
@@ -3,6 +3,10 @@ n_layers = 62
hidden_size = 6144
supports_tensor = true
tasks = ["TextGeneration"]
+family = "qwen"
+quantization = "4bit"
+base_model = "Qwen3 Coder 480B"
+capabilities = ["text", "code"]
[storage_size]
in_bytes = 289910292480
diff --git a/resources/inference_model_cards/mlx-community--Qwen3-Coder-480B-A35B-Instruct-8bit.toml b/resources/inference_model_cards/mlx-community--Qwen3-Coder-480B-A35B-Instruct-8bit.toml
index cdb1d6ec..1c21307c 100644
--- a/resources/inference_model_cards/mlx-community--Qwen3-Coder-480B-A35B-Instruct-8bit.toml
+++ b/resources/inference_model_cards/mlx-community--Qwen3-Coder-480B-A35B-Instruct-8bit.toml
@@ -3,6 +3,10 @@ n_layers = 62
hidden_size = 6144
supports_tensor = true
tasks = ["TextGeneration"]
+family = "qwen"
+quantization = "8bit"
+base_model = "Qwen3 Coder 480B"
+capabilities = ["text", "code"]
[storage_size]
in_bytes = 579820584960
diff --git a/resources/inference_model_cards/mlx-community--Qwen3-Next-80B-A3B-Instruct-4bit.toml b/resources/inference_model_cards/mlx-community--Qwen3-Next-80B-A3B-Instruct-4bit.toml
index db55b7f9..386a3fa1 100644
--- a/resources/inference_model_cards/mlx-community--Qwen3-Next-80B-A3B-Instruct-4bit.toml
+++ b/resources/inference_model_cards/mlx-community--Qwen3-Next-80B-A3B-Instruct-4bit.toml
@@ -3,6 +3,10 @@ n_layers = 48
hidden_size = 2048
supports_tensor = true
tasks = ["TextGeneration"]
+family = "qwen"
+quantization = "4bit"
+base_model = "Qwen3 Next 80B"
+capabilities = ["text"]
[storage_size]
in_bytes = 46976204800
diff --git a/resources/inference_model_cards/mlx-community--Qwen3-Next-80B-A3B-Instruct-8bit.toml b/resources/inference_model_cards/mlx-community--Qwen3-Next-80B-A3B-Instruct-8bit.toml
index e36e24b8..0e2bf2a5 100644
--- a/resources/inference_model_cards/mlx-community--Qwen3-Next-80B-A3B-Instruct-8bit.toml
+++ b/resources/inference_model_cards/mlx-community--Qwen3-Next-80B-A3B-Instruct-8bit.toml
@@ -3,6 +3,10 @@ n_layers = 48
hidden_size = 2048
supports_tensor = true
tasks = ["TextGeneration"]
+family = "qwen"
+quantization = "8bit"
+base_model = "Qwen3 Next 80B"
+capabilities = ["text"]
[storage_size]
in_bytes = 88814387200
diff --git a/resources/inference_model_cards/mlx-community--Qwen3-Next-80B-A3B-Thinking-4bit.toml b/resources/inference_model_cards/mlx-community--Qwen3-Next-80B-A3B-Thinking-4bit.toml
index bc3bdf50..2a3e3c19 100644
--- a/resources/inference_model_cards/mlx-community--Qwen3-Next-80B-A3B-Thinking-4bit.toml
+++ b/resources/inference_model_cards/mlx-community--Qwen3-Next-80B-A3B-Thinking-4bit.toml
@@ -3,6 +3,10 @@ n_layers = 48
hidden_size = 2048
supports_tensor = true
tasks = ["TextGeneration"]
+family = "qwen"
+quantization = "4bit"
+base_model = "Qwen3 Next 80B"
+capabilities = ["text", "thinking"]
[storage_size]
in_bytes = 47080074240
diff --git a/resources/inference_model_cards/mlx-community--Qwen3-Next-80B-A3B-Thinking-8bit.toml b/resources/inference_model_cards/mlx-community--Qwen3-Next-80B-A3B-Thinking-8bit.toml
index dd5512a7..65d33253 100644
--- a/resources/inference_model_cards/mlx-community--Qwen3-Next-80B-A3B-Thinking-8bit.toml
+++ b/resources/inference_model_cards/mlx-community--Qwen3-Next-80B-A3B-Thinking-8bit.toml
@@ -3,6 +3,10 @@ n_layers = 48
hidden_size = 2048
supports_tensor = true
tasks = ["TextGeneration"]
+family = "qwen"
+quantization = "8bit"
+base_model = "Qwen3 Next 80B"
+capabilities = ["text", "thinking"]
[storage_size]
in_bytes = 88814387200
diff --git a/resources/inference_model_cards/mlx-community--gpt-oss-120b-MXFP4-Q8.toml b/resources/inference_model_cards/mlx-community--gpt-oss-120b-MXFP4-Q8.toml
index c725e728..f579c618 100644
--- a/resources/inference_model_cards/mlx-community--gpt-oss-120b-MXFP4-Q8.toml
+++ b/resources/inference_model_cards/mlx-community--gpt-oss-120b-MXFP4-Q8.toml
@@ -3,6 +3,10 @@ n_layers = 36
hidden_size = 2880
supports_tensor = true
tasks = ["TextGeneration"]
+family = "gpt-oss"
+quantization = "MXFP4-Q8"
+base_model = "GPT-OSS 120B"
+capabilities = ["text", "thinking"]
[storage_size]
in_bytes = 70652212224
diff --git a/resources/inference_model_cards/mlx-community--gpt-oss-20b-MXFP4-Q8.toml b/resources/inference_model_cards/mlx-community--gpt-oss-20b-MXFP4-Q8.toml
index bf8f1a60..af1e04ad 100644
--- a/resources/inference_model_cards/mlx-community--gpt-oss-20b-MXFP4-Q8.toml
+++ b/resources/inference_model_cards/mlx-community--gpt-oss-20b-MXFP4-Q8.toml
@@ -3,6 +3,10 @@ n_layers = 24
hidden_size = 2880
supports_tensor = true
tasks = ["TextGeneration"]
+family = "gpt-oss"
+quantization = "MXFP4-Q8"
+base_model = "GPT-OSS 20B"
+capabilities = ["text", "thinking"]
[storage_size]
in_bytes = 12025908224
diff --git a/resources/inference_model_cards/mlx-community--llama-3.3-70b-instruct-fp16.toml b/resources/inference_model_cards/mlx-community--llama-3.3-70b-instruct-fp16.toml
index dd451015..e61660c2 100644
--- a/resources/inference_model_cards/mlx-community--llama-3.3-70b-instruct-fp16.toml
+++ b/resources/inference_model_cards/mlx-community--llama-3.3-70b-instruct-fp16.toml
@@ -3,6 +3,10 @@ n_layers = 80
hidden_size = 8192
supports_tensor = true
tasks = ["TextGeneration"]
+family = "llama"
+quantization = "fp16"
+base_model = "Llama 3.3 70B"
+capabilities = ["text"]
[storage_size]
in_bytes = 144383672320
diff --git a/src/exo/master/api.py b/src/exo/master/api.py
index 01e145b4..282cb0f1 100644
--- a/src/exo/master/api.py
+++ b/src/exo/master/api.py
@@ -74,6 +74,7 @@ from exo.shared.types.api import (
ErrorResponse,
FinishReason,
GenerationStats,
+ HuggingFaceSearchResult,
ImageData,
ImageEditsTaskParams,
ImageGenerationResponse,
@@ -262,6 +263,7 @@ class API:
self.app.get("/v1/models")(self.get_models)
self.app.post("/models/add")(self.add_custom_model)
self.app.delete("/models/custom/{model_id:path}")(self.delete_custom_model)
+ self.app.get("/models/search")(self.search_models)
self.app.post("/v1/chat/completions", response_model=None)(
self.chat_completions
)
@@ -1222,6 +1224,10 @@ class API:
supports_tensor=card.supports_tensor,
tasks=[task.value for task in card.tasks],
is_custom=is_custom_card(card.model_id),
+ family=card.family,
+ quantization=card.quantization,
+ base_model=card.base_model,
+ capabilities=card.capabilities,
)
for card in await get_model_cards()
]
@@ -1257,6 +1263,30 @@ class API:
{"message": "Model card deleted", "model_id": str(model_id)}
)
+ async def search_models(
+ self, query: str = "", limit: int = 20
+ ) -> list[HuggingFaceSearchResult]:
+ """Search HuggingFace Hub for mlx-community models."""
+ from huggingface_hub import list_models
+
+ results = list_models(
+ search=query or None,
+ author="mlx-community",
+ sort="downloads",
+ limit=limit,
+ )
+ return [
+ HuggingFaceSearchResult(
+ id=m.id,
+ author=m.author or "",
+ downloads=m.downloads or 0,
+ likes=m.likes or 0,
+ last_modified=str(m.last_modified or ""),
+ tags=list(m.tags or []),
+ )
+ for m in results
+ ]
+
async def run(self):
cfg = Config()
cfg.bind = f"0.0.0.0:{self.port}"
diff --git a/src/exo/shared/models/model_cards.py b/src/exo/shared/models/model_cards.py
index 47b54d38..79dc95be 100644
--- a/src/exo/shared/models/model_cards.py
+++ b/src/exo/shared/models/model_cards.py
@@ -78,6 +78,10 @@ class ModelCard(CamelCaseModel):
supports_tensor: bool
tasks: list[ModelTask]
components: list[ComponentInfo] | None = None
+ family: str = ""
+ quantization: str = ""
+ base_model: str = ""
+ capabilities: list[str] = []
@field_validator("tasks", mode="before")
@classmethod
diff --git a/src/exo/shared/types/api.py b/src/exo/shared/types/api.py
index e5710014..5a7bae1e 100644
--- a/src/exo/shared/types/api.py
+++ b/src/exo/shared/types/api.py
@@ -43,6 +43,10 @@ class ModelListModel(BaseModel):
supports_tensor: bool = Field(default=False)
tasks: list[str] = Field(default=[])
is_custom: bool = Field(default=False)
+ family: str = Field(default="")
+ quantization: str = Field(default="")
+ base_model: str = Field(default="")
+ capabilities: list[str] = Field(default_factory=list)
class ModelList(BaseModel):
@@ -206,6 +210,15 @@ class AddCustomModelParams(BaseModel):
model_id: ModelId
+class HuggingFaceSearchResult(BaseModel):
+ id: str
+ author: str = ""
+ downloads: int = 0
+ likes: int = 0
+ last_modified: str = ""
+ tags: list[str] = Field(default_factory=list)
+
+
class PlaceInstanceParams(BaseModel):
model_id: ModelId
sharding: Sharding = Sharding.Pipeline
diff --git a/src/exo/worker/engines/mlx/auto_parallel.py b/src/exo/worker/engines/mlx/auto_parallel.py
index 84ed1f29..36589b7a 100644
--- a/src/exo/worker/engines/mlx/auto_parallel.py
+++ b/src/exo/worker/engines/mlx/auto_parallel.py
@@ -164,6 +164,12 @@ def _inner_model(model: nn.Module) -> nn.Module:
if isinstance(inner, nn.Module):
return inner
+ inner = getattr(model, "language_model", None)
+ if isinstance(inner, nn.Module):
+ inner_inner = getattr(inner, "model", None)
+ if isinstance(inner_inner, nn.Module):
+ return inner_inner
+
raise ValueError("Model must either have a 'model' or 'transformer' attribute")
diff --git a/src/exo/worker/engines/mlx/utils_mlx.py b/src/exo/worker/engines/mlx/utils_mlx.py
index d7fb9958..5dae48a4 100644
--- a/src/exo/worker/engines/mlx/utils_mlx.py
+++ b/src/exo/worker/engines/mlx/utils_mlx.py
@@ -384,6 +384,17 @@ def load_tokenizer_for_model_id(
eos_token_ids=eos_token_ids,
)
+ if "gemma-3" in model_id_lower:
+ gemma_3_eos_id = 1
+ gemma_3_end_of_turn_id = 106
+ if tokenizer.eos_token_ids is not None:
+ if gemma_3_end_of_turn_id not in tokenizer.eos_token_ids:
+ tokenizer.eos_token_ids = list(tokenizer.eos_token_ids) + [
+ gemma_3_end_of_turn_id
+ ]
+ else:
+ tokenizer.eos_token_ids = [gemma_3_eos_id, gemma_3_end_of_turn_id]
+
return tokenizer
← 20632789 feat: add custom HuggingFace model support (#1368)
·
back to Exo
·
add resources dir to nix (#1376) 7b6cad94 →