Skip to content

Add any Hugging Face GGUF model, not just the curated catalog - #38

Open
sachin-detrax wants to merge 1 commit into
AtomicBot-ai:mainfrom
sachin-detrax:feat/huggingface-custom-models
Open

Add any Hugging Face GGUF model, not just the curated catalog#38
sachin-detrax wants to merge 1 commit into
AtomicBot-ai:mainfrom
sachin-detrax:feat/huggingface-custom-models

Conversation

@sachin-detrax

Copy link
Copy Markdown
Contributor

Closes #27.

The local model list was a fixed array of nine models compiled into models-catalog.ts. Anything else on Hugging Face was unreachable without running llama-server by hand or waiting for a maintainer to ship a new release. Models are now open-ended.

A user-added entry is a full LocalModelDef with a custom- prefixed id, persisted in localModels.customModels (config v33 → v34; older files inherit []). loadConfig registers them in a catalog registry, so getLocalModelDef / isKnownLocalModelId resolve curated and custom entries through one lookup — daemon start, installer, TUI rows and CLI all picked them up without changes.

Entry points

  • TUI+ Add a model from Hugging Face..., pinned under Local text models. Enter opens a prompt taking either a reference (resolve and add) or free text (search, then pick by digit).
  • CLImodels add <ref> and models search <query>.
  • Prompt/models add and /models search.

References accepted

  • repo URLs, /resolve/ and /blob/ file URLs
  • hf://owner/repo[@rev]/file.gguf
  • a pasted hf download ... command (one- or two-argument, trailing flags dropped)
  • bare owner/name

Naming only a repo picks a 4-bit quant and any mmproj projector, so vision repos come out vision-capable. HF_TOKEN is honoured for gated repos, including on the weights download.

hf:// is parsed by hand rather than with new URL: that puts the owner in the host slot and lowercases it, and HF owners are case-sensitive, so hf://Qwen/... would silently 404.

Two pre-existing rendering defects fixed along the way

Both surfaced while testing this, and both could shred the panel on any daemon failure:

  1. Multi-line daemon errors were written straight into a single-row status line, so a llama-server stack trace rendered 48–66 lines into a 24-row budget. toStatusLine() truncates at every error-message write and at the render point, with wrap="truncate-end" on the Ink Text.
  2. The LlmPanel frame budget did not account for overlay modal / banner rows, so opening a modal during an error overran the frame. estimateOverlayRows() now covers them.

Regression tests pin the frame height for both.

Also

extractLoadFailure pulls the diagnostic reason out of llama-server logs so a failed load reports why rather than a generic message.

Tests

35 files changed, +2104/−52, including new suites for the HF reference parser, panel rendering, frame height, and the reducer.

npx tsc -p tsconfig.json --noEmit          # clean
npx vitest run src/local-llm/ src/tui/ src/cli/
#   before: 6 failed | 797 passed (803)
#   after:  6 failed | 835 passed (841)

The 6 failures are identical by name on main and on this branch (ChatLog/SplashBanner/TuiApp smoke, llm-panel selectors, persistEmbeddingHybridRecall) — pre-existing, untouched by this change. 38 tests added, none broken.

🤖 Generated with Claude Code

The local model list was a fixed array of nine models compiled into
models-catalog.ts. Anything else on Hugging Face was unreachable without
running llama-server by hand or waiting for a maintainer to ship a new
release.

Models are now open-ended. A user-added entry is a full LocalModelDef
with a `custom-` prefixed id, persisted in localModels.customModels
(config v33 -> v34, older files inherit []). loadConfig registers them
in a catalog registry, so getLocalModelDef / isKnownLocalModelId resolve
curated and custom entries through one lookup and every existing
consumer -- daemon start, installer, TUI rows, CLI -- picked them up
without changes.

Entry points:

  TUI  "+ Add a model from Hugging Face..." pinned under Local text
       models. Enter opens a prompt taking either a reference (resolve
       and add) or free text (search, then pick by digit).
  CLI  models add <ref> / models search <query>.
       /models add and /models search from the prompt.

References accepted: repo URLs, /resolve/ and /blob/ file URLs,
hf://owner/repo[@rev]/file.gguf, a pasted `hf download ...` command
(one- or two-argument, trailing flags dropped), and bare owner/name.
Naming only a repo picks a 4-bit quant and any mmproj projector, which
makes vision repos come out vision-capable. HF_TOKEN is honoured for
gated repos, including on the weights download.

hf:// is parsed by hand rather than with `new URL`: that puts the owner
in the host slot and lowercases it, and HF owners are case-sensitive, so
hf://Qwen/... would silently 404.

Two rendering defects surfaced while testing this, both pre-existing and
both able to shred the panel on any daemon failure:

  - daemonError / errorLine are view-model fields rendered as a single
    <Text> in a one-row slot, but a failed llama-server start stored a
    4KB, ~25-line log tail. Ink overlaps lines when a frame outgrows its
    budget rather than clipping: 48 rows into a 24-row budget. Flattened
    on write and on read, with wrap="truncate-end" on the one-row slots
    -- flattening alone still wrapped a long single line.
  - llm-panel computed its list budget without subtracting the modals
    and banners it renders above the list, unlike local-models-panel:
    42 rows into 24, and 66 with both defects together.

Flattening the message hid the real cause behind "did not become
healthy", so extractLoadFailure now lifts the diagnostic line out of the
log tail into the message head. Full detail still reaches the runtime
feed and the L logs tab.

Routing between "add this model" and "set this base URL" goes through
looksLikeHuggingFaceReference, stricter than the parser: a bare
owner/name is ambiguous with 192.168.1.5/api and stays behind the
explicit `add` verb. An inline regex here had already drifted once --
it knew about huggingface.co but not hf://, so /models hf://... was
silently persisted as http://hf://... instead of adding the model.

Tests: HF reference parsing and quant/projector selection, add-row
placement and prompt key handling (including panel hotkeys not stealing
characters mid-URL and digits only picking when results show), rendered
frames, and frame-height regressions asserting height <= budget.

Closes AtomicBot-ai#27

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@sosidudku1

Copy link
Copy Markdown
Collaborator

Strong PR, thank you. The input parser is the best part of it: repo URLs, hf:// refs, direct file links and even a pasted hf download command all resolve correctly, and falling back to a download-sorted live search is the right default. A few things are worth aligning before merge.

Two blocking items, both small:

  1. Sharded repos produce a silently broken model. When the picked file is a -00001-of-NNNNN.gguf shard, only the first part is downloaded and the model will not start, and the user finds out only at llama-server launch. Please either reject sharded picks with a clear error ("sharded models are not supported yet") or put an explicit warning into the "added" message.
  2. Filter out MTP companion files in the default-quant fallback. When none of the preferred quants exist and the code falls back to the smallest GGUF, a speculative-decoding companion file can win the pick and it is not a runnable model. Filtering names matching an MTP/ folder or mtp- / -mtp patterns closes this; it is a few lines in pickDefaultGgufFile.

Nice to have, your call whether in this PR or as follow-ups:

  1. Offer the quant list when a repo is added without an explicit file. listHuggingFaceGgufFiles already returns every file with sizes, and the numbered-pick modal from search results could be reused as a second step, with Enter keeping the recommended quant. Seeing "Q4_K_M 4.2 GB vs Q8_0 8.1 GB" before committing to a download is a real UX win.
  2. Show an approximate size before a search result is committed to the catalog. Right now the user picks by repo id and download count only, and learns the size after the entry is added.

Two things we will split into separate issues rather than grow this PR: memory-requirement estimates read from the GGUF header instead of the file-size heuristic (you left a note about this yourself), and hardware-aware quant selection instead of the constant Q4 preference. This PR does not need to carry them.

With 1 and 2 in, this is good to merge from my side.

@sosidudku1

Copy link
Copy Markdown
Collaborator

We merged #41, #48 and #50 today, which is why this branch now shows conflicts, they touch the same files. Could you rebase onto main? The two blocking items from the review (sharded repos, MTP filter) still apply after the rebase. Everything else is ready on our side.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Feature request: Allow users to point to a model on HuggingFace in the initial set up

2 participants