Concept · Reeve
Where Reeve's work runs
Which accelerator answers each of Reeve's model requests, how requests take turns and fall back, why the NPU, GPU or CPU is busy and then quiet, and how background work is kept gentle.
- Article
- 1207
- Applies to
- Reeve 0.17.2
- Last reviewed
- For
- For developers
Three kinds of model work#
| Kind | Used by | What it is |
|---|---|---|
| Chat | triage, summarize, map, ask, docs, session, the upgrade scout, cluster's group names | A small instruction-tuned model, the same on every accelerator, so answers stay comparable. |
| Embeddings | search, docs' passages, cluster | A small embedding model. Each model gets its own index, so switching back and forth rebuilds nothing, and an index never mixes two models. |
| Vision | triage of a failed UI test, looking at the failing step's screenshot | A small vision model, asked only narrow questions; OCR supplies the words, and rules give the verdict. |
A reranker, when the Smith has set one up ("sharper search"), takes a second look at search's ten best files.
Code does everything else, with no model: outline, changes, refs, history, layout, pr, the history tools, and every pass/fail verdict.
The accelerators#
Reeve uses the accelerators the Smith sets up for every agent on the PC, each behind a model server that speaks the common chat-completions API: the NPU, each GPU, and the CPU as a last resort for short requests. The Smith owns them, their order and their idle times; Reeve uses them as any agent does. See Where the work runs for how the staff share them.
Without the Smith hired, Reeve keeps the model servers itself, and a developer can set up GPUs or the CPU with accelerators setup (see Reeve's command line).
Choosing one, per request#
- Candidates: the accelerators that serve the request's kind, whose request cap fits it, that haven't failed in the last 10 minutes, and, for background work, that no game is using.
- In order: the NPU first, then GPUs with 2 GB or more of their own memory, largest first, then graphics that share the PC's memory, then the CPU.
- The NPU first: when the NPU can do a request, it alone takes it, and the request waits its turn there even while a GPU is free. A GPU or the CPU takes only what the NPU can't do: a longer request, a kind it doesn't serve, or while it has failed.
- The first with a free slot and nobody waiting, else the shortest line. Background work that would be fifth or later in every line waits.
Every answer from a model says where it ran: [reeve: … on <accelerator>, …].
Taking turns#
Every model request waits its turn in a line, shared by every agent on the PC: one at a time on the NPU, and as many at once on a GPU as it has slots. A request someone is waiting on (your assistant's, a terminal's) goes ahead of background work. Each request is refused above its accelerator's cap, and larger inputs are split into chunks (map-reduce), which is why summarize takes about 10 seconds per 120 lines.
When one fails#
When a model server won't start within 30 seconds, refuses the connection, answers with a server error or times out, the request goes once to the next candidate, and the one that failed is skipped for 10 minutes. When every one that would do has failed, Reeve tries them anyway rather than refuse everything.
Embeddings fall back only to an accelerator serving the same model. With none, search ranks by keywords instead, and says so.
Games come first#
While another program keeps a GPU's 3D engine more than 25% busy, background work keeps off that GPU. A request someone is waiting on may still use it. With Castellan's Use the graphics card for models when there's an NPU off, Reeve sends nothing to a GPU at all on a PC with an NPU, not even as the fallback.
Why the NPU, GPU or CPU comes and goes#
- A turn: one request is a few seconds' work, then the accelerator is idle again.
- A model loading: the first request after a model server starts, or after an idle one was stopped to give its memory back.
- Indexing: the first
searchin a repository computes embeddings for every file. Background work like this rests between requests, so a GPU or the CPU is busy at most half the time by default: How hard background work may run the graphics card, in Reeve's settings. - The rounds: fetches, package installs, TypeScript scans and OCR are CPU and disk work, on their schedule.
Reeve's page shows it as it happens, under Accelerators: each one's servers, whether it failed and why, whether a game is holding it back, the slots held, and who's waiting. Your assistant's status tool gives the same.
Related articles
Is this page right?
If something on it is wrong or out of date, tell us and we'll fix the page.
Still stuck? Write to support@castellan-software.com and mention article 1207. Every version of Reeve, and what changed in it, is in its release notes.