Make more of one GPU
Serve several models from one GPU, keep your own work first, and, if you share the machine, choose how much of it other people's requests can use.
Who runs the model decides what Saylek can do
A model on your machine is run either by Saylek itself or by your own runtime (Ollama, LM Studio, llama.cpp or vLLM). The machine's page in Machines labels each model Saylek-managed or with its runtime's name.
| What differs | Saylek runs the model | Your runtime runs the model |
|---|---|---|
| Loading | On the first request for it | Your runtime decides |
| Models loaded at once | At most 2 | Your runtime decides |
| Unloading | After 5 minutes idle, by default | Your runtime decides |
| Checking that a model fits | Not checked when a model loads (the wizard only estimates): a model that does not fit fails to load | Your runtime decides |
| Requests at once | One per model | One shared request at a time, unless the runtime reports more room |
1. Add the models you want
For models Saylek runs, register each model file:
saylek model add /absolute/path/to/model.gguf
Or copy .gguf files into ~/.saylek/models/ and run saylek model rescan. For your own
runtime, add the models in the runtime; if they do not appear, choose Scan again under the
machine's Settings, Manage sources.
You know it worked when saylek models, or the machine's page, lists every model.
2. Know how models load and unload
For models Saylek runs, a request loads its model if it is not loaded yet, so that first
request takes longer. At most two models stay loaded. When a third is requested, the loaded
model used least recently gives way if it is idle; if both are in use, the request gets
all model slots busy. A model unloads after 5 minutes without requests.
To change that, set keep_warm_secs under [daemon] in ~/.saylek/config.toml, in
seconds: higher keeps models ready longer, 0 frees memory as soon as the last request
ends. Saylek reads it when it starts. For your own runtime, use its settings instead.
3. Keep your own work first
- Work sent to Saylek's local API goes first. Requests that applications on this machine
send to
http://127.0.0.1:8443/v1(set up in Use the models you already serve) are taken before waiting shared requests, and stop a shared request that is already running at its next step. That request goes to another machine if one is free; otherwise it waits, or ends with an error the application can retry. A request that has already started streaming ends withowner_preempted. With your own runtime, Saylek stops passing it shared requests; the runtime may finish one it already has. - Requests through the hosted API do not stop running work, even your own requests with your own key on your own machine. When requests are waiting for your machine, Saylek's servers offer your own first.
- Application priority stops nothing. It orders one account's own requests; see Give important applications priority.
- Work sent straight to your runtime's own address bypasses Saylek, so Saylek neither counts nor stops it.
Optional: limit what people you share with can use
These apply only when sharing is on. host models and host slots are terminal only.
| To | In the web | From the terminal |
|---|---|---|
| Offer only some models | Terminal only | saylek host models MODEL_ID; --clear offers all again |
| Take fewer shared requests at once | Terminal only | saylek host slots 1; --clear resets it |
| Share only at set hours | Edit schedule on the machine's page | saylek host hours --set 09:00-17:00 |
| Stop shared requests for now | Turn off Share from this machine on the machine's page | saylek host stop, and saylek host start to share again |
Hours use the machine's own clock. A schedule has at most 5 periods, each a set of days with
one time range, or all day. A slot number above what the machine can take is capped, and
saylek host slots says so.
If something goes wrong
| What you see | What to do |
|---|---|
all model slots busy | Two models Saylek runs are both in use. Wait, or send fewer models at once |
model_load_failed | The model may not fit next to what is loaded. Lower keep_warm_secs (step 2), or use a smaller model |
The wizard says a model may not fit | It is an estimate. Skip it, choose another GPU, or try anyway |
| Shared requests still arrive at hours you set | Hours use the machine's clock, not yours. Check with saylek host hours |
Reference
- Sharing your models: sharing, pausing and stopping.
- Use the models you already serve: connecting your runtime.
- CLI reference: every command.
Last checked 2026-09-29 · read as markdown at /docs/make-more-of-one-gpu.md