# Make more of one GPU

Serve several models from one GPU, keep your own work first, and, if you share the machine,
choose how much of it other people's requests can use.

## Who runs the model decides what Saylek can do

A model on your machine is run either by Saylek itself or by your own runtime
(Ollama, LM Studio, llama.cpp or vLLM). The machine's page in [Machines](/machines) labels each model
**Saylek-managed** or with its runtime's name.

| What differs | Saylek runs the model | Your runtime runs the model |
|---|---|---|
| Loading | On the first request for it | Your runtime decides |
| Models loaded at once | At most 2 | Your runtime decides |
| Unloading | After 5 minutes idle, by default | Your runtime decides |
| Checking that a model fits | Not checked when a model loads (the wizard only estimates): a model that does not fit fails to load | Your runtime decides |
| Requests at once | One per model | One shared request at a time, unless the runtime reports more room |

## 1. Add the models you want

For models Saylek runs, register each model file:

```bash
saylek model add /absolute/path/to/model.gguf
```

Or copy `.gguf` files into `~/.saylek/models/` and run `saylek model rescan`. For your own
runtime, add the models in the runtime; if they do not appear, choose **Scan again** under the
machine's **Settings**, **Manage sources**.

**You know it worked when** `saylek models`, or the machine's page, lists every model.

## 2. Know how models load and unload

For models Saylek runs, a request loads its model if it is not loaded yet, so that first
request takes longer. At most two models stay loaded. When a third is requested, the loaded
model used least recently gives way if it is idle; if both are in use, the request gets
`all model slots busy`. A model unloads after 5 minutes without requests.

To change that, set `keep_warm_secs` under `[daemon]` in `~/.saylek/config.toml`, in
seconds: higher keeps models ready longer, `0` frees memory as soon as the last request
ends. Saylek reads it when it starts. For your own runtime, use its settings instead.

## 3. Keep your own work first

- **Work sent to Saylek's local API goes first.** Requests that applications on this machine
  send to `http://127.0.0.1:8443/v1` (set up in
  [Use the models you already serve](/docs/use-your-existing-runtime)) are taken before
  waiting shared requests, and stop a shared request that is already running at its next
  step. That request goes to another machine if one is free; otherwise it waits, or ends
  with an error the application can retry. A request that has already started streaming
  ends with `owner_preempted`. With your own runtime, Saylek stops passing it shared
  requests; the runtime may finish one it already has.
- **Requests through the hosted API do not stop running work**, even your own requests with
  your own key on your own machine. When requests are waiting for your machine, Saylek's
  servers offer your own first.
- **Application priority stops nothing.** It orders one account's own requests; see
  [Give important applications priority](/docs/prioritize-applications).
- **Work sent straight to your runtime's own address bypasses Saylek**, so Saylek neither
  counts nor stops it.

## Optional: limit what people you share with can use

These apply only when sharing is on. `host models` and `host slots` are terminal only.

| To | In the web | From the terminal |
|---|---|---|
| Offer only some models | Terminal only | `saylek host models MODEL_ID`; `--clear` offers all again |
| Take fewer shared requests at once | Terminal only | `saylek host slots 1`; `--clear` resets it |
| Share only at set hours | **Edit schedule** on the machine's page | `saylek host hours --set 09:00-17:00` |
| Stop shared requests for now | Turn off **Share from this machine** on the machine's page | `saylek host stop`, and `saylek host start` to share again |

Hours use the machine's own clock. A schedule has at most 5 periods, each a set of days with
one time range, or all day. A slot number above what the machine can take is capped, and
`saylek host slots` says so.

## If something goes wrong

| What you see | What to do |
|---|---|
| `all model slots busy` | Two models Saylek runs are both in use. Wait, or send fewer models at once |
| `model_load_failed` | The model may not fit next to what is loaded. Lower `keep_warm_secs` (step 2), or use a smaller model |
| The wizard says a model `may not fit` | It is an estimate. Skip it, choose another GPU, or try anyway |
| Shared requests still arrive at hours you set | Hours use the machine's clock, not yours. Check with `saylek host hours` |

## Reference

- [Sharing your models](/docs/share-your-models): sharing, pausing and stopping.
- [Use the models you already serve](/docs/use-your-existing-runtime): connecting your runtime.
- [CLI reference](/docs/cli-reference): every command.
