Skip to content

Start without a runtime

No model runtime on your machine? Saylek can download one starter model, run it on your GPU, and answer your applications with a Saylek API key.

Before you begin

You needDetail
A machine that can serveLinux x86_64 with an NVIDIA GPU and NVIDIA driver 555.42.02 or newer (CUDA 12.5); Windows with WSL2 (x86_64) and an NVIDIA GPU, with NVIDIA driver 555.85 or newer installed in Windows; or a Mac with Apple Silicon. A machine without one of these cannot run the starter model. It can still serve a runtime you already run (Use the models you already serve) and use models shared with you.
A terminal on that machineThe download asks for your consent there. The browser cannot start it. On Windows, use your WSL2 terminal, not PowerShell.
A Saylek accountCreate one at saylek.com with your email.

1. Connect the machine

curl -fsSL https://saylek.com/install.sh | sh -s -- --connect

This signs you in through your browser and links the machine to your account. It does not download a model. With none on the machine, the installer ends with no model on this machine yet and points you to saylek gpu wizard.

2. Download the starter model

On the machine:

saylek gpu wizard

After its setup steps, and only when nothing on the machine can serve a model yet, the wizard offers one starter model and asks download it now? [y/N]. Answer y. It picks the largest Qwen3.5 model that fits your largest GPU, or a Mac's memory, from 0.8B (about 0.5 GB) to 27B (about 16.7 GB). The file comes from huggingface.co and is checked against a pinned checksum. Saylek downloads nothing without this answer.

In Machines, Add model on the machine's page shows the same command under Starting without a model?.

You know it worked when the wizard prints [ok] downloaded + registered.

3. Check that the model is ready

In Machines, open the machine. The model shows as Ready, labelled Saylek-managed. From the terminal, saylek models lists it.

The first request loads the model, so it takes longer than the rest. On Linux, the first load also downloads Saylek's inference runtime once and verifies its signature.

4. Connect an application and send a request

In Applications, choose Add application, name it, and choose Create key. Copy the key now: it is shown once. Then send a test request, using the model id from the models list:

curl https://api.saylek.com/v1/models -H "Authorization: Bearer YOUR_API_KEY"

curl https://api.saylek.com/v1/chat/completions \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "MODEL_ID", "messages": [{"role": "user", "content": "Say hi in one sentence."}]}'

You know it worked when the response carries an answer. To set up your own client, follow Connect your app. Hosted requests pass through Saylek's servers; Privacy and egress says what that means for your prompt.

Add another model later

Saylek does not download models on its own. To serve another, register a model file you have:

saylek model add /absolute/path/to/model.gguf

Or copy a .gguf file into ~/.saylek/models/ and run saylek model rescan. At most two models stay loaded at once; Make more of one GPU explains.

If something goes wrong

What you seeWhat to do
The wizard does not offer a downloadRun it in an interactive terminal. If you did, something on the machine can already serve: check with saylek models, or follow Use the models you already serve
On Linux without an NVIDIA GPU, the wizard says the machine will serve on its CPU, or offers a model to run on the CPUSaylek cannot run a model on a Linux CPU. Answer N; connect a runtime you already run, or use models shared with you
download did not completeRun saylek gpu wizard again. saylek help models shows how to fetch a model yourself
checksum mismatch or size mismatchThe partial file was removed. Run saylek gpu wizard again
NVIDIA driver not detectedOn a machine with an NVIDIA GPU, install its driver (under WSL2, in Windows), then retry. The message may start with "no free model slot"
no published inference runtime for GPU archEither the driver supports too old a CUDA version, or Saylek has no runtime for this GPU model yet. Update the driver and retry; if the message stays, this GPU cannot serve yet
the inference runtime needs glibcThe machine needs a newer Linux release: Ubuntu 24.04+, Debian 13+ or Fedora 40+
could not fetch the inference runtimeA network problem. Retry
model_load_failedThe model did not load, often because it does not fit. saylek call on the machine, or ~/.saylek/log/, shows the reason

Troubleshooting covers the rest.

Reference

Last checked 2026-09-25 · read as markdown at /docs/start-without-a-runtime.md