Use the models you already serve
Connect the runtime you already run (Ollama, LM Studio, llama.cpp or vLLM), and call its models from compatible applications with a Saylek API key, on this machine or anywhere else. Your runtime keeps loading the models and managing the GPU.
Before you begin
| You need | Detail |
|---|---|
| Your runtime, running | On the machine you are connecting, listening on its own address (127.0.0.1). Default ports: Ollama 11434, LM Studio 1234, llama.cpp 8080, vLLM 8000. |
| A supported machine | Linux x86_64, or macOS on Apple Silicon. On Windows, the Linux build inside WSL2. |
| A Saylek account | Create one at saylek.com with your email. |
1. Connect the machine and its runtime
On the machine that runs your runtime:
curl -fsSL https://saylek.com/install.sh | sh -s -- --connect
This signs you in through your browser, links the machine to your account, and connects a
runtime it finds answering on one of the default ports. It does not share the machine with
anyone. Installed earlier without --connect? Run the line again; it is safe to repeat.
You know it worked when the installer says the machine is connected, and in Machines its Models section lists your runtime's models.
If the installer did not find your runtime, for example because it uses another port, connect it on the machine with its address and port:
saylek model add http://127.0.0.1:11434 # Ollama. vLLM is usually :8000
It prints [ok] register endpoint when it connects. If the runtime needs its own API key,
put the key in an environment variable and add --token-env VARIABLE_NAME. For a runtime
you started after installing, you can also open the machine in Machines, then
Settings, Manage sources, and choose Scan again; the machine's Models section then
shows the same command, with Copy command.
2. Check that your models are ready
In Machines, open the machine. Its Models section shows each of your runtime's models as Ready, and the page says Only you can use its models. Your own keys can use them with sharing off.
From the terminal, saylek models lists the same models, each with your runtime's address.
3. Connect an application
-
Open Applications, choose Add application, name it, and choose Create key. Copy the key now: it is shown once.
-
List the models your key can use, and pick one of your runtime's model ids:
curl https://api.saylek.com/v1/models -H "Authorization: Bearer YOUR_API_KEY"
You know it worked when your runtime's models are in the list.
4. Send your first request
curl https://api.saylek.com/v1/chat/completions \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "MODEL_ID", "messages": [{"role": "user", "content": "Say hi in one sentence."}]}'
You know it worked when the response carries an answer. To set up your own client,
follow Connect your app: OpenAI-compatible clients use
https://api.saylek.com/v1, and Claude Code and Anthropic clients use
https://api.saylek.com with no /v1.
Where your prompt goes. Hosted requests pass through Saylek's servers, even when your own machine serves them. If someone who shares with you serves the same model, their machine can answer instead of yours. Privacy and egress has the details.
If something goes wrong
| What you see | What to do |
|---|---|
The installer ends with waiting for this machine to connect | Run saylek status on the machine; Machines updates when it connects |
saylek models says no models loaded yet | Start the runtime, then saylek model add <address> |
missing port | Add the port to the address, for example :11434 |
connection refused or upstream unreachable | Start the runtime, or give the port it actually listens on |
authentication required | The runtime expects a key: add --token-env VARIABLE_NAME |
| The machine's page says your runtime "is not answering", or no longer lists its models | Start the runtime again, then check with saylek models |
| Your models are missing from the hosted list | Check the machine is online in Machines, or run step 1 again with --connect |
| Only some of them are missing | An earlier selection leaves them out: saylek host models --clear |
| The daemon is not running | Start it as Quickstart step 2 describes |
Troubleshooting covers the rest.
Optional: use Saylek on this machine without a key
Saylek also runs a local API on the machine, for applications running there:
| Setting | Value |
|---|---|
| Base URL | http://127.0.0.1:8443/v1 |
| API key | Any non-empty string. The local API does not check it. |
Because it has no authentication, anything that can reach that port on the machine can use
it. It speaks the OpenAI format only. saylek call sends a one-line test through it and
prints [ok] <model> responded.
Use it instead of your runtime's own address for two things. A request your runtime cannot serve can go to another machine you can use. And your own work goes first: while work sent through the local API needs the runtime, Saylek stops passing hosted requests to it, your own included, and those wait or are stopped part-way. Saylek does not unload the runtime's models or free its memory. Work sent to the runtime's own address bypasses Saylek.
Reference
- Connect your app: OpenAI-compatible and Claude Code setup.
- Sharing your models: sharing, pausing, and stopping.
- Privacy and egress: what leaves your machine.
- CLI reference: every command.
Last checked 2026-09-24 · read as markdown at /docs/use-your-existing-runtime.md