# Use the models you already serve

Connect the runtime you already run (Ollama, LM Studio, llama.cpp or vLLM), and call its models from
compatible applications with a Saylek API key, on this machine or anywhere else. Your runtime keeps loading the
models and managing the GPU.

## Before you begin

| You need | Detail |
|---|---|
| Your runtime, running | On the machine you are connecting, listening on its own address (`127.0.0.1`). Default ports: Ollama 11434, LM Studio 1234, llama.cpp 8080, vLLM 8000. |
| A supported machine | Linux x86_64, or macOS on Apple Silicon. On Windows, the Linux build inside WSL2. |
| A Saylek account | Create one at [saylek.com](https://saylek.com) with your email. |

## 1. Connect the machine and its runtime

On the machine that runs your runtime:

```bash
curl -fsSL https://saylek.com/install.sh | sh -s -- --connect
```

This signs you in through your browser, links the machine to your account, and connects a
runtime it finds answering on one of the default ports. It does not share the machine with
anyone. Installed earlier without `--connect`? Run the line again; it is safe to repeat.

**You know it worked when** the installer says the machine is connected, and in
[Machines](/machines) its Models section lists your runtime's models.

**If the installer did not find your runtime**, for example because it uses another port,
connect it on the machine with its address and port:

```bash
saylek model add http://127.0.0.1:11434   # Ollama. vLLM is usually :8000
```

It prints `[ok] register endpoint` when it connects. If the runtime needs its own API key,
put the key in an environment variable and add `--token-env VARIABLE_NAME`. For a runtime
you started after installing, you can also open the machine in [Machines](/machines), then
**Settings**, **Manage sources**, and choose **Scan again**; the machine's Models section then
shows the same command, with **Copy command**.

## 2. Check that your models are ready

In [Machines](/machines), open the machine. Its Models section shows each of your runtime's
models as **Ready**, and the page says **Only you can use its models.** Your own keys can use
them with sharing off.

From the terminal, `saylek models` lists the same models, each with your runtime's address.

## 3. Connect an application

1. Open [Applications](/applications), choose **Add application**, name it, and choose
   **Create key**. Copy the key now: it is shown once.
2. List the models your key can use, and pick one of your runtime's model ids:

   ```bash
   curl https://api.saylek.com/v1/models -H "Authorization: Bearer YOUR_API_KEY"
   ```

**You know it worked when** your runtime's models are in the list.

## 4. Send your first request

```bash
curl https://api.saylek.com/v1/chat/completions \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "MODEL_ID", "messages": [{"role": "user", "content": "Say hi in one sentence."}]}'
```

**You know it worked when** the response carries an answer. To set up your own client,
follow [Connect your app](/docs/connect-your-app): OpenAI-compatible clients use
`https://api.saylek.com/v1`, and Claude Code and Anthropic clients use
`https://api.saylek.com` with no `/v1`.

**Where your prompt goes.** Hosted requests pass through Saylek's servers, even when your own
machine serves them. If someone who shares with you serves the same model, their machine can
answer instead of yours. [Privacy and egress](/docs/privacy-and-egress) has the details.

## If something goes wrong

| What you see | What to do |
|---|---|
| The installer ends with `waiting for this machine to connect` | Run `saylek status` on the machine; Machines updates when it connects |
| `saylek models` says `no models loaded yet` | Start the runtime, then `saylek model add <address>` |
| `missing port` | Add the port to the address, for example `:11434` |
| `connection refused` or `upstream unreachable` | Start the runtime, or give the port it actually listens on |
| `authentication required` | The runtime expects a key: add `--token-env VARIABLE_NAME` |
| The machine's page says your runtime "is not answering", or no longer lists its models | Start the runtime again, then check with `saylek models` |
| Your models are missing from the hosted list | Check the machine is online in Machines, or run step 1 again with `--connect` |
| Only some of them are missing | An earlier selection leaves them out: `saylek host models --clear` |
| The daemon is not running | Start it as [Quickstart](/docs/quickstart) step 2 describes |

[Troubleshooting](/docs/troubleshooting) covers the rest.

## Optional: use Saylek on this machine without a key

Saylek also runs a local API on the machine, for applications running there:

| Setting | Value |
|---|---|
| Base URL | `http://127.0.0.1:8443/v1` |
| API key | Any non-empty string. The local API does not check it. |

Because it has no authentication, anything that can reach that port on the machine can use
it. It speaks the OpenAI format only. `saylek call` sends a one-line test through it and
prints `[ok] <model> responded`.

Use it instead of your runtime's own address for two things. A request your runtime cannot
serve can go to another machine you can use. And your own work goes first: while work sent
through the local API needs the runtime, Saylek stops passing hosted requests to it, your own
included, and those wait or are stopped part-way. Saylek does not unload the runtime's models
or free its memory. Work sent to the runtime's own address bypasses Saylek.

## Reference

- [Connect your app](/docs/connect-your-app): OpenAI-compatible and Claude Code setup.
- [Sharing your models](/docs/share-your-models): sharing, pausing, and stopping.
- [Privacy and egress](/docs/privacy-and-egress): what leaves your machine.
- [CLI reference](/docs/cli-reference): every command.
