# Models & Launch Options

> List the on-device model catalog, pin a specific GGUF for your integration, and launch the Cephable desktop app with a model and port preselected for automated environments.

Source: https://developers.cephable.com/docs/automate-http-server/models

---
Automate runs on a local GGUF served by a shared llama.cpp server inside the Cephable app. Which model is loaded affects speed, quality, and how many tools the agent can juggle — so an integration that cares about reproducibility should pin one explicitly.

---

## Inspect the catalog

```bash
curl -sS http://127.0.0.1:4317/v1/automate/models \
  -H "Authorization: Bearer $CEPHABLE_AUTOMATE_KEY"
```

```jsonc
{
  "object": "list",
  "data": [
    {
      "name": "Gemma 4 S",
      "family": "Gemma",
      "sizeCode": "S",
      "parameterCountBillions": 1,
      "fileSizeMb": 850,
      "supportsTools": true,
      "availableForDevice": true,
      "downloaded": false,
      "selected": false
    },
    {
      "name": "Gemma 4 M",
      "family": "Gemma",
      "sizeCode": "M",
      "parameterCountBillions": 4,
      "fileSizeMb": 2650,
      "supportsTools": true,
      "availableForDevice": true,
      "downloaded": true,
      "selected": true
    }
  ],
  "catalog": { "status": "ready", "modelCount": 7 }
}
```

A model is **usable** only when all three of these hold:

| Flag | Meaning if false |
|---|---|
| `supportsTools` | Not tool-capable, so Automate cannot use it at all |
| `availableForDevice` | This machine cannot run it — RAM, GPU, or OS below requirement |
| `downloaded` | The GGUF is not on disk. Downloads happen **only in the Cephable UI**; there is no download endpoint |

The catalog syncs from Cephable asynchronously after app launch. At startup, poll until `catalog.status === "ready"` before trying to select anything:

```javascript
async function waitForCatalog(endpoint, token, timeoutMs = 120_000) {
    const deadline = Date.now() + timeoutMs;
    while (Date.now() < deadline) {
        const body = await getJson(endpoint, token, '/v1/automate/models');
        if (body.catalog.status === 'ready' && body.data.length > 0) return body.data;
        await new Promise((r) => setTimeout(r, 1000));
    }
    throw new Error('The Cephable model catalog did not become ready');
}
```

---

## Pin a model for your integration

```bash
curl -sS http://127.0.0.1:4317/v1/automate/models/select \
  -H "Authorization: Bearer $CEPHABLE_AUTOMATE_KEY" \
  -H "Content-Type: application/json" \
  -d '{ "family": "Gemma", "sizeCode": "M" }'
```

The selector takes `modelName`, `family`, and `sizeCode` in any combination, all case-insensitive. It must resolve to **exactly one** catalog entry — an ambiguous selector is an error naming the candidates, which is friendlier than silently picking one:

```jsonc
{ "error": { "message": "Model selector is ambiguous; matches: Gemma 4 M, Gemma 4 M (tools)", "type": "invalid_request_error" } }
```

### What selection actually does

- It sets a **session** preference. The user's synced account setting is untouched, and the pin is dropped when the server stops.
- It **terminates a warm worker** if one is running, so the next run pays a cold start (model load, several seconds to tens of seconds).
- It requires the workflow to be idle. A run in flight gives `409`; a non-idle-but-not-busy workflow gives `400 Cannot switch models while workflow status is …`.
- The first run after a switch verifies that the requested model really became active and fails loudly if it did not, rather than quietly benchmarking or automating against the wrong weights.

**Select once at integration startup**, not per request.

### Unpin

An empty body restores whatever the app itself would have used:

```bash
curl -sS http://127.0.0.1:4317/v1/automate/models/select \
  -H "Authorization: Bearer $CEPHABLE_AUTOMATE_KEY" \
  -H "Content-Type: application/json" -d '{}'
```

### Confirm what is loaded

`/health` reports both layers:

```jsonc
{
  "modelName": "gemma-4-4b-it-Q4_K_M.gguf",
  "automateModelSelection": { "name": "Gemma 4 M", "family": "Gemma", "sizeCode": "M", "parameterCountBillions": 4, "fileSizeMb": 2650 },
  "backend": { "flavorId": "vulkan", "accelerator": "vulkan", "cpuFallback": false },
  "contextSize": 16384
}
```

- `modelName` is the GGUF **file name** actually loaded — that is also what a run record's `model` field contains.
- `automateModelSelection` is the catalog entry you pinned, or `null` when the app's own choice is in effect.
- `backend` is the llama.cpp acceleration in use. `cpuFallback: true` means GPU acceleration failed and the run degraded to CPU, which changes latency by an order of magnitude. Record it before comparing timings.
- `contextSize` is sized from this device's RAM, so it differs machine to machine. Two runs with different `contextSize` are not directly comparable.

---

## Launch options

The Cephable executable accepts process-scoped overrides for the preferred model and the server port — useful for CI, test rigs, and kiosk or lab machines.

```powershell
Cephable.exe --model "Gemma 4 M" --port 4321
Cephable.exe --family Gemma --size M --port 4321
```

```bash
# macOS
/Applications/Cephable.app/Contents/MacOS/Cephable --family Gemma --size M --port 4321
```

| Argument | Aliases | Meaning |
|---|---|---|
| `--model` | | Exact catalog model name |
| `--family` | | Model family, e.g. `Gemma` |
| `--size` | `--size-code` | Size code, e.g. `S`, `M`, `L` |
| `--port` | `--automate-http-port` | Port the Automate server should bind first |

Both `--name value` and `--name=value` forms work. Selectors are case-insensitive and must resolve to exactly one catalog model.

### Rules that matter

- **They do not bypass access control.** A Professional user must still enable **Automate HTTP Server** in Extensions. Launch arguments never enable the server, grant a license, or skip the secure-storage check.
- **The access key is not accepted as a launch argument.** Operating systems expose process command lines to other local processes, so passing it there would leak it. Read it from the app and pass it to your client out of band.
- **They override the process, not the account.** The synced model preference is unchanged; the override lasts for the running process.
- **A second launch reconfigures the running instance.** Cephable is single-instance: launching it again with new arguments applies them to the existing app, including rebinding the server when the port changes.
- **Failures surface on `/health`.** If a model selector could not be applied — misspelled name, model not downloaded, not device-compatible — `launchModelSelectionError` says why, and the app keeps its previous selection.

### Waiting for the server after launch

The app takes a while to start, sync the catalog, and bind. Poll `/health` rather than sleeping a fixed amount:

```javascript
async function waitForServer(endpoint, token, timeoutMs = 180_000) {
    const deadline = Date.now() + timeoutMs;
    let lastError = 'no response';
    while (Date.now() < deadline) {
        try {
            const response = await fetch(`${endpoint}/health`, { headers: { authorization: `Bearer ${token}` } });
            if (response.status === 401) throw new Error('The access key was rejected');
            if (response.ok) return await response.json();
        } catch (error) {
            lastError = error instanceof Error ? error.message : String(error);
        }
        await new Promise((r) => setTimeout(r, 1000));
    }
    throw new Error(`Cephable did not expose the Automate server at ${endpoint}: ${lastError}`);
}
```

---

## Choosing a model

| | Smaller (S) | Larger (M / L) |
|---|---|---|
| Latency | Fast; good for single-tool tasks, rewriting, classification | Slower per token, but usually fewer wasted tool calls |
| Multi-step reliability | Drifts on long chains | Holds a plan across more steps |
| Context | Same device-sized window | Same device-sized window, but consumed faster by longer reasoning |
| Fit | Short prompts, high call volume, low-power machines | Research, multi-file work, anything with branching |

Two practical notes:

1. **Shrink the task before shrinking the model.** A well-scoped prompt on a small model beats a vague prompt on a large one, and costs a fraction of the time.
2. **Pin the model *and* record `contextSize` and `backend`** in anything you intend to compare over time. The model name alone does not determine performance on a given machine.
