# OpenAI Compatibility

> Point an existing OpenAI-compatible client or evaluation framework at Cephable's local Automate assistant using POST /v1/chat/completions — the request mapping, the cephable extension object, and exactly which parameters are ignored.

Source: https://developers.cephable.com/docs/automate-http-server/openai-compatibility

---
The Automate HTTP Server exposes a non-streaming `POST /v1/chat/completions` and a `GET /v1/models`. This lets any tool that accepts an OpenAI base URL and API key drive the full Cephable assistant with no code changes.

```
Base URL   http://127.0.0.1:4317/v1
API key    your Automate access key
Model      cephable-agent
```

> Each chat-completion request runs the **complete assistant loop** — planning, tool calls, file and web work, the lot. It is not a text completion against the GGUF. Expect tens of seconds to minutes per request, and raise your client's timeout accordingly.

---

## Minimal request

```bash
curl -sS http://127.0.0.1:4317/v1/chat/completions \
  -H "Authorization: Bearer $CEPHABLE_AUTOMATE_KEY" \
  -H "Content-Type: application/json" \
  -d '{
        "model": "cephable-agent",
        "messages": [{ "role": "user", "content": "Summarize the newest file in my Downloads folder." }]
      }'
```

```jsonc
{
  "id": "automate-run-9c2e…",
  "object": "chat.completion",
  "created": 1789412568,
  "model": "cephable-agent",
  "choices": [
    { "index": 0, "message": { "role": "assistant", "content": "The newest file is…" }, "finish_reason": "stop" }
  ],
  "usage": { "prompt_tokens": 5120, "completion_tokens": 344, "total_tokens": 5464 },
  "cephable": { "schemaVersion": 1, "status": "completed", "steps": [ ] }
}
```

The `cephable` field carries the **entire native run record** — status, steps, trace, usage, backend, timings. Read it when you need to know what the agent did, not just what it said.

---

## How the request is mapped

| OpenAI field | Mapped to | Notes |
|---|---|---|
| **last** `messages[]` entry with `role: "user"` | `prompt` | Must have **string** content. Content-part arrays are not supported |
| all `messages[]` with `role: "system"` | `additionalWorkflowPrompt` | Joined with a blank line, in order. Advisory preferences, as in the app |
| `timeout_ms` | `timeoutMs` | Non-standard but accepted. Defaults to 900000 |
| `tools` | `clientTools` | **Supported.** `type: "function"` entries become tools you execute yourself; the run parks with `finish_reason: "tool_calls"`. See [Custom tools](/docs/automate-http-server/custom-tools) |
| `messages[]` with `role: "tool"` | tool results | **Supported.** A `tool_call_id` Cephable minted resumes the parked run it belongs to |
| `cephable.answerContract` | `answerContract` | See below |
| `cephable.include` | `include` | `{ steps?, trace?, events? }` |
| `cephable.continuation` | `continuation` | Continue the previous conversation |
| `cephable.thinkingLevel` | `thinkingLevel` | `"low"`/`"medium"`/`"high"`/`"max"`; anything else is ignored |
| `cephable.mcpServers` | `mcpServers` | Inline MCP servers for this run — see [Custom tools](/docs/automate-http-server/custom-tools) |
| `model` | — | **Ignored.** There is one assistant; the value is not validated |

### Earlier `user` and `assistant` messages are dropped

Only the **last** user message becomes the prompt. Prior turns in the `messages` array are not replayed — Cephable owns conversation state itself.

If your client sends chat history, either:

- fold the context you need into the last user message, or
- set `cephable.continuation: true` and send only the new turn, letting Cephable continue its own thread.

The one exception is a **tool-result turn**: when the array carries `role: "tool"` messages whose `tool_call_id`s Cephable minted for a run that is still parked, the request resumes that run rather than starting a new one — which is exactly what a tool-calling client does without knowing anything Cephable-specific. If those ids belong to no live run (it timed out, or was cancelled), the request falls through and starts a fresh run.

A run whose latest user message has non-string content (an array of content parts, an image) is rejected with `400 The latest user message must contain string content`.

---

## The `cephable` extension object

Evaluation frameworks that only speak OpenAI carry Cephable-specific options here:

```jsonc
{
  "model": "cephable-agent",
  "messages": [{ "role": "user", "content": "How many PDFs are in my Documents folder?" }],
  "timeout_ms": 600000,
  "cephable": {
    "answerContract": "Reply with exactly: FINAL ANSWER: <integer>",
    "include": { "trace": false, "events": false },
    "continuation": false
  }
}
```

When `answerContract` is set, `choices[0].message.content` is the **extracted** `finalAnswer` rather than the full prose — a client that declared a contract asked for the value, not the reasoning around it. The untouched answer is still available as `cephable.answer`.

### What the extension object does **not** carry

These native options have no chat-completions equivalent. Use [`POST /v1/runs`](/docs/automate-http-server/runs) if you need them:

- `restrictToWorkspace`
- `selectedSkillIds` / `selectedMcpServerIds`
- `hitlAnswers`
- `allowDestructiveTools`
- `taskId`

---

## Unsupported OpenAI parameters

Everything below is accepted by the JSON parser and then ignored. Nothing errors, so a client that depends on them will silently get default behavior.

| Parameter | Behavior |
|---|---|
| `stream` | **Ignored.** The response is always a single non-streaming `chat.completion` object. There is no SSE endpoint |
| `functions` (the pre-`tools` shape) / `tool_choice` | Ignored. Use `tools`, which **is** supported; the agent decides when to call, so it cannot be forced or restricted per request |
| `temperature`, `top_p`, `top_k`, `seed` | Ignored. Sampling is governed by the app's model profile |
| `max_tokens`, `max_completion_tokens` | Ignored |
| `n` | Ignored — always exactly one choice |
| `stop`, `logit_bias`, `logprobs`, `presence_penalty`, `frequency_penalty` | Ignored |
| `response_format` (JSON mode / schema) | Ignored. Use `cephable.answerContract` instead |
| `user`, `metadata` | Ignored — use `/v1/runs` with `taskId` for correlation |

There are no other OpenAI routes. `/v1/completions`, `/v1/embeddings`, `/v1/responses`, and `/v1/chat/completions/{id}` all return `404`.

---

## Status codes

| Status | Meaning |
|---|---|
| `200` | The run completed — or parked on your tools, with `finish_reason: "tool_calls"` |
| `500` | The run ran and ended `failed`, `canceled`, or `terminated`. The body is a **chat-completion-shaped** object whose `cephable` field holds the full record — check `cephable.status` and `cephable.errorCode` |
| `400` | Invalid request, no string user message, oversized body, or run timeout — an `{ error }` envelope, not a chat completion |
| `401` | Bad or missing key |
| `409` | A run is already in flight. Cancel it or wait — see [Quick start](/docs/automate-http-server/quick-start) |

A `409` is the one most OpenAI clients handle badly, because the OpenAI API never returns it for this reason. Wrap your client or add a readiness gate on `/health`.

---

## Using it from the official OpenAI SDKs

A complete, runnable version of the Python example below — including the tool-calling loop and a working
`@tool`-decorated tool set — is the
[python-langchain-tools sample](https://github.com/Cephable/Cephable-Automate-Agent-Samples/tree/main/samples/python-langchain-tools).
It also shows how to read Cephable's native run record back out, which `langchain-openai` drops.

### Python

```python
import os
from openai import OpenAI

client = OpenAI(
    base_url="http://127.0.0.1:4317/v1",
    api_key=os.environ["CEPHABLE_AUTOMATE_KEY"],
    timeout=900.0,   # agent runs are long; the SDK default is far too short
    max_retries=0,   # never auto-retry a run: a 409 means one is already going
)

completion = client.chat.completions.create(
    model="cephable-agent",
    messages=[{"role": "user", "content": "List the three largest files in my Downloads folder."}],
    extra_body={"cephable": {"include": {"trace": False, "events": False}}},
)

print(completion.choices[0].message.content)

record = completion.model_extra.get("cephable", {})
print(record.get("status"), len(record.get("steps", [])), "steps")
```

### Node.js

```typescript
import OpenAI from 'openai';

const client = new OpenAI({
    baseURL: 'http://127.0.0.1:4317/v1',
    apiKey: process.env.CEPHABLE_AUTOMATE_KEY!,
    timeout: 900_000,
    maxRetries: 0,
});

const completion = await client.chat.completions.create({
    model: 'cephable-agent',
    messages: [{ role: 'user', content: 'Summarize my newest meeting notes file.' }],
    // @ts-expect-error — Cephable extension object
    cephable: { include: { trace: false, events: false } },
});

console.log(completion.choices[0].message.content);
```

Two settings matter more than anything else with the SDKs:

- **`maxRetries: 0`.** The default retry behavior will fire a second request at a busy server and get `409`, or worse, start a duplicate run.
- **A very long timeout.** The SDK defaults (around 10 minutes for Python, 10 for Node) can be shorter than a real research run, and an SDK abort leaves the run executing inside Cephable. Pair a long client timeout with `POST /v1/automate/cancel` in your error path.

---

## When to use `/v1/runs` instead

Reach for the native contract when you want any of:

- `restrictToWorkspace`, `allowDestructiveTools`, `hitlAnswers`, explicit skills or configured MCP servers
- `taskId` correlation
- the run record without unwrapping a chat completion
- clear separation between "the request failed" and "the run failed" — `/v1/runs` always returns `schemaVersion: 1` in the second case

Use `/v1/chat/completions` when the value is *not having to write a client at all*: an existing eval framework, an LLM gateway, a plugin that only accepts an OpenAI base URL.
