> ## Documentation Index
> Fetch the complete documentation index at: https://docs.squarecloud.app/llms.txt
> Use this file to discover all available pages before exploring further.

# Python SDK: AI Gateway chat

> Call the Square Cloud AI Gateway from squarecloud-api with client.ai.chat(): OpenAI-compatible chat completions, without streaming.

`client.ai.chat(request)` calls the [AI Gateway](/en/api-reference/ai-gateway), an **OpenAI-compatible** chat completions endpoint. It needs the `ai:chat` scope and a Standard plan or above.

Examples use the `client` from [Creating the client](/en/sdks/py/client#creating-the-client).

```python theme={"system"}
completion = client.ai.chat({
    "model": "cubic",
    "messages": [
        {"role": "system", "content": "You are a helpful assistant."},
        {"role": "user", "content": "What is Square Cloud?"},
    ],
    "max_tokens": 512,
})

print(completion["choices"][0]["message"]["content"])
print(completion["usage"]["total_tokens"])
```

## Request

`request` is a dict sent as the body, as is, so any other OpenAI parameter passes through. It is typed as `squarecloud.types.ChatRequest`.

| Field | Description |
| - | - |
| `model` | `cubic`. Accepted for compatibility: there is no other model |
| `messages` | `{ role, content?, tool_call_id?, tool_calls? }`, with `role` `system`, `user`, `assistant` or `tool` |
| `tools` / `tool_choice` | OpenAI function calling |
| `max_tokens` | Maximum tokens of the answer |
| `temperature` | Sampling temperature |

The response has the OpenAI shape: `id`, `object`, `created`, `model`, `choices` (`index`, `message`, `finish_reason`) and `usage`.

## No streaming

`ai.chat()` does not stream: it returns the whole completion. `"stream": True` is 400 `stream_not_supported`.

## Timeout

The gateway gives each request **90 seconds** in total, then answers 503 `server_overloaded`. The SDK waits at least 120 s before timing out, so you get the gateway's answer.

## Errors

AI errors use the OpenAI format, so their codes are **lowercase**. They still raise a [`SquareCloudAPIError`](/en/sdks/py/errors), with `code` set to the OpenAI code (or its `type` when there is no code):

```python theme={"system"}
from squarecloud import SquareCloudAPIError

try:
    client.ai.chat({"messages": [{"role": "user", "content": "Hi"}]})
except SquareCloudAPIError as error:
    if error.code == "server_overloaded":
        ...  # safe to retry yourself
    else:
        raise
```

| Status | Code | When |
| - | - | - |
| 400 | `stream_not_supported` | `stream: true` was sent |
| 400 | `invalid_messages`, `invalid_tools`, `invalid_tool_choice`, `invalid_temperature`, `invalid_max_tokens` | A malformed field |
| 400 | `context_length_exceeded` | The conversation exceeds the plan's context window |
| 401 | `access_denied` | Invalid API key |
| 403 | `upgrade_required` | The plan has no AI Gateway access |
| 429 | `rate_limit_exceeded`, `concurrent_limit_reached` | Too fast, or a request already in flight |
| 429 | `daily_limit_reached`, `daily_request_limit_reached`, `daily_spend_limit_reached` | A daily budget is used up (resets at 00:00 UTC) |
| 503 | `server_overloaded` | Capacity is full or the 90 s deadline passed: safe to retry |
| 503 | `daily_capacity_reached` | The platform's daily capacity is used up |

The SDK **never retries** AI errors: the request is a non-idempotent `POST`. See the [AI Gateway reference](/en/api-reference/ai-gateway) for plan limits.

## Next steps

<CardGroup cols={2}>
  <Card title="Errors" icon="triangle-exclamation" href="/en/sdks/py/errors">
    Error class, retries and rate limits.
  </Card>

  <Card title="AI Gateway reference" icon="code" href="/en/api-reference/ai-gateway">
    Models, plan limits and the OpenAI-compatible endpoint.
  </Card>
</CardGroup>
