Skip to main content
string
required
The API key for your account. You can find this in your account settings.
Requires an API key with the ai:chat scope. The AI Gateway lets you use Square Cloud’s hosted AI model inside your own products (chatbots, assistants, automations) through an OpenAI-compatible chat completions endpoint. If your code already speaks the OpenAI API, it speaks the AI Gateway too: point the SDK at our base URL, use your account API key, and set the model to cubic.
The AI Gateway is in beta with free early access: during the beta there is no extra charge, and usage counts only against the gateway’s own token limits. Limits and plan availability may change when the beta ends.

How to connect

  • Base URL: https://api.squarecloud.app/v2/ai
  • API key: your account API key (the same one used by the Square Cloud API and CLI, from the account page). The Authorization header is accepted with or without the Bearer prefix.
  • Model: cubic, Square Cloud’s hosted model, the same one powering the dashboard AI assistant.

Plans and limits

Every request bills its real tokens against the gateway’s own token limits: a daily limit, reset at 00:00 UTC, and a weekly limit of 4x the daily one, reset on Monday at 00:00 UTC. They are separate from the dashboard AI assistant’s limits, so gateway traffic never uses up the assistant’s allowance, and vice versa. Standard and Pro also have a daily request ceiling, and every plan has a daily fair-use allowance on the gateway, sized for moderate use (a guard for unattended keys, based on what the requests cost to serve); both reset at 00:00 UTC. Every Standard plan size has the same gateway limits, and so does every Pro size. Enterprise token limits grow with the plan: 10M tokens per day on Enterprise 32, 12M on Enterprise 48, 14M on Enterprise 64, 16M on Enterprise 96 and 20M on Enterprise 128 and above, with 4x that per week.
Hobby plans (and accounts without a plan) do not have AI Gateway access: upgrading to Standard or above enables it.

Request contract

string
Accepted for SDK compatibility; the gateway always answers as cubic.
array
required
OpenAI-style messages with roles system, user, assistant and tool. Content must be a string: this endpoint is text only, so a multi-part array (image_url, input_audio, file) is rejected with 400 multimodal_not_supported. Up to 100 messages per request.
number
OpenAI’s current name for the answer’s token limit, the one the OpenAI SDKs send. It wins over max_tokens when both are present. Silently clamped to your plan’s output ceiling.
number
The legacy name, still accepted. Silently clamped to your plan’s output ceiling.
number
From 0 to 2.
array
Function calling in the standard OpenAI format, up to 32 tool definitions. Tool calls come back as finish_reason: "tool_calls", and you send results back as role: "tool" messages.
string | object
Standard OpenAI tool_choice values.
Unknown parameters are ignored, with two deliberate exceptions: audio and a modalities value other than ["text"] are rejected rather than ignored, so a request for spoken output never comes back as silent text. Text in, text out. Images, audio and files are not accepted on any plan, and this is a product decision rather than a temporary gap. Send text and read text back. Not supported yet: streaming (stream: true returns 400 stream_not_supported). The model decides on its own to search the web when the conversation asks for current or external information. Searches run server-side, and only the final answer is returned, citing result URLs. If you declare your own tool named web_search, yours replaces the built-in one.

Errors

Every error uses the OpenAI shape { "error": { "message", "type", "param", "code" } }, including authentication, scope, rate-limit, invalid-body, 500 and 503 errors, and the code is always lowercase. Three of these look alike but mean different things:
  • 429 daily_spend_limit_reached: the daily limit of your key, set by your plan.
  • 503 daily_capacity_reached: the daily capacity of the platform for the gateway, shared by every customer.
  • 503 server_overloaded: a momentary overload, or a request that took longer than 90 seconds in total (provider queue, retries and search rounds included). Retry in a few seconds.

Frequently asked questions

No, the API is stateless, exactly like OpenAI’s: send the conversation history in messages on every request.
cubic, Square Cloud’s hosted model. There is no model list to choose from; the model field is accepted for SDK compatibility.
Yes, that is the main use case. It is a normal HTTPS API, so it works from anywhere.
No. The gateway has its own daily and weekly token limits, separate from the dashboard assistant’s, so heavy gateway usage never uses up the assistant’s allowance, and vice versa.