OpenAI-compatible AI Gateway (Beta)
Use Square Cloud’s hosted cubic model in your own products through the OpenAI-compatible POST /v2/ai/chat/completions, with tool calling and web search.
string
required
The API key for your account. You can find this in your account settings.
ai:chat scope.
The AI Gateway lets you use Square Cloud’s hosted AI model inside your own products (chatbots, assistants, automations) through an OpenAI-compatible chat completions endpoint. If your code already speaks the OpenAI API, it speaks the AI Gateway too: point the SDK at our base URL, use your account API key, and set the model to cubic.
The AI Gateway is in beta with free early access: during the beta there is no extra charge, and usage counts only against the gateway’s own token limits. Limits and plan availability may change when the beta ends.
How to connect
- Base URL:
https://api.squarecloud.app/v2/ai - API key: your account API key (the same one used by the Square Cloud API and CLI, from the account page). The
Authorizationheader is accepted with or without theBearerprefix. - Model:
cubic, Square Cloud’s hosted model, the same one powering the dashboard AI assistant.
Plans and limits
Every request bills its real tokens against the gateway’s own token limits: a daily limit, reset at 00:00 UTC, and a weekly limit of 4x the daily one, reset on Monday at 00:00 UTC. They are separate from the dashboard AI assistant’s limits, so gateway traffic never uses up the assistant’s allowance, and vice versa. Standard and Pro also have a daily request ceiling, and every plan has a daily fair-use allowance on the gateway, sized for moderate use (a guard for unattended keys, based on what the requests cost to serve); both reset at 00:00 UTC.
Every Standard plan size has the same gateway limits, and so does every Pro size. Enterprise token limits grow with the plan: 10M tokens per day on Enterprise 32, 12M on Enterprise 48, 14M on Enterprise 64, 16M on Enterprise 96 and 20M on Enterprise 128 and above, with 4x that per week.
Hobby plans (and accounts without a plan) do not have AI Gateway access: upgrading to Standard or above enables it.
Request contract
string
Accepted for SDK compatibility; the gateway always answers as
cubic.array
required
OpenAI-style messages with roles
system, user, assistant and tool. Content must be a string: this endpoint is text only, so a multi-part array (image_url, input_audio, file) is rejected with 400 multimodal_not_supported. Up to 100 messages per request.number
OpenAI’s current name for the answer’s token limit, the one the OpenAI SDKs send. It wins over
max_tokens when both are present. Silently clamped to your plan’s output ceiling.number
The legacy name, still accepted. Silently clamped to your plan’s output ceiling.
number
From 0 to 2.
array
Function calling in the standard OpenAI format, up to 32 tool definitions. Tool calls come back as
finish_reason: "tool_calls", and you send results back as role: "tool" messages.string | object
Standard OpenAI
tool_choice values.audio and a modalities value other than ["text"] are rejected rather than ignored, so a request for spoken output never comes back as silent text.
Text in, text out. Images, audio and files are not accepted on any plan, and this is a product decision rather than a temporary gap. Send text and read text back.
Not supported yet: streaming (stream: true returns 400 stream_not_supported).
Built-in web search
The model decides on its own to search the web when the conversation asks for current or external information. Searches run server-side, and only the final answer is returned, citing result URLs. If you declare your own tool namedweb_search, yours replaces the built-in one.
Errors
Every error uses the OpenAI shape{ "error": { "message", "type", "param", "code" } }, including authentication, scope, rate-limit, invalid-body, 500 and 503 errors, and the code is always lowercase.
Three of these look alike but mean different things:
429 daily_spend_limit_reached: the daily limit of your key, set by your plan.503 daily_capacity_reached: the daily capacity of the platform for the gateway, shared by every customer.503 server_overloaded: a momentary overload, or a request that took longer than 90 seconds in total (provider queue, retries and search rounds included). Retry in a few seconds.
Frequently asked questions
Does it remember conversations?
Does it remember conversations?
No, the API is stateless, exactly like OpenAI’s: send the conversation history in
messages on every request.Which model is it?
Which model is it?
cubic, Square Cloud’s hosted model. There is no model list to choose from; the model field is accepted for SDK compatibility.Can I call it from an app hosted on Square Cloud?
Can I call it from an app hosted on Square Cloud?
Yes, that is the main use case. It is a normal HTTPS API, so it works from anywhere.
Does gateway usage affect my dashboard AI assistant?
Does gateway usage affect my dashboard AI assistant?
No. The gateway has its own daily and weekly token limits, separate from the dashboard assistant’s, so heavy gateway usage never uses up the assistant’s allowance, and vice versa.
Related
- SDKs:
api.ai.chat()(JavaScript),client.ai.chat()(Python),c.AI.Chat()(Go)

