Documentation
Three models, one endpoint, OpenAI-compatible format. Start with free test credit, create your key and integrate over HTTP using the examples below.
https://api.nexichat.ai/v1Introduction
#The Nexi API exposes the same three models that power the chat. A single endpoint, POST https://api.nexichat.ai/v1/chat/{model}, where the model is air, prime or nova.
Requests and responses follow OpenAI's chat completions format. NexiChat uses its own path, such as /chat/prime. Use the HTTP examples below; changing only base_url in an SDK that calls /chat/completions is not sufficient.
The OpenAPI 3.1 specification — operations, schemas, error codes and how to resolve each — is at https://api.nexichat.ai/v1/openapi.json. Use it to generate a client, or let an agent read the API without this page.
model field in the body. Send that field and it is ignored — no error, no effect.First request
#Create a key in the console without contacting sales. The first account receives 1.00 € of free test credit, with no initial top-up. Create a key in the console and make the call. Testing uses the production API and consumes that credit; there is no separate sandbox.
curl https://api.nexichat.ai/v1/chat/prime \
-H "Authorization: Bearer nexi_live_..." \
-H "Content-Type: application/json" \
-d '{
"messages": [{"role": "user", "content": "Olá!"}]
}'The response:
{
"id": "chatcmpl-...",
"object": "chat.completion",
"created": 1786500000,
"model": "nexi-prime-v1",
"choices": [{
"index": 0,
"message": { "role": "assistant", "content": "Olá! Em que posso ajudar?" },
"finish_reason": "stop"
}],
"usage": { "prompt_tokens": 12, "completion_tokens": 8, "total_tokens": 20 }
}Authentication
#Send Authorization: Bearer nexi_live_… with every request.
The key is shown once, when you create it. After that only the prefix is visible — lose it and your only option is to revoke and create another. You can hold several at a time, which makes rotation painless: create the new one, swap it in your code, revoke the old one.
Nexi Code CLI
#Nexi Code is the official NexiChat CLI: work on your local project from the terminal. Install from the website and run nexicode login to connect your account in the browser. This uses your account plan; API keys and API balance are separate.
curl -fsSL https://nexichat.ai/install.sh | bash
nexicode login
nexicode --helpirm https://nexichat.ai/install.ps1 | iex
nexicode login
nexicode --helpModels and pricing
#Prices in euros per million tokens. No subscription — you pay for what you use.
| Model | URL | Input | Output |
|---|---|---|---|
Nexi Air 2 Fast answers, classification, simple tasks. nexi-air-v1 | /chat/air | 0,26 € | 3,10 € |
Nexi Prime 1 The balanced one. Long text, analysis, coding. nexi-prime-v1 | /chat/prime | 1,80 € | 6,20 € |
Nexi Nova 1 Reasoning. Hard problems, maths, planning. nexi-nova-v1 | /chat/nova | 3,20 € | 9,50 € |
Nova's reasoning tokens are billed as output. That is how usage is reported, and it is also why it is the most expensive of the three.
The request
#POST https://api.nexichat.ai/v1/chat/{air|prime|nova} with a JSON body.
| Field | Notes |
|---|---|
messagesarrayrequired | 1 to 200 messages. Each has a role and content. |
streamboolean | Defaults to false. Set true for an SSE response. |
toolsarray | Up to 64 functions the model may ask you to run. |
max_tokensinteger | Cap on generated tokens. Omit to use the model maximum. |
temperaturenumber | 0 to 2. Lower is more deterministic; use 0 for classification and extraction. |
Messages
#Four roles, each with its own shape.
| role | Shape | What for |
|---|---|---|
system | content | Your instructions. The Nexi identity always comes before them. |
user | content | What the person wrote. |
assistant | content and/or tool_calls | Previous replies. This is how you give the model history. |
tool | content + tool_call_id | The result of a function you ran. |
Each message accepts up to 200,000 characters. The conversation is stateless: you send the full history with every request, and nothing is stored on Nexi's side.
The response
#The text is in choices[0].message.content. usage carries the token counts — that is what the cost is computed from, and it matches what you see in the console.
| finish_reason | Means |
|---|---|
stop | The model finished on its own. |
length | Hit max_tokens. The reply is cut off mid-way. |
tool_calls | Waiting for you to run a function. |
x-request-id header. Log it: it is how support finds your exact request, and without it "it failed yesterday afternoon" is impossible to investigate.Streaming
#With "stream": true the response is SSE: data: lines in OpenAI's format, ending with data: [DONE]. The text arrives in choices[0].delta.content.
data: {"object":"chat.completion.chunk",
"choices":[{"index":0,"delta":{"content":"Olá"}}]}
data: {"object":"chat.completion.chunk",
"choices":[{"index":0,"delta":{"content":"!"}}]}
data: [DONE]curl https://api.nexichat.ai/v1/chat/nova \
-H "Authorization: Bearer nexi_live_..." \
-H "Content-Type: application/json" \
-d '{
"messages": [{"role": "user", "content": "Explica a relatividade"}],
"stream": true
}' --no-buffer--no-buffer with curl, stream=True with requests). Without it everything arrives at once at the end and streaming buys you nothing — the single most common complaint from people wiring up SSE for the first time.Tools
#The model can ask you to run your own functions. It takes three steps, and the middle one is yours: the API never runs your code — it returns the request, you run the function, you send the result back.
curl https://api.nexichat.ai/v1/chat/prime \
-H "Authorization: Bearer nexi_live_..." \
-H "Content-Type: application/json" \
-d '{
"messages": [{"role": "user", "content": "Que tempo faz no Porto?"}],
"tools": [{
"type": "function",
"function": {
"name": "meteorologia",
"description": "Tempo atual numa cidade",
"parameters": {
"type": "object",
"properties": {"cidade": {"type": "string"}},
"required": ["cidade"]
}
}
}]
}'tool_call must get its matching role: "tool" reply, with the tool_call_id lining up. A history containing the call but not the result is rejected — and from then on the whole conversation fails, not just that message.List models
#GET https://api.nexichat.ai/v1/models returns the list in OpenAI's format. Handy for populating a picker without hard-coding names in your code.
{
"object": "list",
"data": [
{ "id": "nexi-air-v1", "object": "model", "owned_by": "nexichat" },
{ "id": "nexi-prime-v1", "object": "model", "owned_by": "nexichat" },
{ "id": "nexi-nova-v1", "object": "model", "owned_by": "nexichat" }
]
}Errors
#Always the same shape. code is stable and meant for branching in your code; message is for humans and may change without notice.
{
"error": {
"code": "rate_limit_rpm",
"message": "Rate limit exceeded for your tier."
}
}| HTTP | code | When it happens |
|---|---|---|
| 400 | invalid_request | The body failed validation. The message names the field. |
| 400 | tools_unsupported | You sent tools to a model that does not support them. |
| 401 | invalid_api_key | Key missing, malformed or revoked. |
| 402 | insufficient_balance | Out of balance. Top up in the console. |
| 404 | unknown_model | The model in the URL is not air, prime or nova. |
| 429 | rate_limit_rpm | Requests per minute above your tier. |
| 429 | rate_limit_tpm | Tokens per minute above your tier. |
| 429 | daily_spend_limit_reached | You hit the daily cap you set in the console. |
| 502 | upstream_error | The model failed. Retrying usually works. |
| 503 | model_unavailable | Model temporarily unavailable. |
429 and 502 are worth retrying with exponential backoff. The other 4xx are not — retrying an invalid_request gives the same answer every time.
Rate limits
#Your tier moves up on its own with what the account holds: the higher of your total top-ups and your current balance. No form, no wait. If your integration needs more from day one, talk to us: we can set a minimum tier on the account.
| Tier | Requirement | Requests/min | Tokens/min |
|---|---|---|---|
| 1 | below 10 € | 10 | 10 000 |
| 2 | 10 € to 100 € | 100 | 200 000 |
| 3 | above 100 € | 500 | 1 000 000 |
Every request counts towards the per-minute request limit, including those rejected with 429: retrying without waiting only extends the wait. A request rejected for tokens spends no tokens. Before calling the model, the API reserves an estimate of the input plus max_tokens; only real usage counts at the end. A max_tokens sized to the answer you expect fits more requests into the same minute.
You can also set a daily spend cap in the console. Once hit, the API returns daily_spend_limit_reached until the next day — it is a safety net against a runaway loop in your code, not a commercial limit.
Authenticated chat responses include RateLimit-Policy and RateLimit: the enforced request and token quotas, remaining allowance and seconds to reset. These are a snapshot; concurrent requests may consume the allowance. Streaming headers include the token reservation for the request, reconciled when it ends.
RateLimit-Policy: "requests";q=10;w=60;qu="requests", "tokens";q=10000;w=60;qu="tokens"
RateLimit: "requests";r=9;t=30, "tokens";r=9000;t=30A 429 error includes Retry-After in seconds and the same value in error.retry_after. Respect that delay for automatic retries. For the daily spend ceiling, you can lower max_tokens, adjust the ceiling in the console, or wait for the next UTC day. For older clients, X-RateLimit-Limit, X-RateLimit-Remaining and X-RateLimit-Reset describe the request window; reset is a Unix timestamp in seconds.
Versions and changes
#The public API uses URL versioning: /v1/chat/prime and /v1/models. The current version is v1. Compatible additions, such as optional fields or new models, may ship in the same version. Clients should ignore unknown response fields. Breaking contract changes require a new major version.
There is no announced retirement date for v1. A migration will be documented here with the new contract, differences and dates before removal. Only endpoints with an announced deprecation will send Deprecation (RFC 9745), and Sunset(RFC 8594) when a retirement date exists. The Link; rel="service-doc" header points to this policy.
Best practices
#| Do this | Why |
|---|---|
| Log the x-request-id | Without it, investigating a request that failed two days ago is impossible. |
| Exponential backoff on 429 | Retrying straight away only makes it worse — and counts against the limit again. |
| Set a daily spend cap | A runaway loop in your code costs real money until somebody notices. |
| One key per environment | Revoking the test key must not take production down. |
| Air for the simple things | Classifying, extracting and summarising do not need Nova. The price gap is 6×. |
Ready to start
Create your key and make the first request in under a minute.