Docs · Integrate

Chat completions

POST /v1/chat/completions sends a list of messages and returns the model's reply. Every chat model in the catalog uses this endpoint.

curl https://<your-endpoint>/v1/chat/completions \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-4.1-mini",
    "messages": [{"role": "user", "content": "Hello!"}]
  }'

Multi-turn conversations and system messages

The API is stateless: send the full conversation history with every request. Put a system message first to set the role and rules.

resp = client.chat.completions.create(
    model="gpt-4.1-mini",
    messages=[
        {"role": "system", "content": "You are a concise assistant."},
        {"role": "user", "content": "What is a vector database?"},
        {"role": "assistant", "content": "A database optimized for similarity search over embeddings."},
        {"role": "user", "content": "Give me two use cases."},
    ],
    temperature=0.3,
    max_tokens=800,
)
print(resp.choices[0].message.content)
print(resp.usage)

Common parameters

ParameterDescription
modelRequired. A model ID from the catalog (case-sensitive)
messagesRequired. Array of messages with role system / user / assistant / tool
max_tokens / max_completion_tokensOutput token cap; reasoning models take max_completion_tokens
temperature / top_pSampling randomness; usually not supported by reasoning models
stopStop sequences
streamtrue to stream over SSE — see Streaming
tools / tool_choiceFunction calling — see Tool calling
response_formatJSON output — see Structured output
reasoning_effortReasoning depth — see Reasoning
Parameters are passed through to the model. Unsupported parameters may be ignored or rejected with a 400, depending on the model.

Response

{
  "id": "chatcmpl-...",
  "object": "chat.completion",
  "model": "gpt-4.1-mini",
  "choices": [{
    "index": 0,
    "message": { "role": "assistant", "content": "Hello! How can I help you today?" },
    "finish_reason": "stop"
  }],
  "usage": { "prompt_tokens": 9, "completion_tokens": 10, "total_tokens": 19 }
}
finish_reasonMeaning
stopFinished normally
lengthHit max_tokens; the output is truncated
tool_callsThe model wants to call tools
content_filterBlocked by a safety policy

usage reports input and output tokens; cost follows the model price.