All models

gpt-5.6-luna

gpt-5.6-luna
StreamingTool callingVisionReasoningJSON modeWeb search
Integration guide

Overview

gpt-5.6-luna is available on MAX API through the OpenAI-compatible API.

Pricing

Official list prices in USD. Token prices are per 1M tokens.
Input ≤ 272,000 tokens
Input
$0.2
Output
$1.2
Cache read
$0.02
Cache write
$0.25
Input > 271,999 tokens
Input
$0.4
Output
$1.8
Cache read
$0.04
Cache write
$0.5

This model uses tiered pricing: the tier is chosen by the total input length of each request.

Cache read applies to prompt tokens served from the prompt cache; cache write applies to tokens written into it.

Example request

curl https://<your-endpoint>/v1/chat/completions \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-5.6-luna",
    "messages": [{"role": "user", "content": "Hello!"}]
  }'

https://<your-endpoint> is a placeholder. Your endpoint is shown in the console after you sign in.