Docs · Capabilities

Reasoning

Reasoning models think before they answer, which helps with math, coding, planning and multi-step analysis. You control how hard they think, trading quality against latency and cost.

OpenAI format: reasoning_effort

resp = client.chat.completions.create(
    model="o4-mini",
    reasoning_effort="low",            # minimal / low / medium / high (model-dependent)
    max_completion_tokens=8000,        # reasoning tokens count toward this cap
    messages=[{"role": "user", "content": "How many prime numbers are below 100?"}],
)
print(resp.choices[0].message.content)
print(resp.usage.completion_tokens_details)   # reasoning_tokens are billed as output
  • Reasoning tokens are not shown in the reply but count toward completion_tokens and are billed as output.
  • Cap output (including reasoning) with max_completion_tokens; reasoning models usually reject max_tokens.
  • Reasoning models usually do not support sampling parameters such as temperature or top_p; sending them may return a 400.

Anthropic format: thinking

# Newer Claude models: adaptive thinking, depth set by effort
msg = client.messages.create(
    model="claude-sonnet-4-5",
    max_tokens=16000,
    thinking={"type": "adaptive"},
    output_config={"effort": "medium"},    # low / medium / high …
    messages=[{"role": "user", "content": "Plan a 3-step database migration."}],
)

# Older Claude models: fixed budget (>= 1024 and < max_tokens)
# thinking={"type": "enabled", "budget_tokens": 4000}

Claude generations differ: newer models use thinking: {"type": "adaptive"} with output_config.effort, older ones use budget_tokens. Parameters are passed through as-is; Anthropic's official documentation is the reference.

Other vendors

Whether reasoning models from Gemini, DeepSeek, Qwen and others honor reasoning_effort in the OpenAI format depends on the model. Some (for example, versions named "thinking") always reason.

Tips

  • Use low effort or a non-reasoning model for simple tasks — faster and cheaper.
  • Reasoning requests can take minutes: use streaming or set the client timeout to 300 seconds or more.
  • Watch reasoning tokens in usage to keep costs in check.