문서 · 연동
스트리밍
Set stream: true to receive tokens as server-sent events (SSE) while they are generated. You get the first token sooner and avoid client timeouts on long outputs.
본문은 현재 영어로만 제공됩니다.
curl https://<your-endpoint>/v1/chat/completions \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-4.1-mini",
"messages": [{"role": "user", "content": "Hello!"}],
"stream": true
}'Event format
Each event is a data: {JSON} line with incremental text in choices[0].delta.content. The stream ends with data: [DONE].
data: {"id":"chatcmpl-...","choices":[{"index":0,"delta":{"role":"assistant","content":""}}]}
data: {"id":"chatcmpl-...","choices":[{"index":0,"delta":{"content":"Hel"}}]}
data: {"id":"chatcmpl-...","choices":[{"index":0,"delta":{"content":"lo!"},"finish_reason":"stop"}]}
data: [DONE]Usage in streams
Set stream_options.include_usage and the final chunk carries the usage for the whole request.
stream = client.chat.completions.create(
model="gpt-4.1-mini",
messages=[{"role": "user", "content": "Write a haiku about routers."}],
stream=True,
stream_options={"include_usage": True},
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)
if chunk.usage: # final chunk: empty choices, carries usage
print("\n", chunk.usage)Things to know
- If you relay streams through your own reverse proxy (e.g. Nginx), disable response buffering (
proxy_buffering off), or the content arrives all at once at the end. - Set the read timeout for streams as the gap between chunks, not the total request duration.
- Reasoning models may think for a while before the first token. That is expected.
- In the Anthropic format, stream events match Anthropic's (
message_start,content_block_delta, …) — see Anthropic format.