文档 · 接入

流式输出

设置 stream: true,以服务器推送事件(SSE)边生成边返回。首字更快,也能避免长输出触发客户端超时。

curl https://<your-endpoint>/v1/chat/completions \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-4.1-mini",
    "messages": [{"role": "user", "content": "Hello!"}],
    "stream": true
  }'

事件格式

每个事件是一行 data: {JSON},增量文本在 choices[0].delta.content;流以 data: [DONE] 结束。

data: {"id":"chatcmpl-...","choices":[{"index":0,"delta":{"role":"assistant","content":""}}]}

data: {"id":"chatcmpl-...","choices":[{"index":0,"delta":{"content":"Hel"}}]}

data: {"id":"chatcmpl-...","choices":[{"index":0,"delta":{"content":"lo!"},"finish_reason":"stop"}]}

data: [DONE]

在流式中获取用量

设置 stream_options.include_usage,最后一个 chunk 会携带整次请求的 usage。

stream = client.chat.completions.create(
    model="gpt-4.1-mini",
    messages=[{"role": "user", "content": "Write a haiku about routers."}],
    stream=True,
    stream_options={"include_usage": True},
)
for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)
    if chunk.usage:            # final chunk: empty choices, carries usage
        print("\n", chunk.usage)

注意事项

  • 如果你在自己的 Nginx 等反向代理后转发流式响应,请关闭响应缓冲(如 proxy_buffering off),否则内容会攒到最后一次性到达。
  • 流式请求的读超时按「两次数据之间的间隔」设置,而不是整次请求时长。
  • 推理模型在输出第一个字之前可能先思考较长时间,这是正常现象。
  • Anthropic 格式的流式事件与 Anthropic 官方一致(message_start、content_block_delta 等),见Anthropic 格式。