Documentation
Documentation

Streaming

Server-sent events, stream options, and disconnects.

Display text as it arrives

Set stream to true to receive server-sent events. SDKs expose these events as an iterable stream. With cURL, -N disables output buffering so you can see events as they arrive.

Set ARNICT_API_KEY and install your SDK as shown in the quickstart. These examples include the client configuration.

Node.js
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://api.arnict.com/v1",
  apiKey: process.env.ARNICT_API_KEY,
});

const stream = await client.chat.completions.create({
  model: "zai/glm-5.3-flash-uncensored",
  messages: [{ role: "user", content: "Write a short welcome message." }],
  max_tokens: 256,
  stream: true,
  stream_options: { include_usage: true },
});

for await (const chunk of stream) {
  process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
  if (chunk.usage) console.error("Usage:", chunk.usage);
}

Handle the final usage event

Request token counts with stream_options.include_usage. A usage event can contain an empty choices array. Check that a choice exists before reading its delta, and handle usage separately from generated text.

When reading raw HTTP events, stop at data: [DONE]. Do not try to parse [DONE] as JSON. A lost connection before this marker can mean the response is incomplete.

Not every event contains text

An event may carry a role, tool-call update, finish reason or usage rather than content. Do not assume every event has choices[0].delta.content.

Handle interrupted streams

Treat text received before a disconnection as a partial result. Let your user retry or recover the task deliberately. Restarting a request creates a new generation and may produce different text.

Requests with no confirmed token usage are not charged. An interrupted request may still have confirmed usage. Check the request record in Usage if you need to understand its status or cost.