Display text as it arrives
Set stream to true to receive server-sent events. SDKs expose these events as an iterable stream. With cURL, -N disables output buffering so you can see events as they arrive.
Set ARNICT_API_KEY and install your SDK as shown in the quickstart. These examples include the client configuration.
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://api.arnict.com/v1",
apiKey: process.env.ARNICT_API_KEY,
});
const stream = await client.chat.completions.create({
model: "zai/glm-5.3-flash-uncensored",
messages: [{ role: "user", content: "Write a short welcome message." }],
max_tokens: 256,
stream: true,
stream_options: { include_usage: true },
});
for await (const chunk of stream) {
process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
if (chunk.usage) console.error("Usage:", chunk.usage);
}Handle the final usage event
Request token counts with stream_options.include_usage. A usage event can contain an empty choices array. Check that a choice exists before reading its delta, and handle usage separately from generated text.
When reading raw HTTP events, stop at data: [DONE]. Do not try to parse [DONE] as JSON. A lost connection before this marker can mean the response is incomplete.
Not every event contains text
Handle interrupted streams
Treat text received before a disconnection as a partial result. Let your user retry or recover the task deliberately. Restarting a request creates a new generation and may produce different text.
Requests with no confirmed token usage are not charged. An interrupted request may still have confirmed usage. Check the request record in Usage if you need to understand its status or cost.
- Retry and rate-limit guidanceAvoid duplicate work and repeated requests during an outage.
