stream_struct 0.3.1 copy "stream_struct: ^0.3.1" to clipboard
stream_struct: ^0.3.1 copied to clipboard

Stream an LLM's structured output as it arrives: tolerant partial JSON that fills the typed object token by token. Decodes SSE, with OpenAI, Anthropic and Gemini extractors.

stream_struct: a token stream parsed into the object as it fills in

stream_struct #

Turn a language model's token stream into a stream of the structured object as it fills in.

stream_struct parses a JSON object and fills it in token by token as it streams

A model asked for JSON emits it one token at a time. Mid-stream you are holding something like {"title": "The quick bro , which jsonDecode throws on until the very last token lands. So the usual choices are to wait for the whole response before showing anything, or to hand-roll a fragile parser. stream_struct is the parser, done once and tested, plus the provider glue.

dart pub add stream_struct

Parse one partial buffer #

parsePartialJson decodes a truncated buffer into the value it holds so far. It closes an open string value and any open array or object, and drops a dangling key, colon, or comma.

import 'package:stream_struct/stream_struct.dart';

parsePartialJson('{"title": "The quick bro');   // {title: The quick bro}
parsePartialJson('{"a": 1, "tags": ["x"');       // {a: 1, tags: [x]}
parsePartialJson('{"a": 1, "colo');              // {a: 1}   (partial key dropped)

It returns null while nothing is decodable yet (an empty buffer, a lone { with only a partial key, a half-written tr). Treat null as "no update this frame" and keep the previous value; the next token resolves it.

Stream the object as it grows #

streamPartialJson accumulates a delta stream and emits the value after each token, skipping frames that do not parse yet or that did not change.

await for (final partial in streamPartialJson(modelDeltas)) {
  final map = partial as Map<String, dynamic>;
  setState(() => _draft = map);   // render the object filling in
}

Plug in your provider #

sseJson decodes the Server-Sent Events body providers stream, and an adapter pulls the text fragment out of each event. OpenAI, Anthropic, and Gemini shapes are built in, so a response goes end to end with no line handling of your own:

final response = await request.close();

streamPartialJsonFrom(sseJson(response), openAiDelta)  // choices[0].delta.content
    .listen((partial) => print(partial));

streamPartialJsonFrom(sseJson(response), anthropicDelta); // delta.text / delta.partial_json
streamPartialJsonFrom(sseJson(response), geminiDelta);    // candidates[0].content.parts[0].text

sseJson takes the raw byte stream, so chunk boundaries falling inside a line or an event are handled for you. It follows the event-stream format: several data: lines in one event are joined with newlines, one leading space after the colon is stripped, : comments and the event:/id:/retry: fields are ignored, and the [DONE] sentinel ends the stream rather than being parsed. If you want the payloads without the JSON decode, use sseData, and if your transport already gives you lines, sseDataFromLines.

Type it #

streamPartial<T> maps each growing object through a builder. Write the builder to tolerate a half-filled map and you get a typed value on every step.

final titles = streamPartial<String>(
  modelDeltas,
  (m) => (m['title'] as String?) ?? '',
);

From an HTTP response, streamPartialFrom is the same thing over a provider's chunks, which is the whole path in one call:

streamPartialFrom(sseJson(response), openAiDelta, Recipe.fromPartial)
    .listen((recipe) => setState(() => _recipe = recipe));
text fragments a provider's chunks
Object? frames streamPartialJson streamPartialJsonFrom
your type streamPartial streamPartialFrom

example/openai_end_to_end.dart runs that path end to end with no API key, on bytes chopped at arbitrary boundaries the way a socket delivers them.

What it handles #

  • open string values are kept and closed, so partial text shows as it types
  • open objects and arrays are closed to any depth
  • dangling keys, colons, and commas are dropped
  • braces and quotes inside strings, and escaped quotes, do not confuse it
  • a valid partial number is kept; an unresolved literal skips that one frame

It does the retrieval-of-structure, not generation. It never calls a model; you bring the stream. A frame that still will not parse yields null rather than a guess, which keeps a wrong intermediate value off your screen.

Roadmap #

Typed streaming today needs a small hand-written builder. Generated builders, so streamPartial<T> needs no mapping for your own classes, are planned next.

License #

MIT.

0
likes
160
points
25
downloads
screenshot

Documentation

API reference

Publisher

verified publisherdeveloperyusuf.com

Weekly Downloads

Stream an LLM's structured output as it arrives: tolerant partial JSON that fills the typed object token by token. Decodes SSE, with OpenAI, Anthropic and Gemini extractors.

Repository (GitHub)
View/report issues

Topics

#llm #ai #json #streaming #structured-output

License

MIT (license)

Dependencies

meta

More

Packages that depend on stream_struct