llm_sdk 0.7.0 copy "llm_sdk: ^0.7.0" to clipboard
llm_sdk: ^0.7.0 copied to clipboard

Unified interface for talking to LLMs (Claude, OpenAI, Gemini) from Dart/Flutter: multi-provider, streaming, tool calling, structured outputs.

0.7.0 #

The breaking decisions, grouped into one release so there is a single migration rather than three. The only break is on the LlmProvider contract — code that uses LlmClient needs no change at all.

  • BREAKING — ToolChoice replaces forceTool. String? forceTool could express exactly one of the four cases every provider supports: you could force a named tool, but not say "no tools this turn" or "you must call something, your pick". ToolChoice covers all four and maps to each dialect:

    ToolChoice Claude OpenAI Gemini
    auto {type: auto} "auto" {mode: AUTO}
    none {type: none} "none" {mode: NONE}
    any {type: any} "required" {mode: ANY}
    named('x') {type: tool, name: x} {type: function, function: {name: x}} {mode: ANY, allowed_function_names: [x]}
    • generateStream gains toolChoice too — it previously had no way to constrain tools at all.
    • Omitting it sends nothing, leaving the provider's own default in place.
    • The constructor is private, so an incoherent state like "named mode with no name" cannot be built.
    • In the tool loop, toolChoice applies to the first step only. Otherwise any or a named tool would force a call on every turn, the loop would never reach a final answer, and it would run straight into maxSteps.
    • Migration: forceTool: 'x' becomes toolChoice: ToolChoice.named('x'). If you implement LlmProvider yourself, both methods change signature — the README has the diff.
    • Noted while reading the spec: forced tool use returns a 400 on Claude Fable 5.1, which limits generateObject there, since it is built on a forced tool. Documented rather than worked around.
  • maxTokens is now per request. It moves into GenerationOptions, where the provider's own value stays the default. It is the cap most often adjusted between calls, and it was reachable only at provider construction.

  • Three escape hatches, so a field the SDK doesn't model no longer means forking the package. All three are opt-in and deliberately not portable — they speak the provider's dialect, not the SDK's:

    • GenerationOptions.providerOptions is deep-merged into the request body, so {'generationConfig': {'seed': 42}} adds to Gemini's generationConfig instead of replacing what the SDK put there. At an equal key with a non-map value, the caller wins — including over the SDK's own values.
    • extraHeaders on every provider, applied last so it can override the SDK's headers. This is what makes it possible to route through your own backend rather than ship an API key inside a published app.
    • LlmResponse.raw carries the provider's decoded JSON body — citations, logprobs, safetyRatings, cache counters. null on the streaming path, where there is no single body to expose.
  • The sealed-type contract is now written down. Part and LlmStreamEvent stay sealed so adapters must handle every variant, but new variants will appear and would break an exhaustive switch at compile time. Both types now document it: put a _ case in your switches and your code survives minor releases. LlmResponse.text / .toolCalls don't switch at all.

  • +21 tests (112 total): all four ToolChoice modes on all three dialects plus their omission, the first-step-only behaviour proven by a loop that would otherwise hit maxSteps, per-request maxTokens on each provider, deep-merge semantics including a caller value overriding the SDK's, extraHeaders override, and raw present on the buffered path and absent on the streamed one.

0.6.0 #

  • Streaming and tools now work together. The automatic tool loop used to live only on generate (the Future path); streamText had no tools parameter at all and streamEvents forwarded tool calls without ever running them. Both now run the same loop as generate, built once in LlmClient on top of the unchanged two-method provider contract — no adapter changed.

    await for (final chunk in client.streamText(
      [Message.user('Compare the weather in Douala and Yaoundé.')],
      tools: [weather], // ← runs automatically, mid-stream
    )) {
      stdout.write(chunk);
    }
    
    • streamText gains an optional tools parameter.
    • TextDelta and ToolCallDelta are forwarded for every step, so a UI can show "running getWeather…" while it happens.
    • Exactly one StreamDone is emitted, at the very end: per-step StreamDones are absorbed, so "done" never means "done with this step".
    • Its usage is the sum across all steps — a tool round-trip no longer goes uncounted.
    • Bounded by maxSteps, and an unknown tool raises, exactly as on generate. A third-party provider that ends a stream without a StreamDone now raises instead of looping silently.
    • Tool execution is shared between both loops, so they cannot drift apart.
    • Need the old raw single-step behaviour? Call provider.generateStream(...) directly — LlmProvider is public.
  • Retries now obey rate limits. RetryPolicy ignored the Retry-After header that 429s carry, so a rate-limited request was replayed 400 ms later and simply took another 429. The header now wins over the backoff, in both spec forms — a number of seconds, or an HTTP date. Hand-rolled date parsing, because HttpDate.parse lives in dart:io and this package must run on the web too.

    • New maxRetryAfter (default 60 s): beyond it the SDK stops retrying and returns the response, instead of blocking the request for minutes.
    • New respectRetryAfter (default true) to ignore the header entirely.
    • RetryPolicy.parseRetryAfter is public, for callers who read it themselves.
  • The backoff is now jittered. Every client that hit the same 429 used to retry on the same millisecond and re-saturate the service. Each delay is now drawn uniformly from [nominal × (1 - jitter), nominal] — the "equal jitter" recipe, with jitter defaulting to 0.5. Never longer than the nominal backoff, so the worst case stays predictable. jitter: 0 restores a strictly deterministic schedule, and random: takes a seeded Random for tests.

    • delayFor keeps returning the nominal backoff; the new jitteredDelayFor is what the retry loop actually waits.
  • Refreshed default models. The previous defaults had aged out; each provider now points at a current, generally available model (no previews):

    Provider Before Now Why this one
    Claude claude-opus-4-8 claude-opus-5 Same tier as before, current id.
    OpenAI gpt-4o gpt-5.6-terra The balanced intelligence/cost tier, the role gpt-4o played.
    Gemini gemini-1.5-pro gemini-3.8-flash Current stable flagship; the only current-generation Pro is preview-only.

    Defaults are a convenience, not a contract: the model field now documents that it tracks each provider's current general-purpose model and may change in a minor release. Pin it explicitly in production.

  • Note for OpenAI users who set maxTokens: current OpenAI models expect max_completion_tokens rather than the max_tokens this adapter still sends, and the reasoning tiers also reject temperature / topP. Leaving maxTokens at its default (null) avoids the issue; a proper fix is tracked for a later release, since max_completion_tokens is not understood by every OpenAI-compatible local server.

  • GeminiProvider's URL test no longer asserts on the default model id, so refreshing a default can't break the suite.

  • A failing tool no longer aborts the run. tool.run throwing used to propagate straight to the caller: the model never got a chance to recover, and the turn you had already paid for was lost. The error is now handed back as a tool result flagged isError — which Claude carries on the wire as is_error — so the model can retry differently or say the information is unavailable. Same for a tool the model asks for but you never registered.

    • New ToolErrorPolicy: feedToModel (the new default) or throwToCaller, which restores the previous behaviour and rethrows the original exception with its stack trace intact.
    • New onToolError callback on LlmClient, fired in both policies — a swallowed exception would otherwise be exactly the kind of silent failure 0.5.1 set out to remove.
    • ToolResultPart gains an optional isError flag (defaults to false, so existing call sites are unaffected).
    • Behaviour change: code that relied on a throwing tool aborting the run must now pass toolErrors: ToolErrorPolicy.throwToCaller.
  • Tools requested in the same turn now run concurrently. Providers emit parallel tool calls on purpose; running them in series added up the latencies. Results are still re-injected in call order. parallelTools: false restores sequential execution for tools that share non-concurrent state.

  • generate now sums usage across every step, like the streamed loop. A two-step tool round-trip that really cost 360 tokens used to be reported as 210 — every cost estimate built on it was low. usage stays null when no provider reported any.

  • +39 tests (91 total): the streamed tool loop against a scripted provider (re-injection, single StreamDone, summed usage, maxSteps, unknown tool, broken provider contract) plus one end-to-end test per provider driving real SSE through each dialect's very different tool-result encoding — a user message with tool_result blocks on Claude, a tool-role message on OpenAI, a functionResponse matched by name on Gemini. Plus tool-error handling on both paths and in both policies, is_error on the wire, usage summing, and concurrency ordering proven without timing-dependent assertions. Plus both Retry-After forms, the maxRetryAfter cut-off, and jitter bounds.

0.5.1 #

  • Fix: streaming no longer fails silently. A mid-stream error or a dropped connection used to end the stream quietly, leaving a truncated answer with no signal at all — and on OpenAI it even emitted a StreamDone that looked like a clean success. The contract is now explicit: a stream always ends with either a StreamDone or an error.
    • A provider error announced during generation (Claude's error SSE event, an {"error": ...} payload on OpenAI / Gemini) throws the new LlmStreamErrorException, carrying the provider's raw payload.
    • A stream that ends before its terminal marker throws the new LlmStreamInterruptedException, carrying partial — the response assembled so far. Terminal markers: message_stop (Claude), [DONE] (OpenAI), and the last chunk's finishReason (Gemini, which has no dedicated marker).
    • Both extend LlmException, so existing on LlmException catch handlers keep working unchanged.
    • Malformed tool-call JSON on a complete stream now throws a readable LlmException naming the tool, instead of a bare FormatException. On an interrupted stream, a tool call cut mid-JSON is reported with empty arguments rather than hiding the interruption.
  • Lower SDK floor: ^3.11.5 → ^3.4.0. The code only needs Dart 3 features; the old bound locked out most Flutter stable users for no reason.
  • +10 tests (52 total) covering every path above on all three providers.
  • .gitignore: ignore .DS_Store.

No breaking change. Note that code which previously relied on a truncated stream ending quietly will now see an exception — that is the fix.

0.5.0 #

  • Sampling options: new GenerationOptions (temperature, topP, stopSequences) accepted by every LlmClient method (generate, generateText, streamText, streamEvents, generateObject) and mapped by each adapter to its own dialect (temperature/top_p/stop on OpenAI, stop_sequences on Claude, generationConfig on Gemini). A null field is omitted, so the model's default applies.
  • Network resilience: new RetryPolicy (configurable per provider via the retry: constructor argument). The generate path now retries transient failures — HTTP 408/429/5xx, timeouts, and dropped connections — with exponential backoff, and every request is bounded by a timeout. Streaming applies the connection timeout only (replaying a started stream is unsafe). RetryPolicy.none opts out.
  • +12 tests (42 total): options mapping per provider, backoff timing, retry / exhaustion / non-retryable-status / network-drop / timeout, and client-level options passthrough. No breaking change: the new parameters are optional.

0.4.0 #

  • Local models via OpenAIProvider: apiKey is now optional (defaults to empty) and the Authorization header is omitted when empty, for self-hosted OpenAI-compatible servers (Ollama, LM Studio, llama.cpp, vLLM). baseUrl was already overridable — no other API change.
  • Added example/ollama_example.dart (fully local streaming + tool calling).
  • README: new "Local models" section.

0.3.1 #

  • README: demo GIFs ("switch provider = 1 line", word-by-word streaming).
  • pubspec: screenshots: (shown on pub.dev).
  • No code or API change.

0.3.0 #

  • GeminiProvider adapter (Generative Language generateContent API): assistant role → model, system_instruction, tool results re-attached by name as functionResponse (Gemini has no call ID — we map ToolCallPart.id == name), forcing via tool_config mode ANY, SSE streaming (streamGenerateContent?alt=sse), usage and finishReason. API key in the x-goog-api-key header, configurable baseUrl.
  • +7 Gemini tests (29 total). No change to the core: all 3 providers share the same abstraction.

0.2.0 #

  • OpenAIProvider adapter (Chat Completions API): encoding/decoding, tool calling (tool_calls + JSON-string arguments), structured outputs via forced tool, SSE streaming (assembling tool calls by index), usage and finishReason. Configurable baseUrl (compatible with OpenAI-like endpoints). No change to the core: the abstraction holds as-is.
  • +8 OpenAI tests (22 total).

0.1.0 #

First slice: provider-agnostic core + complete Claude adapter.

  • Core: types (Message/Part/Tool/LlmResponse/LlmStreamEvent), LlmProvider contract (generate + generateStream), LlmClient carrying the tool loop, streamText, generateText and generateObject<T>.
  • ClaudeProvider adapter (Anthropic Messages API): encoding/decoding, tool calling, structured outputs via forced tool, SSE streaming (TextDelta / ToolCallDelta / StreamDone), usage and finishReason.
  • 20 tests (client logic on a mocked provider + Claude round-trip/SSE).
2
likes
160
points
124
downloads
screenshot

Documentation

API reference

Publisher

unverified uploader

Weekly Downloads

Unified interface for talking to LLMs (Claude, OpenAI, Gemini) from Dart/Flutter: multi-provider, streaming, tool calling, structured outputs.

Repository (GitHub)
View/report issues

Topics

#llm #ai #claude #openai #gemini

License

MIT (license)

Dependencies

http

More

Packages that depend on llm_sdk