llm_sdk 0.7.0
llm_sdk: ^0.7.0 copied to clipboard
Unified interface for talking to LLMs (Claude, OpenAI, Gemini) from Dart/Flutter: multi-provider, streaming, tool calling, structured outputs.
0.7.0 #
The breaking decisions, grouped into one release so there is a single migration
rather than three. The only break is on the LlmProvider contract — code
that uses LlmClient needs no change at all.
-
BREAKING —
ToolChoicereplacesforceTool.String? forceToolcould express exactly one of the four cases every provider supports: you could force a named tool, but not say "no tools this turn" or "you must call something, your pick".ToolChoicecovers all four and maps to each dialect:ToolChoiceClaude OpenAI Gemini auto{type: auto}"auto"{mode: AUTO}none{type: none}"none"{mode: NONE}any{type: any}"required"{mode: ANY}named('x'){type: tool, name: x}{type: function, function: {name: x}}{mode: ANY, allowed_function_names: [x]}generateStreamgainstoolChoicetoo — it previously had no way to constrain tools at all.- Omitting it sends nothing, leaving the provider's own default in place.
- The constructor is private, so an incoherent state like "named mode with no name" cannot be built.
- In the tool loop,
toolChoiceapplies to the first step only. Otherwiseanyor a named tool would force a call on every turn, the loop would never reach a final answer, and it would run straight intomaxSteps. - Migration:
forceTool: 'x'becomestoolChoice: ToolChoice.named('x'). If you implementLlmProvideryourself, both methods change signature — the README has the diff. - Noted while reading the spec: forced tool use returns a 400 on Claude
Fable 5.1, which limits
generateObjectthere, since it is built on a forced tool. Documented rather than worked around.
-
maxTokensis now per request. It moves intoGenerationOptions, where the provider's own value stays the default. It is the cap most often adjusted between calls, and it was reachable only at provider construction. -
Three escape hatches, so a field the SDK doesn't model no longer means forking the package. All three are opt-in and deliberately not portable — they speak the provider's dialect, not the SDK's:
GenerationOptions.providerOptionsis deep-merged into the request body, so{'generationConfig': {'seed': 42}}adds to Gemini'sgenerationConfiginstead of replacing what the SDK put there. At an equal key with a non-map value, the caller wins — including over the SDK's own values.extraHeaderson every provider, applied last so it can override the SDK's headers. This is what makes it possible to route through your own backend rather than ship an API key inside a published app.LlmResponse.rawcarries the provider's decoded JSON body — citations, logprobs,safetyRatings, cache counters.nullon the streaming path, where there is no single body to expose.
-
The sealed-type contract is now written down.
PartandLlmStreamEventstaysealedso adapters must handle every variant, but new variants will appear and would break an exhaustive switch at compile time. Both types now document it: put a_case in your switches and your code survives minor releases.LlmResponse.text/.toolCallsdon't switch at all. -
+21 tests (112 total): all four
ToolChoicemodes on all three dialects plus their omission, the first-step-only behaviour proven by a loop that would otherwise hitmaxSteps, per-requestmaxTokenson each provider, deep-merge semantics including a caller value overriding the SDK's,extraHeadersoverride, andrawpresent on the buffered path and absent on the streamed one.
0.6.0 #
-
Streaming and tools now work together. The automatic tool loop used to live only on
generate(theFuturepath);streamTexthad notoolsparameter at all andstreamEventsforwarded tool calls without ever running them. Both now run the same loop asgenerate, built once inLlmClienton top of the unchanged two-method provider contract — no adapter changed.await for (final chunk in client.streamText( [Message.user('Compare the weather in Douala and Yaoundé.')], tools: [weather], // ← runs automatically, mid-stream )) { stdout.write(chunk); }streamTextgains an optionaltoolsparameter.TextDeltaandToolCallDeltaare forwarded for every step, so a UI can show "running getWeather…" while it happens.- Exactly one
StreamDoneis emitted, at the very end: per-stepStreamDones are absorbed, so "done" never means "done with this step". - Its
usageis the sum across all steps — a tool round-trip no longer goes uncounted. - Bounded by
maxSteps, and an unknown tool raises, exactly as ongenerate. A third-party provider that ends a stream without aStreamDonenow raises instead of looping silently. - Tool execution is shared between both loops, so they cannot drift apart.
- Need the old raw single-step behaviour? Call
provider.generateStream(...)directly —LlmProvideris public.
-
Retries now obey rate limits.
RetryPolicyignored theRetry-Afterheader that 429s carry, so a rate-limited request was replayed 400 ms later and simply took another 429. The header now wins over the backoff, in both spec forms — a number of seconds, or an HTTP date. Hand-rolled date parsing, becauseHttpDate.parselives indart:ioand this package must run on the web too.- New
maxRetryAfter(default60 s): beyond it the SDK stops retrying and returns the response, instead of blocking the request for minutes. - New
respectRetryAfter(defaulttrue) to ignore the header entirely. RetryPolicy.parseRetryAfteris public, for callers who read it themselves.
- New
-
The backoff is now jittered. Every client that hit the same 429 used to retry on the same millisecond and re-saturate the service. Each delay is now drawn uniformly from
[nominal × (1 - jitter), nominal]— the "equal jitter" recipe, withjitterdefaulting to0.5. Never longer than the nominal backoff, so the worst case stays predictable.jitter: 0restores a strictly deterministic schedule, andrandom:takes a seededRandomfor tests.delayForkeeps returning the nominal backoff; the newjitteredDelayForis what the retry loop actually waits.
-
Refreshed default models. The previous defaults had aged out; each provider now points at a current, generally available model (no previews):
Provider Before Now Why this one Claude claude-opus-4-8claude-opus-5Same tier as before, current id. OpenAI gpt-4ogpt-5.6-terraThe balanced intelligence/cost tier, the role gpt-4oplayed.Gemini gemini-1.5-progemini-3.8-flashCurrent stable flagship; the only current-generation Pro is preview-only. Defaults are a convenience, not a contract: the
modelfield now documents that it tracks each provider's current general-purpose model and may change in a minor release. Pin it explicitly in production. -
Note for OpenAI users who set
maxTokens: current OpenAI models expectmax_completion_tokensrather than themax_tokensthis adapter still sends, and the reasoning tiers also rejecttemperature/topP. LeavingmaxTokensat its default (null) avoids the issue; a proper fix is tracked for a later release, sincemax_completion_tokensis not understood by every OpenAI-compatible local server. -
GeminiProvider's URL test no longer asserts on the default model id, so refreshing a default can't break the suite. -
A failing tool no longer aborts the run.
tool.runthrowing used to propagate straight to the caller: the model never got a chance to recover, and the turn you had already paid for was lost. The error is now handed back as a tool result flaggedisError— which Claude carries on the wire asis_error— so the model can retry differently or say the information is unavailable. Same for a tool the model asks for but you never registered.- New
ToolErrorPolicy:feedToModel(the new default) orthrowToCaller, which restores the previous behaviour and rethrows the original exception with its stack trace intact. - New
onToolErrorcallback onLlmClient, fired in both policies — a swallowed exception would otherwise be exactly the kind of silent failure 0.5.1 set out to remove. ToolResultPartgains an optionalisErrorflag (defaults tofalse, so existing call sites are unaffected).- Behaviour change: code that relied on a throwing tool aborting the run
must now pass
toolErrors: ToolErrorPolicy.throwToCaller.
- New
-
Tools requested in the same turn now run concurrently. Providers emit parallel tool calls on purpose; running them in series added up the latencies. Results are still re-injected in call order.
parallelTools: falserestores sequential execution for tools that share non-concurrent state. -
generatenow sumsusageacross every step, like the streamed loop. A two-step tool round-trip that really cost 360 tokens used to be reported as 210 — every cost estimate built on it was low.usagestaysnullwhen no provider reported any. -
+39 tests (91 total): the streamed tool loop against a scripted provider (re-injection, single
StreamDone, summed usage,maxSteps, unknown tool, broken provider contract) plus one end-to-end test per provider driving real SSE through each dialect's very different tool-result encoding — ausermessage withtool_resultblocks on Claude, atool-role message on OpenAI, afunctionResponsematched by name on Gemini. Plus tool-error handling on both paths and in both policies,is_erroron the wire, usage summing, and concurrency ordering proven without timing-dependent assertions. Plus bothRetry-Afterforms, themaxRetryAftercut-off, and jitter bounds.
0.5.1 #
- Fix: streaming no longer fails silently. A mid-stream error or a dropped
connection used to end the stream quietly, leaving a truncated answer with no
signal at all — and on OpenAI it even emitted a
StreamDonethat looked like a clean success. The contract is now explicit: a stream always ends with either aStreamDoneor an error.- A provider error announced during generation (Claude's
errorSSE event, an{"error": ...}payload on OpenAI / Gemini) throws the newLlmStreamErrorException, carrying the provider's raw payload. - A stream that ends before its terminal marker throws the new
LlmStreamInterruptedException, carryingpartial— the response assembled so far. Terminal markers:message_stop(Claude),[DONE](OpenAI), and the last chunk'sfinishReason(Gemini, which has no dedicated marker). - Both extend
LlmException, so existingon LlmException catchhandlers keep working unchanged. - Malformed tool-call JSON on a complete stream now throws a readable
LlmExceptionnaming the tool, instead of a bareFormatException. On an interrupted stream, a tool call cut mid-JSON is reported with empty arguments rather than hiding the interruption.
- A provider error announced during generation (Claude's
- Lower SDK floor:
^3.11.5→^3.4.0. The code only needs Dart 3 features; the old bound locked out most Flutter stable users for no reason. - +10 tests (52 total) covering every path above on all three providers.
.gitignore: ignore.DS_Store.
No breaking change. Note that code which previously relied on a truncated stream ending quietly will now see an exception — that is the fix.
0.5.0 #
- Sampling options: new
GenerationOptions(temperature,topP,stopSequences) accepted by everyLlmClientmethod (generate,generateText,streamText,streamEvents,generateObject) and mapped by each adapter to its own dialect (temperature/top_p/stopon OpenAI,stop_sequenceson Claude,generationConfigon Gemini). Anullfield is omitted, so the model's default applies. - Network resilience: new
RetryPolicy(configurable per provider via theretry:constructor argument). Thegeneratepath now retries transient failures — HTTP 408/429/5xx, timeouts, and dropped connections — with exponential backoff, and every request is bounded by atimeout. Streaming applies the connection timeout only (replaying a started stream is unsafe).RetryPolicy.noneopts out. - +12 tests (42 total): options mapping per provider, backoff timing, retry / exhaustion / non-retryable-status / network-drop / timeout, and client-level options passthrough. No breaking change: the new parameters are optional.
0.4.0 #
- Local models via
OpenAIProvider:apiKeyis now optional (defaults to empty) and theAuthorizationheader is omitted when empty, for self-hosted OpenAI-compatible servers (Ollama, LM Studio, llama.cpp, vLLM).baseUrlwas already overridable — no other API change. - Added
example/ollama_example.dart(fully local streaming + tool calling). - README: new "Local models" section.
0.3.1 #
- README: demo GIFs ("switch provider = 1 line", word-by-word streaming).
- pubspec:
screenshots:(shown on pub.dev). - No code or API change.
0.3.0 #
GeminiProvideradapter (Generative LanguagegenerateContentAPI): assistant role →model,system_instruction, tool results re-attached by name asfunctionResponse(Gemini has no call ID — we mapToolCallPart.id == name), forcing viatool_configmodeANY, SSE streaming (streamGenerateContent?alt=sse),usageandfinishReason. API key in thex-goog-api-keyheader, configurablebaseUrl.- +7 Gemini tests (29 total). No change to the core: all 3 providers share the same abstraction.
0.2.0 #
OpenAIProvideradapter (Chat Completions API): encoding/decoding, tool calling (tool_calls+ JSON-string arguments), structured outputs via forced tool, SSE streaming (assembling tool calls byindex),usageandfinishReason. ConfigurablebaseUrl(compatible with OpenAI-like endpoints). No change to the core: the abstraction holds as-is.- +8 OpenAI tests (22 total).
0.1.0 #
First slice: provider-agnostic core + complete Claude adapter.
- Core: types (
Message/Part/Tool/LlmResponse/LlmStreamEvent),LlmProvidercontract (generate+generateStream),LlmClientcarrying the tool loop,streamText,generateTextandgenerateObject<T>. ClaudeProvideradapter (Anthropic Messages API): encoding/decoding, tool calling, structured outputs via forced tool, SSE streaming (TextDelta/ToolCallDelta/StreamDone),usageandfinishReason.- 20 tests (client logic on a mocked provider + Claude round-trip/SSE).
