flutter_local_ai 0.1.1
flutter_local_ai: ^0.1.1 copied to clipboard
On-device AI for Flutter: Apple Foundation Models, Android Gemini Nano, Windows AI Foundry, Chrome Prompt API. Text, streaming, tool calling, JSON schemas, genUI specs.
0.1.1 #
Windows #
- Zero-config build.
flutter build windowsnow resolves the Windows App SDK's C++/WinRT projection on its own:Microsoft.WindowsAppSDK.AI2.5.5 andMicrosoft.Windows.CppWinRT2.0.250303.1 from the NuGet cache, or downloaded from nuget.org into the build tree, then projected withcppwinrt.exeonce per build tree. No CMake edits, no NuGet in Visual Studio. When that cannot happen (offline, no cache) the build warns and falls back to the unconfigured plugin exactly as before, so nothing that built yesterday stops building.FLUTTER_LOCAL_AI_WINDOWS_AI=ON|OFF,FLUTTER_LOCAL_AI_NUGET_DOWNLOAD=OFFand the pre-existingFLUTTER_LOCAL_AI_WINRT_INCLUDE_DIRsteer it, as environment variables or cache entries. Seedoc/platform-support.md. - Structured output, natively, through
LanguageModel.GenerateStructuredJsonResponseAsync(App SDK 2.0+).supportsStructuredOutputis nowtrueon a configured Windows build. A response that completes but strays from the schema throwsSTRUCTURED_OUTPUT_INVALID, with the model's text in the error details. - Diagnosable availability.
LocalAi.availabilityReason()reports the actual WinRT activation failure ("Class not registered (0x80040154)") and what it means — no Windows App Runtime, or no package identity — instead of a generic hint. Generation statuses map toGENERATION_BLOCKEDandPROMPT_TOO_LONGrather than a bare status number; an older runtime under a newer build surfaces as such instead of as a generic failure. - CI compiles the arm. A
windows-latestjob builds the example twice, with the projection resolved and with it forced off. Previously nothing compiled the Windows C++ at all. It is still not a device pass: the runtime needs a Copilot+ PC or supported GPU, Windows 11 25H2+, and an MSIX-packaged app with thesystemAIModelscapability, none of which a plugin can supply.
Documentation #
- A known limitations and fallbacks section in the README spells out
what each ❌ in the platform table is blocked on: Foundation Models exist
only from OS 26 (older OSes get
unavailableOsTooOld, not a crash); ML Kit's structured output is compile-time KSP with no runtime schema to bridge and has no function calling for Gemini Nano; Windows AI has no tool API; bring-your-own models belong to flutter_gemma and itsflutter_gemma_builtin_aibridge, not this package. - The Windows setup section describes the Flutter-side flow — env vars,
msixpackaging withsystemAIModels— and Microsoft's runtime requirements, instead of a generic CMake/NuGet recipe. - The example's "Windows AI setup required" dialog no longer describes steps that stopped existing in 0.1.0.
Tool declarations carry their schema #
LocalAiTool.parameterSchema. A tool can now declare its parameters as a JSON Schema object instead of a flat list of scalars, so a constraint the flat form cannot express — a string enum, a list, a nested object — reaches the model intact. It is the same subsetGenerationConfig.schemaaccepts, validated in Dart with a path-qualified error and translated on Apple by the sameSchemaBuilderthat backs structured output, so the model is constrained to the declaration rather than asked to respect it. The flatparameterslist still works and now lowers to the same schema; supply one or the other, not both.
Tool errors are an answer, not a failed turn #
- A tool body that throws no longer aborts generation. The error is encoded
as a
{"error": "..."}tool result the model reads and can respond to, which is what a declined confirmation or a denied permission looks like in an agent loop.LocalAiToolException(message, {details})is the deliberate spelling; any other exception is reported the same way, as is a result that will not JSON-encode. - A suspended tool call is cancellable.
stopGeneration()now unwinds a turn waiting ononCallinstead of waiting for it to answer. There is no timeout on the native side, by design: a tool may sit for as long as a person takes to approve an action.
Testing #
FakeLocalAiHost.invokeToolplays the model's half of a tool call, through the same registry the native host dispatches with — so an adapter's tool loop, including its refusal and suspension paths, is testable off device.FakeLocalAiSessionnow exposes thetoolsit was opened with and theirtoolSchemas, and a fake configured withoutsupportsToolCallingrejects a session that binds tools, as a real host does.
Behaviour and wire changes #
No source change is needed to move from 0.1.0: LocalAiTool.parameters is
now optional rather than required, and ToolParameter / ToolArgumentType
are untouched. Two things behind the API did change.
- A tool body that throws used to fail the turn and now answers the model
instead. Code relying on an exception to abort generation should call
stopGeneration()explicitly. - Wire:
ToolSpeccarriesparametersSchemaJsoninstead of aList<ToolParameterSpec>, andToolParameterSpec/ToolArgumentKindare gone frompigeon.dart. Nothing outside this package's own native hosts reads the wire, but a hot restart across this upgrade needs a full rebuild rather than a reload.
0.1.0 #
A rewrite onto one architecture, exposing a standalone OS-model layer that external adapters can use through the public session API.
One implementation instead of two #
Every platform previously carried two paths: a hand-rolled method channel with its own generation logic, and nothing else. Both Dart surfaces now run on a single pigeon-typed session host per platform.
FlutterLocalAiis now a thin facade over the session layer rather than a parallel implementation. Its API is unchanged; there is simply one place where generation happens.- The Android plugin drops from 515 to 39 lines, the Apple plugin from 932 to 38, and Windows from 373 to 48. What is left is registration.
- The native contract lives in
pigeon.dart, generated for Dart, Kotlin, Swift and C++. Its session half is shape-compatible withflutter_gemma_builtin_ai'sBuiltInAiService.
New capability #
- Session API.
LocalAiModelmintsLocalAiSessions that buffer a turn (addQueryChunk/addImage) then generate it, withstopGeneration,sizeInTokensandclose. Several conversations can be open at once. - Availability facade.
LocalAi.availability(),LocalAi.ensureReady()with download progress, andLocalAi.capabilities(). Capabilities are reported by the running host, not assumed per platform — the same binary answers differently across OS versions. - Web. A Chrome Prompt API arm, including schema-constrained output via
responseConstraint, alongside Apple's native schema path. - Images.
addImageon Android. Apple needs OS 27 and reportssupportsVision: falseuntil then. - Exact token counts. Native on Android and Apple 26.4+; elsewhere
sizeInTokensfalls back to a documentedlength / 4estimate rather than failing. - Cancellation.
stopGenerationreaches the model, not just the Dart subscription. - Per-call sampling on the wire.
LocalAiGenerationOverrideslets one call vary temperature or length without disturbing the conversation, which is how both Apple'srespond(options:)and ML Kit's request builder already work.
Platform backends #
- Android. Prompt API beta4 / Kotlin 2.3.21, with multiple images, native system instructions where AICore supports them, an expanded output budget and cancellable, serialized generation.
- Android. A real download total from
DownloadStarted.bytesToDownload, soensureReady(onProgress:)reports a percentage rather than an unknown. - Android. The transcript is settled when a turn fails or is cancelled:
partial streamed text is committed, otherwise the abandoned prompt is
dropped instead of being left for the next turn to resend.
stopGenerationjoins the cancelled job before it answers Dart. - Apple. Fixed service access control; full and structured responses are cancellable; an OS 27 SDK-gated image attachment path (awaiting OS 27 validation).
- Apple. Tool calling and structured output are reported only on iOS/macOS 26+ rather than unconditionally, so the capability gate is usable on an older OS instead of promising what the runtime cannot do.
- Windows. The App SDK 2.0 Text namespace, readiness/preparation, async generation/cancellation and explicit CMake setup (awaiting Windows build/device validation).
- Native model ownership is shared across the facade, the session API and genUI: session IDs no longer collide, in-flight session creation is awaited on shutdown, and closing one owner can no longer destroy another's live model.
- genUI instructions stay isolated from ongoing conversations, so generating a module no longer rewrites the chat the user is in.
Documentation #
- Per-platform coverage, build requirements and what remains unverified are
documented in
doc/platform-support.md.
Breaking #
- Drop the
genuidependency.GenUiModuleSpec.toComponents()becomestoComponentMaps(), returning plain maps instead of genui's typedComponent, so the package no longer pulls a renderer — and its native plugins — into apps that only generate text. The doc comment carries the three-line adaptation for genui users. FlutterLocalAiPlatformandMethodChannelFlutterLocalAiare removed. The platform seam is nowLocalAiHost; tests substitute one withdebugLocalAiHost.LocalAiPlatformInfo.fromMapis replaced byLocalAiPlatformInfo.fromCapabilities. The type is now a narrowed view ofLocalAiBackendCapabilities, which is the single source of truth.LocalAiBackendgainschromePromptApi, so exhaustive switches over it need a new case.- Behaviour:
generateTextwithoutinstructionsnow continues one shared conversation on every platform. Apple already did; Android silently discarded history. Passinstructions:for a stateless call, as the docs have always described. registerToolsnow restarts the shared conversation, because Apple's FoundationModels binds tools when a session is constructed and cannot add them to a live one.- The SDK floor moves to Dart 3.8 / Flutter 3.32, set by the
flutter_lints6 development tooling, whose own pubspec declaressdk: ^3.8.0— 3.8 being what Flutter 3.32 ships. The web arm does not push it higher: itsextension type/dart:js_interopinterop has been stable since Dart 3.3, and its null-aware elements land exactly on 3.8.
Preserved #
Native Apple tool calling, schema-constrained output with Dart-side schema validation, Windows AI Foundry, genUI module specs, AICore availability reasons and the Play Store redirect all carry over unchanged, and are now reachable from the session API as well.
0.0.16 #
Hardening of the structured-output API #
- A
schemanow implies JSON mode consistently: you no longer have to also setresponseFormat: ResponseFormat.json, andGenerationConfig.toMap()never sends a schema paired with atextformat (effectiveResponseFormat/requestsStructuredOutputexpose this). - Schemas are validated in Dart before the platform channel
(
GenerationConfig.validateSchema()), so unsupported constructs fail fast with a path-qualifiedArgumentErrorinstead of an opaque native error. AiResponse.decodedJsondecodes a root array or scalar;AiResponse.jsonstays object-only for backwards compatibility.generateTextStreamnow rejects schema-constrained requests up front (no backend can constrain streamed output yet) instead of silently returning free-form text — the returned stream errors immediately. UsegenerateTextfor structured output.
0.0.15 #
Structured (JSON-schema) outputs #
- Apple (iOS 26 / macOS 26): native schema-constrained generation. A JSON
Schema passed through
GenerationConfig.schemais translated into a FoundationModelsGenerationSchema, so the model is forced to emit matching JSON. Supported constructs: nested objects (withrequired), arrays (withminItems/maxItems), string enums, and the scalar types. Read the result withAiResponse.json(object roots) orAiResponse.decodedJson(any root). - Android / Windows: unchanged — these backends are text-out only and report
supportsStructuredOutput: false. Passing aschema(orResponseFormat.json) throwsSTRUCTURED_OUTPUT_UNSUPPORTED. The ML Kit GenAI on-device Prompt API does not currently expose aresponseSchema/responseMimeType; gate ongetPlatformInfo().supportsStructuredOutput.
0.0.14 #
Android: tool calling is explicitly unsupported #
registerTools()on Android now fails with a clearUNSUPPORTEDerror instead of pretending to work. The ML Kit GenAI Prompt API has no function calling — text/image in, text out only (verified against thegenai-prompt:1.0.0-beta2API surface) — and a prompt-emulated JSON call protocol proved too unreliable on Gemini Nano to ship: the model answers in prose instead of performing the call. Gate ongetPlatformInfo().supportsToolCalling(false on Android); revisit when Gemma 4 / Agent Mode reaches the Prompt API surface.- The example app shows the tool-calls toggle only on iOS/macOS, and
debugPrints every request and response — including the genUI tab's raw model output, so generation issues are inspectable from the console.
Android: widest available device support + robust model download #
- Bumped
com.google.mlkit:genai-promptfrom1.0.0-alpha1to1.0.0-beta2. alpha1 only resolved the Prompt API feature on Pixel 9 devices, socheckStatus()never reportedDOWNLOADABLEon anything else; beta2 carries the full current supported-device list — Pixel 9/10 series plus Samsung (Galaxy Z Fold7 / Z TriFold, S26 series), Honor, iQOO, Lenovo, Motorola, OnePlus, OPPO, POCO, realme, vivo and Xiaomi flagships (nano-v2 / nano-v3). - Aligned
com.google.mlkit:genai-commonwith the1.0.0-beta3version thatgenai-prompt:1.0.0-beta2declares in its POM (it's the artifact resolving per-device feature configs; the previous explicitbeta2pin understated the actually-resolved version) andplay-services-taskswith the POM's18.2.0floor. - Documented the supported-device matrix in the README, including that any
API 26+ device can install the app and gracefully degrades via
isAvailable()/getModelStatus()on unsupported hardware. downloadModel()failures now always surface on the download status stream: the Android implementation emits afailedstatus (with the error message) before propagating the exception, and the Dart layer converts a faileddownloadModelmethod call into afailedstatus as a last resort. UIs watching the stream can no longer hang on a silent native error.downloadModel()is idempotent: when the model is already downloaded it emitscompletedimmediately instead of erroring.
0.0.13 #
Stateless one-shot generation #
generateTextandgenerateTextStreamgain an optionalinstructionsparameter. When given, the call runs in a throwaway session with exactly those instructions: nothing accumulates in the shared session created byinitialize, and its instructions stay untouched. Apple'sLanguageModelSessionkeeps every prompt + response of its transcript counting toward the 4096-token context window, so stateless callers that carry their own context in the prompt would otherwise hitexceededContextWindowSizeafter a few calls. On Android ML Kit (already stateless per call) the one-shot instructions simply replace the session-level ones for that prompt.LocalAiUiGeneratornow generates every module one-shot, so repeated generations no longer fill the shared session's context window.
0.0.12 #
Streaming text generation #
- New
FlutterLocalAi.generateTextStream(prompt:, config:)returns aStream<String>of delta chunks as the on-device model decodes, closing when the generation completes. Implemented on Apple FoundationModels (LanguageModelSession.streamResponse, cumulative snapshots converted to deltas) and Android ML Kit GenAI (generateContentwith aStreamingCallback). Backends without a streaming implementation surface an error on the stream so callers can fall back togenerateText. LocalAiUiGenerator.generateModulegains an optionalonTextcallback that receives the cumulative raw model output while the module is being generated (for live progress/preview UI). When the platform cannot stream, generation silently degrades to the previous blocking call.
0.0.11 #
genUI — localized generation #
LocalAiUiGenerator.generateModulegains an optionallanguageparameter (an English language name such as "Italian" or "German"). When set, the prompt instructs the on-device model to write all user-facing copy — title, blurb and every block label, item and note — in that language, so generated modules match the app's locale. Omitting it preserves the previous behaviour.
0.0.10 #
genUI — reliable generation on Android (Gemini Nano) #
- Fixed the "model did not return JSON" failure on Android. Gemini Nano caps output at 256 tokens, so a verbose module was being truncated mid-JSON. The genUI prompt is now backend-aware: on Android it asks for compact, minified JSON with at most 3 blocks and shorter strings, and runs at a lower temperature (0.2) for more reliable structure — keeping output well within the token budget.
- Made
LocalAiUiGenerator's JSON extraction tolerant of small-model quirks: it now repairs JSON truncated by the output cap (cutting at the last complete block and closing the open brackets) and strips trailing commas (viareplaceAllMapped—replaceAlldoes not expand$1), so a partial response still renders instead of being discarded. - Verified end-to-end on a Google Pixel 10 (Android 16, Gemini Nano via AICore): multiple goals now produce valid, rendered modules.
Tests #
- Fixed the test suite: updated the platform-interface fake to implement the new
availabilityReason()member, and added regression coverage forLocalAiUiGenerator.parseModelOutput(clean / fenced / prose / trailing-comma / truncated-and-repaired / invalid inputs).
Generative UI in the example app #
- The example app now has a Generative UI tab (alongside Text) that turns a goal into a rendered on-device module, with backend labelling and a JSON view.
Platform & docs #
- Raised the macOS deployment target (Podfile
10.15→12.0) and refreshed dependencies; pins the project to Flutter 3.41.9 via FVM. - Documented genUI in the README (Platform Support table, a Generative UI usage
guide, and
LocalAiUiGenerator/GenUiModuleSpecAPI reference).
0.0.9 #
genUI reuse #
- Exposed
LocalAiUiGenerator.genUiInstructions(the module/block schema) andLocalAiUiGenerator.parseModelOutput(text)so any on-device backend (e.g. a downloaded Gemma model via flutter_gemma) can drive the same genUI generation.
0.0.8 #
genUI on Android (Pixel) #
LocalAiUiGeneratoris now backend-aware and works on Android ML Kit GenAI / Gemini Nano (e.g. Google Pixel) as well as Apple FoundationModels — the genUI schema is delivered via instructions (prepended to the prompt natively on Android), and the output-token budget is tuned per backend.- Added
availabilityReasonon Android (mapsFeatureStatus→ available/downloadable/downloading/unavailable) for accurate UI status.
0.0.7 #
genUI integration #
- Added a
genuiintegration so the package can turn a natural-language goal into a renderable module spec:LocalAiUiGenerator(on-device, via FoundationModels) andGenUiModuleSpec(typed blocks →genuicomponents). - Added
availabilityReason()(Dart + Apple native) to report exactly why the model is unavailable (e.g.deviceNotEligible, Apple Intelligence disabled).
Apple platforms - Generation fix #
- Fixed a FoundationModels
GenerationErrorcaused by combining.greedysampling with a temperature; options are now chosen exclusively. Generation failures now surface a fully-reflected, diagnosable error description.
0.0.6 #
Android - Availability & Dependencies #
- Added
genai-commondependency to align with updated ML Kit GenAI APIs. (Contributed by kaitotokyo) - Fixed availability checks by using
Generation.getClient()andFeatureStatus/GenAiExceptionfromgenai-common, with improved AICore incompatible handling. (Contributed by kaitotokyo) - Made generation config parsing safer and only apply
maxOutputTokens/temperaturewhen provided. (Contributed by kaitotokyo)
0.0.5 #
Apple platforms - Thread Safety #
- Introduced a
ModelManageractor for thread-safe FoundationModels access, session initialization, tool registration, and text generation. (Contributed by kaitotokyo)
Android - Availability Check #
- Improved availability checks using
FeatureStatus, betterGenAiExceptionhandling, and ensured the model client is closed. (Contributed by kaitotokyo) - Updated AICore/MLKit incompatibility error message for clarity. (Contributed by kaitotokyo)
0.0.4 #
Apple platforms - Improvements #
- Lowered iOS and macOS deployment targets to allow plugin compilation on older OS versions; runtime still reports unsupported below 26.0.
0.0.3 #
Apple platforms - Tool Support #
- ✅ Tools API support (iOS & macOS only) - Added support for tool execution on Apple platforms; Android and Windows tooling support is planned
0.0.2 #
Windows - Initial Support #
- ✅ Added Windows platform support - Initial implementation structure for Windows AI APIs (Windows AI Foundry)
- ✅ Windows plugin structure - Created C++/WinRT plugin implementation with method channel handlers
- ✅ Windows version checking - Added availability check for Windows 11 22H2 (build 22621) or later
- ✅ CMake build configuration - Added Windows CMakeLists.txt for plugin compilation
- ✅ Example app Windows support - Added Windows platform to example app
- ✅ Documentation updates - Added Windows setup instructions and platform-specific notes to README
Improvements #
- Updated package description to include Windows AI APIs
- Added Windows to platform support table
- Comprehensive Windows implementation documentation
Status #
- Windows AI API integration structure is in place and ready for full implementation
- Plugin provides availability checking, initialization flow, and error handling
- Ready for Windows AI Foundry API integration when APIs become available
0.0.1-dev.9 #
Android - Complete Implementation #
- ✅ Completed Android support - Full working implementation using ML Kit GenAI (Gemini Nano)
- ✅ Improved FlutterLocalAiPlugin.kt - Enhanced with proper context management, coroutine scope handling, and error detection
- ✅ Java 11 support - Updated build.gradle to require Java 11 (required for ML Kit GenAI)
- ✅ AICore integration - Added proper AICore library declaration and error handling
- ✅ Play Store integration - Added
openAICorePlayStore()method to help users install AICore - ✅ Enhanced error handling - Improved error code -101 detection and user-friendly error messages
- ✅ Dependencies - Added
play-services-tasks:18.0.2dependency - ✅ Example app updates - Added AICore library declaration and dependencies to example app
- ✅ Comprehensive documentation - Updated README.md with complete Android setup instructions, AICore handling guide, and code examples
Improvements #
- Better token counting (filtering empty strings)
- Improved Play Store opening logic with proper activity resolution
- Proper cleanup in
onDetachedFromEngine(canceling coroutine scope) - Enhanced logging for debugging AICore issues
Documentation #
- Complete Android setup guide with step-by-step instructions
- AICore requirement explanation and handling examples
- Platform-specific usage examples
- Error handling best practices
- Updated platform support table (Android now shows ✅)
0.0.1-dev.8 #
- Wip on Android
- First usage of AICore
0.0.1-dev.7 #
- Added support to macOS
- Migrated to Swift Package Manager
0.0.1-dev.6 #
- Enhanced documentation
0.0.1-dev.5 #
- Improved Android logic
0.0.1-dev.4 #
- Improved iOS logic
0.0.1-dev.3 #
- Enhanced documentation
0.0.1-dev.2 #
- Added development warning
0.0.1-dev.1 #
- Initial beta release
- Android implementation using ML Kit GenAI
- iOS implementation structure (placeholder for Apple GenAI API)
- Dart API for text generation
- Example app included
- Comprehensive test suite