Flutter Edge AI
On-device AI for Flutter on Android, iOS, Web, macOS, Windows and Linux. Run Gemma and other open models, or the model the operating system already ships, with text, vision, audio, function calling, embeddings, RAG and speech β no server, no cloud. The core is small: you add only the runtimes and stores your app uses.
Formerly
flutter_gemma. Through 1.11.3 this package shipped asflutter_gemma. Version 2.0 also moves RAG out of core intoflutter_edge_ai_rag. See the migration guide.
π Full documentation: flutteredge.ai
Features
- Pluggable engines: LiteRT-LM (
.litertlm), MediaPipe (.task), ONNX Runtime, and the OS's own models (Gemini Nano, Apple Foundation Models, Windows AI Foundry, Chrome Prompt API). - Multimodal: image and audio input with Gemma 4, Gemma 3n and other vision models.
- Function calling and thinking mode on the models that support them.
- CPU, GPU and NPU backends β Qualcomm NPU on Android, Intel NPU on Windows.
- Embeddings and on-device RAG with Qdrant Edge or SQLite +
sqlite-vec, payload filters included. - Speech: on-device STT, TTS and a push-to-talk voice loop.
- Agent skills:
SKILL.mdtools the model invokes through function calling. - Model management: installs from the network, assets, bundled files or local paths; Hugging Face installs in one call; retries, typed download errors, LoRA weights.
- Genkit integration for on-device and hybrid cloud flows.
Packages
Core registers no engine. Add flutter_edge_ai plus what your app needs:
| You want to⦠| Add |
|---|---|
Run .litertlm models (Gemma 4, Qwen3, FastVLM; all desktop) and LiteRT embeddings |
flutter_edge_ai_litertlm |
Run .task / .bin models (MediaPipe; mobile and web) |
flutter_edge_ai_mediapipe |
| Run ONNX models β text generation and embeddings | flutter_edge_ai_onnx |
| Use the model built into the OS or browser | flutter_edge_ai_builtin_ai |
| Tokenizers for text embeddings (needed with the LiteRT or ONNX embedding backend) | flutter_edge_ai_embeddings |
| On-device RAG | flutter_edge_ai_rag + flutter_edge_ai_qdrant (native) or flutter_edge_ai_sqlite (all six platforms) |
| Speech-to-text, text-to-speech, voice loop | flutter_edge_ai_speech |
| Agent skills over function calling | flutter_edge_ai_agent |
| Memory a model costs, read from the OS | flutter_edge_ai_diagnostics |
| Genkit, and on-device/cloud routing | genkit_flutter_edge_ai, genkit_hybrid |
Supported platforms
| Engine | Android | iOS | Web | macOS | Windows | Linux |
|---|---|---|---|---|---|---|
LiteRT-LM (.litertlm) |
β | β | β οΈ preview ΒΉ | β | β | β |
MediaPipe (.task) |
β | β | β | β | β | β |
| ONNX Runtime | β | β | β | β | β | β |
| Built-in AI | β | β | β | β | β | β |
ΒΉ Web .litertlm is text and function calling only β no vision, audio or LoRA.
LiteRT-LM ships arm64 on Android, iOS and macOS, x86_64 on Windows and x86_64 / arm64
on Linux; on Android, MediaPipe .task also runs on x86_64 / armeabi-v7a, and on Linux
ONNX is x86_64 only. Details, GPU backends and per-feature limits:
installation.
Installation
dependencies:
flutter_edge_ai: ^2.1.0
flutter_edge_ai_litertlm: ^1.9.0 # or any other engine from the table
Then complete the platform setup below.
Platform setup
iOS
Set the deployment target to 15.0 (16.0 if you use
flutter_edge_ai_mediapipe) and link pods statically in ios/Podfile:
platform :ios, '15.0'
use_frameworks! :linkage => :static
For large models add com.apple.developer.kernel.extended-virtual-addressing,
com.apple.developer.kernel.increased-memory-limit and
com.apple.developer.kernel.increased-debugging-memory-limit to
Runner.entitlements. The iOS Simulator runs on CPU only.
Android
Release builds need <uses-permission android:name="android.permission.INTERNET"/>
in android/app/src/main/AndroidManifest.xml to download models (Flutter's
template adds it only to the debug and profile manifests). Anything backed by
LiteRT β .litertlm, LiteRT embeddings, speech β needs minSdk 30 and is
arm64-v8a only. flutter_edge_ai_builtin_ai needs minSdk 26 (and macOS
12.0). The GPU and NPU uses-native-library entries still merge in from the
plugin manifest automatically.
Web
Copy cache_api.js and opfs_helper.js from this package's web/ folder into
your app's web/ and load them in index.html, then add the script for each
engine you use. For MediaPipe:
<script src="cache_api.js"></script>
<script src="opfs_helper.js"></script>
<script type="module">
import { FilesetResolver, LlmInference } from 'https://cdn.jsdelivr.net/npm/@mediapipe/tasks-genai@0.10.29';
window.FilesetResolver = FilesetResolver;
window.LlmInference = LlmInference;
</script>
The LiteRT-LM, ONNX, web embeddings and SQLite snippets are in the web setup guide. MediaPipe and LiteRT-LM run on WebGPU; ONNX also runs on CPU (WASM).
macOS
macOS needs a build step that stages the LiteRT-LM companion libraries into
the app (why), and it runs from a
CocoaPods Podfile. Flutter 3.44+ uses Swift Package Manager by default, so a new
app has no macos/Podfile: turn SPM off for the app in pubspec.yaml:
flutter:
config:
enable-swift-package-manager: false
Run flutter pub get (it writes macos/Podfile), then replace that Podfile's
post_install block with the one below; the next build runs pod install
itself. A green build proves nothing: without the block the first model load
fails with Library not loaded: @rpath/libGemmaModelConstraintProvider.dylib. If your app already has a macos/Podfile (another plugin
needs CocoaPods), keep SPM on and just paste the block. Prefer the pubspec
setting to flutter config --no-enable-swift-package-manager, which changes
only your machine β not your teammates' or CI.
post_install do |installer|
installer.pods_project.targets.each do |target|
flutter_additional_macos_build_settings(target)
end
# flutter_gemma: stage the upstream Apple companion dylibs into the built
# .app. `hook/build.dart` deliberately skips them from Native Assets on macOS
# (#247 β Google ships them without `-Wl,-headerpad_max_install_names`, so the
# JIT bundling path cannot rewrite their install_name), which leaves this
# build phase to stage them.
#
# The phase only LOCATES and RUNS a script; the staging logic itself lives in
# flutter_gemma_litertlm and is delivered next to the dylibs it stages. That
# is deliberate: this block is frozen into your Xcode project, and a copy of
# the logic frozen there cannot be fixed by upgrading the package.
installer.aggregate_targets.each do |aggregate_target|
aggregate_target.user_targets.each do |user_target|
phase_name = '[flutter_gemma] Setup LiteRT-LM macOS'
# Only the app target embeds the Frameworks/ this phase patches.
# RunnerTests inherits Runner's framework search paths and has no
# Contents/Frameworks of its own β having the phase there creates a
# cross-target dependency on Runner's framework output that Xcode reports
# as "Cycle inside Flutter Assemble" (#300). Remove any stale copy from
# non-app targets and skip them.
unless user_target.name == 'Runner'
user_target.build_phases
.select { |p| p.respond_to?(:name) && p.name == phase_name }
.each { |p| user_target.build_phases.delete(p) }
next
end
existing = user_target.shell_script_build_phases.find { |p| p.name == phase_name }
phase = existing || user_target.new_shell_script_build_phase(phase_name)
# The embedded LiteRtLm binary is an INPUT so the phase re-runs whenever
# Flutter's always-out-of-date `embed` phase re-copies the raw, unpatched
# binary over the patched one. Without it Xcode caches the phase after the
# first build and the second incremental build ships an unpatched
# LiteRtLm that fails dlopen at runtime (#368).
phase.input_paths = [
'$(BUILT_PRODUCTS_DIR)/$(PRODUCT_NAME).app/Contents/Frameworks/LiteRtLm.framework/Versions/A/LiteRtLm',
]
# A declared output lets Xcode order the phase in its dependency graph
# instead of treating it as "runs every build with no outputs" β the other
# half of the cycle warning (#300). The script touches this file.
phase.output_paths = ['$(DERIVED_FILE_DIR)/flutter_gemma_litertlm_macos.stamp']
phase.shell_script = <<~SHELL
set -e
STAGER="${HOME}/Library/Caches/flutter_gemma/native/macos_arm64/stage_macos_companions.sh"
if [ ! -f "${STAGER}" ]; then
echo "[flutter_gemma] ERROR: ${STAGER} not found." >&2
echo " flutter_gemma_litertlm 1.6.2+ installs it there from its build hook." >&2
echo " Upgrade the package, then: flutter clean && flutter pub get" >&2
exit 1
fi
sh "${STAGER}" "${BUILT_PRODUCTS_DIR}/${PRODUCT_NAME}.app/Contents/Frameworks"
mkdir -p "$(dirname "${SCRIPT_OUTPUT_FILE_0}")"
touch "${SCRIPT_OUTPUT_FILE_0}"
SHELL
end
end
end
Add to macos/Runner/DebugProfile.entitlements and Release.entitlements:
<key>com.apple.security.cs.disable-library-validation</key>
<true/>
<key>com.apple.security.network.client</key>
<true/>
Windows and Linux
Nothing to add: the native libraries are downloaded and bundled at build time.
GPU on Linux needs a vendor Vulkan driver (NVIDIA, AMD or Intel) β Mesa's
llvmpipe is not enough for Gemma 4.
Quick start
Initialization
Register the engines you added, once, in main():
import 'package:flutter/widgets.dart';
import 'package:flutter_edge_ai/flutter_edge_ai.dart';
import 'package:flutter_edge_ai_litertlm/flutter_edge_ai_litertlm.dart';
Future<void> main() async {
WidgetsFlutterBinding.ensureInitialized();
await FlutterEdgeAi.initialize(
inferenceEngines: const [LiteRtLmEngine()],
// Gated Hugging Face models need a token; never hard-code it.
huggingFaceToken: const String.fromEnvironment('HUGGINGFACE_TOKEN').isEmpty
? null
: const String.fromEnvironment('HUGGINGFACE_TOKEN'),
);
runApp(const MyApp());
}
Install a model
Once per device; installed models survive restarts.
await FlutterEdgeAi.installModel(
modelType: ModelType.gemmaIt,
fileType: ModelFileType.litertlm, // selects the engine β declare it
).fromNetwork(
'https://huggingface.co/litert-community/Gemma3-1B-IT/resolve/main/Gemma3-1B-IT_multi-prefill-seq_q4_ekv4096.litertlm',
).withProgress((progress) => print('Downloading: $progress%')).install();
litert-community/Gemma3-1B-IT is gated: accept its license on Hugging Face, then pass your token with --dart-define=HUGGINGFACE_TOKEN=β¦.
Chat
final model = await FlutterEdgeAi.getActiveModel(maxTokens: 2048);
final chat = await model.createChat();
await chat.addQueryChunk(Message.text(text: 'Explain quantum computing', isUser: true));
// Stream the reply token by token:
final reply = StringBuffer();
await for (final response in chat.generateChatResponseAsync()) {
if (response is TextResponse) reply.write(response.token);
}
await model.close();
Three things that fail quietly
maxTokensis the context window, not the reply length: prompt, history and answer together..litertlmneeds at least 1024. Cap the reply withcreateChat(maxOutputTokens: β¦).Message.isUserdefaults tofalse. A user message withoutisUser: truegets an empty answer.fileTypepicks the engine, not the file name.installModeldefaults toModelFileType.task, so a.litertlmmust sayModelFileType.litertlm.
Supported models
| Model | ModelType |
Function calling | Thinking | Vision / audio |
|---|---|---|---|---|
| Gemma 4 E2B / E4B Β² | gemma4 |
β | β | β / β |
| Gemma 3n E2B / E4B Β² | gemmaIt |
β ΒΉ | β | β / β |
| Gemma 3 1B, Gemma 3 270M | gemmaIt |
β | β | β |
| FunctionGemma 270M | functionGemma |
β | β | β |
| Qwen3 0.6B | qwen3 |
β | β | β |
| Qwen3.5 / 3.6 / 3.8 | qwen35 |
β | β | β |
| Qwen 2.5 0.5B / 1.5B | qwen |
β | β | β |
| DeepSeek R1 | deepSeek |
β | β | β |
| Phi-4 Mini | phi |
β | β | β |
| FastVLM, Qwen2-VL, SmolVLM2, LLaVA-OneVision | general |
β | β | β / β |
| SmolLM, SmolLM3, LFM2.5, Phi-4 Mini Reasoning | general |
β | β | β |
ΒΉ The downloadable E4B .litertlm.
Β² On Web: no audio input; Gemma 3n vision is native-only; Gemma 4 thinking needs the .litertlm web build.
Download links, sizes, formats per platform, embedding and speech models: models.
Going further
- Getting started β sessions, system instructions, removing models
- Models β model sources, Hugging Face installs
- Function calling Β· Multimodal Β· Thinking mode
- Embeddings & RAG Β· Speech Β· Agent skills
- Built-in AI Β· ONNX Runtime Β· Genkit Β· Memory diagnostics
- Desktop Β· Troubleshooting
- Fine-tune a small model and convert it to
.litertlmwith litetune (alpha).
Teach your AI assistant this package
flutter_edge_ai ships agent skills that teach Claude Code, Codex, Cursor and
other coding assistants this API:
dart run skills@ get --all
Links
- Migration guide Β· Changelog Β· Desktop support
- Example app Β· Issues Β· Contributing
Libraries
- core/api/embedding_installation_builder
- core/api/flutter_edge_ai
- core/api/inference_installation_builder
- core/api/stt_installation_builder
- core/api/tts_asset_listing
- core/api/tts_installation_builder
- core/chat
- core/chat_event
- core/di/download_hub/configure_download_updates_stream_mobile
- core/di/download_hub/configure_download_updates_stream_stub
- core/di/platform/mobile_service_factory
- Mobile-platform service factory. This file is only compiled on iOS/Android platforms. Uses background_downloader for model downloads.
- core/di/platform/web_service_factory
- Web-platform service factory. This file is only compiled on web platform. Uses dart:js_interop for browser-based downloads.
- core/di/service_registry
- core/domain/cache_metadata
- core/domain/download_error
- core/domain/download_exception
- core/domain/model_source
- core/domain/platform_types
- Shared platform value types for flutter_edge_ai.
- core/domain/web_storage_mode
- core/embedding/common_embedding_model
- core/embedding/common_embedding_model_stub
- core/embedding/embedder_backend_notice
- core/embedding/embedder_cache
- core/embedding/embedding_worker
- core/embedding/forward_pass
- core/embedding/pooling
- core/embedding/tokenizer_adapter
- core/extensions
- core/function_call_parser
- core/genai/genai_chat_extension
- core/genai/genai_input_converter
- core/genai/genai_output_converter
- core/handlers/asset_source_handler
- core/handlers/bundled_source_handler
- core/handlers/file_source_handler
- core/handlers/network_source_handler
- core/handlers/source_handler
- core/handlers/source_handler_registry
- core/handlers/web_asset_source_handler
- core/handlers/web_asset_source_handler_stub
- Stub implementation for non-web platforms This file is used when dart:js_interop is not available
- core/handlers/web_bundled_source_handler
- core/handlers/web_bundled_source_handler_stub
- core/handlers/web_file_source_handler
- core/handlers/web_file_source_handler_stub
- core/handlers/web_network_source_handler
- core/handlers/web_network_source_handler_stub
- core/image_error_handler
- core/image_processor
- core/image_tokenizer
- core/infrastructure/background_downloader_service
- core/infrastructure/blob_url_manager
- core/infrastructure/blob_url_manager_stub
- core/infrastructure/flutter_asset_loader
- core/infrastructure/flutter_asset_loader_stub
- Stub implementation for platforms where dart:io is not available (web) This file is used when large_file_handler cannot be imported
- core/infrastructure/in_memory_model_repository
- core/infrastructure/platform_file_system_service
- core/infrastructure/platform_file_system_service_stub
- core/infrastructure/url_utils
- core/infrastructure/web_cache_interop
- JavaScript interop for Cache API
- core/infrastructure/web_cache_interop_stub
- Stub implementation for non-web platforms
- core/infrastructure/web_cache_service
- Web cache service for persistent model storage
- core/infrastructure/web_cache_service_stub
- core/infrastructure/web_download_service
- core/infrastructure/web_download_service_stub
- core/infrastructure/web_file_system_service
- core/infrastructure/web_js_interop
- core/infrastructure/web_js_interop_stub
- core/infrastructure/web_opfs_interop
- JavaScript interop for OPFS (Origin Private File System)
- core/infrastructure/web_opfs_interop_stub
- Stub for OPFS interop on non-web platforms
- core/infrastructure/web_opfs_service
- Dart service wrapper for OPFS (Origin Private File System)
- core/lifecycle/close_notifier
- core/message
- core/model
- core/model_management/active_embedding_identity
- core/model_management/cancel_token
- core/model_management/constants/preferences_keys
- core/model_management/managers/web_model_manager
- core/model_management/model_activation
- core/model_management/model_specs
- Model specification value types β
dart:io-free, shared across all platforms (mobile, desktop, web). - core/model_management/utils/download_temp_reclaim
- Decision logic + filesystem walk for reclaiming orphaned
background_downloaderpartial temp files (#383), split out so both the pure predicate and the directory sweep are testable β the predicate as a pure unit test, the sweep as an on-device integration test β without the surroundingFileDownloader/path_provider/ gate machinery. - core/model_response
- core/multimodal_image_handler
- core/parsing/deepseek_function_call_format
- core/parsing/function_call_format
- core/parsing/function_call_format_factory
- core/parsing/function_gemma_format
- core/parsing/function_gemma_wire
- FunctionGemma's wire format, in one place.
- core/parsing/json_function_call_format
- core/parsing/json_parsing_utils
- core/parsing/llama_function_call_format
- core/parsing/phi_function_call_format
- core/parsing/qwen_function_call_format
- core/parsing/sdk_passthrough_function_call_format
- core/parsing/sdk_response_parser
- core/parsing/sdk_text_extractor
- core/registry/embedding_backend_provider
- core/registry/embedding_registry
- core/registry/embedding_tokenizer_provider
- core/registry/embedding_tokenizer_registry
- core/registry/engine_registry
- core/registry/hugging_face_resolver
- core/registry/hugging_face_resolver_registry
- core/registry/hugging_face_resolver_source
- core/registry/inference_engine_provider
- core/registry/runtime_config
- core/registry/skill_executor_provider
- core/registry/skill_executor_registry
- core/registry/stt_backend_provider
- core/registry/stt_registry
- core/registry/tts_backend_provider
- core/registry/tts_registry
- core/services/asset_loader
- core/services/download_service
- core/services/file_system_service
- core/services/model_repository
- core/services/protected_files_registry
- core/tool
- core/utils/edge_ai_log
- core/utils/file_name_utils
- core/utils/host_native_library
- core/vision_encoder_validator
- desktop/flutter_edge_ai_desktop
- desktop/flutter_edge_ai_desktop_stub
- flutter_edge_ai
- flutter_edge_ai_default
- flutter_edge_ai_default_web
- flutter_edge_ai_interface
- genai
- genai_primitives adoption surface for flutter_edge_ai (#181).
- mobile/flutter_edge_ai_mobile
- mobile/smart_downloader
- model_file_manager_interface
- web/flutter_edge_ai_web
- web/web_image_format
- web/web_model_source