llama_cpp_flutter 0.6.1
llama_cpp_flutter: ^0.6.1 copied to clipboard
On-device LLM inference for Flutter: llama.cpp with Metal on iOS/macOS, wllama on web. Streaming sessions, chat formats, model downloads, and memory-aware multi-agent orchestration.
llama_cpp_flutter #
On-device LLM inference for Flutter, backed by llama.cpp: load a local GGUF model and stream generated text through one cross-platform API. Chat formats, model downloading, KV-cache snapshots, multimodal input, and memory-aware multi-agent orchestration are layered on top — use as much or as little as you need.
Platform support #
| Platform | Status | Backend |
|---|---|---|
| iOS | Supported | llama.cpp xcframework, Metal |
| macOS | Supported | llama.cpp xcframework, Metal |
| Web | Supported | wllama, Wasm |
| Android | Not yet | see Roadmap |
| Windows | Not yet | — |
| Linux | Not yet | — |
On unsupported platforms createLlamaRuntime() throws an
UnsupportedError naming the platform, rather than failing later with a
missing-plugin error.
Quick start #
createLlamaRuntime() returns the platform's LlamaRuntime — the
llama.cpp plugin on iOS/macOS, wllama on web. Describe the model with a
ModelSpec, load it, and stream:
import 'dart:io';
import 'package:llama_cpp_flutter/llama_cpp_flutter.dart';
Future<void> main() async {
final runtime = createLlamaRuntime();
final spec = ModelSpec(
id: 'gemma-3-4b-it-q4km',
displayName: 'Gemma 3 4B IT',
modelUrl: huggingFaceModelUri(
repo: 'unsloth/gemma-3-4b-it-GGUF',
file: 'gemma-3-4b-it-Q4_K_M.gguf',
),
contextSize: 4096,
format: resolveChatFormat('gemma')!,
);
final session = await runtime.loadModel(
spec,
localPath: '/path/to/gemma-3-4b-it-Q4_K_M.gguf',
);
await for (final text in session.generate('Why is the sky blue?')) {
stdout.write(text);
}
await session.dispose();
}
generate consumes a raw prompt string. For multi-turn conversations,
render the prompt through the spec's ChatFormat — or skip prompt handling
entirely and use the ChatClient adapter.
Where the model file comes from #
-
Native, already on device: pass
localPathas above (e.g. a user-picked file, or a path your app downloaded earlier). -
Native, downloaded for you: give the runtime an artifact cache directory and omit
localPath— the spec's URLs (model, optional multimodal projector, optional draft model) are fetched into it first, skipping files already present:final runtime = createLlamaRuntime( artifactCacheDirectory: cacheDir.path, ); final session = await runtime.loadModel( spec, onProgress: (progress) => debugPrint('load: $progress'), );For manual control over downloads (auth tokens, separate progress per artifact), use
ModelDownloader.downloadSpecArtifactsand pass the resulting paths toloadModelyourself. -
Web: omit
localPathand the runtime streamsspec.modelUrldirectly (Hugging FaceresolveURLs send the CORS headers browsers need).huggingFaceModelUribuilds those URLs on every platform. -
Any platform, app-managed: build an
ArtifactStoreand let it own the bytes. Same API everywhere — a directory you name on iOS/macOS, the origin-private file system in the browser — so an app that shows users what is on their device, how much space it takes, and lets them delete it needs one code path:final store = createArtifactStore(directory: cacheDir.path); // web: ignored final model = await store.fetch(spec.modelUrl, onProgress: onProgress); final mmproj = await store.fetch(spec.mmprojUrl!); final paths = await store.resolve( modelKey: model.key, mmprojKey: mmproj.key, ); final session = await runtime.loadModel( spec, localPath: paths.modelPath, localMmprojPath: paths.mmprojPath, );Downloads resume, imports (
importFileon native,importStreamanywhere) land atomically, andlist,totalSizeBytes, anddeletecover the rest. On the web the resolved paths are opaque handles the runtime dereferences straight to their stored files, so a model past the ~2 GiB wasm32 per-file limit is still split and staged transparently instead of being read back through the network stack. Speculative-decoding draft models remain native-only; don't passdraftKeyon the web.
Web notes #
The wllama Wasm binary ships as a Flutter asset — no extra setup. For
multi-threaded inference the page must be cross-origin isolated (served
with Cross-Origin-Opener-Policy: same-origin and
Cross-Origin-Embedder-Policy: require-corp); otherwise wllama silently
falls back to a single thread, which is slow enough to look like a hang for
multi-billion-parameter models. Check
runtime.supportsMultiThreading and surface a warning to users when it is
false.
Chat formats #
A ChatFormat renders conversation turns into a model family's wire format
and decodes its output (including streamed tool calls). Built-ins cover
Gemma, Llama 3, Mistral, Qwen, LFM2/LFM2.5, and generic ChatML; resolve one
by name with resolveChatFormat, register your own with
registerChatFormat.
You can also detect the right format from the model file itself —
detectChatFormatNameForGguf reads the GGUF's architecture and embedded
chat template:
final name = detectChatFormatNameForGguf(headerBytes);
final format = resolveChatFormat(name) ?? resolveChatFormat('chatml')!;
Using with an agent framework #
createLlamaChatClient wraps a loaded session in a ChatClient, handling
prompt rendering, sampling defaults, tool-call decoding, and multimodal
turns:
final client = createLlamaChatClient(
spec: spec,
sessionProvider: () async => session,
);
final response = await client.getResponse(
messages: [ChatMessage.fromText(ChatRole.user, 'Hello!')],
);
print(response.text);
The client plugs into anything that accepts a ChatClient from
package:extensions/ai.dart. This package depends only on extensions, so
the agent framework — if any — is entirely your choice; wrap the client in
whatever agent abstraction it provides.
Structured generation events #
generateEvents is the typed alternative to the raw string stream: text
arrives as events, and a successful run always ends with a completed event
carrying the engine's stats — so the finish reason and token accounting
cannot be missed. Failures surface as stream errors.
await for (final event in session.generateEvents('Why is the sky blue?')) {
switch (event) {
case LlamaTextEvent(:final text):
stdout.write(text);
case LlamaCompletedEvent(:final stats):
debugPrint('finished: ${stats?.finishReason}');
case LlamaWarningEvent(:final message):
debugPrint('warning: $message');
}
}
Smoothing the stream for display #
Tokens arrive in bursts — a prompt-processing pause, then several tokens in
one frame, then nothing while the next batch decodes — which makes streamed
text jump and stutter on screen. TokenSmoother (from chat.dart) re-paces
a Stream<String> into a steady, typewriter-style grapheme stream.
import 'package:llama_cpp_flutter/chat.dart';
import 'package:llama_cpp_flutter/llama_cpp_flutter.dart';
final display = session.generate('Why is the sky blue?').smoothed();
The release rate tracks the arrival rate rather than draining the buffer
on a fixed schedule. The smoother keeps a running estimate of how fast text
is arriving and releases at that rate, holding roughly window of text in
reserve and using the backlog only as a correction; the rate itself is
low-pass filtered over smoothing, so a change in generation speed comes out
as a ramp instead of a jump. Because the rate is fractional and carried
across frames, it can emit slower than one grapheme per frame — which is
what lets it match a slow model instead of outrunning it.
The difference, measured against a 5 tok/s source (4 graphemes every 200ms), looking at the gap between consecutive graphemes:
| Pacing | median gap | worst gap |
|---|---|---|
| drain the backlog over a fixed window | 16 ms | 152 ms |
| track the arrival rate | 48 ms | 64 ms |
The fixed-window version empties its buffer in four frames and then stalls for the rest of the token interval — the stutter is still there, just moved. Tracking the rate spreads the same four graphemes evenly across the interval. At 60 tok/s the same code emits several graphemes every frame; only the speed changes.
Grapheme clusters stay intact (emoji ZWJ sequences are never split
mid-render), and an atomic predicate releases matching chunks whole — use
it for in-band markers that must not appear half-formed:
stream.smoothed(atomic: (chunk) => chunk.startsWith('<tool_call>'));
Smoothing is opt-in and never applied inside generation: it trails the source
by about window and runs a periodic timer, which headless, batch, and test
callers don't want. Apply it at the widget that renders the text. When the
source ends, the remaining backlog drains within window. Cancelling the
smoothed subscription cancels the upstream one, so stopping a generation
still propagates back to the runtime.
Concurrency and lifecycle #
- A session runs one generation at a time. Calling
generatewhile a run is in flight supersedes it: the engine cancels the current run and starts the newest call once cancellation completes. Prefer awaiting or cancelling the previous run explicitly. session.cancel()stops the in-flight run; the stream completes normally and the stats callback (when the engine reports one) carriesLlamaFinishReason.cancelled. Cancelling the stream subscription has the same effect.session.dispose()is safe during generation — the stream closes, though a stats callback may not be delivered.maxSequences > 1gives a session independent KV-cache sequences (separate conversations sharing one loaded model), not parallel decode lanes;sequenceIdselects which one a run extends.
See the LlamaSession API docs for the full contract, including KV-cache
persistence (saveState/loadState) and in-memory stashing
(stashState/restoreStashedState).
Multi-agent orchestration #
The orchestration layer runs several logical agents over one loaded model:
per-agent KV-cache swapping via sequences and stashes, GGUF-driven memory
estimation, and dynamic context budgeting. It is entirely optional and
lives in its own entrypoint — see doc/orchestration.md and SPEC.md:
import 'package:llama_cpp_flutter/orchestration.dart';
Package layout #
The main entrypoint stays deliberately small; deeper layers have their own imports:
| Entrypoint | Contents |
|---|---|
llama_cpp_flutter.dart |
Runtime, sessions, ModelSpec, downloads, format resolution, ChatClient adapter |
chat.dart |
Concrete chat formats, templates, stream decoders, LlamaChatClient, prompt diagnostics, token-stream transformers |
gguf.dart |
GGUF metadata reading, artifact cache naming |
orchestration.dart |
Multi-agent orchestration over one loaded model |
bridge.dart |
Low-level iOS/macOS plugin bridge |
Setup: the vendored xcframework (iOS/macOS) #
The xcframework is large and is not committed. It is downloaded from the
official llama.cpp releases
automatically during pod install (the podspec runs the fetch script, which
is a no-op once the installed framework matches the pin in
tool/versions.env). Manual install / options:
./scripts/fetch_llama_xcframework.sh
# Try a different upstream release; its zip checksum is required:
LLAMA_CPP_TAG_OVERRIDE=<tag> \
LLAMA_XCFRAMEWORK_ZIP_SHA256_OVERRIDE=<sha256> \
./scripts/fetch_llama_xcframework.sh
# Development-only escape hatch when you don't have the checksum yet
# (prints a loud warning; never ship a build produced this way):
LLAMA_CPP_TAG_OVERRIDE=<tag> LLAMA_CPP_ALLOW_UNVERIFIED=1 \
./scripts/fetch_llama_xcframework.sh
# Skip the automatic fetch during pod install (offline/lint environments):
LLAMA_CPP_FLUTTER_SKIP_FETCH=1 pod install
Downloads are always verified against the sha256 pinned in
tool/versions.env (or the explicit override above) before unpacking.
Both backends track upstream automatically: a weekly workflow re-pins the
native xcframework to the latest llama.cpp release and rebuilds the wllama
wasm against that same tag, opening a PR gated by CI (ABI check + macOS
link). Manual equivalents: tool/update_deps.sh and
tool/update_wllama.sh.
This writes darwin/Frameworks/llama.xcframework. As a fallback, build from
source with ./scripts/build_llama_xcframework.sh (requires Xcode + CMake;
LLAMA_REF=<tag-or-commit> to pin, LLAMA_ALL_PLATFORMS=1 for every Apple
slice).
The example/ app is a minimal harness that links the plugin; CI builds it on
macOS so every PR exercises compile + link against the pinned framework. Run a
real model through it with:
cd example
flutter test integration_test/model_smoke_test.dart -d macos \
--dart-define=MODEL_PATH=/absolute/path/to/model.gguf
Advanced: the native bridge #
package:llama_cpp_flutter/bridge.dart is the low-level iOS/macOS plugin
bridge underneath the neutral API. Most apps never need it; reach for it
only for Apple-specific control the neutral API doesn't expose. Its
session type is LlamaBridgeSession, so it coexists cleanly with the
neutral LlamaSession.
import 'package:llama_cpp_flutter/bridge.dart';
final llama = LlamaCppFlutter();
final session = await llama.loadModel('/path/to/model.gguf');
await for (final token in session.generate('Hello, world!')) {
stdout.write(token);
}
await session.dispose();
await llama.shutdown();
How it works:
- Native bridge: Pigeon — a typed
@HostApifor control and an@EventChannelApitoken stream. - Threading: all native calls run from a dedicated Dart worker
isolate (bound with
BackgroundIsolateBinaryMessenger), so model loading and streaming never block the UI. - Backend: a vendored
llama.xcframework(Metal-enabled).
Roadmap #
- Android is the most-requested missing platform and the next target
for a native backend; there is no committed timeline yet. Windows and
Linux are open beyond that. The neutral
LlamaRuntimeAPI is designed so new backends slot in without app-code changes. - The API is pre-1.0 and may still see breaking changes; they are called out per release in the changelog.
Requirements & notes #
- SDK: Dart ^3.9.0 / Flutter >=3.35.0.
- Deployment targets: iOS 16.4 / macOS 13.3 (Metal build minimums).
- Models are loaded from a runtime file path — nothing is bundled. On
sandboxed macOS the app needs
com.apple.security.files.user-selected.read-only. - The iOS Simulator has limited Metal support; pass
gpuLayers: 0to force CPU there. - Model weights have their own licenses — shipping a model in your app means complying with that model's license, not this package's.
- Regenerate the Pigeon bridge after editing
pigeons/messages.dart:dart run pigeon --input pigeons/messages.dart.