llm_llamacpp 0.3.2
llm_llamacpp: ^0.3.2 copied to clipboard
llama.cpp backend implementation for LLM interactions. Enables local on-device inference with GGUF models on Android, iOS, macOS, Windows, and Linux.
llm_llamacpp #
Local LLM inference via llama.cpp for Dart and Flutter.
Available on pub.dev.
Part of the dart-llm ecosystem.
Features #
- Local on-device inference with GGUF models
- Streaming token generation
- Non-streaming responses - Get complete responses with
chatResponse() - Chat templates read from the GGUF itself and applied by llama.cpp
- Model-aware tool calling - Definitions advertised in the format the loaded model's family expects, calls parsed back out of the raw token stream
- Advanced generation options - Temperature, top-p, top-k, repeat penalty, frequency/presence penalties
- GPU acceleration support (CUDA, Metal, Vulkan)
- Cross-platform: Android, iOS, macOS, Windows, Linux
- Isolate-based inference (non-blocking UI)
- Model management - Discover, load, pool, and download models
- GGUF metadata - Read model info without loading
- Improved error handling - Specific exception types with detailed error messages
- Vision model support - Load and use vision models (image input processing coming soon)
Installation #
dependencies:
llm_llamacpp: ^0.3.2
Prerequisites #
GGUF model #
Download a model in GGUF format. Important: use properly converted GGUF files from trusted sources — see Model Compatibility.
Native library #
Nothing to do. The native library is a
native asset: hook/build.dart
resolves it during dart pub get / flutter build, and there is no manual step,
no jniLibs to copy and no library path to set.
How it resolves, in order:
- ABI fingerprint. A SHA-256 over
lib/src/bindings/llama_bindings.dartyields a 12-character fingerprint identifying the exact FFI surface this Dart code expects. - Prebuilt download. The hook fetches
llm_llamacpp-v<version>-abi<fingerprint>-<os>-<arch>.zipfrom this repo's GitHub release tagged<version>(a bare version, novprefix), where<version>is read from this package'spubspec.yaml. Because the fingerprint is part of the filename, a binding change can never be paired with a mismatched binary — the URL simply stops resolving. - Cache. Bundles are cached per
(version, fingerprint, os, arch), so a stale extract is never reused and one download serves every target. - Source build. On a 404 the hook configures and builds llama.cpp itself
with CMake. This requires the vendored submodule, so it works in a checkout
(
git submodule update --init) but not from a pub.dev install — the published archive excludesllamacpp/. For published versions the prebuilt is the only path.
The primary library and every ggml* library it links against are bundled
together, so a built app keeps working when moved off the build machine.
Prebuilts are produced by .github/workflows/build-release.yaml, which reads the
same version: field the hook does, so bumping the package version is what
triggers a new native release. Release tags are bare versions — 0.3.2, not
v0.3.2 — because that is the tag the hook's download URL is built from.
Two escape hatches, both optional:
| Variable | Effect |
|---|---|
LLM_LLAMACPP_LIB_DIR |
Directory searched first when loading the library. Useful for running pure-Dart scripts and tests against a local build under .dart_tool/. |
LLM_LLAMACPP_ANDROID_VULKAN |
Tri-state override for the Android arm64-v8a Vulkan backend, which is otherwise enabled automatically whenever glslc is on PATH. 0 forces CPU-only; 1 forces Vulkan and turns a missing glslc into a hard build error. |
Model Compatibility #
⚠️ Important: GGUF Source Matters #
Not all GGUF files are created equal. Different converters include different metadata, and llama.cpp requires specific keys for certain architectures.
Recommended Model Sources #
| Source | Compatibility | Notes |
|---|---|---|
| HuggingFace Official | ✅ Excellent | Models converted by the model authors (e.g., Qwen, Meta) |
| Unsloth | ✅ Excellent | High-quality conversions with imatrix quantization |
| TheBloke | ✅ Good | Wide variety of models |
| QuantFactory | ✅ Good | Many model options |
| Ollama model blobs | ⚠️ Limited | May be missing required metadata (see below) |
Ollama Models Compatibility #
Ollama stores models in /root/.ollama/models/blobs/ (or ~/.ollama/) as raw GGUF files. However, Ollama's converter may not include all metadata that llama.cpp requires.
| Model Type | Ollama Blob | HuggingFace GGUF |
|---|---|---|
| Standard LLMs (Llama, Mistral, Phi) | ✅ Works | ✅ Works |
| Qwen2/Qwen2.5 | ✅ Works | ✅ Works |
| Qwen3-VL (vision) | ❌ Missing rope.dimension_sections |
✅ Works |
| Gemma3 (vision) | ❌ Missing attention.layer_norm_rms_epsilon |
✅ Works |
| Other vision models | ⚠️ May have issues | ✅ Recommended |
If you encounter errors like:
error loading model hyperparameters: key not found in model: <arch>.rope.dimension_sections
Download the model directly from HuggingFace instead of using Ollama's blob.
Recommended Models #
Small & Fast (< 1GB)
- Qwen/Qwen2.5-0.5B-Instruct-GGUF (~400MB)
Balanced (1-5GB)
- unsloth/Llama-3.2-1B-Instruct-GGUF (~800MB)
- unsloth/Phi-3.5-mini-instruct-GGUF (~2GB)
- unsloth/Qwen2.5-3B-Instruct-GGUF (~2GB)
High Quality (5-20GB)
- unsloth/Llama-3.1-8B-Instruct-GGUF (~5GB Q4)
- unsloth/Qwen2.5-14B-Instruct-GGUF (~9GB Q4)
Vision Models
- unsloth/Qwen3-VL-8B-Instruct-GGUF (~5GB Q4)
- Qwen/Qwen3-VL-4B-Instruct-GGUF (~3GB Q4)
Note: Vision model support is for text inference only. Image input requires additional multimodal bindings (not yet implemented).
Usage #
Example App (Recommended) #
For a complete, production-ready example demonstrating real-world usage, see the example_app:
- Full Flutter app with chat interface
- Model download from HuggingFace with progress tracking
- Tool calling demonstration (calculator tool)
- Mobile platform support (Android/iOS)
- Offline inference after model download
cd example_app
flutter run
The example app is the recommended starting point for understanding how to integrate llm_llamacpp into a Flutter application.
CLI Example (Simple) #
For a minimal command-line example, see example/cli_example.dart:
dart run example/cli_example.dart /path/to/model.gguf
Simplified Model Acquisition (getModel) #
The easiest way to get models from HuggingFace - deterministic, no guessing:
import 'package:llm_llamacpp/llm_llamacpp.dart';
final repo = LlamaCppRepository();
// GGUF repo - auto-downloads Q4_K_M variant
final path = await repo.getModel(
'Qwen/Qwen2.5-0.5B-Instruct-GGUF',
outputDir: '/models/',
);
print('Model ready: $path');
// Specific quantization
final path = await repo.getModel(
'unsloth/Llama-3.2-1B-Instruct-GGUF',
outputDir: '/models/',
quantization: QuantizationType.q5_k_m, // Q5_K_M variant
);
// Safetensors repo - MUST specify quantization
final path = await repo.getModel(
'meta-llama/Llama-3.2-1B',
outputDir: '/models/',
quantization: QuantizationType.q4_k_m, // Required for conversion
);
// Specific file (bypass matching)
final path = await repo.getModel(
'Qwen/Qwen2.5-0.5B-Instruct-GGUF',
outputDir: '/models/',
preferredFile: 'qwen2.5-0.5b-instruct-q8_0.gguf',
);
repo.dispose();
With Progress Updates
await for (final status in repo.getModelStream(
'Qwen/Qwen2.5-0.5B-Instruct-GGUF',
outputDir: '/models/',
)) {
print('${status.stage.name}: ${status.message}');
if (status.progress != null) {
print(' Progress: ${status.progressPercent}');
}
if (status.isComplete) {
print('Ready: ${status.modelPath}');
}
}
Error Handling #
The API is deterministic - it throws clear errors instead of guessing:
try {
final path = await repo.getModel(
'some/model-repo',
outputDir: '/models/',
);
} on ModelNotFoundException catch (e) {
// No exact quantization match found
// e.availableFiles lists what's available
print(e);
// "ModelNotFoundException: No Q4_K_M GGUF found in 'some/model-repo'.
// Available GGUF files:
// - model-q5_k_m.gguf (1.2 GB)
// - model-q8_0.gguf (2.1 GB)
// Specify file: getModel('some/model-repo', preferredFile: 'model-q5_k_m.gguf')"
} on AmbiguousModelException catch (e) {
// Multiple files match the quantization
// e.matchingFiles lists all matches
print(e);
// "AmbiguousModelException: Multiple Q4_K_M files found in 'some/model-repo':
// - model-q4_k_m.gguf
// - model-v2-q4_k_m.gguf
// Specify file: getModel('some/model-repo', preferredFile: 'model-v2-q4_k_m.gguf')"
} on ConversionRequiredException catch (e) {
// Only safetensors, must specify quantization
print(e);
// "ConversionRequiredException: Quantization required for safetensors conversion.
// Repository: 'some/model-repo' only has safetensors (no GGUF).
// Example: getModel('some/model-repo', quantization: QuantizationType.q4_k_m)"
} on UnsupportedModelException catch (e) {
// No GGUF or safetensors at all
print(e);
}
Model Management (LlamaCppRepository) #
The LlamaCppRepository also provides discovery, loading/pooling, and low-level downloads:
import 'package:llm_llamacpp/llm_llamacpp.dart';
final repo = LlamaCppRepository();
// Discover models in a directory
final models = await repo.discoverModels('/path/to/models');
for (final model in models) {
print('${model.name}: ${model.metadata?.sizeLabel} (${model.fileSizeLabel})');
}
// Read GGUF metadata without loading
final metadata = await GgufMetadata.fromFile('/path/to/model.gguf');
print('Architecture: ${metadata.architecture}');
print('Parameters: ${metadata.sizeLabel}');
print('Quantization: ${metadata.quantizationType}');
print('Context length: ${metadata.contextLength}');
// Load models with pooling (reference counting)
final model = await repo.loadModel('/path/to/model.gguf');
print('Loaded, ref count: ${repo.getModelRefCount(model.path)}');
// Load same model again (reuses existing, increments ref count)
final model2 = await repo.loadModel('/path/to/model.gguf');
print('Ref count now: ${repo.getModelRefCount(model.path)}'); // 2
// Unload (decrements ref count, disposes when 0)
repo.unloadModel('/path/to/model.gguf');
// Check system capabilities
for (final backend in repo.getAvailableBackends()) {
print('${backend.name}: ${backend.isAvailable ? "✓" : "✗"} ${backend.deviceName ?? ""}');
}
print('Recommended GPU layers: ${repo.recommendedGpuLayers}');
// Low-level download (specific file)
await for (final progress in repo.downloadModel(
'Qwen/Qwen2.5-0.5B-Instruct-GGUF',
'qwen2.5-0.5b-instruct-q4_k_m.gguf',
'/path/to/models/',
)) {
print('${progress.progressPercent} - ${progress.status}');
}
// Check what's available for a model
final plan = await repo.planModelAcquisition('Qwen/Qwen2.5-0.5B-Instruct');
print(plan); // "Download GGUF: ..." or "Convert safetensors → GGUF"
repo.dispose();
Converting Safetensors to GGUF (Manual) #
For manual conversion with full control:
final repo = LlamaCppRepository();
// Check if conversion is needed
final plan = await repo.planModelAcquisition('meta-llama/Llama-3.2-1B');
if (plan.method == AcquisitionMethod.convertFromSafetensors) {
// Convert safetensors → GGUF with Q4_K_M quantization
await for (final progress in repo.convertModel(
repoId: 'meta-llama/Llama-3.2-1B',
outputPath: '/path/to/llama-3.2-1b-q4.gguf',
quantization: QuantizationType.q4_k_m,
llamaCppPath: '/path/to/llama.cpp', // Optional, will auto-detect
)) {
print('${progress.stage}: ${progress.message}');
}
}
Requirements for conversion:
- Python 3.8+ with:
pip install transformers torch safetensors sentencepiece - A llama.cpp checkout. In this repository it is already vendored at
packages/llm_llamacpp/llamacpp(git submodule update --init); otherwisegit clone https://github.com/ggml-org/llama.cpp - Build the quantize tool (llama.cpp is CMake-only; the old
maketargets are gone):cmake -B build -DLLAMA_BUILD_COMMON=ON cmake --build build --target llama-quantize --config Release
Available quantization types:
| Type | Size | Quality | Use Case |
|---|---|---|---|
q4_k_m |
~4.0x smaller | Good | Recommended default |
q5_k_m |
~3.2x smaller | Better | Balanced quality/size |
q6_k |
~2.7x smaller | High | Near-original quality |
q8_0 |
~2x smaller | Excellent | Minimal quality loss |
q3_k_m |
~5.3x smaller | Lower | Memory constrained |
q2_k |
~8x smaller | Lowest | Extreme compression |
Basic Chat (Streaming) #
import 'package:llm_llamacpp/llm_llamacpp.dart';
final repo = LlamaCppChatRepository(
contextSize: 2048,
nGpuLayers: 0, // Set > 0 for GPU acceleration
);
try {
await repo.loadModel('/path/to/model.gguf');
final stream = repo.streamChat('model', messages: [
LLMMessage(role: LLMRole.system, content: 'You are helpful.'),
LLMMessage(role: LLMRole.user, content: 'Hello!'),
]);
await for (final chunk in stream) {
print(chunk.message?.content ?? '');
}
} finally {
repo.dispose();
}
Conversation Continuity #
Maintain conversation history by passing all previous messages:
// First message
final messages = [
LLMMessage(role: LLMRole.user, content: 'What is 2+2?'),
];
var stream = repo.streamChat('model', messages: messages);
String response1 = '';
await for (final chunk in stream) {
response1 += chunk.message?.content ?? '';
}
// Continue conversation - add assistant response and new user message
messages.add(LLMMessage(role: LLMRole.assistant, content: response1));
messages.add(LLMMessage(role: LLMRole.user, content: 'What about 3+3?'));
stream = repo.streamChat('model', messages: messages);
String response2 = '';
await for (final chunk in stream) {
response2 += chunk.message?.content ?? '';
}
// The model sees the full conversation history
print('Response 1: $response1');
print('Response 2: $response2');
Non-Streaming Chat #
Get a complete response without streaming:
final repo = LlamaCppChatRepository();
await repo.loadModel('/path/to/model.gguf');
final response = await repo.chatResponse('model', messages: [
LLMMessage(role: LLMRole.user, content: 'What is 2+2?'),
]);
print(response.content); // Complete response
print('Tokens used: ${response.evalCount}');
repo.dispose();
Chat Templates #
There is nothing to configure. Every GGUF embeds its own chat template, and the
package applies it through llama.cpp's llama_chat_apply_template(). The
template classes this package used to expose (ChatMLTemplate, Llama3Template,
getTemplateForModel, …) were removed in 0.1.5 — hand-picking a template was a
reliable way to disagree with what the model was actually trained on.
GPU Acceleration #
final repo = LlamaCppChatRepository(
nGpuLayers: 99, // Offload all layers to GPU
);
await repo.loadModel('/path/to/model.gguf', options: ModelLoadOptions(
nGpuLayers: 99, // Use GPU for all layers
useMemoryMap: true,
));
CUDA Setup (NVIDIA)
Requires CUDA toolkit 12.4+ for modern GPUs:
# Ubuntu/Debian
wget https://developer.download.nvidia.com/compute/cuda/repos/ubuntu2404/x86_64/cuda-keyring_1.1-1_all.deb
sudo dpkg -i cuda-keyring_1.1-1_all.deb
sudo apt-get update
sudo apt-get install cuda-toolkit-12-8
# Set environment
export PATH=/usr/local/cuda/bin:$PATH
export LD_LIBRARY_PATH=/usr/local/cuda/lib64:$LD_LIBRARY_PATH
Tool Calling #
Pass tools: and the package handles the rest — do not hand-write tool
syntax into the system prompt, which fights the format the model was trained on:
final stream = repo.streamChat(modelPath,
messages: messages,
tools: [MyTool()],
);
The package detects the model's tool-call family from its GGUF chat template,
falling back to probing the tokenizer vocabulary for the family's opening
delimiter (necessary because GGUF conversions often ship a template with the
tools branch stripped). It then advertises the tool definitions in that
family's format and parses calls back out of the raw token stream. Supported
families: LFM2/LFM2.5 (<|tool_call_start|> with Pythonic calls), Hermes/Qwen
(<tool_call>), Mistral ([TOOL_CALLS]), Llama 3.x (<|python_tag|>), plus
bare Pythonic call lists and bare JSON. Add one in
lib/src/tool_calls/tool_call_syntax.dart.
Tools are executed internally and their results are not surfaced as
role: tool chunks. A UI that wants to show them needs the tool to report its
own invocations — see example_app's CalculatorTool(onInvoke: ...).
Replay rawContent across turns
If you maintain conversation history yourself, append the assistant turn from
chunk.message.rawContent, not from the visible text:
String? rawTurn;
await for (final chunk in stream) {
rawTurn = chunk.message?.rawContent ?? rawTurn;
}
messages.add(LLMMessage(role: LLMRole.assistant, content: rawTurn ?? visible));
rawContent keeps the tool-call markup that is stripped from content. Replay
only the visible text and history ends up showing the assistant announcing a
tool and then answering without calling one — the model copies that pattern and
stops calling tools after the first turn.
Platform Support #
| Platform | Architecture | GPU Support |
|---|---|---|
| Linux | x86_64 | CUDA, Vulkan |
| macOS | arm64, x86_64 | Metal |
| Windows | x86_64 | CUDA, Vulkan |
| Android | arm64-v8a | Vulkan (auto when glslc is available) |
| Android | x86_64 | - |
| iOS | arm64 | Metal |
Configuration #
LlamaCppChatRepository(
contextSize: 4096, // Token context window
batchSize: 512, // Batch size for processing
threads: null, // null = auto-detect
nGpuLayers: 0, // Layers to offload to GPU (99 = all)
maxToolAttempts: 90, // Max tool calling iterations (default)
);
Troubleshooting #
Library not found #
Under Flutter the library ships inside the app bundle (on macOS/iOS as
llama.framework/llama, on Android as a JNI library) and this should not happen
— if it does, the build hook failed; check the pub get / build output for its
messages.
For pure-Dart programs the loader searches, in order:
LLM_LLAMACPP_LIB_DIR, if set- The current directory
- The usual system locations
The hook's output lives under .dart_tool/, so point the override at it:
export LLM_LLAMACPP_LIB_DIR=$(dirname $(find .dart_tool/hooks_runner \
\( -name 'libllama.*' -o -name 'llama.dll' \) | head -1))
If the hook could not download a prebuilt and you are working from a checkout,
make sure the submodule is present (git submodule update --init) and CMake is
installed so the source build can run.
Model loading errors #
"key not found in model" - The GGUF is missing required metadata. Download from a different source (see Model Compatibility above).
"Failed to load model" - Check file path, permissions, and that it's a valid GGUF file.
Out of memory #
- Use a smaller model (Q4_K_M or Q4_0 quantization)
- Reduce context size
- Offload layers to GPU with
nGpuLayers
Slow inference #
- Enable GPU acceleration (
nGpuLayers: 99) - Use a more aggressively quantized model (Q4_0 vs Q8_0)
- Reduce context size
- Increase batch size
CUDA errors #
- Ensure CUDA toolkit version matches your driver
- Check
nvidia-smifor driver CUDA version - Set
LD_LIBRARY_PATHto include CUDA libs
Error Handling #
The package provides specific exception types for better error handling:
try {
final stream = chatRepo.streamChat('model', messages: messages);
await for (final chunk in stream) {
print(chunk.message?.content ?? '');
}
} on ModelLoadException catch (e) {
print('Failed to load model: ${e.message}');
if (e.modelPath != null) {
print('Model path: ${e.modelPath}');
}
} on TokenizationException catch (e) {
print('Tokenization failed: ${e.message}');
if (e.prompt != null) {
print('Problematic prompt: ${e.prompt}');
}
} on ContextCreationException catch (e) {
print('Context creation failed: ${e.message}');
print('Requested contextSize: ${e.contextSize}');
print('Requested batchSize: ${e.batchSize}');
} on InferenceException catch (e) {
print('Inference error: ${e.message}');
if (e.details != null) {
print('Details: ${e.details}');
}
} on VisionNotSupportedException catch (e) {
print('Vision not supported: ${e.message}');
}
Generation Options #
Fine-tune generation behavior with GenerationOptions:
final options = GenerationOptions(
temperature: 0.8, // Higher = more creative
topP: 0.95, // Nucleus sampling threshold
topK: 40, // Top-K sampling limit
maxTokens: 1024, // Maximum tokens to generate
seed: 42, // For reproducible outputs
repeatPenalty: 1.1, // Penalty for repetition (>1.0 discourages)
frequencyPenalty: 0.5, // Penalty based on token frequency
presencePenalty: 0.3, // Penalty for token presence
);
final stream = chatRepo.streamChatWithGenerationOptions(
'model',
messages: messages,
generationOptions: options,
);
Performance Tips #
-
GPU Acceleration: Always enable GPU layers when available:
final repo = LlamaCppChatRepository(nGpuLayers: 99); -
Context Size: Use the minimum context size needed:
final repo = LlamaCppChatRepository(contextSize: 2048); // Instead of 4096 -
Batch Size: Increase batch size for faster processing:
final repo = LlamaCppChatRepository(batchSize: 1024); -
Model Quantization: Use Q4_K_M for best balance of size and quality.
-
Memory Mapping: Enable memory mapping for large models:
final model = await repo.loadModel( '/path/to/model.gguf', options: ModelLoadOptions(useMemoryMap: true), );