Xybrid Flutter SDK

Run LLMs, ASR, and TTS natively in Flutter apps — private, offline, no cloud required.

pub package License: Apache 2.0

Installation

flutter pub add xybrid_flutter

Or add to your pubspec.yaml:

dependencies:
  xybrid_flutter: ^0.6.0
Alternative installation (git / local path)

From git (unreleased changes):

dependencies:
  xybrid_flutter:
    git:
      url: https://github.com/xybrid-ai/xybrid.git
      ref: main
      path: bindings/flutter

Quick Start

import 'package:xybrid_flutter/xybrid_flutter.dart';

Future<void> main() async {
  WidgetsFlutterBinding.ensureInitialized();

  // Runs locally with no key. Pass an apiKey to light up the dashboard:
  //   await Xybrid.init(apiKey: const String.fromEnvironment('XYBRID_API_KEY'));
  await Xybrid.init();

  // Load a TTS model from the registry
  final model = await XybridModelLoader.fromRegistry('kokoro-82m').load();

  // Run text-to-speech
  final result = await model.run(XybridEnvelope.text('Hello from Xybrid!'));
  print('Audio: ${result.audioBytes?.length} bytes');
}

Inference runs entirely on-device whether or not you authenticate. Without an apiKey, telemetry is disabled and the first inference logs a one-shot hint pointing at the dashboard (suppress with XYBRID_QUIET=1). Get a free key at dashboard.xybrid.dev.

Features

Model Loading

Load models from the Xybrid registry or local bundles:

// From registry (downloads + caches automatically)
final model = await XybridModelLoader.fromRegistry('kokoro-82m').load();

// From local bundle
final model = await XybridModelLoader.fromBundle('path/to/model.xyb').load();

// Check if already cached
if (Xybrid.isModelCached('kokoro-82m')) {
  print('Model ready, no download needed');
}

Download Progress

Track model downloads with progress events:

final loader = XybridModelLoader.fromRegistry('kokoro-82m');

await for (final event in loader.loadWithProgress()) {
  switch (event) {
    case LoadProgress(:final progress, :final downloadedBytes, :final totalBytes):
      print('Downloading: ${(progress * 100).toInt()}% '
          '($downloadedBytes / ${totalBytes ?? '?'} bytes)');
    case LoadComplete():
      print('Model ready!');
    case LoadError(:final message):
      print('Error: $message');
  }
}

progress spans every file the model needs and never moves backwards. totalBytes is null when the source publishes no size; downloadedBytes is exact either way, so megabytes, speed and time remaining are all derivable.

Input Envelopes

Type-safe inputs for different model types:

// Text (for TTS or LLM)
final textInput = XybridEnvelope.text('Hello world');

// Text with TTS voice selection
final ttsInput = XybridEnvelope.text('Hello', voiceId: 'af_heart', speed: 1.0);

// Audio (for ASR / speech-to-text)
final audioInput = XybridEnvelope.audio(
  bytes: wavBytes,
  sampleRate: 16000,
  channels: 1,
);

// Embedding vector
final embeddingInput = XybridEnvelope.embedding([0.1, 0.2, 0.3]);

// Vision-language prompt with an encoded image
final image = XybridEnvelope.image(bytes: pngBytes, format: 'png');
final visionInput = XybridEnvelope.userMessage(
  text: 'Describe this image.',
  images: [image],
);

Inference Results

final result = await model.run(envelope);

if (result.success) {
  // Text output (ASR transcription or LLM response)
  print(result.text);

  // Audio output (TTS) — get as WAV for playback
  final wav = result.audioAsWav(sampleRate: 24000, channels: 1);

  // Embedding output
  print(result.embedding);

  // Inference timing
  print('Latency: ${result.latencyMs}ms');
}

Reasoning (thinking models)

Reasoning models (metadata reasoning: true, e.g. lfm2.5-1.2b-thinking) produce a chain-of-thought before their answer. Xybrid keeps it out of the answer text and surfaces it on reasoningContent — null for non-thinking models. Nothing to enable; just read it if you want it.

final result = await model.run(XybridEnvelope.text(
    'Is 97 a prime number? Reason, then answer.'));

if (result.text != null) print('Answer: ${result.text}');
if (result.reasoningContent != null) print('Reasoning: ${result.reasoningContent}');

Structured Output (JSON Schema)

Constrain a local llama model so its output is always schema-valid — no retry loop, no parse failures. jsonSchemaToGbnf turns a JSON Schema into the GBNF grammar GenerationConfig.grammar expects:

final grammar = jsonSchemaToGbnf(
  schemaJson: '{"type":"object","properties":{"city":{"type":"string"}},"required":["city"]}',
);

final result = await model.run(
  XybridEnvelope.text('Extract the city: "I flew into Paris last night."'),
  config: GenerationConfig.greedy(grammar: grammar),
);
// result.text is guaranteed to parse against the schema

GenerationConfig.greedy is the usual pairing — deterministic decoding plus a grammar is the standard extraction shape. Raw GBNF works too: pass it to grammar directly.

Tool Calling

One run is one model turn, so the loop lives in your code: run a request carrying tools, execute the calls the model asks for, then run a continuation envelope that feeds the outcomes back.

final tools = [
  ToolDefinition(
    name: 'get_weather',
    description: 'Current weather for a city',
    parametersJson: '{"type":"object","properties":{"city":{"type":"string"}}}',
  ),
];
final config = GenerationConfig.greedy(tools: tools);

const question = 'Weather in Paris?';
final first = await model.run(XybridEnvelope.text(question), config: config);

if (first.hasToolCalls) {
  final results = [
    for (final call in first.toolCalls)
      ToolResult(callId: call.id, name: call.name, contentJson: runTool(call)),
  ];

  final answer = await model.run(
    XybridEnvelope.toolResults(
      userText: question,
      priorAssistantText: first.text!,  // raw output, tool-call block included
      results: results,
    ),
    config: config,  // same tools as the first turn
  );
}

Two rules the loop depends on: run the continuation with the same tools as the original turn so the executor rebuilds an identical chat prefix, and pass priorAssistantText verbatim — tool-call block included.

Gate your tool UI on the bundle's advisory flag:

if (model.supportsToolCalling == true) showToolsToggle();

null means the bundle says nothing and never implies support. Tool calling is llama.cpp-only; unsupported paths reject tool-bearing requests rather than silently generating without them.

LLM Streaming

Stream tokens in real-time as the LLM generates:

final model = await XybridModelLoader.fromRegistry('qwen-2.5-0.5b').load();

await for (final token in model.runStreaming(XybridEnvelope.text('What is ML?'))) {
  stdout.write(token.token);

  if (token.isFinal) {
    print('\n--- Done (${token.finishReason}) ---');
  }
}

Conversation Memory

Multi-turn LLM conversations with automatic history management:

final model = await XybridModelLoader.fromRegistry('qwen-2.5-0.5b').load();
final context = ConversationContext();
context.setSystem('You are a helpful assistant.');

// Turn 1
context.pushText('What is Rust?', MessageRole.user);
final result = await model.runWithContext(
  XybridEnvelope.text('What is Rust?'),
  context,
);
context.pushText(result.text!, MessageRole.assistant);

// Turn 2 — the model remembers Turn 1
context.pushText('How does it compare to Go?', MessageRole.user);
final result2 = await model.runWithContext(
  XybridEnvelope.text('How does it compare to Go?'),
  context,
);

Streaming with context also supported:

await for (final token in model.runStreamingWithContext(envelope, context)) {
  stdout.write(token.token);
}

Platform Support

Platform ONNX Runtime Candle LLM (llama.cpp) Notes
macOS ✅ ✅ Metal ✅ Apple Silicon only (M1+)
iOS ✅ CoreML ✅ Metal ✅ arm64, downloads ORT from HuggingFace
Android ✅ — ✅ arm64-v8a, x86_64; ORT from Maven Central
Linux ✅ ✅ CPU ✅ x86_64
Windows ✅ ✅ CPU ✅ x86_64

Model Support

Model Type All Platforms
Kokoro 82M TTS ✅
KittenTTS Nano TTS ✅
Whisper Tiny (Candle) ASR ✅
Wav2Vec2 (ONNX) ASR ✅
SmolLM2 360M LLM ✅
Qwen 2.5 0.5B LLM ✅
Qwen 3.5 0.8B LLM ✅
Qwen 3.5 2B LLM ✅
Gemma 3 1B LLM ✅
Llama 3.2 1B LLM ✅

Platform Requirements

  • macOS: 13.3+, Xcode 15+, Apple Silicon
  • iOS: 13.0+, Xcode 15+
  • Android: minSdk 21, NDK r25+, 64-bit only
  • Linux/Windows: x86_64

Native Libraries

Native ML runtimes are resolved automatically at build time:

  • Android: ONNX Runtime pulled from Maven Central (com.microsoft.onnxruntime:onnxruntime-android)
  • iOS: ONNX Runtime is part of the precompiled library, so nothing extra is downloaded or installed. (Monorepo source builds fetch an ONNX Runtime xcframework from HuggingFace into ~/.xybrid/cache/ort-ios/; simulator source builds also need xz.)
  • macOS/Linux/Windows: ONNX Runtime downloaded by the ort Rust crate at compile time

The Rust library itself ships as a precompiled, signature-verified binary for every supported platform via cargokit, downloaded at build time. No Rust toolchain is required, and having one installed does not change anything — the published package is precompiled-only and cannot be built from source, because its Rust crate lives in the xybrid monorepo workspace.

The first build of an app downloads that binary, gzip-compressed: roughly 10 MB per Android ABI and 30–50 MB for iOS and macOS, where it is a static library of over 100 MB once unpacked. A slow first build is this download, not a Rust compile.

The download is kept in a shared cache at ~/.xybrid/cache/precompiled/, so flutter clean and other projects on the same machine reuse it instead of downloading again. Every reuse re-verifies the binary's signature against the key pinned in this package, and entries unused for 90 days are removed. Set XYBRID_PRECOMPILED_CACHE_DIR to move the cache — for example into a directory your CI persists between runs — or to an empty value to turn it off.

Flutter hides native build-step output unless you pass -v; with it (or in the Xcode / Android Studio build log) cargokit reports what it is doing:

INFO: Downloading precompiled aarch64-apple-ios_libxybrid_flutter_ffi.a.gz (33.6 MB) from https://github.com/…
INFO: aarch64-apple-ios_libxybrid_flutter_ffi.a.gz: 14.2 MB of 33.6 MB (42%)
INFO: Downloaded aarch64-apple-ios_libxybrid_flutter_ffi.a.gz: 33.6 MB in 21s (1.6 MB/s)
INFO: Unpacked aarch64-apple-ios_libxybrid_flutter_ffi.a.gz to 119.2 MB
INFO: Using precompiled xybrid_flutter for aarch64-apple-ios (downloaded)

Later builds print (cached); after flutter clean, or in another project, (shared cache). A line starting Building xybrid_flutter for means a source build, which only happens inside the monorepo.

Building from source is for monorepo development, where the workspace root and the xybrid-* crates are present. There it is the default whenever a Rust toolchain is installed, because cargokit's crate hash only covers bindings/flutter/rust — a precompiled binary would not pick up edits to xybrid-core, xybrid-sdk or xybrid-ffi-facade.

Either default can be overridden with a cargokit_options.yaml next to your app's pubspec.yaml:

use_precompiled_binaries: false

Example App

A full example app with 8 demo screens (TTS, ASR, LLM chat, pipelines, device info) is available:

https://github.com/xybrid-ai/xybrid/tree/main/examples/flutter

API Reference

Class Purpose
Xybrid SDK initialization, cache checking, factory methods
XybridModelLoader Load models from registry or local bundle
XybridModel Run inference (batch, streaming, with context)
XybridEnvelope Type-safe inputs: audio, text, embedding
XybridResult Inference output: text, audio, embedding, latency
StreamToken Individual LLM token during streaming
ConversationContext Multi-turn conversation history with FIFO pruning
XybridPipeline Multi-stage pipeline execution from YAML
MessageRole Enum: system, user, assistant
LoadEvent Download progress events: LoadProgress, LoadComplete, LoadError

Full API documentation: pub.dev/documentation/xybrid_flutter

Telemetry

The plugin reports binding=flutter in a small X-Xybrid-Client header attached to registry metadata calls. See docs/telemetry/registry.md for the exact wire format and the opt-out switch (XYBRID_TELEMETRY_OPTOUT=1).

License

Apache 2.0 — see LICENSE

Libraries

xybrid
Xybrid - Hybrid cloud-edge ML inference orchestrator.
xybrid_flutter
Xybrid Flutter SDK for hybrid cloud-edge ML inference.