ollama_dart 4.0.0 copy "ollama_dart: ^4.0.0" to clipboard
ollama_dart: ^4.0.0 copied to clipboard

Type-safe Dart client for Ollama chat, generation, embeddings, System One decisions, model management, and cloud web search.

Ollama Dart Client #

tests ollama_dart Discord MIT

Dart client for the Ollama API — chat, streaming, tool calling, embeddings, System One decisions, model management, and cloud web search. Connect to local, self-hosted, or Ollama Cloud models from Dart and Flutter across iOS, Android, macOS, Windows, Linux, Web, and server-side Dart.

Tip

Coding agents: start with llms.txt. It links to the package docs, examples, and optional references in a compact format.

Table of Contents

Features #

Generation and streaming #

  • Chat completions with context memory and multimodal inputs
  • Text generation for prompt-style completions
  • Embeddings for semantic search and retrieval
  • NDJSON streaming for chat and completions
  • Tool calling, thinking mode, and structured output
  • Cached prompt token metrics and model-defined thinking controls
  • System One classification, yes/no probabilities, and ordered scoring

Local model operations #

  • Pull, push, copy, create, delete, and inspect models
  • Upload binary blobs and import GGUF or Safetensors model files
  • List running models and query server version
  • Connect to local or remote Ollama instances with optional auth

Cloud web tools #

  • Search the web and fetch page content with an Ollama API key
  • Explicit cloud configuration using the same authentication and transport settings

Why choose this client? #

  • Pure Dart with no Flutter dependency — works in mobile apps, backends, and CLIs.
  • Type-safe request and response models with minimal dependencies (http, logging, meta).
  • Streaming, retries, interceptors, and error handling built into the client.
  • Mirrors the Ollama API closely, including model management endpoints most wrappers skip.
  • Strict semver versioning so downstream packages can depend on stable, predictable version ranges.

Quickstart #

Requires Dart 3.12 or later. See the Migration Guide for upgrade instructions.

dependencies:
  ollama_dart: ^4.0.0
import 'package:ollama_dart/ollama_dart.dart';

Future<void> main() async {
  final client = OllamaClient();

  try {
    final response = await client.chat.create(
      request: ChatRequest(
        model: 'gpt-oss',
        messages: [ChatMessage.user('Explain what Dart isolates do.')],
      ),
    );

    print(response.message?.content);
  } finally {
    client.close();
  }
}

Configuration #

Configure local hosts, remote servers, and retries

Use OllamaClient() for the default local daemon at http://localhost:11434, or OllamaClient.fromEnvironment() to read OLLAMA_HOST. Use OllamaConfig when you need a remote host, bearer auth, or a different timeout policy.

import 'package:ollama_dart/ollama_dart.dart';

Future<void> main() async {
  final client = OllamaClient(
    config: OllamaConfig(
      baseUrl: 'http://localhost:11434',
      timeout: const Duration(minutes: 5),
      retryPolicy: RetryPolicy(
        maxRetries: 3,
        initialDelay: Duration(seconds: 1),
      ),
    ),
  );

  client.close();
}

Environment variable:

  • OLLAMA_HOST

Use BearerTokenProvider when the Ollama server is exposed behind an authenticated reverse proxy or remote deployment.

For direct Ollama Cloud inference or hosted web tools, pass the cloud host explicitly:

final client = OllamaClient.withApiKey(
  apiKey,
  baseUrl: 'https://ollama.com',
);

Here apiKey is your Ollama API key. Configure the host without an /api suffix; resources append their endpoint paths. Web search and fetch use the configured host and require Ollama Cloud or a proxy exposing those endpoints. Local model creation, blob uploads, and System One require a local Ollama server.

By default the client does not send an X-Request-ID header — Ollama's CORS allow-list excludes it, so sending it breaks the preflight in browser targets (Flutter Web / dart2wasm). A request ID is still generated internally for logging and error correlation. Set OllamaConfig(sendRequestIdHeader: true) to emit the header when talking to an intermediary (e.g. a reverse proxy) you've configured to accept it.

Usage #

How do I run a chat completion? #

Show example

Use client.chat.create(...) for conversational flows. The chat response exposes message?.content, which keeps simple completions ergonomic in Dart and Flutter UIs.

import 'package:ollama_dart/ollama_dart.dart';

Future<void> main() async {
  final client = OllamaClient();

  try {
    final response = await client.chat.create(
      request: ChatRequest(
        model: 'gpt-oss',
        messages: [
          ChatMessage.system('You are a concise assistant.'),
          ChatMessage.user('What is hot reload?'),
        ],
      ),
    );

    print(response.message?.content);
  } finally {
    client.close();
  }
}

For structured output, set format to constrain the response to valid JSON:

final response = await client.chat.create(
  request: ChatRequest(
    model: 'gpt-oss',
    messages: [ChatMessage.user('List 3 colors as JSON')],
    format: ResponseFormat.json(),
  ),
);

→ Full example

How do I discover a model's thinking controls? #

/api/show advertises supported values and the model default. Keep think unset to use that default. Existing ThinkValue.enabled(...) and ThinkValue.level(...) constructors remain available; use ThinkValue.string(...) for names advertised by a model.

final details = await client.models.show(
  request: const ShowRequest(model: 'gpt-oss'),
);
final thinking = details.thinking;
if (thinking != null) {
  print(thinking.values.map((value) => value.toJson()).toList());
  print('Default: ${thinking.defaultValue.toJson()}');
}

final response = await client.chat.create(
  request: const ChatRequest(
    model: 'gpt-oss',
    messages: [ChatMessage.user('What is 15 * 7?')],
    think: ThinkValue.string('high'),
  ),
);
print('Cached prompt tokens: ${response.promptEvalCachedCount}');

Thinking values are model-defined; use a value returned by thinking.values. promptEvalCount includes cached prompt tokens, while promptEvalDuration measures uncached prompt evaluation. The cached count is nullable because older servers omit it.

→ Full model-inspection example

How do I stream local model output? #

Show example

Streaming uses Ollama's NDJSON response format and works well for terminals and live Flutter widgets. This is the fastest way to surface partial output from a local model.

The final event (done == true) carries token and timing metrics, including nullable promptEvalCachedCount. Chunks can contain multiple tokens; use the final evalCount for generated token usage. promptEvalDuration measures uncached prompt evaluation.

import 'dart:io';

import 'package:ollama_dart/ollama_dart.dart';

Future<void> main() async {
  final client = OllamaClient();

  try {
    final stream = client.chat.createStream(
      request: ChatRequest(
        model: 'gpt-oss',
        messages: [ChatMessage.user('Write a haiku about local models.')],
      ),
    );

    await for (final chunk in stream) {
      stdout.write(chunk.message?.content ?? '');
    }
  } finally {
    client.close();
  }
}

→ Full example

How do I use tool calling? #

Show example

Tool calling is declared on the request with typed ToolDefinition objects. This makes local agent-style workflows possible without switching to another API format.

For a follow-up request, replay the assistant's content, thinking, and tool calls together, then identify each tool result with toolName and toolCallId when the model supplies an ID. See the full example for a complete round trip.

import 'package:ollama_dart/ollama_dart.dart';

Future<void> main() async {
  final client = OllamaClient();

  try {
    final response = await client.chat.create(
      request: ChatRequest(
        model: 'gpt-oss',
        messages: [ChatMessage.user('What is the weather in Paris?')],
        tools: [
          ToolDefinition(
            type: ToolType.function,
            function: ToolFunction(
              name: 'get_weather',
              description: 'Get the current weather for a location',
              parameters: {
                'type': 'object',
                'properties': {
                  'location': {'type': 'string'},
                },
                'required': ['location'],
              },
            ),
          ),
        ],
      ),
    );

    print(response.message?.toolCalls?.length ?? 0);
  } finally {
    client.close();
  }
}

→ Full example

How do I generate plain text? #

Show example

Use the completions resource when you want prompt-style generation instead of chat messages. This is useful for legacy templates, code infill helpers, or smaller server utilities.

import 'package:ollama_dart/ollama_dart.dart';

Future<void> main() async {
  final client = OllamaClient();

  try {
    final result = await client.completions.generate(
      request: GenerateRequest(
        model: 'gpt-oss',
        prompt: 'Complete this sentence: Dart is great for',
      ),
    );

    print(result.response);
  } finally {
    client.close();
  }
}

→ Full example

How do I generate images (experimental)? #

Show example

Ollama 0.35.0 rejects image generation with HTTP 400. The feature was temporarily removed in 0.32.6; upstream identifies 0.32.5 as the last release with support. The client retains the experimental fields for compatible servers.

On a server that supports image generation, pass width/height/steps to /api/generate; the response carries a base64 string in image (decode it before writing bytes). These fields may change or be removed in a future Ollama release.

import 'dart:convert';
import 'dart:io';

import 'package:ollama_dart/ollama_dart.dart';

Future<void> main() async {
  final client = OllamaClient();

  try {
    final result = await client.completions.generate(
      request: const GenerateRequest(
        model: 'x/z-image-turbo',
        prompt: 'a sunset over mountains',
        width: 1024,
        height: 768,
        steps: 20,
      ),
    );

    final image = result.image;
    if (image != null) {
      File('generated_image.png').writeAsBytesSync(base64Decode(image));
    }
  } finally {
    client.close();
  }
}

→ Full example

How do I create embeddings? #

Show example

Embeddings are exposed as a first-class resource, so semantic search or retrieval code can stay inside the same Ollama client. This is useful for local RAG pipelines in Dart.

import 'package:ollama_dart/ollama_dart.dart';

Future<void> main() async {
  final client = OllamaClient();

  try {
    final response = await client.embeddings.create(
      request: const EmbedRequest(
        model: 'nomic-embed-text',
        input: EmbedInput.list(['Dart', 'Flutter']),
      ),
    );

    print(response.embeddings?.length ?? 0);
  } finally {
    client.close();
  }
}

→ Full example

How do I manage local models? #

Show example

Model management is part of the same client, which means pull, inspect, and runtime checks do not require a separate admin tool. That is useful for installers, desktop apps, and local dev tooling.

import 'package:ollama_dart/ollama_dart.dart';

Future<void> main() async {
  final client = OllamaClient();

  try {
    final models = await client.models.list();
    print(models.models?.length ?? 0);
  } finally {
    client.close();
  }
}

→ Full example

How do I make System One decisions? #

System One requires Ollama v0.35.0 or later and a compatible local decision model such as nimble. Its single JSON response answers named questions against a shared state; it does not stream. It supports choice questions, yes/no probabilities (noul), and scores over ordered criteria.

final response = await client.systemOne.create(
  request: SystemOneRequest(
    model: 'nimble',
    state: const SystemOneContent.string('The customer asks for a refund.'),
    questions: {
      'intent': SystemOneQuestion.choice(
        instructions: const SystemOneContent.string('Classify the request.'),
        criteria: {
          'refund': 'A request to return money',
          'other': 'Any other request',
        },
      ),
    },
  ),
);
print(response.answers['intent']);

State and instructions also accept SystemOneContent.object(...) and SystemOneContent.array(...), serialized as JSON text. Send 1–64 named questions; choice and score questions need 2–26 criteria. A noul answer is a probability of true. A score is the probability-weighted average of zero-based criterion indices. Confidence measures probability concentration, not calibrated correctness. Questions are evaluated independently against the shared state. Ollama Cloud and MLX/Safetensors runners do not support this endpoint.

→ Full example with all question types

How do I upload files to create a model? #

Use client.blobs.exists(digest: ...) and client.blobs.create(digest: ..., bytes: ...) against your local server, then pass original file names and their SHA256 digests through CreateRequest.files. Blob methods return no JSON payload; an existing blob is a successful upload, and only a missing blob returns false from exists.

import 'package:ollama_dart/ollama_dart.dart';

Future<void> importModelFile({
  required OllamaClient client,
  required String fileName,
  required String digest,
  required List<int> bytes,
  required String modelName,
}) async {
  if (!await client.blobs.exists(digest: digest)) {
    await client.blobs.create(digest: digest, bytes: bytes);
  }
  await client.models.create(
    request: CreateRequest(model: modelName, files: {fileName: digest}),
  );
}

Supply the SHA256 digest of the same bytes passed to this function.

Keep split GGUF shard names intact and upload every shard. GGUF weights must be quantized before import; quantize and draftQuantize apply during Safetensors import. LoRA adapters are no longer supported by current Ollama servers; the existing adapters field remains for older-server compatibility.

→ Full file-upload and model-import example

How do I search and fetch the web? #

Use an explicit cloud client with your Ollama API key. Local Ollama's experimental web proxy paths are outside these methods.

final client = OllamaClient.withApiKey(
  apiKey,
  baseUrl: 'https://ollama.com',
);
try {
  final results = await client.web.search(
    request: const WebSearchRequest(query: 'Dart isolates', maxResults: 3),
  );
  for (final result in results.results ?? <WebSearchResult>[]) {
    print('${result.title}: ${result.url}');
  }
  final page = await client.web.fetch(
    request: const WebFetchRequest(url: 'https://dart.dev'),
  );
  print(page.content);
} finally {
  client.close();
}

Omit maxResults to use the service default of 5; the maximum is 10. Web fetch accepts a single URL and returns its title, extracted content, and links.

→ Full example

Error Handling #

Handle local daemon failures, retries, and streaming issues

ollama_dart throws typed exceptions so you can distinguish between API failures, timeouts, aborts, and streaming problems. Catch ApiException first for HTTP errors, then fall back to OllamaException for everything else.

import 'dart:io';

import 'package:ollama_dart/ollama_dart.dart';

Future<void> main() async {
  final client = OllamaClient();

  try {
    await client.version.get();
  } on ApiException catch (error) {
    stderr.writeln('Ollama API error ${error.statusCode}: ${error.message}');
  } on OllamaException catch (error) {
    stderr.writeln('Ollama client error: $error');
  } finally {
    client.close();
  }
}

→ Full example

Examples #

See the example/ directory for complete examples:

Example Description
chat_example.dart Chat completions
streaming_example.dart Streaming responses
tool_calling_example.dart Tool calling
completions_example.dart Plain text generation
embeddings_example.dart Text embeddings
models_example.dart Model management
system_one_example.dart Choice, yes/no, and score decisions
blobs_example.dart Binary upload and file-based model creation
web_example.dart Hosted web search and page fetch
image_generation_example.dart Experimental image generation (unavailable in Ollama 0.35.0)
version_example.dart Server version
error_handling_example.dart Exception handling patterns
ollama_dart_example.dart Quick-start overview

API Coverage #

API Status
Chat ✅ Full
Completions ✅ Full
Embeddings ✅ Full
Models ✅ Full
Blobs ✅ Existence checks and binary uploads
System One ✅ Local decision API; requires Ollama v0.35.0+
Web search / fetch ✅ Hosted endpoints with explicit cloud configuration
Version ✅ Full

Coverage targets Ollama's public native API and hosted web tools. Existing experimental image-generation fields remain available for compatible servers; Ollama 0.35.0 rejects image generation. OpenAI and Anthropic compatibility endpoints can be used through openai_dart and anthropic_sdk_dart; experimental server-control routes and debug fields are excluded.

Official Documentation #

If these packages are useful to you or your company, please consider sponsoring the project. Development and maintenance are provided to the community for free, but integration tests against real APIs and the tooling required to build and verify releases still have real costs. Your support, at any level, helps keep these packages maintained and free for the Dart & Flutter community.

License #

This package is licensed under the MIT License.

This is a community-maintained package and is not affiliated with or endorsed by Ollama.

92
likes
160
points
14.3k
downloads

Documentation

API reference

Publisher

verified publisherdavidmiguel.com

Weekly Downloads

Type-safe Dart client for Ollama chat, generation, embeddings, System One decisions, model management, and cloud web search.

Homepage
Repository (GitHub)
View/report issues

Topics

#nlp #gen-ai #llms #ollama

Funding

Consider supporting this project:

github.com

License

MIT (license)

Dependencies

http, logging, meta

More

Packages that depend on ollama_dart