backend library

The service-provider interface for custom llamadart backends.

Import this library next to package:llamadart/llamadart.dart to:

Apps that only load models and generate need just package:llamadart/llamadart.dart. This library follows the package's semantic versioning.

import 'package:llamadart/backend.dart';
import 'package:llamadart/llamadart.dart';

final class FakeSpeechBackend implements LlamaBackend, BackendTextToSpeech {
  // ...
}

Classes

BackendAvailability
Optional backend capability for exposing selectable backend options.
BackendBatchEmbeddings
Optional backend capability for batching embedding requests.
BackendDartLogLevel
Optional backend capability for backends whose worker isolate keeps its own Dart logger and forwards its records to the main isolate.
BackendDecision
Optional backend capability for encoder-plus-head decision models.
BackendDecisionCapabilities
Runtime support for an optional backend decision-model path.
BackendDecisionHeadInfo
A decision head loaded by a backend.
BackendDecisionOutput
Raw head outputs for one sequence.
BackendDecisionSequence
Encoder input for one question.
BackendEmbeddings
Optional backend capability for generating text embeddings.
BackendEmbeddingsSupport
Optional backend capability for reporting whether BackendEmbeddings is actually available for the active runtime.
BackendGenerationCapabilities
Optional GenerationParams controls that the loaded model's runtime applies.
BackendGenerationCapabilitiesSupport
Optional backend capability reporting BackendGenerationCapabilities.
BackendGenerationLimitSupport
Optional declaration that a runtime cannot reliably report generation limits.
BackendGpuEnumeration
Optional backend capability for enumerating GPU-class devices for offload selection (e.g. pinning to the discrete GPU on a laptop, or surfacing which device is in use).
BackendGrammarConstraintsSupport
Optional backend capability for reporting grammar-constrained decoding.
BackendLazyGrammarSupport
Optional backend capability for reporting lazy grammar activation.
BackendModelFileTypeDiagnostics
Optional backend capability for exposing loaded model file type metadata.
BackendNativeChatGeneration
Optional backend capability for native structured chat generation.
BackendNextTokenScoring
Optional backend capability for scoring the token that would follow a prompt, without sampling.
BackendNextTokenScoringSupport
Optional backend capability for reporting whether BackendNextTokenScoring is actually available for the active runtime.
BackendPerfContextData
Native performance timings reported by llama.cpp for the active context.
BackendPerformanceDiagnostics
Optional backend capability for exposing llama.cpp perf timings.
BackendPromptSpeechToTextSupport
Optional backend capability for prompt-adapted speech recognition.
BackendRuntimeDiagnostics
Optional backend capability for exposing resolved runtime diagnostics.
BackendStatePersistence
Optional backend capability for persisting the KV cache to disk and restoring it later, mirroring llama_state_save_file / llama_state_load_file in llama.cpp. Saving captures the native runtime state of contextHandle together with the token sequence that produced it. Loading restores the native KV cache and returns the saved token sequence; subsequent inference can skip prompt evaluation when callers re-issue a prompt with the restored token prefix and prompt-prefix reuse enabled.
BackendStatePersistenceSupport
Optional backend capability for reporting whether BackendStatePersistence is actually available for the active runtime.
BackendTextToSpeech
Optional backend capability for dedicated text-to-speech synthesis.
BackendTextToSpeechCapabilities
Runtime capabilities for an optional backend text-to-speech path.
BackendTextToSpeechProgress
Progress reported by a backend TTS implementation.
BackendTextToSpeechRequest
Sampling and input values passed to a backend TTS implementation.
BackendTextToSpeechResult
Complete PCM output returned by a backend TTS implementation.
LiteRtLmAsrProcessResult
One native ASR processing result.
LiteRtLmAsrPushResult
Result of pushing PCM into a bounded native ASR session.
LiteRtLmAsrRuntimeSession
Synchronous low-level session implemented by the native runtime.
LiteRtLmBackend
Web-safe placeholder for the native-only LiteRT-LM backend.
LiteRtLmRuntimeClient
Web-safe placeholder for the native-only runtime client.
LiteRtLmRuntimeMetrics
Runtime metrics shape shared with the native LiteRT-LM implementation.
LiteRtLmRuntimeResult
Generated text and runtime metrics from a LiteRT-LM run.
LlamaBackend
Platform-agnostic interface for local model inference.
StateLoadResult
Result of BackendStatePersistence.stateLoadFile. Contains the token sequence saved alongside the native KV-cache state.

Enums

BackendTextToSpeechModel
Model family reported by a backend text-to-speech implementation.
BackendTextToSpeechPhase
Phase reported while a backend synthesizes speech.
LiteRtLmAsrProcessState
Native ASR processing state.

Extensions

LlamaEngineBackendHooks on LlamaEngine
Low-level hooks that TextToSpeechEngine, DecisionEngine and backend integrations use on a LlamaEngine.