backend library
The service-provider interface for custom llamadart backends.
Import this library next to package:llamadart/llamadart.dart to:
- implement a custom LlamaBackend, or fake one in tests, through the
optional
Backend*interfaces such as BackendTextToSpeech and BackendDecision; - drive the LiteRT-LM runtime directly through LiteRtLmBackend and LiteRtLmRuntimeClient;
- call the low-level LlamaEngineBackendHooks that
TextToSpeechEngineandDecisionEnginebuild on.
Apps that only load models and generate need just
package:llamadart/llamadart.dart. This library follows the package's
semantic versioning.
import 'package:llamadart/backend.dart';
import 'package:llamadart/llamadart.dart';
final class FakeSpeechBackend implements LlamaBackend, BackendTextToSpeech {
// ...
}
Classes
- BackendAvailability
- Optional backend capability for exposing selectable backend options.
- BackendBatchEmbeddings
- Optional backend capability for batching embedding requests.
- BackendDartLogLevel
- Optional backend capability for backends whose worker isolate keeps its own Dart logger and forwards its records to the main isolate.
- BackendDecision
- Optional backend capability for encoder-plus-head decision models.
- BackendDecisionCapabilities
- Runtime support for an optional backend decision-model path.
- BackendDecisionHeadInfo
- A decision head loaded by a backend.
- BackendDecisionOutput
- Raw head outputs for one sequence.
- BackendDecisionSequence
- Encoder input for one question.
- BackendEmbeddings
- Optional backend capability for generating text embeddings.
- BackendEmbeddingsSupport
- Optional backend capability for reporting whether BackendEmbeddings is actually available for the active runtime.
- BackendGenerationCapabilities
- Optional GenerationParams controls that the loaded model's runtime applies.
- BackendGenerationCapabilitiesSupport
- Optional backend capability reporting BackendGenerationCapabilities.
- BackendGenerationLimitSupport
- Optional declaration that a runtime cannot reliably report generation limits.
- BackendGpuEnumeration
- Optional backend capability for enumerating GPU-class devices for offload selection (e.g. pinning to the discrete GPU on a laptop, or surfacing which device is in use).
- BackendGrammarConstraintsSupport
- Optional backend capability for reporting grammar-constrained decoding.
- BackendLazyGrammarSupport
- Optional backend capability for reporting lazy grammar activation.
- BackendModelFileTypeDiagnostics
- Optional backend capability for exposing loaded model file type metadata.
- BackendNativeChatGeneration
- Optional backend capability for native structured chat generation.
- BackendNextTokenScoring
- Optional backend capability for scoring the token that would follow a prompt, without sampling.
- BackendNextTokenScoringSupport
- Optional backend capability for reporting whether BackendNextTokenScoring is actually available for the active runtime.
- BackendPerfContextData
- Native performance timings reported by llama.cpp for the active context.
- BackendPerformanceDiagnostics
- Optional backend capability for exposing llama.cpp perf timings.
- BackendPromptSpeechToTextSupport
- Optional backend capability for prompt-adapted speech recognition.
- BackendRuntimeDiagnostics
- Optional backend capability for exposing resolved runtime diagnostics.
- BackendStatePersistence
-
Optional backend capability for persisting the KV cache to disk and
restoring it later, mirroring
llama_state_save_file/llama_state_load_filein llama.cpp. Saving captures the native runtime state ofcontextHandletogether with the token sequence that produced it. Loading restores the native KV cache and returns the saved token sequence; subsequent inference can skip prompt evaluation when callers re-issue a prompt with the restored token prefix and prompt-prefix reuse enabled. - BackendStatePersistenceSupport
- Optional backend capability for reporting whether BackendStatePersistence is actually available for the active runtime.
- BackendTextToSpeech
- Optional backend capability for dedicated text-to-speech synthesis.
- BackendTextToSpeechCapabilities
- Runtime capabilities for an optional backend text-to-speech path.
- BackendTextToSpeechProgress
- Progress reported by a backend TTS implementation.
- BackendTextToSpeechRequest
- Sampling and input values passed to a backend TTS implementation.
- BackendTextToSpeechResult
- Complete PCM output returned by a backend TTS implementation.
- LiteRtLmAsrProcessResult
- One native ASR processing result.
- LiteRtLmAsrPushResult
- Result of pushing PCM into a bounded native ASR session.
- LiteRtLmAsrRuntimeSession
- Synchronous low-level session implemented by the native runtime.
- LiteRtLmBackend
- Web-safe placeholder for the native-only LiteRT-LM backend.
- LiteRtLmRuntimeClient
- Web-safe placeholder for the native-only runtime client.
- LiteRtLmRuntimeMetrics
- Runtime metrics shape shared with the native LiteRT-LM implementation.
- LiteRtLmRuntimeResult
- Generated text and runtime metrics from a LiteRT-LM run.
- LlamaBackend
- Platform-agnostic interface for local model inference.
- StateLoadResult
- Result of BackendStatePersistence.stateLoadFile. Contains the token sequence saved alongside the native KV-cache state.
Enums
- BackendTextToSpeechModel
- Model family reported by a backend text-to-speech implementation.
- BackendTextToSpeechPhase
- Phase reported while a backend synthesizes speech.
- LiteRtLmAsrProcessState
- Native ASR processing state.
Extensions
- LlamaEngineBackendHooks on LlamaEngine
-
Low-level hooks that
TextToSpeechEngine,DecisionEngineand backend integrations use on a LlamaEngine.