local_ai_llama_cpp 0.0.3
local_ai_llama_cpp: ^0.0.3 copied to clipboard
llama.cpp adapter for LocalAI Kit: runs any GGUF model as a LocalLlm (streaming, GBNF-constrained structured output) and as a LocalEmbedding, with all FFI work isolated in worker isolates.
Changelog #
0.0.3 #
- First pub.dev release of the llama.cpp adapter, with streaming GGUF chat, embeddings, persistent KV-cache planning and GBNF-constrained output.
- Add model-family chat templates and bring-your-own native library support.
0.0.2 #
- Initial release of the llama.cpp adapter:
LlamaCppLlmAdapter(any GGUF chat model, streaming, GBNF-constrained structured output),LlamaCppEmbeddingAdapter(the firstLocalEmbeddingimplementation in the kit) andLlamaCppAdapterPluginunder provider keyllama-cpp.