llm_vllm 0.3.0 copy "llm_vllm: ^0.3.0" to clipboard
llm_vllm: ^0.3.0 copied to clipboard

VLLM backend implementation for LLM interactions. Provides streaming chat, embeddings, tool calling, vision support, and model management.

example/README.md

llm_vllm examples #

CLI #

dart run example/cli_example.dart
dart run example/cli_example.dart Qwen/Qwen3-0.6B
dart run example/cli_example.dart Qwen/Qwen3-0.6B http://localhost:8000

If your vLLM server was started with --api-key, pass the key as the third argument or set VLLM_API_KEY.

Minimal Usage #

import 'package:llm_vllm/llm_vllm.dart';

final repo = VLLMChatRepository(baseUrl: 'http://localhost:8000');

final stream = repo.streamChat('Qwen/Qwen3-0.6B', messages: [
  LLMMessage(role: LLMRole.user, content: 'Hello!'),
]);

await for (final chunk in stream) {
  print(chunk.message?.content ?? '');
}

Model Listing #

final vllmRepo = VLLMRepository(baseUrl: 'http://localhost:8000');
final models = await vllmRepo.models();

Integration Environment #

export VLLM_BASE_URL=http://localhost:8000
export VLLM_CHAT_MODEL=Qwen/Qwen3-0.6B
export VLLM_EMBEDDING_MODEL=BAAI/bge-small-en-v1.5
export VLLM_API_KEY=optional-key
0
likes
0
points
453
downloads

Publisher

unverified uploader

Weekly Downloads

VLLM backend implementation for LLM interactions. Provides streaming chat, embeddings, tool calling, vision support, and model management.

Repository (GitHub)
View/report issues

Topics

#vllm #llm #ai #chat #embeddings

License

unknown (license)

Dependencies

http, llm_core

More

Packages that depend on llm_vllm