llm_chatgpt 0.3.2
llm_chatgpt: ^0.3.2 copied to clipboard
OpenAI/ChatGPT backend implementation for LLM interactions. Provides streaming chat, embeddings, and tool calling via the OpenAI API.
llm_chatgpt #
OpenAI/ChatGPT backend implementation for LLM interactions in Dart.
Available on pub.dev.
Features #
- Streaming chat responses
- Tool/function calling
- Vision (image) support
- Embeddings
- Reasoning models: per-model detection,
reasoning_effortmapping, and streaming usage - Structured output (
json_objectand nativejson_schemawithstrict) - Configurable base URL for OpenAI-compatible servers
Installation #
dependencies:
llm_chatgpt: ^0.3.2
Prerequisites #
You need an OpenAI API key. Get one from platform.openai.com.
Important: Never commit your API key to version control. Use environment variables or a .env file.
Usage #
Basic Chat #
import 'package:llm_chatgpt/llm_chatgpt.dart';
final repo = ChatGPTChatRepository(apiKey: 'your-api-key');
final stream = repo.streamChat('gpt-5.4-nano', messages: [
LLMMessage(role: LLMRole.user, content: 'Hello!'),
]);
await for (final chunk in stream) {
print(chunk.message?.content ?? '');
}
Tool Calling #
final stream = repo.streamChat('gpt-5.4-nano',
messages: messages,
tools: [MyTool()],
);
Structured Output #
Use LLMChatOptions.responseFormat to enforce JSON output natively via the OpenAI response_format API:
import 'package:llm_core/llm_core.dart';
// Simple JSON mode
final stream = repo.streamChat(
'gpt-5.4-nano',
messages: [LLMMessage(role: LLMRole.user, content: 'List three fruits as JSON.')],
options: const LLMChatOptions(responseFormat: JsonFormat()),
);
// JSON Schema mode (strict schema enforcement)
const schema = {
'type': 'object',
'properties': {
'name': {'type': 'string'},
'age': {'type': 'integer'},
},
'required': ['name', 'age'],
'additionalProperties': false,
};
final stream = repo.streamChat(
'gpt-5.4-nano',
messages: [LLMMessage(role: LLMRole.user, content: 'Return a person object.')],
options: const LLMChatOptions(
responseFormat: JsonSchemaFormat(name: 'Person', schema: schema),
),
);
Embeddings #
final embeddings = await repo.embed(
model: 'text-embedding-3-small',
messages: ['Hello world', 'Goodbye world'],
);
Non-Streaming Response #
Get a complete response without streaming:
final response = await repo.chatResponse('gpt-5.4-nano', messages: [
LLMMessage(role: LLMRole.user, content: 'Hello!'),
]);
print(response.content);
print('Tokens: ${response.evalCount}');
Using LLMChatOptions #
Encapsulate all options in a single object:
import 'package:llm_core/llm_core.dart';
final options = LLMChatOptions(
tools: [MyTool()],
toolAttempts: 5,
timeout: Duration(minutes: 5),
retryConfig: RetryConfig(maxAttempts: 3),
);
final stream = repo.streamChat('gpt-5.4-nano', messages: messages, options: options);
Reasoning models #
Reasoning models (o-series, gpt-5 family) are detected by model id
(gptIsReasoningModel) and handled differently from conventional models:
temperature/top_pare dropped on reasoning models — the API rejects them with a400.reasoningEffortmaps toreasoning_effort, clamped to what the family accepts (gptEffortWireValue): o-serieslow/medium/high; gpt-5minimal/low/medium/high; gpt-5.1+none/low/medium/high(plusxhighon codex-max ids). Never sent to conventional models or too1-mini/o1-preview, which predate the parameter.- OpenAI has no exact reasoning-token budget, so
reasoningBudgetis honored as a derived effort level (an explicitreasoningEffortwins). - Reasoning models always reason, so the knobs apply regardless of
think. - Streaming requests set
stream_options: {include_usage: true}; reasoning-token usage surfaces asLLMUsage.reasoningTokensfromcompletion_tokens_details.reasoning_tokens.
final options = LLMChatOptions(reasoningEffort: ReasoningEffort.high);
final stream = repo.streamChat('gpt-5.4', messages: messages, options: options);
Vision #
Attach images to a message; they are sent as OpenAI image_url content parts:
final stream = repo.streamChat(
'gpt-5.4-nano',
messages: [
LLMMessage(
role: LLMRole.user,
content: 'What is in this image?',
images: [base64EncodedImage],
),
],
);
OpenAI-compatible servers #
baseUrl points the client at any server exposing OpenAI's
/v1/chat/completions with Authorization: Bearer:
final repo = ChatGPTChatRepository(
apiKey: 'your-key',
baseUrl: 'https://my-openai-compatible-host',
);
Azure OpenAI is not supported. It needs a
/openai/deployments/{deployment}/chat/completions?api-version=... path and an
api-key header; this package always builds $baseUrl/v1/chat/completions with
a bearer token. For a self-hosted OpenAI-compatible server, llm_vllm is
usually the better fit — it probes the deployment's real capabilities.
Advanced Configuration #
Builder Pattern #
Use the builder for complex configurations:
import 'package:llm_core/llm_core.dart';
// Standard OpenAI
final repo = ChatGPTChatRepository.builder()
.apiKey('your-api-key')
.baseUrl('https://api.openai.com')
.maxToolAttempts(10)
.retryConfig(RetryConfig(
maxAttempts: 5,
initialDelay: Duration(seconds: 1),
maxDelay: Duration(seconds: 30),
))
.timeoutConfig(TimeoutConfig(
connectionTimeout: Duration(seconds: 10),
readTimeout: Duration(minutes: 5),
totalTimeout: Duration(minutes: 10),
))
.build();
// An OpenAI-compatible server
final compatRepo = ChatGPTChatRepository.builder()
.apiKey('your-key')
.baseUrl('https://my-openai-compatible-host')
.maxToolAttempts(10)
.retryConfig(RetryConfig(maxAttempts: 3))
.timeoutConfig(TimeoutConfig(readTimeout: Duration(minutes: 5)))
.build();
Retry Configuration #
Configure automatic retries for failed requests:
import 'package:llm_core/llm_core.dart';
final repo = ChatGPTChatRepository(
apiKey: 'your-api-key',
retryConfig: RetryConfig(
maxAttempts: 3,
initialDelay: Duration(seconds: 1),
maxDelay: Duration(seconds: 30),
retryableStatusCodes: [429, 500, 502, 503, 504],
),
);
Timeout Configuration #
Configure timeouts for different scenarios:
import 'package:llm_core/llm_core.dart';
final repo = ChatGPTChatRepository(
apiKey: 'your-api-key',
timeoutConfig: TimeoutConfig(
connectionTimeout: Duration(seconds: 10),
readTimeout: Duration(minutes: 5),
totalTimeout: Duration(minutes: 10),
),
);
Models #
See OpenAI Models for available models:
gpt-5.4-nano- Low-cost current-generation chat model used by live testsgpt-5.4-mini- Larger current-generation small modelgpt-5.4- More capable current-generation modeltext-embedding-3-small- Low-cost embeddings used by live teststext-embedding-3-large- Higher quality embeddings