LlamaGenerationStats class
Token accounting and timing for one completed generation run.
Constructors
- LlamaGenerationStats({required int promptTokenCount, required int cachedTokenCount, required int generatedTokenCount, required LlamaFinishReason finishReason, required Duration prefillDuration, required Duration decodeDuration, int? draftedTokenCount, int? acceptedTokenCount})
-
Creates a LlamaGenerationStats.
const
Properties
- acceptedTokenCount → int?
-
Drafted tokens the main model accepted; null when the session runs
without speculative decoding.
final
- cachedTokenCount → int
-
Prompt tokens served from the reused KV-cache prefix.
final
- decodeDuration → Duration
-
Wall-clock time spent generating tokens.
final
- decodeTokensPerSecond → double
-
Generation speed, in tokens per second.
no setter
- draftedTokenCount → int?
-
Tokens proposed by the draft model; null when the session runs without
speculative decoding.
final
- finishReason → LlamaFinishReason
-
Why the run stopped.
final
- generatedTokenCount → int
-
Tokens generated.
final
- hashCode → int
-
The hash code for this object.
no setterinherited
- prefillDuration → Duration
-
Wall-clock time spent ingesting the prompt.
final
- prefillTokensPerSecond → double
-
Prompt-ingestion speed over the tokens that were actually decoded
(prompt minus cached), in tokens per second.
no setter
- promptTokenCount → int
-
Prompt tokens fed to the model (including the reused prefix).
final
- runtimeType → Type
-
A representation of the runtime type of the object.
no setterinherited
Methods
-
noSuchMethod(
Invocation invocation) → dynamic -
Invoked when a nonexistent method or property is accessed.
inherited
-
toString(
) → String -
A string representation of this object.
inherited
Operators
-
operator ==(
Object other) → bool -
The equality operator.
inherited