TokenSmoother class
Re-paces a token stream into a steady, typewriter-style character stream.
Token streams arrive in bursts — a prompt-processing pause, then several tokens in one frame, then nothing while the next batch decodes — which makes streamed text jump and stutter on screen. TokenSmoother buffers incoming chunks and releases them a grapheme at a time on a frame timer, at a rate that continuously tracks how fast text is actually arriving.
The release rate is not derived from the backlog alone. Draining whatever is buffered over a fixed interval empties the buffer during a decode pause and then stalls, which just relocates the stutter. Instead the smoother estimates the arrival rate and releases at that rate, using the backlog only as a correction:
arrival ← arrival + α·(unitsThisTick − arrival) // rate estimate
lead ← arrival · framesPerWindow // buffer to hold
target ← arrival + (backlog − lead) / correctionFrames
release ← release + α·(target − release) // rate is smoothed
In steady state backlog == lead and the output rate equals the input
rate, so a 5 tok/s model and a 60 tok/s model both render evenly — only
the speed differs. When generation speeds up or slows down, the backlog
term pulls the release rate toward the new arrival rate over roughly
smoothing, so the change is a ramp rather than a jump. Lag behind the
source settles around window.
Because the rate is fractional and accumulated across ticks, the smoother can emit slower than one grapheme per tick — that is what lets it match a slow model instead of outrunning it.
This is a presentation concern, so it is opt-in: apply it at the widget that renders the text, not inside generation. Headless callers (batch runs, orchestration, tests) should leave the stream unpaced.
final display = session.generate(prompt).smoothed();
A chunk for which atomic returns true is never split; it is released whole in a single emission. Use this for in-band markers or control sequences that must not appear half-rendered.
The transformer is single-subscription. Cancelling the output subscription cancels the upstream subscription, so stopping a generation propagates back to the runtime.
- Inheritance
-
- Object
- StreamTransformerBase<
String, String> - TokenSmoother
Constructors
- TokenSmoother({Duration tick = const Duration(milliseconds: 16), Duration window = const Duration(milliseconds: 250), Duration smoothing = const Duration(milliseconds: 150), AtomicChunkPredicate atomic = _neverAtomic})
-
Creates a token stream smoother.
const
Properties
- atomic → AtomicChunkPredicate
-
Decides which chunks are emitted whole instead of paced.
final
- hashCode → int
-
The hash code for this object.
no setterinherited
- runtimeType → Type
-
A representation of the runtime type of the object.
no setterinherited
- smoothing → Duration
-
Time constant of the rate filter: how quickly the release rate follows
a change in generation speed. Larger values are steadier but slower to
react to the model speeding up or slowing down.
final
- tick → Duration
-
How often the release budget is recomputed. One frame by default.
final
- window → Duration
-
How much text the smoother aims to keep buffered, expressed as time at
the current arrival rate — and so also the steady-state lag behind the
source. Larger values ride out longer decode pauses at the cost of the
text trailing further behind generation.
final
Methods
-
bind(
Stream< String> stream) → Stream<String> -
Transforms the provided
stream.override -
cast<
RS, RT> () → StreamTransformer< RS, RT> -
Provides a
StreamTransformer<RS, RT>view of this stream transformer.inherited -
noSuchMethod(
Invocation invocation) → dynamic -
Invoked when a nonexistent method or property is accessed.
inherited
-
toString(
) → String -
A string representation of this object.
inherited
Operators
-
operator ==(
Object other) → bool -
The equality operator.
inherited