Skip to content

Deliver streamed tokens live instead of after generation completes - #195

Open
james-333i wants to merge 3 commits into
huggingface:mainfrom
james-333i:fix/llama-token-streaming
Open

Deliver streamed tokens live instead of after generation completes#195
james-333i wants to merge 3 commits into
huggingface:mainfrom
james-333i:fix/llama-token-streaming

Conversation

@james-333i

@james-333i james-333i commented Aug 27, 2026

Copy link
Copy Markdown

streamResponse ran the whole generation loop before the consuming task received anything, so tokens arrived in one burst at the end. This yields each token as it is sampled.

Includes the chunked prompt ingestion commit it builds on, and the build-fix commit from #193 as its base.

The open-ended dependency range resolves llama.swift to releases
wrapping current llama.cpp builds, where the Llama trait no longer
compiles: llama_sampler_init_penalties regained its leading n_vocab
parameter, and llama_model_params replaced use_mmap and use_mlock
with a llama_load_mode enum.

Pass the vocabulary size at all three penalties call sites and set
load_mode to LLAMA_LOAD_MODE_MMAP, matching the previous mmap-only
behavior. Verified against llama.swift 2.10549.0 with the full live
test suite.
Prompts longer than the batch capacity (512 tokens by default) threw
insufficientMemory before generation started, so multi-turn
conversations failed as soon as the rendered chat history crossed the
batch size, regardless of how much memory was actually available.

Feed decoder-only prompts through llama_decode in batch-sized chunks
with absolute positions, requesting logits only for the final token.
Generation positions now derive from the full prompt length rather
than the last batch's token count. Encoder models keep the
single-batch requirement, and a prompt that cannot fit in the context
window now fails with a new promptExceedsContextWindow error instead
of a misleading memory error.

Adds a live test generating from a prompt several times the batch
size.
streamResponse consumed an inner AsyncThrowingStream whose builder ran
the entire generation loop synchronously on the consuming task, so
every snapshot buffered and arrived in one burst after generation
finished.

Yield snapshots directly from the generation loop on the streaming
task, and check for task cancellation between tokens so an abandoned
stream stops decoding promptly.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant