Deliver streamed tokens live instead of after generation completes - #195
Open
james-333i wants to merge 3 commits into
Open
Deliver streamed tokens live instead of after generation completes#195james-333i wants to merge 3 commits into
james-333i wants to merge 3 commits into
Conversation
The open-ended dependency range resolves llama.swift to releases wrapping current llama.cpp builds, where the Llama trait no longer compiles: llama_sampler_init_penalties regained its leading n_vocab parameter, and llama_model_params replaced use_mmap and use_mlock with a llama_load_mode enum. Pass the vocabulary size at all three penalties call sites and set load_mode to LLAMA_LOAD_MODE_MMAP, matching the previous mmap-only behavior. Verified against llama.swift 2.10549.0 with the full live test suite.
Prompts longer than the batch capacity (512 tokens by default) threw insufficientMemory before generation started, so multi-turn conversations failed as soon as the rendered chat history crossed the batch size, regardless of how much memory was actually available. Feed decoder-only prompts through llama_decode in batch-sized chunks with absolute positions, requesting logits only for the final token. Generation positions now derive from the full prompt length rather than the last batch's token count. Encoder models keep the single-batch requirement, and a prompt that cannot fit in the context window now fails with a new promptExceedsContextWindow error instead of a misleading memory error. Adds a live test generating from a prompt several times the batch size.
streamResponse consumed an inner AsyncThrowingStream whose builder ran the entire generation loop synchronously on the consuming task, so every snapshot buffered and arrived in one burst after generation finished. Yield snapshots directly from the generation loop on the streaming task, and check for task cancellation between tokens so an abandoned stream stops decoding promptly.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
streamResponseran the whole generation loop before the consuming task received anything, so tokens arrived in one burst at the end. This yields each token as it is sampled.Includes the chunked prompt ingestion commit it builds on, and the build-fix commit from #193 as its base.