Skip to content

Ingest decoder prompts in batch-sized chunks - #196

Open
james-333i wants to merge 2 commits into
huggingface:mainfrom
james-333i:fix/llama-chunked-prompt
Open

Ingest decoder prompts in batch-sized chunks#196
james-333i wants to merge 2 commits into
huggingface:mainfrom
james-333i:fix/llama-chunked-prompt

Conversation

@james-333i

@james-333i james-333i commented Aug 27, 2026

Copy link
Copy Markdown

Prompts longer than n_batch failed to decode, which surfaced as generation errors after a few exchanges of context growth. This feeds the prompt to llama_decode in batch-sized chunks and reports a prompt that exceeds the context window as a typed error.

Includes the build-fix commit from #193 as its base.

The open-ended dependency range resolves llama.swift to releases
wrapping current llama.cpp builds, where the Llama trait no longer
compiles: llama_sampler_init_penalties regained its leading n_vocab
parameter, and llama_model_params replaced use_mmap and use_mlock
with a llama_load_mode enum.

Pass the vocabulary size at all three penalties call sites and set
load_mode to LLAMA_LOAD_MODE_MMAP, matching the previous mmap-only
behavior. Verified against llama.swift 2.10549.0 with the full live
test suite.
Prompts longer than the batch capacity (512 tokens by default) threw
insufficientMemory before generation started, so multi-turn
conversations failed as soon as the rendered chat history crossed the
batch size, regardless of how much memory was actually available.

Feed decoder-only prompts through llama_decode in batch-sized chunks
with absolute positions, requesting logits only for the final token.
Generation positions now derive from the full prompt length rather
than the last batch's token count. Encoder models keep the
single-batch requirement, and a prompt that cannot fit in the context
window now fails with a new promptExceedsContextWindow error instead
of a misleading memory error.

Adds a live test generating from a prompt several times the batch
size.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant