Skip to content

Add an assistant prefill option to LlamaLanguageModel - #198

Open
james-333i wants to merge 2 commits into
huggingface:mainfrom
james-333i:feat/llama-assistant-prefill
Open

Add an assistant prefill option to LlamaLanguageModel#198
james-333i wants to merge 2 commits into
huggingface:mainfrom
james-333i:feat/llama-assistant-prefill

Conversation

@james-333i

@james-333i james-333i commented Aug 27, 2026

Copy link
Copy Markdown

Adds assistantPrefill to the custom generation options, appended after the rendered assistant header so the model continues from supplied text. The motivating case is suppressing default reasoning output on models whose chat template offers no switch for it, such as prefilling an empty think block on Qwen, while leaving thinking available to callers who want it.

Includes the build-fix commit from #193 as its base.

The open-ended dependency range resolves llama.swift to releases
wrapping current llama.cpp builds, where the Llama trait no longer
compiles: llama_sampler_init_penalties regained its leading n_vocab
parameter, and llama_model_params replaced use_mmap and use_mlock
with a llama_load_mode enum.

Pass the vocabulary size at all three penalties call sites and set
load_mode to LLAMA_LOAD_MODE_MMAP, matching the previous mmap-only
behavior. Verified against llama.swift 2.10549.0 with the full live
test suite.
llama_chat_apply_template cannot express chat template kwargs, so
there is no way to pass switches like Qwen's enable_thinking. Models
that reason by default then spend their entire token budget inside a
think block with no way to turn it off.

Add assistantPrefill to CustomGenerationOptions. The text is appended
after the assistant header of the rendered prompt and the model
continues from it, so prefilling an empty think block suppresses
reasoning output, matching how llama.cpp itself implements a zero
reasoning budget.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant