Skip to content

Fix constrained sampling ignoring the allowed-token set - #197

Open
james-333i wants to merge 2 commits into
huggingface:mainfrom
james-333i:fix/llama-constrained-sampling
Open

Fix constrained sampling ignoring the allowed-token set#197
james-333i wants to merge 2 commits into
huggingface:mainfrom
james-333i:fix/llama-constrained-sampling

Conversation

@james-333i

@james-333i james-333i commented Aug 27, 2026

Copy link
Copy Markdown

Structured generation masked logits in the buffer returned by llama_get_logits, but the context refreshes that buffer, so the sampler drew from the full vocabulary and produced invalid JSON. This applies the constraint through llama_token_data_array and llama_sampler_apply so masked tokens cannot be sampled.

Includes the build-fix commit from #193 as its base.

The open-ended dependency range resolves llama.swift to releases
wrapping current llama.cpp builds, where the Llama trait no longer
compiles: llama_sampler_init_penalties regained its leading n_vocab
parameter, and llama_model_params replaced use_mmap and use_mlock
with a llama_load_mode enum.

Pass the vocabulary size at all three penalties call sites and set
load_mode to LLAMA_LOAD_MODE_MMAP, matching the previous mmap-only
behavior. Verified against llama.swift 2.10549.0 with the full live
test suite.
LlamaTokenBackend.sample(from:) masked disallowed tokens by writing
negative infinity into the buffer returned by llama_get_logits, then
called llama_sampler_sample. The context can refresh that buffer when
the sampler fetches logits, so the mask is dropped and sampling runs
over the full vocabulary. Constrained JSON generation then emits
tokens outside the allowed set and fails.

Build a llama_token_data_array containing only the allowed tokens and
their logits, apply the sampler chain to it with llama_sampler_apply,
and return the selected candidate.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant