-
Notifications
You must be signed in to change notification settings - Fork 39
All issues
Issue creation is restricted in this repository
Issues
is:issue state:open
is:issue state:open
Search results
perf(cache): speculative rollback rewrites the whole KV buffer every round, so its cost scales with context instead of with the block
area:coremlxcel-core: MLX FFI, primitives, KV cache, layersmlxcel-core: MLX FFI, primitives, KV cache, layerspriority:mediumMedium priorityMedium prioritystatus:readyReady to be worked onReady to be worked ontype:performancePerformance improvementsPerformance improvementsStatus: Open.#1209 In lablup/mlxcel;perf(speculative): the adaptive block-size controller holds Gemma 4 MTP 5% below its own optimum
area:inferenceGeneration, sampling, decoding (incl. speculative, DRY)Generation, sampling, decoding (incl. speculative, DRY)priority:mediumMedium priorityMedium prioritystatus:readyReady to be worked onReady to be worked ontype:performancePerformance improvementsPerformance improvementsStatus: Open.#1207 In lablup/mlxcel;fix(speculative): Gemma 4 MTP breaks byte-identity on generation 15+, and its gate returns unconditionally true
area:inferenceGeneration, sampling, decoding (incl. speculative, DRY)Generation, sampling, decoding (incl. speculative, DRY)platform:macosmacOS (Apple Silicon) specificmacOS (Apple Silicon) specificpriority:mediumMedium priorityMedium prioritystatus:readyReady to be worked onReady to be worked ontype:bugBug fixes, error corrections, or issue resolutionsBug fixes, error corrections, or issue resolutionsStatus: Open.#1188 In lablup/mlxcel;perf(speculative): cut the MTP drafter's per-step cost, now the binding constraint on M5-class hardware
area:coremlxcel-core: MLX FFI, primitives, KV cache, layersmlxcel-core: MLX FFI, primitives, KV cache, layersarea:inferenceGeneration, sampling, decoding (incl. speculative, DRY)Generation, sampling, decoding (incl. speculative, DRY)priority:highHigh priorityHigh prioritystatus:readyReady to be worked onReady to be worked ontype:performancePerformance improvementsPerformance improvementsStatus: Open.#1185 In lablup/mlxcel;test: e2e_crossover_larger_models_benefit_more asserts on wall-clock noise and flakes under load
area:benchmarkBenchmark harness and performance measurement (bench_*.sh, /update-benchmarks)Benchmark harness and performance measurement (bench_*.sh, /update-benchmarks)priority:mediumMedium priorityMedium prioritystatus:readyReady to be worked onReady to be worked ontype:testTest related changesTest related changesStatus: Open.#1184 In lablup/mlxcel;perf(inference): decide by measurement whether a device-side early-exit MTP verify walk beats the batched full-logits verifier
area:coremlxcel-core: MLX FFI, primitives, KV cache, layersmlxcel-core: MLX FFI, primitives, KV cache, layersarea:inferenceGeneration, sampling, decoding (incl. speculative, DRY)Generation, sampling, decoding (incl. speculative, DRY)priority:lowLow priorityLow prioritystatus:readyReady to be worked onReady to be worked ontype:performancePerformance improvementsPerformance improvementsStatus: Open.#1179 In lablup/mlxcel;fix(server): /v1/responses and the offline CLI still swallow template rejections that /v1/chat/completions now returns 400 for
area:cliCommand-line interface / CLI flagsCommand-line interface / CLI flagsarea:docsUser and developer documentationUser and developer documentationarea:inferenceGeneration, sampling, decoding (incl. speculative, DRY)Generation, sampling, decoding (incl. speculative, DRY)priority:mediumMedium priorityMedium prioritystatus:readyReady to be worked onReady to be worked ontype:bugBug fixes, error corrections, or issue resolutionsBug fixes, error corrections, or issue resolutionsStatus: Open.#1176 In lablup/mlxcel;fix(server): model-aware snapshot capacity default: 27B hybrid snapshots thrash the 512 MiB store to a 0% multi-turn hit rate
area:coremlxcel-core: MLX FFI, primitives, KV cache, layersmlxcel-core: MLX FFI, primitives, KV cache, layersarea:inferenceGeneration, sampling, decoding (incl. speculative, DRY)Generation, sampling, decoding (incl. speculative, DRY)priority:highHigh priorityHigh prioritystatus:readyReady to be worked onReady to be worked ontype:bugBug fixes, error corrections, or issue resolutionsBug fixes, error corrections, or issue resolutionsStatus: Open.#1167 In lablup/mlxcel;feat(models): video input for the Qwen VL families (Qwen3.5-VL / Qwen3.8, then the rest)
area:modelsModel architectures, weights, loading, metadataModel architectures, weights, loading, metadatamodeltype:vlmVision-language modelVision-language modelpriority:lowLow priorityLow prioritystatus:readyReady to be worked onReady to be worked ontype:enhancementNew features, capabilities, or significant additionsNew features, capabilities, or significant additionsStatus: Open.#1166 In lablup/mlxcel;perf(server): gate the history-boundary render on model snapshot-reuse capability
area:inferenceGeneration, sampling, decoding (incl. speculative, DRY)Generation, sampling, decoding (incl. speculative, DRY)priority:mediumMedium priorityMedium prioritystatus:readyReady to be worked onReady to be worked ontype:performancePerformance improvementsPerformance improvementsStatus: Open.#1153 In lablup/mlxcel;chore: Enforce the
// Used by:convention in CI, staleness includedpriority:lowLow priorityLow prioritystatus:readyReady to be worked onReady to be worked ontype:choreMaintenance tasks (build, CI, etc.)Maintenance tasks (build, CI, etc.)Status: Open.#1141 In lablup/mlxcel;chore: Add a CI check keeping environment-variables.md in sync with src
priority:lowLow priorityLow prioritystatus:readyReady to be worked onReady to be worked ontype:choreMaintenance tasks (build, CI, etc.)Maintenance tasks (build, CI, etc.)Status: Open.#1140 In lablup/mlxcel;