-
Notifications
You must be signed in to change notification settings - Fork 3k
Pull requests: JustVugg/colibri
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
serve: add opt-in assistant continuation for GLM-5.3
#1402
opened Sep 8, 2026 by
enitimeago
Loading…
5 tasks done
perf: COALESCE=1 multi-expert coalesced pread for contiguous shard runs
#1398
opened Sep 8, 2026 by
KyleSanderson
•
Draft
5 tasks done
perf: IDOT_TEAM=1 fuses three OpenMP regions into one team per expert…
#1397
opened Sep 8, 2026 by
KyleSanderson
Loading…
5 tasks done
feat(web): drop redundant Medium from GLM 5.3 reasoning selector
#1396
opened Sep 7, 2026 by
dmoraesrs
Loading…
Opt-in exact verify mode for speculative batches (COLI_EXACT_VERIFY=1)
#1395
opened Sep 7, 2026 by
rybruscoe
Loading…
fix(qwen36 tier): qt_fill_wait returns only when the last upload is resident
#1390
opened Sep 7, 2026 by
crichalchemist
Contributor
Loading…
qwen36 tier: QT_MAX_ROWS, placement pin, ownership-based int8 free (follow-up to #1344)
#1388
opened Sep 7, 2026 by
crichalchemist
Contributor
Loading…
test(tools): gate the logprob gap check on the engine preamble; add the native-MTP witness
#1357
opened Sep 5, 2026 by
monotophic
Contributor
Loading…
feat(engine): replace the ablation scoring mode with a checked, digest-bound one
#1356
opened Sep 5, 2026 by
monotophic
Contributor
Loading…
feat(tools): offline ablation-evidence checker and an evidence-bound eval harness
#1355
opened Sep 5, 2026 by
monotophic
Contributor
Loading…
OpenAI API Instrumentation (e.g. seed, logprobs, echo, an array prompt)
#1353
opened Sep 5, 2026 by
monotophic
Contributor
Loading…
feat(qwen36): load MLX-affine dense weights and norms from a qpack container
#1343
opened Sep 4, 2026 by
Avicennasis
Contributor
Loading…
perf: O(1) LRU victim selection via intrusive recency lists (#1050)
#1342
opened Sep 4, 2026 by
Petsku01
Contributor
Loading…
qwen36: Vulkan expert tier, and staged device-local uploads for cards without Resizable BAR
#1338
opened Sep 4, 2026 by
crichalchemist
Contributor
Loading…
docs: report GLM-5.3-Flash PRO 6000 speed and quality findings
#1337
opened Sep 4, 2026 by
lEWFkRAD
Contributor
Loading…
2 of 5 tasks
feat(qwen36): name a checkpoint by its geometry, and a converter that refuses what it cannot place
#1326
opened Sep 3, 2026 by
kreuzzelg
Contributor
Loading…
feat(qwen36): routed experts from a qpack container through bounded Metal slots
#1323
opened Sep 2, 2026 by
Avicennasis
Contributor
Loading…
glm53: size the expert cache around the model, not around MemAvailable
#1321
opened Sep 2, 2026 by
ronaldcklomp
Contributor
Loading…
fix(planner): honor cgroup memory limits
#1316
opened Sep 1, 2026 by
Avicennasis
Contributor
Loading…
3 of 5 tasks
Add resumable, hash-verified qpack installers for Hugging Face and static mirrors
#1315
opened Sep 1, 2026 by
Avicennasis
Contributor
Loading…
perf(quant): compute four output rows per pass in matmul_fp8 (+21.34% tok/s, bit-exact)
#1313
opened Sep 1, 2026 by
BrianHeeseIs
Contributor
Loading…
fix(coli): discover serve processes where there is no /proc
#1307
opened Aug 31, 2026 by
BrianHeeseIs
Contributor
Loading…
docs: add a reproducible benchmarking protocol
#1294
opened Aug 31, 2026 by
ZacharyZcR
Contributor
Loading…
Add MLX affine qpack reader, inspector, and packed affine Metal GEMV
#1290
opened Aug 31, 2026 by
Avicennasis
Contributor
Loading…
Previous Next
ProTip!
Mix and match filters to narrow down what you’re looking for.