Skip to content
#

pallas

Here are 33 public repositories matching this topic...

Frontier-class open models on a free Kaggle TPU v5e-8: GLM-5.3-Flash 320B MoE (~165 tok/s, our own JAX engine) and Qwen3.8-27B bf16 (~130 tok/s), 262k context, prefix caching. Works with Claude Code, Codex, opencode and pi.

  • Updated Oct 6, 2026
  • Python

Measurement harness for the sliding window attention premium in the vLLM TPU Ragged Paged Attention v3 kernel: per layer decode cost, block size control, throughput, and goodput for Gemma 4 31B on TPU v6e.

  • Updated Jul 29, 2026
  • Python

Add this topic to your repo

To associate your repository with the pallas topic, visit your repo's landing page and select "manage topics."

Learn more