Skip to content

Download Ollaya

A desktop app and a command line for macOS, Windows and Linux, or a Docker image.

Command line

curl -fsSL https://ollaya.dev/install.sh | sh

The script detects your CPU and NVIDIA GPU, downloads the release from GitHub, checks its sha256 and, where systemd runs, sets up the ollaya service. With a GPU it also fetches the CUDA libraries (1 to 1.6 GB). It never installs drivers.

Run a model

ollaya run winnow:e4b

The recommended model: close to Jev's accuracy, in about 90 ms on an NVIDIA GPU with 10 GB or more (an 8 GB download). Otherwise, ollaya run laya answers in a fraction of a second, on the CPU too.

Requirements

  • x86-64 or ARM64 with glibc 2.38 or newer: Ubuntu 24.04, Debian 13, Fedora 39, RHEL 10 or newer.
  • Runs on the CPU. An NVIDIA GPU is optional: driver R525 or newer, on x86-64 (CUDA 13 from R580, CUDA 12 before and for GTX 10-series and Volta cards).

Desktop app

Start and stop the server, download models and try them, in one window. Your code talks to the same local API; for the ollaya command, install the command line too.

On its own the app runs models on the CPU. With an NVIDIA GPU, install the command line as well: the app then starts the server from it, on the GPU, and its status line says which one is running.

Prefer a tarball? Every release on GitHub Releases has the archives and a sha256sum.txt.