Download Ollaya
A desktop app and a command line for macOS, Windows and Linux, or a Docker image.
Command line
curl -fsSL https://ollaya.dev/install.sh | shThe script detects your CPU and NVIDIA GPU, downloads the release from GitHub, checks its sha256 and, where systemd runs, sets up the ollaya service. With a GPU it also fetches the CUDA libraries (1 to 1.6 GB). It never installs drivers.
Run a model
ollaya run winnow:e4bThe recommended model: close to Jev's accuracy, in about 90 ms on an NVIDIA GPU with 10 GB or more (an 8 GB download). Otherwise, ollaya run laya answers in a fraction of a second, on the CPU too.
Requirements
- x86-64 or ARM64 with glibc 2.38 or newer: Ubuntu 24.04, Debian 13, Fedora 39, RHEL 10 or newer.
- Runs on the CPU. An NVIDIA GPU is optional: driver R525 or newer, on x86-64 (CUDA 13 from R580, CUDA 12 before and for GTX 10-series and Volta cards).
Desktop app
Start and stop the server, download models and try them, in one window. Your code talks to the same local API; for the ollaya command, install the command line too.
On its own the app runs models on the CPU. With an NVIDIA GPU, install the command line as well: the app then starts the server from it, on the GPU, and its status line says which one is running.
Desktop app
Start and stop the server, download models and try them, in one window. Your code talks to the same local API; for the ollaya command, install the command line too.
On its own the app runs models on the CPU. With an NVIDIA GPU, install the command line as well: the app then starts the server from it, on the GPU, and its status line says which one is running.
The Windows app is not code-signed yet. If SmartScreen stops the installer, choose More info, then Run anyway.
Command line
irm https://ollaya.dev/install.ps1 | iexIn PowerShell. It installs ollaya for your user, with no administrator rights, and puts it on your PATH. The archives are checked against the release's sha256. With an NVIDIA GPU it also fetches the CUDA libraries (about 1 GB). It never installs drivers.
Run a model
ollaya run winnow:e4bThe recommended model: close to Jev's accuracy, in about 90 ms on an NVIDIA GPU with 10 GB or more (an 8 GB download). Otherwise, ollaya run laya answers in a fraction of a second, on the CPU too.
Requirements
- Windows 10 or 11 on a 64-bit x86 PC. Runs on the CPU.
- An NVIDIA GPU is optional: driver R527 or newer (CUDA 13 from R580, CUDA 12 before and for GTX 10-series and Volta cards). The command line uses it; the desktop app uses it when the command line is installed too, and the CPU otherwise.
- WSL 2 with the Linux installer works too. The server in WSL answers Windows programs at
localhost:11435.
Desktop app
Start and stop the server, download models and try them, in one window. Your code talks to the same local API; for the ollaya command, install the command line too.
Command line
curl -fsSL https://ollaya.dev/install.sh | shThe script downloads the release from GitHub and checks its sha256. Start the server with ollaya serve, or let ollaya run start it for you.
Run a model
ollaya run layaRequirements
- A Mac with Apple silicon (arm64) and macOS 14 or newer.
CPU
docker run -d --name ollaya -p 11435:11435 -v ollaya:/home/ollaya/.ollaya ghcr.io/ollaya-dev/ollayaNVIDIA GPU
docker run -d --name ollaya --gpus=all -p 11435:11435 -v ollaya:/home/ollaya/.ollaya ghcr.io/ollaya-dev/ollaya:cudaNeeds the NVIDIA Container Toolkit and a host driver with CUDA 13 support (R580 or newer). For drivers with CUDA 12 (R525 to R575), or a GTX 10-series or Volta GPU, use :cuda12.
Run a model
docker exec -it ollaya ollaya run layaThe CPU image is built for linux/amd64 and linux/arm64, the :cuda and :cuda12 images for linux/amd64. Models are kept in the ollaya volume.
Prefer a tarball? Every release on GitHub Releases has the archives and a sha256sum.txt.