PIRAMID

Install

Install piramid from a binary or cargo, fetch a model and start the server.

This covers running piramid serve, configuring it, and the HTTP API. config.example.yaml, attached to each release, lists every setting at its default.

Binary

Linux and macOS, x86_64 or arm64:

curl -fsSL https://piramiddb.com/install.sh | sh

The script installs piramid to ~/.local/bin. PIRAMID_INSTALL_DIR changes the directory and PIRAMID_VERSION picks a release tag. The binaries run models on the CPU. Windows builds and every archive are on the releases page.

Cargo

Rust 1.87 or newer. inference-candle runs models; add gpu-cuda for CUDA, which needs the CUDA toolkit at build time:

cargo install piramid --locked --features inference-candle
cargo install piramid --locked --features inference-candle,gpu-cuda

Without inference-candle the server stores and searches collections but cannot generate.

Model

Piramid runs Qwen2 and Qwen3 checkpoints (Qwen2.5 included). The model directory holds config.json, tokenizer.json, tokenizer_config.json and the .safetensors weights:

hf download Qwen/Qwen2.5-0.5B-Instruct --local-dir ./models/Qwen2.5-0.5B-Instruct

Run

piramid.yaml
startup:
  embedding:
    provider: openai
    model: text-embedding-3-small

runtime:
  inference:
    enabled: true
    model_path: ./models/Qwen2.5-0.5B-Instruct
    kv_cache:
      max_bytes: 2147483648
export OPENAI_API_KEY=...
piramid serve --config piramid.yaml

The server listens on 127.0.0.1:6333. Configuration covers GPU settings and every other block; HTTP API and Python client show how to use it.

On this page