Install
Install piramid from a binary or cargo, fetch a model and start the server.
This covers running piramid serve, configuring it, and the HTTP API. config.example.yaml, attached
to each release, lists every setting at its default.
Binary
Linux and macOS, x86_64 or arm64:
curl -fsSL https://piramiddb.com/install.sh | shThe script installs piramid to ~/.local/bin. PIRAMID_INSTALL_DIR changes the directory and
PIRAMID_VERSION picks a release tag. The binaries run models on the CPU. Windows builds and every
archive are on the releases page.
Cargo
Rust 1.87 or newer. inference-candle runs models; add gpu-cuda for CUDA, which needs the CUDA
toolkit at build time:
cargo install piramid --locked --features inference-candle
cargo install piramid --locked --features inference-candle,gpu-cudaWithout inference-candle the server stores and searches collections but cannot generate.
Model
Piramid runs Qwen2 and Qwen3 checkpoints (Qwen2.5 included). The model directory holds
config.json, tokenizer.json, tokenizer_config.json and the .safetensors weights:
hf download Qwen/Qwen2.5-0.5B-Instruct --local-dir ./models/Qwen2.5-0.5B-InstructRun
startup:
embedding:
provider: openai
model: text-embedding-3-small
runtime:
inference:
enabled: true
model_path: ./models/Qwen2.5-0.5B-Instruct
kv_cache:
max_bytes: 2147483648export OPENAI_API_KEY=...
piramid serve --config piramid.yamlThe server listens on 127.0.0.1:6333. Configuration covers GPU settings
and every other block; HTTP API and Python client show how to use it.