S3
Every acknowledged write is durable here.
Details
Immutable logs acknowledge writes; a published root selects sealed packs, indexes and manifests.
Open-source vector database
Every write is durable in object storage before it is acknowledged. A restart recovers your data; RAM and SSD are disposable caches. Built to go far on less.
from glider_client import Client
client = Client("http://localhost:8080")
client.create_collection("demo", dimensions=3)
docs = client.collection("demo")
docs.upsert([{"id": 1, "vector": [0, 0, 0],
"metadata": {"color": "red"}}])
hits = docs.query([1, 1, 0.9], k=10)# After creating and writing the demo collection
curl -sS localhost:8080/v1/collections/demo/query \
-H 'content-type: application/json' \
-d '{"vector":[1,1,0.9],"k":10}'claude mcp add glider \
-e GLIDER_COLLECTION=memory -- glider-mcpSIFT1M · 128 dimensions · k=10 · S3 Standard · c7g.2xlarge, eu-central-1. Recall is the static mean; latency is from concurrent readers and writers. Methodology & raw results →
10,000 collections · 10M vectors · one server: warm queries over 64 active collections reached 2,457 queries/s at 18.5 ms p95. All 10,000 collections were verified after kill -9 during ingest. Multi-tenant results →
Storage playground
Pick a spot, follow the blocks, then pull the plug.
Route and select in RAM → fetch from RAM, SSD or S3 → rerank in RAM on the CPU.
Fetch: RAM hot cache → SSD cache → S3 range read. Then CPU rerank in RAM: fetched blocks + unsealed tail.
Click the map to route a query through centroids and cached blocks.
Acknowledged writes lost: 0
Top-10 neighbours appear as connected points after a query.
| Rank | ID | Distance |
|---|
Latency = 0.25 ms RAM routing + 0.25 ms RAM selection + 0.5 ms RAM CPU reranking + 0.1 ms per RAM hit + 0.5 ms per SSD hit + 20 ms per batch of four parallel S3 reads. Illustrative, not measured.
Every acknowledged write is durable here.
Immutable logs acknowledge writes; a published root selects sealed packs, indexes and manifests.
Fetched blocks stay here for the next query.
Slots show capacity, not physical blocks. Least recently used blocks are evicted when full; losing this disk loses no acknowledged writes.
Routing and recent data live in process memory.
Full vectors enter RAM only as fetched blocks or unsealed writes. The hot block cache is bounded to 4 MiB; a crash clears all process memory.
30,000 sampled dots stand for the selected dataset, already sealed and clustered (including the 100K example). Seeded Gaussian-mixture islands with Poisson-disc centres have varied populations averaging about 4,000 vectors per centroid. Projected islands can overlap; cluster membership represents the original embedding. Routing uses the M31 profile's 32 probes, scaled to 4 for 100K. Reranking is exact within selected blocks and the tail; the overall search is approximate. The projection and five-bit scores use two dimensions, while sizes assume 128-dimensional float32 vectors: 24 B per directory slot, 80 B of five-bit codes + 4 B ID offset + 8 B posting sequence per sealed row, 512 B per centroid, and 640 B per tail/log row. Blocks are sized proportionally up to 120 KiB; packs hold up to eight blocks (960 KiB) plus sketches. Compression, allocator overhead, routing cache files and retired objects are omitted. RAM's hot cache is 4 MiB; SSD defaults to 256 MiB. Latency = 0.25 ms RAM routing + 0.25 ms RAM selection + 0.5 ms RAM CPU reranking + 0.1 ms per RAM hit + 0.5 ms per SSD hit + 20 ms per batch of four parallel S3 range GETs. Trace widths follow these simulated phase times with a 14% minimum per phase; the animation stretches the whole query to 1.2 seconds. Each batch adds 100 new IDs; 32 log objects trigger sealing. Restart compresses the real lease wait and recovery into five visual stages.
How it works
A write is acknowledged once it is on S3; queries read clustered packs through a local SSD cache.
Why Glider
A restart fences the old writer at the store and replays the log. No operator step.
Reuse a write’s request ID within the retry window to get its original outcome without applying it twice.
A local SSD cache serves warm queries. Lose it and you lose no data.
The server builds its clustered index on its own and rebuilds it as data grows.
Create and delete collections over HTTP, each with its own dimension and metric.
Equality, sets, numeric ranges and nested logic, with an exact mode for every match.
Filters
Filtered search is approximate by default. Add "exact": true and you get every match, up to k.
{
"vector": [1, 1, 0.9],
"k": 10,
"filter": {
"color": "red",
"price": {"$gt": 1.5, "$lte": 10},
"tag": {"$in": ["a", "b"]}
},
"exact": true
}
Agent memory · MCP
glider-mcp gives MCP clients three tools. Memories survive restarts, and request IDs make writes safe to retry within the retained window.
docker run -d --name glider -p 8080:8080 \
-v glider-data:/var/lib/glider \
-e GLIDER_DATA_DIR=/var/lib/glider/data \
ghcr.io/omerfeyzioglu/glider:latest
pip install \
"glider-client[mcp] @ git+https://github.com/omerfeyzioglu/glider#subdirectory=clients/python"
claude mcp add glider \
-e GLIDER_COLLECTION=memory -- glider-mcp
Quickstart
Only Docker and curl. This demo uses a local Docker volume that survives container removal. For object storage, run it on S3.
Once it is running, open localhost:8080/console to browse points and query your server.
Keep the server running. Use another terminal for the next steps.
docker run --rm -p 8080:8080 -e GLIDER_DIMENSIONS=3 \
-v glider-quickstart-data:/var/lib/glider \
-e GLIDER_DATA_DIR=/var/lib/glider/data \
ghcr.io/omerfeyzioglu/glider:latest
curl -sS localhost:8080/v1/write \
-H 'content-type: application/json' \
-d '{"upsert":[
{"id":1,"vector":[0,0,0],"metadata":{"color":"red"}},
{"id":2,"vector":[1,1,1]}]}'
curl -sS localhost:8080/v1/query \
-H 'content-type: application/json' \
-d '{"vector":[1,1,0.9],"k":2,"include_metadata":true}'
docker run --rm -p 8080:8080 -e GLIDER_DIMENSIONS=3 \
-v glider-quickstart-data:/var/lib/glider \
-e GLIDER_DATA_DIR=/var/lib/glider/data \
ghcr.io/omerfeyzioglu/glider:latest
pip install "git+https://github.com/omerfeyzioglu/glider#subdirectory=clients/python"
from glider_client import Client
client = Client("http://localhost:8080")
client.upsert([
{"id": 1, "vector": [0, 0, 0], "metadata": {"color": "red"}},
{"id": 2, "vector": [1, 1, 1]},
])
from glider_client import Client
client = Client("http://localhost:8080")
for hit in client.query([1, 1, 0.9], k=2):
print(hit.id, hit.distance)
Good to know