Benchmark methodology

The numbers. And everything behind them.

The front-page RSS result comes from a defined core-engine test that anyone with the source tree can run.

This page documents exactly what that test does, what it measures, what it deliberately does not measure, and how to reproduce it. The test ships with Meivo itself: crates/gm-server/tests/ram_regression.rs.

Resident memory across ingestion
Footprint and regression ceiling
01

What is measured

Peak RSS(resident set size, the physical RAM a process actually occupies) of a single process running Meivo's full ingestion path, while it ingests 51.2 MB of text. RSS is read with the memory-stats crate.

02

The corpus: deterministic by construction

The test generates 1,000 documents of roughly 50 KB each, 51.2 MB total, from a fixed 115-word vocabulary. Randomness comes from a seeded xorshift64* generator with the constant seed 0x6772_6561_746D_656D ("greatmem" in ASCII), so every run ingests the byte-identical corpus.

03

The system under test: the real pipeline

Nothing is mocked. The test wires up the same components the server uses.

Storage: the real SQLite store with sqlite-vec and FTS5.
Chunking: the production chunker, roughly 400-token chunks with 15% overlap.
Embeddings: real BGE-small-en-v1.5 via fastembed / ONNX Runtime, one model instance behind the bounded batch queue.
Fact extraction: off. The test isolates the embed-and-store write path.
04

Sampling and assertions

RSS is sampled every 50 documents and once more after the pipeline drains. The test then asserts three things:

all 1,000 documents were stored;
chunking really happened;
peak RSS < 1 GB.

The same test runs in CI on every push to main. If that tested workload crosses the threshold, the build breaks.

05

The published run

The numbers on the front page come from a release-mode run on 2026-06-12: Apple M1 Pro, 16 GB RAM, macOS, total duration 2,620 seconds. Peak RSS was 588.1 MB. From 2.6 MB ingested to 51.2 MB ingested, sampled RSS ranged from 494.9 MB to 588.1 MB. These measurements describe this run and workload.

06

The idle figure

The 19 MB idle server number is rss_bytes reported by GET /v1/stats right after booting the published 158 MB Docker image, before the embedding model is loaded.

07

Honest limitations

This is a memory benchmark, not a throughput benchmark.
It does not prove a corpus-size-independent ceiling for arbitrary data, fact extraction, concurrent queries, other storage backends, or the three-service enterprise stack.
RSS includes allocator slack and varies a few percent between platforms and runs.
Absolute numbers differ with the embedding model you configure.
We publish no measurements of other tools; measure anything you like with the same corpus generator.
08

Reproduce it

# from the Meivo source tree
	cargo test -p gm-server --release -- --ignored ram_regression --nocapture

Requirements: Rust stable, roughly 150 MB of disk for the first model download, and patience. Expect tens of minutes on CPU.

Raw output of the published run
[ram_regression] docs=  50  ingested=  2.6 MB  rss= 535.9 MB  peak= 535.9 MB
[ram_regression] docs= 100  ingested=  5.1 MB  rss= 530.5 MB  peak= 535.9 MB
[ram_regression] docs= 150  ingested=  7.7 MB  rss= 529.3 MB  peak= 535.9 MB
[ram_regression] docs= 200  ingested= 10.2 MB  rss= 535.8 MB  peak= 535.9 MB
[ram_regression] docs= 250  ingested= 12.8 MB  rss= 512.6 MB  peak= 535.9 MB
[ram_regression] docs= 300  ingested= 15.4 MB  rss= 578.1 MB  peak= 578.1 MB
[ram_regression] docs= 350  ingested= 17.9 MB  rss= 545.5 MB  peak= 578.1 MB
[ram_regression] docs= 400  ingested= 20.5 MB  rss= 547.7 MB  peak= 578.1 MB
[ram_regression] docs= 450  ingested= 23.1 MB  rss= 520.9 MB  peak= 578.1 MB
[ram_regression] docs= 500  ingested= 25.6 MB  rss= 545.8 MB  peak= 578.1 MB
[ram_regression] docs= 550  ingested= 28.2 MB  rss= 550.6 MB  peak= 578.1 MB
[ram_regression] docs= 600  ingested= 30.7 MB  rss= 543.4 MB  peak= 578.1 MB
[ram_regression] docs= 650  ingested= 33.3 MB  rss= 556.5 MB  peak= 578.1 MB
[ram_regression] docs= 700  ingested= 35.9 MB  rss= 543.7 MB  peak= 578.1 MB
[ram_regression] docs= 750  ingested= 38.4 MB  rss= 530.4 MB  peak= 578.1 MB
[ram_regression] docs= 800  ingested= 41.0 MB  rss= 526.9 MB  peak= 578.1 MB
[ram_regression] docs= 850  ingested= 43.6 MB  rss= 494.9 MB  peak= 578.1 MB
[ram_regression] docs= 900  ingested= 46.1 MB  rss= 560.0 MB  peak= 578.1 MB
[ram_regression] docs= 950  ingested= 48.7 MB  rss= 588.1 MB  peak= 588.1 MB
[ram_regression] docs=1000  ingested= 51.2 MB  rss= 541.3 MB  peak= 588.1 MB
[ram_regression] SUMMARY: ingested 51.2 MB across 1000 docs -> 31000 chunks; peak RSS 588.1 MB (limit 1024 MB)
test ram_regression_50mb_ingest_stays_under_1gb ... ok
test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 2619.96s