What we measured.

One synthetic company, a frozen benchmark, and a conventional retrieval baseline answering the identical questions with the same grader. Every figure below comes from a version-pinned artifact in our repository. Ask, and we will send you the file.

Answer grade vs paired raw-RAG baseline 0.694against 0.639
Retrieval recall@5 0.972against 0.856
Mean reciprocal rank 0.958against 0.870
Permission leaks 0against the baseline’s 1
Running cost per company ~$2.58per month, model spend
Built in 8 sessions~390 tests · $4.26 total API spend

Behaviours we can demonstrate on a call.

What we have not shown.

We would rather you hear this from us than find it in week three. None of it is hedging: each line is a thing we tried to establish and did not.