Counts you can trust
“How many incidents in April?” is answered by counting, not by guessing from whatever surfaced. Complete result sets, provable absence.
metor memory is a memory engine for company knowledge: a second brain your AI agents can rely on. It answers with verified facts and exact counts instead of plausible guesses, and it plugs into whatever agent harness you already use.
“How many incidents in April?” is answered by counting, not by guessing from whatever surfaced. Complete result sets, provable absence.
Real corpora are full of look-alikes: similar incidents, near-identical playbooks. metor memory pins the attributes that tell them apart, so agents answer from the right document, not the loudest one.
When wording and vocabulary diverge, semantic matching steps in, as one tool among several, not as the single point of failure.
Every answer carries the documents it stands on. Agents are made to read before they claim.
metor memory speaks MCP and ships a CLI, so it works with any agent runtime: today, and with whatever you switch to next.
We proved the point by running the identical memory under two harnesses and four frontier models. The quality lives in the memory engine, not in the model on top. For our own benchmark runs we settled on Claude Code with Opus 5: a choice, not a dependency.
metor memory is a product of its own. It is not part of the metor agent platform, and it does not need it.
EnterpriseRAG-Bench, built by Onyx, is a synthetic company of about 512,000 documents across nine source systems and 500 adversarial questions, scored as correctness × completeness by an LLM judge. Confident nonsense scores zero. The benchmark story explains why it is the one to watch.
| # | System | Overall | Correctness | Completeness | Doc. recall | Invalid extra |
|---|---|---|---|---|---|---|
| 1 | metor memory · listed as metor.com | 80.34 | 82.0 | 86.22 | 85.53 | 4.96 |
| 2 | CDL (Causal Dynamics Lab) | 78.95 | 82.0 | 85.16 | 80.54 | 14.08 |
| 3 | Troml | 76.79 | 83.8 | 81.84 | 86.55 | 12.65 |
| 4 | Skyller | 71.93 | 77.0 | 79.14 | 81.60 | 8.86 |
| 5 | OpenClaw | 68.22 | 81.6 | 72.86 | 79.02 | 0.47 |
| 6 | SovraRAG.ch | 65.61 | 74.6 | 72.37 | 78.80 | 8.87 |
| 7 | fgroo | 63.27 | 71.0 | 71.03 | 72.50 | 0.63 |
| 8 | OpenAI File Search | 61.03 | 69.8 | 67.87 | 71.65 | 15.70 |
Official leaderboard standings, 28 August 2026. Single pass: one agent, one attempt per question. See the leaderboard.
Interested in the benchmark results in detail, a demo, or metor memory on your own data? Write to us.