metor memory · #1 on EnterpriseRAG-Bench

Your agents are brilliant.
Their memory isn't.

metor memory is a memory engine for company knowledge: a second brain your AI agents can rely on. It answers with verified facts and exact counts instead of plausible guesses, and it plugs into whatever agent harness you already use.

#1on EnterpriseRAG-Bench, verified and rescored by the benchmark maintainers
80.34official overall score, single pass: one agent, one attempt per question
86.2answer completeness, the highest on the board
500 / 500questions answered across ~512,000 workplace documents from nine source systems
What makes it different

Most retrieval hands your agent a pile of plausible text. metor memory gives it instruments.

Counts you can trust

“How many incidents in April?” is answered by counting, not by guessing from whatever surfaced. Complete result sets, provable absence.

Disambiguation built in

Real corpora are full of look-alikes: similar incidents, near-identical playbooks. metor memory pins the attributes that tell them apart, so agents answer from the right document, not the loudest one.

Semantic when it helps

When wording and vocabulary diverge, semantic matching steps in, as one tool among several, not as the single point of failure.

Verified, cited answers

Every answer carries the documents it stands on. Agents are made to read before they claim.

Bring your own harness

The memory stays. The harness and the model remain your choice.

metor memory speaks MCP and ships a CLI, so it works with any agent runtime: today, and with whatever you switch to next.

Claude CodeOpenCodeany MCP-capable harness

We proved the point by running the identical memory under two harnesses and four frontier models. The quality lives in the memory engine, not in the model on top. For our own benchmark runs we settled on Claude Code with Opus 5: a choice, not a dependency.

metor memory is a product of its own. It is not part of the metor agent platform, and it does not need it.

Measured, not promised

The hardest public benchmark for company knowledge. We hold the top spot.

EnterpriseRAG-Bench, built by Onyx, is a synthetic company of about 512,000 documents across nine source systems and 500 adversarial questions, scored as correctness × completeness by an LLM judge. Confident nonsense scores zero. The benchmark story explains why it is the one to watch.

#SystemOverallCorrectnessCompletenessDoc. recallInvalid extra
1metor memory · listed as metor.com80.3482.086.2285.534.96
2CDL (Causal Dynamics Lab)78.9582.085.1680.5414.08
3Troml76.7983.881.8486.5512.65
4Skyller71.9377.079.1481.608.86
5OpenClaw68.2281.672.8679.020.47
6SovraRAG.ch65.6174.672.3778.808.87
7fgroo63.2771.071.0372.500.63
8OpenAI File Search61.0369.867.8771.6515.70

Official leaderboard standings, 28 August 2026. Single pass: one agent, one attempt per question. See the leaderboard.

See it on your own knowledge.

Interested in the benchmark results in detail, a demo, or metor memory on your own data? Write to us.