Self-hosted memory engine

Associative memory for AI agents, modelled on the human mind.

Named for engraphy, the old memory-science term for laying down an engram, the trace a memory leaves in the brain. Engraphy does that for your agents.

72.9% Best-in-class temporal reasoning and open-domain knowledge on LoCoMo, both by wide margins over the next-best published system. Matches the leading memory systems on LoCoMo overall. See the full comparison

Where it runs

Self-hosted, by you

What it needs

Postgres and pgvector

How agents connect

Any MCP client

Install it in your editor, or host the whole engine.

Engraphy is a server you host. Point any MCP client at it. Embeddings run in-process, so there is no API key and no data leaves the box.

The VS Code extension talks to an Engraphy server you run. Self-hosting needs Docker, which is the supported runner: one docker compose up brings up Postgres, pgvector, the migrations and the server.

Engraphy holds everything your agent must remember.

Your agent writes what it learns. Engraphy keeps it clean, findable, and connected over time.

No duplicates

Your agent writes the same fact twice. Engraphy sees the match and stores it once.

Search by meaning

Ask in your own words. Engraphy matches on meaning and on exact terms together.

Linked facts

Engraphy links related memories into a graph. An agent walks from one fact to the next.

Private by space

Engraphy keeps every user, project, and tenant apart. The database enforces the split.

History you keep

A fact changes. Engraphy keeps the old version and links it, so you can trace the change.

Runs on your database

Run Engraphy on Postgres with pgvector. Use your own machine, or your own cloud.

Why it's different

Most agent memory is a pile of text with a search box.

Engraphy works like a database, not a scratchpad. Four things make it different.

Non-destructive dedup

Merge, but never delete

A restatement merges into the memory it repeats. A related but different fact stays its own memory, linked to the first. When a fact changes, the old one is retired and kept, not overwritten.

Universal search

Every fact is findable

Engraphy stores a detail in a typed field, such as a role, a date, or a place. Search still reaches it.

Isolation in the database

Tenants cannot leak

Postgres row-level security separates every space. An agent cannot read a memory from another space.

Supersede, not overwrite

History you can walk

Facts change over time. Engraphy records the new version and keeps the old one linked. You can always ask what was true, and when.

Benchmark

Leads on temporal reasoning and open-domain knowledge. Matches the leading memory systems on LoCoMo.

Engraphy leads on temporal reasoning and open-domain knowledge under the strict judge, both by wide margins, and matches the leading memory systems overall at 66.8% excluding the adversarial category. Temporal reasoning: 72.9%, against the next-best published figure of 58.13%. Open-domain knowledge: 69.2%, against the next-best published figure of 51.15%.

72.9%

Leads temporal reasoning

vs 58.13% next-best (Mem0-graph)

69.2%

Leads open-domain knowledge

vs 51.15% next-best (Mem0)

99.4%

of the facts the extractor proposes get stored

346 of 348 nodes.

77.1%

of answerable questions had the answer in the retrieved context

300 of 389, counting the 17 with no lexical handle on the gold answer as misses.

LoCoMo accuracy with the adversarial category excluded, the denominator every published figure here uses. Engraphy matches Mem0 and clears Zep on this number; Mem0-graph leads it by 1.6 points, and Engraphy leads Mem0-graph outright on temporal reasoning and open-domain knowledge.

By question type, against Mem0, Mem0-graph and Zep.

Question typeEngraphyMem0Mem0-graphZep
Excluding adversarial66.8%66.88%68.44%65.99%
Single-hop69.0%72.93%75.71%76.60%
Multi-hop53.8%67.13%65.71%61.70%
Temporal72.9%55.51%58.13%49.31%
Open-domain69.2%51.15%47.19%41.35%
Adversarial89.2%not reportednot reportednot reported

Engraphy: run locomo-definitive-20260917, LoCoMo's standard benchmark conversations, 500 questions. Competitor figures from Chhikara et al., Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory, arXiv:2504.19413v1, 28 April 2025. That paper's Table 1 column headers do not follow LoCoMo's category order; the single-hop, multi-hop and open-domain figures above are realigned to the one mapping that reproduces the paper's own published overall for all five systems it reports. Full methodology, configuration and confidence intervals below.

Under the grading conventions used to report Mem0's own published score, Engraphy scores 86.1% excluding adversarial, ahead of Mem0's 66.88%. See how this is measured.

Methodology, configuration and limits

Configuration

66.8% is measured at a search width of 20 results, requestable on any Engraphy deployment through the search tool's own limit parameter. The engine's default width is 10, which measures 65.0% on the same questions; raising it is a one-line change on the caller's side, not a different build.

Confidence intervals

Every figure above is a point estimate from 389 non-adversarial questions (500 including adversarial). The 95% interval on each: excluding adversarial 62 to 71, single-hop 62 to 75, multi-hop 43 to 64, temporal 63 to 81, open-domain 50 to 84, adversarial 82 to 94. All three published competitor figures on the aggregate row fall inside Engraphy's interval; Engraphy's own interval on temporal reasoning clears the best published figure outright.

What is being compared

Both sides exclude the adversarial category. The Mem0 paper drops it for want of ground truth, and our harness scores it separately by an abstention rule rather than by the judge. Including it would raise Engraphy to 71.8% across all five categories. That figure is not comparable to anything on the chart, so 66.8% is the one quoted.

A second figure, under the reference rules

The 86.1% figure above is the same run, read again under the reader prompt and judge rules a reference benchmark harness uses rather than Engraphy's own: the reader must commit to an answer instead of declining, and the judge accepts one gold item from a list rather than requiring all of them. Range 76% to 89%, depending on how a disputed edge case in those rules is resolved.

Why three conversations

Ingest cannot be subsetted, because a question is unanswerable until its whole conversation is stored. Three whole conversations buys a real cost reduction, and the price is the wider intervals above than a ten-conversation run would carry.

One run, not ten

The published figures are the mean of ten runs. This is one run of one configuration; a second run of the same configuration answered 19.6% of questions differently. The intervals above are the honest width of a single measurement, not a claim that a repeat run lands on the same point.

The reader model, and the Zep column

The published figures use a GPT-4o-mini reader. Engraphy's reader here is Claude Opus 4.8, a stronger model than what the published systems were measured with, so part of any gap in either direction is the reader rather than the engine. Separately, the Zep column above is not Zep's own measurement: it is Mem0's measurement of Zep, which Zep disputes.

Single-hop and multi-hop

Mem0, Mem0-graph and Zep all lead here under the strict judge: Zep leads single-hop by 7.6 points at the widest, and Mem0 leads multi-hop by 13.3 points at the widest. Temporal reasoning and open-domain knowledge, tracking when something changed and pulling in world knowledge the conversation itself never stated, are where Engraphy leads instead.

The competitor column realignment

Table 1 of the Mem0 paper does not label its per-category columns in LoCoMo's own category order. Checked by trying every plausible assignment of the paper's columns to LoCoMo's four non-adversarial categories and computing each system's overall from its own per-category figures: only one assignment, out of 24 possible, reproduces the paper's published overall for all five systems it reports, Mem0, Mem0-graph and Zep included. That is the mapping used throughout this page.

Judge stability

Grading the same 40 answers a second time changed the verdict once. The judge and the engine are the same vendor, which is a known weakness of this measurement and is recorded here.

Contradictions merge, and the merge says so

A negation sits close in meaning to the fact it overturns, so it scores above the merge threshold and is absorbed into it. Measured at 28 of 36 pairs in our own contradiction set. The merge result names what it merged into, so the caller can repair it with supersede. The engine will not guess, because agreement is a judgment call and no model runs inside the engine.

A near-verbatim restatement can lose its text

If an incoming body overlaps the stored one at a Jaccard of 0.8 or above, it is not added as an addendum. You get addendum_added: false. Nothing is deleted, and not everything written is kept as new text.

Ranking, and how far time travel goes

The cross-encoder slot is specified and stays disabled until there is evidence it earns its cost, so ranking today is reciprocal rank fusion. Facts are retired and linked rather than overwritten, so history is walkable; full bi-temporal querying is not built.

Dataset licence

LoCoMo is published under CC BY-NC 4.0 by Maharana et al. Engraphy is not sold, so this use is non-commercial.

Accuracy measures one axis. The difference worth reading is what a memory store can do at all, next to a plain append-and-search store.

Engraphy against a typical append-and-search memory

CapabilityAppend and searchEngraphy
Duplicate handlingOn read, if at allOn write, before the row exists
The ambiguous caseGuessed, silentlyHanded back to the caller
SchemaApplication-side, if at allEnforced in Postgres
Tenant isolationA WHERE clauseRow-level security
Superseded factsOverwritten or duplicatedRetired, linked, still walkable
Worst-case readThe whole store25 capped nodes
ContradictionsAbsorbed silentlyAbsorbed, and the merge names what it hit

Against named systems rather than a shape: Engraphy compared to Mem0, Zep, Letta, Cognee and the reference MCP memory server. Each page says where the other one wins.

How it works, in two paths.

Writes check for duplicates before they land. Reads fuse two searches into one ranked result.

The write pathCheck before it lands

Writetitle, body, fields
→
Embedand index
→
Decidemerge, hold, insert

Engraphy compares each write against what it already holds. A clear duplicate merges. A clear new fact inserts. The uncertain middle is parked for a decision rather than guessed at.

The read pathTwo searches, one result

Meaningvector
+
Keywordsfull-text
→
Fuse and rankgraph on top

Semantic search and keyword search run together. Engraphy fuses them into one ranked list and removes duplicates. You can follow the graph out from any hit for more context.

Running in a few minutes.

Self-hosting runs on Docker. One command brings up Postgres, migrates the schema, and starts the server.

1
Get the code and set two secrets
git clone https://github.com/devon-clarkk/engraphy.git && cd engraphy

# .env is git-ignored; these two passwords are all it needs
cp deploy/.env.example .env   # then edit it, or generate them:
printf 'POSTGRES_PASSWORD=%s\nENGRAPHY_APP_ROLE_PASSWORD=%s\n' \
  "$(openssl rand -hex 16)" "$(openssl rand -hex 16)" > .env
2
Bring the whole stack up
docker compose up -d

# Postgres + pgvector, then migrations, the app role, and the server.
→ MCP on http://127.0.0.1:8000/mcp/ (first run builds the images; the embedding model ships inside them)
3
Create a space, apply a pack, mint a token
docker compose --profile admin run --rm admin \
  engraphy-admin space create --id personal --display-name "My Memory" --principal me

docker compose --profile admin run --rm admin \
  engraphy-admin pack apply packs/starter/pack.yaml --space personal

docker compose --profile admin run --rm admin \
  engraphy-admin token create --space personal --principal me \
    --client-name my-editor --role readwrite
→ the bearer token prints once; paste it into your MCP client

See the full walkthrough in the developer docs: setup, building a pack, the tool reference, and deployment.