Engraphy Compare Docs ← Back to site

Engraphy › Compare › Zep

Engraphy vs Zep

Checked 24 August 2026 · Engraphy v0.1.0

The short answer

Zep is the better choice if you genuinely need bi-temporal querying: asking what the store believed at an arbitrary past instant, not just what it believes now. Graphiti gives every edge a validity window, and Engraphy does not have that. Engraphy is the better choice if you want the whole thing on one Postgres instance you already know how to run, with isolation enforced by the database and embeddings that never leave the box.

The operational difference is large. Zep today is either Zep Cloud (managed, credit-based) or standing up Graphiti yourself against Neo4j, FalkorDB or Kuzu. The Zep Community Edition was deprecated. Engraphy is docker compose up and a Postgres database.

At a glance

CapabilityEngraphyZep
DeploymentSelf-hosted only, on Postgres + pgvector. One Compose file.Zep Cloud (managed, credit-based), or run the Graphiti library yourself against a graph database.
Where your data livesYour machine or your cloud. Embeddings in-process, no egress.Zep's cloud, or your own graph database with Graphiti. LLM extraction calls a provider either way.
LicenceBusiness Source License 1.1, converting to Apache-2.0 at the Change Date.Graphiti is Apache-2.0. Zep Cloud is commercial. Zep Community Edition is deprecated.
StoragePostgres 16 + pgvector, one database you probably already operate.A temporal knowledge graph on Neo4j, FalkorDB or Kuzu, plus a vector index.
Duplicate handlingBanded on write before the row exists: merge / merge-link / pending / new.Entity resolution at ingest, driven by an LLM.
When a fact changesRetired and linked by a supersedes edge. History is walkable.Bi-temporal validity windows. Edges carry valid-from and valid-to, so you can query a past state directly.
IsolationPostgres row-level security under a NOBYPASSRLS role.User and session scoping in Zep Cloud; on raw Graphiti, whatever you build.
SchemaTyped nodes and edges with attribute schemas, enforced in Postgres.Custom entity and edge types defined in Python.
RetrievalVector plus Postgres full-text fused with reciprocal rank fusion, plus graph traversal.Hybrid semantic and BM25 search with graph traversal and rerankers.
How agents connectMCP over HTTP with a bearer token.SDK and REST API; a Graphiti MCP server exists.
ComplianceNone claimed. You are the operator, so the posture is yours.Zep Cloud carries SOC 2 Type 2 and HIPAA.

Which one to choose

Choose Zep when

  • You need real bi-temporal queries. This is Zep's sharpest advantage and Engraphy does not match it. Engraphy retires and links superseded facts, which makes history walkable, but it is not the same as querying an arbitrary past instant.
  • You want a managed service with SOC 2 Type 2 and HIPAA behind it. Engraphy has no hosted offering at all.
  • You are already running a Neo4j-class graph database and want to use it.
  • You want Apache-2.0 with no field-of-use restriction, and Graphiti as a library rather than a server.
  • You want a much larger community and a longer track record than a v0.1.0 engine has.

Choose Engraphy when

  • You would rather operate one Postgres than a graph database plus a vector index. Engraphy's whole store is Postgres 16 with pgvector.
  • You want no outbound calls on the write path. Engraphy's embedding model runs in-process; there is no API key.
  • You need isolation the database enforces rather than the application. Row-level security under a NOBYPASSRLS role, per space.
  • You want the borderline duplicate handed back to you rather than resolved by a model.
  • You want a memory shape you declare and Postgres enforces, rather than one the ingest pipeline infers.

Bi-temporal versus supersession: the difference matters

Both systems refuse to overwrite history, but they do it differently, and Zep's version is more powerful.

Graphiti gives each edge a validity window. A fact is not just true or retired; it is true between these two instants, and the graph separately records when the system learned it. That lets you reconstruct the state of the world at an arbitrary past moment.

Engraphy records the new version and keeps the old one linked by a supersedes edge. You can always ask what was true and when, by walking the chain. What you cannot do is issue a single query for “the state of this space as of March”. Full bi-temporal querying is not built. If your problem is auditing a changing world, such as contracts, prices or org charts, that gap is the reason to choose Zep.

The operational bill is very different

Since the Community Edition was deprecated, self-hosting Zep means running Graphiti directly: standing up and operating Neo4j, FalkorDB or Kuzu, handling schema migrations, and keeping a graph database healthy in production. That is a real skill set and a real cost.

Engraphy is a Postgres application. docker compose up -d brings up Postgres and pgvector, runs the migrations, provisions a non-superuser app role, and starts the MCP server. If your team already backs up and monitors a Postgres, Engraphy adds a schema to something you already run rather than a new class of database to learn.

What the benchmark says, and what it does not

On LoCoMo, Engraphy scores 66.8%, just ahead of Zep's 65.99% and just behind Mem0's 66.88%, close enough that all three read as level. Temporal reasoning and open-domain knowledge separate them clearly instead: Engraphy leads outright on temporal reasoning, 72.9% against a best published figure of 58.13% held by Mem0-graph rather than Zep, and on open-domain knowledge, 69.2% against Zep's 41.35%.

Zep leads on single-hop and multi-hop questions instead, at 76.60% and 61.70% against Engraphy's 69.0% and 53.8%. Competitor per-category figures are realigned from the Mem0 paper's Table 1, whose column headers do not follow LoCoMo's category order; see the methodology detail below the benchmark bars for how the mapping was verified, and for the confidence interval on every figure here.

What the benchmark says

Mem0-graph68.44%
Mem066.88%
Engraphy66.8%
Zep65.99%
LangMem58.1%
OpenAI memory52.9%
A-Mem48.4%

LoCoMo accuracy with the adversarial category excluded, the denominator every published figure here uses. Engraphy matches Mem0 and clears Zep on this number; Mem0-graph leads it by 1.6 points, and Engraphy leads Mem0-graph outright on temporal reasoning and open-domain knowledge.

Engraphy: run locomo-definitive-20260917, LoCoMo's standard benchmark conversations, 500 questions. Competitor figures from Chhikara et al., Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory, arXiv:2504.19413v1, 28 April 2025. That paper's Table 1 column headers do not follow LoCoMo's category order; the single-hop, multi-hop and open-domain figures above are realigned to the one mapping that reproduces the paper's own published overall for all five systems it reports. Full methodology, configuration and confidence intervals below.

Under the grading conventions used to report Mem0's own published score, Engraphy scores 86.1% excluding adversarial, ahead of Mem0's 66.88%. See how this is measured.

Methodology, configuration and limits

Configuration

66.8% is measured at a search width of 20 results, requestable on any Engraphy deployment through the search tool's own limit parameter. The engine's default width is 10, which measures 65.0% on the same questions; raising it is a one-line change on the caller's side, not a different build.

Confidence intervals

Every figure is a point estimate from 389 non-adversarial questions (500 including adversarial). The 95% interval on each: excluding adversarial 62 to 71, single-hop 62 to 75, multi-hop 43 to 64, temporal 63 to 81, open-domain 50 to 84, adversarial 82 to 94. All three published competitor figures on the aggregate fall inside Engraphy's interval; Engraphy's own interval on temporal reasoning and on open-domain knowledge each clear the best published figure outright.

What is being compared

Both sides exclude the adversarial category. The Mem0 paper drops it for want of ground truth, and our harness scores it separately by an abstention rule rather than by the judge. Including it would raise Engraphy to 71.8% across all five categories. That figure is not comparable to anything on this page, so 66.8% is the one quoted.

A second figure, under the reference rules

The 86.1% figure above is the same run, read again under the reader prompt and judge rules a reference benchmark harness uses rather than Engraphy's own: the reader must commit to an answer instead of declining, and the judge accepts one gold item from a list rather than requiring all of them. Range 76% to 89%, depending on how a disputed edge case in those rules is resolved.

Why three conversations

Ingest cannot be subsetted, because a question is unanswerable until its whole conversation is stored. Three whole conversations buys a real cost reduction, and the price is the wider intervals above than a ten-conversation run would carry.

One run, not ten

The published figures are the mean of ten runs. This is one run of one configuration; a second run of the same configuration answered 19.6% of questions differently. The intervals above are the honest width of a single measurement, not a claim that a repeat run lands on the same point.

The reader model, and the Zep column

The published figures use a GPT-4o-mini reader. Engraphy's reader here is Claude Opus 4.8, a stronger model than what the published systems were measured with, so part of any gap in either direction is the reader rather than the engine. Separately, the Zep column above is not Zep's own measurement: it is Mem0's measurement of Zep, which Zep disputes.

Single-hop and multi-hop

Mem0, Mem0-graph and Zep all lead here under the strict judge: Zep leads single-hop by 7.6 points at the widest, and Mem0 leads multi-hop by 13.3 points at the widest. Temporal reasoning and open-domain knowledge, tracking when something changed and pulling in world knowledge the conversation itself never stated, are where Engraphy leads instead.

The competitor column realignment

Table 1 of the Mem0 paper does not label its per-category columns in LoCoMo's own category order. Checked by trying every plausible assignment of the paper's columns to LoCoMo's four non-adversarial categories and computing each system's overall from its own per-category figures: only one assignment, out of 24 possible, reproduces the paper's published overall for all five systems it reports, Mem0, Mem0-graph and Zep included. That is the mapping used throughout this page.

Judge stability

Grading the same 40 answers a second time changed the verdict once. The judge and the engine are the same vendor, which is a known weakness of this measurement and is recorded here.

What Engraphy does not do

Stated plainly, so that nothing here has to be walked back:

Questions people ask

Does Engraphy support bi-temporal queries like Zep?

No. Engraphy retires superseded facts and links them with a supersedes edge, so history is walkable and you can ask what was true and when. Full bi-temporal querying, asking for the state of the store at an arbitrary past instant in one query, is not built. Zep's Graphiti does have validity windows on edges.

Can I still self-host Zep?

The Zep Community Edition was deprecated and its code moved to a legacy folder. Self-hosting today means running the Graphiti library yourself against Neo4j, FalkorDB or Kuzu, and operating that graph database in production. Engraphy self-hosts on Postgres with pgvector.

Which needs less infrastructure, Engraphy or Zep?

Engraphy. Its entire store is Postgres 16 with the pgvector extension, brought up by one docker compose command. Graphiti needs a graph database plus a vector index, and Zep Cloud needs no infrastructure at all but is a managed service you send data to.

Does either one keep my data on my own machine?

Engraphy always does. It is self-hosted only and its embedding model runs in-process, so there is no API key and no egress on the write path. Graphiti can run on your infrastructure but calls an LLM provider for extraction. Zep Cloud is a hosted service.

Is Zep or Engraphy better for compliance?

Zep Cloud carries SOC 2 Type 2 and HIPAA, which Engraphy has no equivalent of because there is no hosted Engraphy. What Engraphy gives you instead is that the data never leaves your infrastructure and tenant isolation is enforced by Postgres row-level security.

Sources

Try it against your own memory

Engraphy is a server you host. Point any MCP client at it. Embeddings run in-process, so there is no API key and no data leaves the box.

Install for VS Code Self-host quickstart Developer docs