Back to all posts
Guide
10 min read

AI Agent Memory Platforms Compared: Mem0, Zep, Letta and Open Source

DevToolLab Team

DevToolLab Team

August 24, 2026

Diagram of an agent memory pipeline: session events, fact extraction, vector and temporal graph storage, and a retrieved context block

An agent that runs across more than one session hits the same wall every time: the model remembers nothing between calls, so you either replay the entire history into the context window or the agent forgets what the user told it yesterday. Replaying gets expensive on every call, and the workarounds (summarize, truncate, hope) quietly lose the one fact that mattered.

That gap is the whole market for agent memory platforms. Zep publishes 94.7% accuracy on the LoCoMo benchmark at 155ms and a 5,760-token context block. Mem0 publishes 92.5 on LoCoMo and 94.4 on LongMemEval at roughly 6,900 tokens per query, against 25,000-plus for stuffing the full history. Both are self-run on different harnesses, so the interesting question is not which vendor scores highest. It is which memory model fits your data, what it costs at your volume, and what happens when you need to run it yourself.

What a Memory Layer Actually Does

Every platform here sits between your agent and the model and does the same four things: ingest episodes (messages, tool results, JSON records), extract facts and entities with an LLM call, store them as vectors or graph edges, and retrieve a compact block for the current turn.

The differences that matter show up in the last two steps. A vector-first system is cheaper to run and degrades predictably: it returns semantically similar text, including stale text. A graph-first system tracks relationships between entities and, in Zep's case, attaches a validity window to every fact so a superseded fact is marked invalid rather than overwritten. That costs more per write and pays off when your questions are relational ("who approved the last renewal") or temporal ("what did they use before Postgres").

Two things are worth knowing before you shop. A bigger context window does not remove the need for a memory layer, which is the core finding of the BEAM benchmark at 1M and 10M token scales. And Anthropic now ships a client-side memory tool (memory_20250818, on all Claude 4 and later models) where Claude reads and writes files in a /memories directory backed by your own storage. That is a free baseline, so any paid platform has to beat a directory of markdown files.

If you want the underlying architecture with code for each tier, we covered it in AI Agent Memory in 2026. This post is about picking a product.

Head-to-Head Comparison

The short version, with the detail on each one below it.

PlatformMemory modelFree tierPaid entrySelf-hostBest for
Mem0Vector, optional graph10k adds, 1k retrievals$19/mo, $249/mo for graphFull OSS core, Apache-2.0Fastest prototype to production
ZepTemporal knowledge graph10k credits/mo$125/mo, reads unmeteredGraphiti plus a graph DB, Apache-2.0Relational and temporal questions
LettaAgent-owned memory blocks$0, 3 agents, bring your own keys$20/moDocker server, Apache-2.0Stateful agents with an identity
CogneeGraph plus vector, pluggable$0 cloud, 1M tokens, or self-host$2.50 per 1M tokensFree forever, Apache-2.0A graph without a vendor
HindsightMulti-strategy retrievalLocal, bring your own LLM keyCloud, usage-basedDocker, Helm or pip, MITLong-horizon recall quality
MemoriStructured agent state$0 cloud, 5k memories, or self-hostTeam from $60,000/yearBring your own DB, Apache-2.0Regulated environments
HonchoVector plus user modelingSelf-hosted, or $100 cloud creditsCloud, usage-basedYes, AGPL-3.0Personalization, non-commercial
LangMemStore primitives for LangGraphLibrary onlyn/aYes, MITExisting LangGraph stacks

Mem0

Mem0 homepage showing drop-in memory infrastructure for AI agents
Mem0 homepage showing drop-in memory infrastructure for AI agents

Mem0 is the default starting point for most teams: 63,900-plus stars on the Apache-2.0 repository, a $24M raise in October 2025, and an SDK that drops into an existing app in about six lines. The current PyPI release is mem0ai 2.0.19.

The fast path is what it does well. You add messages, search with a user filter, and get back scored memories with categories attached. Dream, announced August 24, 2026, adds background consolidation: it runs weekly per project, marks superseded memories as merged or superseded with a pointer to their replacement, and synthesizes summaries when a pattern emerges. None of that runs on the request path, so add and search latency stay flat.

The gotcha is pricing shape. Graph memory and Dream both sit on the $249/month Pro plan, and the cheap tiers are metered generously for writes but stingily for reads: Starter includes 50,000 adds and only 5,000 retrievals per month, and a chat agent retrieves at least once per turn.

Pricing: Hobby free (10,000 adds, 1,000 retrievals/month) · Starter $19/month (50,000 adds, 5,000 retrievals) · Pro $249/month (500,000 adds, 50,000 retrievals, graph memory, Dream) · Enterprise custom

The open-source package is the same core engine running against your own vector store and LLM keys, defaulting to Qdrant plus OpenAI, so the self-hosted path is real rather than a teaser. Mem0 also publishes OpenMemory, a local-first MCP memory server shared across Claude Desktop, Cursor and other MCP clients.

Python
# pip install mem0ai && export OPENAI_API_KEY=sk-...
from mem0 import Memory

m = Memory()

messages = [
    {"role": "user", "content": "I ship to Austin, TX and prefer Postgres over MySQL."},
    {"role": "assistant", "content": "Noted: Austin shipping address, Postgres preference."},
]
m.add(messages, user_id="u_1042")

hits = m.search(
    "which database does this user prefer?",
    filters={"user_id": "u_1042"},
    top_k=5,
)
for hit in hits["results"]:
    print(round(hit["score"], 3), hit["memory"])

Note filters={"user_id": ...} on search. Older tutorials pass user_id= directly and that no longer matches the 2.x signature, which is the most common paste-and-fail in Mem0 code today.

Zep

Zep homepage showing agent memory at enterprise scale with a context graph dashboard
Zep homepage showing agent memory at enterprise scale with a context graph dashboard

Zep is the graph-native option, built on Graphiti, its Apache-2.0 temporal knowledge graph engine with more than 30,000 stars (currently graphiti-core 0.29.3). Every fact becomes an edge with a validity window, so contradicting information invalidates the old edge with its history intact.

Its published numbers are the strongest on latency at scale: 94.7% on LoCoMo at 155ms and 5,760 tokens, 90.2% on LongMemEval at 162ms and 4,408 tokens, and p95 retrieval under 200ms from small graphs up to more than 100 million nodes. It also does something the others do not: retrieval, storage, threads, users and graph storage are all unmetered. You pay for what you write, not what you read.

Model that billing carefully, because it is unusual. Zep bills credits by ingested bytes: one credit per episode up to 350 bytes, plus one more per additional 350 bytes or part thereof. A 640-byte message is 2 credits. A 1,200-byte JSON record is 4.

Pricing: Free (10,000 credits/month, 2 projects) · Flex $125/month or $1,250/year, 50,000 credits then $25 per 10,000 · Flex Plus $375/month or $3,750/year, 200,000 credits then $75 per 40,000 · Enterprise custom (SOC 2 Type II, HIPAA BAA, BYOK and BYOC)

The catch is self-hosting. Zep Community Edition was discontinued in April 2025; the repository stays Apache-2.0 but gets no updates. Self-hosting today means running Graphiti plus a compatible graph database (Neo4j, FalkorDB or Kuzu) plus your own service layer, which is a different operational commitment than docker run. Graphiti's MCP server (now at mcp-v1.0.2) is the shortcut if all you want is memory for your coding agents.

Python
# pip install zep-cloud
from zep_cloud.client import Zep
from zep_cloud.types import Message

client = Zep(api_key="z_...")

client.user.add(user_id="u_1042", email="dana@example.com", first_name="Dana")
client.thread.create(thread_id="t_88", user_id="u_1042")

client.thread.add_messages(
    "t_88",
    messages=[
        Message(
            role="user",
            name="Dana Reyes",
            content="We moved the deploy to us-east-1 last Tuesday.",
        )
    ],
)

# One string, already assembled, ready to drop into your system prompt.
context = client.thread.get_user_context(thread_id="t_88").context

The get_user_context call is worth copying even if you pick a different vendor. Zep hands back an assembled context block rather than a list of chunks you rank and format yourself, which removes a surprising amount of glue code.

Letta

Letta homepage describing an AI research lab building machines that learn
Letta homepage describing an AI research lab building machines that learn

Letta comes from the Berkeley team behind MemGPT, raised a $10M seed, and its Apache-2.0 server has more than 24,000 stars. Its model is different in a way that matters: Letta does not sell you a memory API to call from your agent, it gives you the agent. State lives on the agent in editable memory blocks the model can rewrite with its own tools, and enable_sleeptime=True lets a background agent reorganize those blocks between turns.

Be clear-eyed about the direction. Letta's own "next phase" post describes a shift to Letta Code, with legacy server-side memory tools, templates, identities and tool rules deprecated in favor of client-side filesystem access and git-backed context repositories. The new Agent SDK (August 17, 2026) is TypeScript-first, the Python letta-client package has not seen a release since June 2, 2026, and letta.com/pricing now 301s straight to the Letta Code pricing docs. If you want a stable memory-as-a-service API for something shipping this quarter, Letta is the riskiest of the three. If you want a stateful coding agent with real memory, it is the most interesting.

Pricing: Free $0 (3 stateful agents, bring your own model keys) · Personal Pro $20/month (up to 20 stateful agents) · API plan $20/month plus $0.10 per active agent per month and $0.00015 per second of tool execution · Teams Pro $20/seat/month · Enterprise custom

Python
# pip install letta-client
from letta_client import Letta

client = Letta(api_key="sk-let-...")

agent = client.agents.create(
    model="openai/gpt-4.1",
    memory_blocks=[
        {"label": "human", "value": "Name: Dana. Works on payments at a Chicago fintech."},
        {"label": "persona", "value": "Concise engineering assistant. Cites files."},
    ],
    enable_sleeptime=True,
)

reply = client.agents.messages.create(
    agent_id=agent.id,
    messages=[{"role": "user", "content": "What did we decide about the deploy region?"}],
)

Self-hosting is a Docker server with your own model provider keys, and the whole stack including Letta Code (npm i -g @letta-ai/letta-code) is Apache-2.0.

Open Source Options Worth Self-Hosting

Cognee homepage, an open-source agent memory platform
Cognee homepage, an open-source agent memory platform

Cognee is the most complete self-hosted option: Apache-2.0, more than 30,000 stars, currently 1.5.3. The 1.x API is four verbs (remember, recall, forget, improve, with the older add/cognify/search still working) and the storage layer is genuinely pluggable, with Kuzu by default for the graph (Neo4j, Postgres, Turso, Neptune supported) and LanceDB for vectors (PGVector, Qdrant, Chroma, Weaviate, Milvus available). Self-hosting is free forever; the hosted option is metered on tokens rather than seats.

Pricing: Free $0/month (1M tokens, 1 workspace, unlimited API calls) · Standard $2.50 per 1M tokens plus $5 per additional workspace · Enterprise custom (BYO cloud, support SLA)

Python
# pip install cognee
import asyncio
import cognee

async def main():
    await cognee.remember("Dana approved the Q3 checkout rewrite and prefers Postgres.")
    print(await cognee.recall("who approved the checkout rewrite?"))

asyncio.run(main())
Hindsight documentation showing memory types and multi-strategy retrieval
Hindsight documentation showing memory types and multi-strategy retrieval

Hindsight from Vectorize is the benchmark story of 2026. MIT licensed, more than 21,000 stars, and it reports 64.1% on the hardest BEAM tier at 10M tokens against 40.6% for Honcho and 24.9% for a RAG baseline. It self-hosts through Docker Compose, Helm or pip install hindsight-all with no Hindsight account, though retain and recall still call an LLM, so you supply a provider key either way. If your workload is long-horizon and retrieval quality is the thing you are optimizing, it belongs on the shortlist on measured results alone.

Pricing: Self-hosted free · Hindsight Cloud usage-based, with free credits to start and no fixed monthly or per-seat fee

Memori homepage describing agent-native memory infrastructure
Memori homepage describing agent-native memory infrastructure

Memori from MemoriLabs takes the no-rip-and-replace angle: Apache-2.0, more than 16,000 stars, turning agent execution into structured state on top of databases you already run, with managed cloud, single-tenant, VPC and on-premises deployments. It is the option when the blocker is a security review rather than a benchmark, but check the quote before you get attached: the open-source build is free on your own database, and the managed tiers start at $60,000/year.

Pricing: Open source self-hosted free (bring your own database) · Cloud Free $0/month (5,000 memories created, 15,000 recalled) · Team from $60,000/year (single production agent, multi-tenant cloud) · Business from $150,000/year (single-tenant) · Enterprise custom (customer VPC or on-premises)

Two more to know about. Honcho from Plastic Labs is well regarded but AGPL-3.0, which your legal team will want to hear about before you embed it in a commercial product. The hosted service gives each new account a dedicated instance and $100 in credits. LangMem is MIT and small, the pragmatic pick if you are already all-in on LangGraph.

What the Benchmark Numbers Mean

LoCoMo and LongMemEval are multi-session conversation benchmarks: given a long synthetic dialogue, can the system answer questions about it. These are the numbers on every pricing page, and vendor runs use different harnesses, models and retrieval budgets, so a two-point gap is not a ranking. BEAM pushes the same idea to 1M and 10M tokens, where context stuffing stops being an option at all.

MemoryArena, a 2026 benchmark from researchers at Stanford, UCSD, UIUC, Princeton and Pitt, is the one that should change how you evaluate. It puts the agent in a continuous loop across web navigation, constrained planning and formal reasoning, where later steps depend on what it learned earlier, and finds that systems near saturation on LoCoMo perform poorly there. Recall and behavior are different capabilities.

The practical version: run a small eval on your own transcripts before you commit. Fifty real conversations with twenty questions you care about will tell you more than any published score, especially about the failure mode that hurts most in production, which is confidently returning a fact the user already corrected.

Cost Math on a Real Workload

Take a support agent handling 2,000 conversations a month, twelve messages each, so 24,000 messages at roughly 600 bytes apiece, with one context retrieval per turn.

On Zep, each 600-byte message is 2 credits, so ingestion is about 48,000 credits and all 24,000 retrievals are free. That fits inside Flex's 50,000 included credits at $125/month, or $104/month billed annually, and read volume never changes the bill.

On Mem0, writes are cheap and reads are the constraint. 24,000 adds sits comfortably inside Starter's 50,000, but 24,000 retrievals blows through Starter's 5,000 limit and lands you on Pro at $249/month. Batching adds per conversation drops you to 2,000 adds, but retrieval count is set by your turn count, not your batching.

The tradeoff is not "Zep is cheaper," it is that the two meter different things. Model your own read-to-write ratio before comparing sticker prices, and sanity check the token side with the AI Token Counter and the LLM Token Cost Calculator, because a layer that returns 6,000 tokens per turn instead of 25,000 usually saves more on inference than it costs in subscription.

Which One Should You Pick

Prototype or side project: Mem0's Hobby tier or Cognee locally. Both get you working memory in an afternoon with no credit card.

Shipping a SaaS or consumer chat product: Mem0 Starter at $19/month if retrieval volume is low, Pro at $249/month once you need graph memory, Dream or real read headroom. Least intrusive SDK of the three, and the OSS escape hatch is real.

Relational or temporal questions, or an audit requirement: Zep. Validity windows on facts, provenance back to the source episode, unmetered reads, and SOC 2 Type II plus a HIPAA BAA at Enterprise. Check your byte volume against the credit math first.

Data cannot leave your infrastructure: Cognee or self-hosted Memori, or Graphiti with Neo4j if you want Zep's temporal model without Zep's cloud. Memori's managed and VPC tiers start at $60,000/year, so price the open-source build first. Budget for the graph database as an operational dependency, not a footnote.

Long-horizon agents where recall quality is the product: Hindsight, on measured BEAM results, with the caveat that it is younger than the rest.

A stateful agent rather than a memory API: Letta, with eyes open about the shift toward Letta Code.

Already using Claude for everything: try the memory tool first. A /memories directory mapped onto your own storage covers a real share of "the agent should remember this" at zero incremental cost. Reach for a platform when you need cross-session search over thousands of users, relational queries, or a story for your security reviewer.

Conclusion

All three commercial platforms work. The differences that will decide your project are commercial and operational rather than technical: what your read-to-write ratio costs, whether stale facts have to be invalidated or merely outranked, and how much of the stack you can run yourself. The open-source side has closed most of the gap, so "we do not want to build this" is a weaker reason to sign than it was a year ago.

Whichever you pick, wire it behind a small interface of your own (remember(user_id, messages) and recall(user_id, query) is enough), then run fifty real conversations through it including a few where the user changes their mind. That test takes a day and will tell you more than every benchmark table in this post, including the one above.

Pricing, benchmark claims and license terms in this guide were verified against each vendor's own site on August 26, 2026. All of them move, so check the source before you commit budget.

Related Posts

Best AI Video Editors for Developers

Descript, DaVinci Resolve, Shotstack, Remotion and auto-editor compared for product demos and code-driven video, with prices, licenses and a tested auto-cut.

By DevToolLab Team•

Best HubSpot Alternatives for Developers

HubSpot, Attio, Twenty, Close and EspoCRM compared on published API limits, webhooks, licenses and per-seat prices, plus how long a 50,000-record sync takes.

By DevToolLab Team•

Best Bolt.new Alternatives in 2026

Lovable, v0, Replit, Base44, Dyad, bolt.diy compared on October 2026 prices: tokens versus credits, per-seat versus flat plans, and what a 4-person team pays.

By DevToolLab Team•