Agent Memory in 2026: Vendors Claim 92.5 on LoCoMo, a Plain Filesystem Scored 74, and Every Memory Product Lost to BM25
Every microVM I run forgets everything when it is torn down, and that is a feature I sell. Snapshot restore gives an agent back its process state in 179 milliseconds, but process state is not memory: it knows where it was in a loop, not that this user hates Terraform modules or that the staging database moved last Tuesday. The question I now get most from teams putting agents on PandaStack (my Firecracker cloud, so weigh my sandbox opinions accordingly) is not "how do I isolate this" but "where does the agent keep what it learned". So I spent the week on the memory layer: the vendors, the benchmarks they quote and the ones they do not, what the labs ship first-party, and what it costs against re-reading the transcript at cached prices. The products are real, the numbers are mostly not comparable, and the unsolved problem is not recall but trust.
Four tiers and a twenty-year-old idea
The architecture every vendor converges on is the one the MemGPT paper laid out in October 2023: treat the context window as RAM and page facts in and out of slower storage the way an operating system does. The 2026 vocabulary is four tiers. Working context is the prompt now. Episodic memory is what happened, with a timestamp. Semantic memory is what is true, extracted from episodes and liable to be superseded. Procedural memory is how to do things, in practice markdown skills or rewritten system prompts. Vendors differ in what they store (facts, triples, raw sessions), how they retrieve (vectors, BM25, graph traversal, or all three), and who writes (an extraction LLM per message, a background job, or the agent through tools).
Anthropic's September 2025 context engineering guidance frames the problem from the other side: compaction, note-taking outside the window, and sub-agents returning 1,000 to 2,000 token summaries. Letta's Sarah Wooders argued in April 2026 that memory isn't a plugin, because the harness already decides what loads, what compacts and what the filesystem exposes. The benchmark fight below is largely about which of them is right.
The benchmark everyone quotes, and the ones they do not
LoCoMo, from Snap and UNC researchers in February 2024, is ten long multi-session dialogues with question-answer pairs. Mem0's April 2025 paper used it to claim a 26 percent relative improvement over OpenAI's memory, 91 percent lower p95 latency and over 90 percent token savings versus full context. Inside, with gpt-4o-mini answering, Mem0 scored 66.88 and its graph variant 68.44, Zep 65.99, LangMem 58.10, OpenAI memory 52.90 and A-Mem 48.38, while the whole conversation in context scored 72.90. Mem0's own paper says full context beats every memory system on LoCoMo; the pitch is that it does so at 26,031 tokens and 17.1 seconds p95 instead of 1,764 tokens and 1.44 seconds.
Zep's rebuttal of 6 May 2025 found three errors in how Mem0 had run Zep: both speakers ingested as the user, timestamps appended to message text instead of the created_at field, and sequential rather than parallel searches. Rerun, Zep reported 75.14 at 0.632 seconds p95, above the full-context baseline. The same post is the best public critique of LoCoMo: category 5 has no ground truth and is unusable, some questions ask about content absent from the image descriptions, actions are attributed to the wrong speaker, and conversations average 16,000 to 26,000 tokens, inside any 2026 context window. Then in August 2025 Letta scored 74.0 with gpt-4o-mini and no memory product at all, just filesystem tools, and argued that the benchmark measures how a harness manages context, not the retrieval mechanism.
A year later the numbers have run away from the baseline. Mem0's research page and README report 92.5 on LoCoMo and 94.4 on LongMemEval for an April 2026 algorithm (single-pass extraction, entity linking, BM25 alongside vectors, temporal reasoning), up from 71.4 and 67.8, at about 6,956 tokens per query, with a caveat I respect: the scores reflect the managed platform, and open-source users should expect "directionally similar gains but not identical numbers". Vectorize's Hindsight paper from December 2025 reports 89.61 against a prior open best of 75.78. MemOS's README says 88.83. None used the same answer model, judge or question subset as the 2025 rows, and only Zep's and Letta's numbers were produced to check someone else's. The spread is the finding.
LongMemEval, published at ICLR 2025, is the benchmark I would read instead: 500 questions across five abilities (information extraction, multi-session reasoning, temporal reasoning, knowledge updates, and abstention), embedded in histories of about 115k tokens for the S variant and roughly 500 sessions for M. Its headline finding is that commercial assistants and long-context models drop about 30 percent in accuracy over sustained interaction. Zep's paper used it honestly: with gpt-4o, full context scored 60.2 and Zep 71.2, an 18.5 percent relative gain, latency down from 28.9 to 2.58 seconds and the prompt from about 115k tokens to 1.6k. The same table shows where graph memory loses: single-session-assistant questions fell from 94.6 to 80.4 percent, because a retriever can miss what the assistant itself said. Hindsight's 83.6 percent with a 20B open-weight model, against 39 percent for the same model given full context, is the year's most interesting result because it is a small-model story.
Then there is MemoryAgentBench from UC San Diego, revised June 2026, which feeds information incrementally and tests four competencies. Its tables are unkind to every product. On accurate retrieval Mem0 scored 25, 32 and 36 percent across three subsets, Zep 44, 25 and 38, MemGPT 41, 38 and 32, Cognee 31, 26 and 29; a plain BM25 baseline averaged 60.5. Long-range understanding landed between 16 and 22 percent for everyone, and selective forgetting never beat 28. Mem0's answer in September 2026 was its own DolphinBench: 600 tool-based tasks over three personas with up to 5,128 messages each, graded on whether the agent called the right tool with the right content. The right direction, and the wrong party to build it.
The systems
| System | Architecture | Licence and status | Scores it reports | Source |
|---|---|---|---|---|
| Mem0 | LLM extracts facts per message; vector plus optional graph; hybrid retrieval since April 2026 | Apache-2.0 SDK v2.2.1 (25 Sep 2026); managed platform; $24M from YC, Peak XV, Basis Set, Oct 2025 | LoCoMo 92.5, LongMemEval 94.4 (platform); 66.9 in own 2025 paper | paper, research, TechCrunch |
| Letta | MemGPT lineage; agent rewrites its own memory blocks; archival search; Letta Code tracks context in git (MemFS) and syncs it to a "context repository" | Apache-2.0; letta 0.34.4 (4 Oct 2026); $10M seed, Felicis, Sep 2024 | LoCoMo 74.0 with filesystem tools; top model-agnostic open harness on Terminal-Bench | letta-code, blog, TechCrunch |
| Zep / Graphiti | Bi-temporal knowledge graph with fact invalidation; Neo4j 5.26+, FalkorDB, Neptune | Graphiti Apache-2.0; Zep cloud free 10k credits, Flex $125/mo, no community server | LongMemEval 71.2 vs 60.2 full context (gpt-4o); LoCoMo 75.1 | paper, repo, pricing |
| LangMem | Hot-path memory tools plus background extractor over LangGraph's store | MIT; 0.0.30 (27 Oct 2025) | 58.1 on LoCoMo as run by Mem0; none self-reported | PyPI, docs |
| Cognee | Knowledge graph plus vectors, local extraction models optional | Apache-2.0; v1.6.1 (24 Sep 2026) | BEAM 0.79 at 100k tokens, 0.67 at 10M | repo |
| Supermemory | "learner-1" extraction model, vector-graph store, MCP and plugins | Managed; $3M raised Oct 2025 | SWE-ContextBench 55.95% fail-to-pass, 30.3% resolved (Feb 2026); claims first on LongMemEval and LoCoMo | site, blog |
| Hindsight (Vectorize) | Four networks (world facts, experiences, observations, mental models); retain, recall, reflect; Postgres backend | MIT | LongMemEval 83.6 with a 20B model, 91.4 scaled; LoCoMo 89.6 | paper, repo |
| MemOS (MemTensor) | "MemCube" units with provenance and versioning across plaintext, activation and parameter memory | Apache-2.0; 2.0 "Stardust" | LoCoMo 88.83 | paper, repo |
Two things stand out. Zep is the only one without a self-hostable server, which matters if your erasure obligations are real. And the two that most changed in 2026 are Letta, whose development moved to Letta Code, a coding harness whose memory is a git repository, and Mem0, which added memory decay in May and background consolidation ("Dream") in September: the ideas the 2023 MemoryBank paper took from Ebbinghaus and the 2025 sleep-time compute paper from Letta's founders quantified (about five times less test-time compute at equal accuracy, 2.5 times lower cost per query amortised).
What the labs ship first-party
The consumer products made memory a default, which changes the threat model more than the architecture. ChatGPT has saved memories and, since 10 April 2025, "reference chat history", a continuously updated summary of all past conversations injected into every new chat for Plus and Pro; saved memories reached free users on 3 June 2025. Claude's memory launched for Team and Enterprise on 11 September 2025 and for Pro and Max on 23 October 2025, scoped per project, with an editable memory summary, an incognito mode that writes nothing, and import and export. Gemini's "personal context" went live on 13 August 2025, on by default, alongside temporary chats retained for 72 hours.
For developers, the memory is files. Claude Code's auto memory writes four kinds of notes (user, feedback, project, reference) into ~/.claude/projects/<project>/memory/, with a MEMORY.md index of which only the first 200 lines or 25 KB load at session start, and it compacts at 200k tokens or about 967k on native 1M models. The API memory tool (memory_20250818) is client-side: Claude asks for file operations under /memories and your code executes them, hence the path-traversal warning in the docs. Anthropic's evaluation for the September 2025 launch put memory plus context editing at a 39 percent improvement on agentic search, editing alone at 29 percent, and an 84 percent token reduction on a 100-turn task. Context editing clears old tool results server-side (default trigger 100k input tokens, keep three); compaction (beta compact-2026-09-04) replaces older turns with a summary Claude writes. OpenAI's Responses API has context_management with a compact_threshold, a /responses/compact endpoint and an opaque encrypted compaction item; state otherwise rides on previous_response_id, retained 30 days unless store: false.
Neither lab ships a hosted semantic memory store: compaction and a file protocol, with extraction, supersession and deletion left to you or the vendors. Given what follows on security, I think that is deliberate.
Memory versus re-reading at cached prices
The economic case for a memory layer used to be that context was small and expensive. Neither holds the same way now. Claude 4.6 and later bill a 1M-token request at the same per-token rate as a 9k one, and cache reads cost a tenth of base input, a twentieth on Opus 5.5. OpenAI's cached input on gpt-5.2 is $0.175 per million against $1.75. Take LongMemEval's 115k-token history. On Sonnet 5.5 at $2 per million, re-reading it costs $0.23 a turn cold and $0.023 from cache. Zep's 1.6k retrieved tokens cost $0.003 and Mem0's roughly 6,956 about $0.014, before counting the extraction calls that built the store. Inside a live session, caching wins; memory earns its keep when the cache is cold (five minutes by default, an hour at 2x write cost), when history exceeds the window, or when the same user returns next week.
The quality argument is stronger. RULER found in 2024 that only half of 17 models claiming 32k or more held up at 32k. On MRCR v2, Google's model page lists Gemini 3.1 Pro at 84.9 percent at 128k and 26.3 percent at 1M, with Claude Sonnet 4.6 also at 84.9 and Opus 4.6 at 84.0 at 128k, and no competitor scored at 1M. A million tokens of window is not a million tokens of recall; that is the "context rot" Anthropic's guidance describes and the reason compaction exists. I could not verify current Fiction.LiveBench figures (the page renders client-side and would not fetch), so they are left out.
A memory layer you can audit
Most of what the vendors sell is a few hundred lines around a database. Here is the version I would run on Postgres with pgvector 0.8.7, keeping two properties most products hide: a timestamp for when a fact was true separate from when it was recorded, and supersession instead of in-place update, so you can answer "what did the agent believe on 3 March" and delete by provenance.
-- memory.sql: pgvector 0.8.7. Bi-temporal rows; never UPDATE content, supersede it.
CREATE EXTENSION IF NOT EXISTS vector;
CREATE TABLE memory (
id bigserial PRIMARY KEY,
user_id text NOT NULL,
kind text NOT NULL CHECK (kind IN ('semantic','episodic','procedural')),
content text NOT NULL,
embedding vector(1536) NOT NULL, -- text-embedding-3-small
observed_at timestamptz NOT NULL, -- when it was true or said
recorded_at timestamptz NOT NULL DEFAULT now(), -- when we wrote it
superseded_by bigint REFERENCES memory(id), -- NULL means current
source_msg text -- provenance: session/message id, for erasure
);
CREATE INDEX ON memory USING hnsw (embedding vector_cosine_ops);
CREATE INDEX ON memory (user_id) WHERE superseded_by IS NULL;
# memory.py: psycopg 3.3.6, pgvector 0.5.0, openai SDK. Extract -> store -> retrieve.
import json, psycopg
from pgvector.psycopg import register_vector
from openai import OpenAI
ai = OpenAI()
db = psycopg.connect("dbname=agent", autocommit=True); register_vector(db)
def embed(text): # 1536 dims; swap for your own embedder
return ai.embeddings.create(model="text-embedding-3-small", input=text).data[0].embedding
EXTRACT = ("From the message, list durable facts about the user or their project as JSON: "
'[{"kind":"semantic|episodic|procedural","content":"...","supersedes":"<old fact text or null>"}]. '
"Return [] if nothing is worth remembering. Never store secrets or credentials.")
def remember(user_id, msg, msg_id, observed_at):
r = ai.chat.completions.create(model="gpt-5-mini",
messages=[{"role": "system", "content": EXTRACT}, {"role": "user", "content": msg}],
response_format={"type": "json_object"})
facts = json.loads(r.choices[0].message.content).get("facts", [])
with db.cursor() as cur:
for f in facts:
cur.execute("INSERT INTO memory (user_id, kind, content, embedding, observed_at, source_msg) "
"VALUES (%s,%s,%s,%s,%s,%s) RETURNING id",
(user_id, f["kind"], f["content"], embed(f["content"]), observed_at, msg_id))
new_id = cur.fetchone()[0]
if f.get("supersedes"): # retire the nearest current fact instead of overwriting it
cur.execute("UPDATE memory SET superseded_by=%s WHERE id = (SELECT id FROM memory "
"WHERE user_id=%s AND superseded_by IS NULL AND id<>%s "
"ORDER BY embedding <=> %s LIMIT 1)",
(new_id, user_id, new_id, embed(f["supersedes"])))
def recall(user_id, query, k=8, half_life_days=30):
# cosine similarity decayed by age of the observation, current facts only
with db.cursor() as cur:
cur.execute("SELECT content, observed_at FROM memory WHERE user_id=%s AND superseded_by IS NULL "
"ORDER BY (1 - (embedding <=> %s)) * "
"exp(-extract(epoch FROM now()-observed_at)/86400.0/%s) DESC LIMIT %s",
(user_id, embed(query), half_life_days, k))
return cur.fetchall()
def forget(source_msg): # GDPR Art. 17: erase by provenance, not by similarity
with db.cursor() as cur:
cur.execute("DELETE FROM memory WHERE source_msg=%s", (source_msg,))
I compiled this in the scratchpad but did not run it against a database, so it is a sketch with the right shape, not tested code. It is also the design MemoryAgentBench says loses to BM25, which is why I would add a tsvector column and union the two result sets before adding a graph.
Poisoning, exfiltration, erasure
Memory turns a one-shot prompt injection into a persistent one, and the demonstrations are old enough that the missing fix is the story. Johann Rehberger showed on 22 May 2024 that a connected document, an uploaded image or a browsed page could write arbitrary memories into ChatGPT, and that OpenAI closed the report as a "Model Safety Issue". On 20 September 2024 he chained it into SpAIware: an injected memory that made every future conversation render an image whose URL carried the chat to his server. OpenAI fixed the exfiltration channel in the macOS app (1.2024.247); the injection remained. The MINJA paper, revised February 2026, generalises this to any agent with a memory bank: an attacker with query-only access plants records that later victims' queries retrieve, with bridging steps that make the malicious reasoning look like the agent's own.
The engineering consequences are unglamorous. Memory writes need the provenance and review of a config change, which is why Anthropic's tool is client-side and its docs lead with path traversal, size caps and expiry, and why Hindsight ships a secrets scanner. Retrieval must be scoped per tenant at the database, not in the prompt. And erasure has to be real: GDPR Article 17 requires deletion "without undue delay" once consent is withdrawn, which a vector store cannot do by similarity and a graph with inferred edges cannot do cleanly at all. Deleting by source_msg is the minimum; what the fact was consolidated into is the hard part, and no vendor documents it well. The consumer products at least offer a switch: Claude's incognito, Gemini's 72-hour temporary chats, ChatGPT's archive and off switch. For agents that act, I treat memory like the 2 AM problem: a memory store is a persistent prompt the attacker gets to edit.
Where the research is going
The research splits into two camps. One puts memory inside the model: Google's Titans (December 2024) adds a neural long-term memory that learns at test time and claims recall past 2M tokens; its successor Hope, built on Nested Learning at NeurIPS 2025, treats memory as modules updating at different frequencies and beats Titans on language modelling and needle-in-haystack tests; Mem-alpha (September 2025) trains the memory-construction policy with reinforcement learning on 30k-token episodes and generalises past 400k. The other keeps memory outside and makes it smarter: A-MEM (NeurIPS 2025) links notes Zettelkasten-style and lets new memories revise old ones, MemOS versions memory like a filesystem, and the sleep-time and "Dream" work moves consolidation off the request path. The in-model camp will not reach you until a frontier lab ships it; what you can use today is the vendor table.
What I take from the week. Memory is a database problem with an LLM on the write path, and the discipline that matters is the one platform teams already have: provenance on every write, bi-temporal rows, deletion by source, tenant isolation at the store, and an eval that is not the vendor's. The benchmark numbers compare harnesses, judges and ingestion settings dressed up as products; the only rows I would act on came from a competitor or a university. This quarter I will add a memory tier to the PandaStack agent templates as a Postgres schema customers own, not a hosted store I run, because erasure and poisoning belong next to the data and not in my blast radius. I will benchmark it against BM25 and a cached full-context baseline before anyone's LoCoMo score, and treat every row the agent writes as untrusted input, because two years of demos say it is.
Related: Agent Observability in 2026, MCP Security in 2026 and It's 2 AM. Do You Know What Your AI Agent Is Doing?.
I'm Ajay Kumar — I build and operate PandaStack, an open-source Firecracker microVM cloud for AI agents. Everything above comes from running it in production.
Need this kind of infrastructure work? See what I do or email hello@ajayk.sh.
Related
Browser Agents in 2026: 90.2% on OSWorld-Verified, 20.6% on OSWorld 2.0, and Three Agent Browsers Already Shut Down
Computer-use agents passed the 72.36 percent human baseline on OSWorld-Verified this year (90.2 percent for the top framework, 86.0 for Claude Fable 5), then scored 20.6 percent on the 108-task OSWorld 2.0 that replaced it; Operator, Project Mariner and ChatGPT Atlas were all discontinued within 19 months; prompt-injection success in Anthropic's own tests fell from 23.6 to 11.2 percent and no vendor claims zero; Cloudflare sells HTTP 402 access to agents while 52 percent of crawling is for training; and hosting settled on one headless Chrome per microVM at $0.02 to $0.12 a browser-hour.
13 minOct 5, 2026Infrastructure as Code in 2026: Two Forks at 1.16 and 1.13, a $6.4B Owner, 912 Public State Files, and One Agent That Ran terraform destroy
Infrastructure as code three years after the BSL relicence: Terraform 1.16.5 under IBM versus OpenTofu 1.13.1 under the Linux Foundation and who shipped what first, HCP Terraform at $0.10 to $0.99 per resource with the legacy free plan gone, CDKTF and System Initiative archived, Pulumi 3.267 and Crossplane 2.4, a Terraform MCP server that grew from registry lookups to workspace administration in 15 months, 44 percent running AI for infrastructure but 34 percent trusting it, 912 exposed state files with 41 live AWS keys, and the agent that ran terraform destroy on 2.5 years of production.
13 minSep 3, 2026Top AI Agent Sandbox Providers: An Engineer's 2026 Roundup
A working engineer's roundup of the top AI agent sandbox providers: isolation, hosting, persistence and pricing compared — including the one I built myself.
10 min