📌 Highlights
Mem++ reframes organizational memory as an evidence-retention problem. Documents remain whole and historical versions remain available, while the system postpones the choice among conflicting records until a question supplies temporal and semantic context.
The strongest reported result is on OrgMemBench with gpt-4.1-mini: Mem++ scores 57.6 overall, compared with 55.0 for RAG and 44.5 for A-Mem. Its largest ablation signal is semantic retrieval: removing the vector leg reduces the claude-sonnet-4-6 score from 57.0 to 21.3 on OrgMemBench and from 87.9 to 11.9 on LongMemEvalS.
- Stores full documents, dates, authors, embeddings, and supersession state without deleting earlier versions.
- Uses temporal filtering before ranking, multi-index fusion, and a three-slot latest-date reserve.
- Leads average LLM-judge performance on LoCoMo with both answerers.
- Organizational evaluation is promising but limited to one synthetic organization and 73 questions.
🎯 Introduction
The paper distinguishes conversational memory from organizational memory. In a single dialogue, later statements often update the same narrator's previous statements. In an organization, independent authors write emails, tickets, meeting notes, and shared documents; a policy revision is a new artifact rather than an edit, and an earlier decision may still be needed for historical questions.
Prior systems either preserve text without explicitly organizing version relationships or compress records into facts, notes, episodes, or graph edges at ingest. The paper argues that this fixes the answerable information before the query is known: reasons, conditions, approvals, and source wording can disappear, while overwriting can make an earlier approved decision unrecoverable. Mem++ therefore aims to preserve the complete record and move version selection to read time.
Figure 1 frames the core problem as a difference in authorship and update semantics rather than merely a longer context window.
🔬 Methodology
Mem++ is modeled as a store R plus optional consolidation C and read operator Phi. Each incoming document is stored once as a full-text row with author, event date, embedding, and state. The write path makes no generative-model call, skips empty or exact duplicate active rows within the same author scope, and never deletes prior rows. Consolidation can group near-duplicates and mark older conflicts as superseded, but it is disabled in headline runs.
At query time, the system constructs an active, date-bounded candidate set when an as-of date is provided. Lexical retrieval targets exact content terms, the tag index targets author labels, and vector retrieval uses cosine similarity between query and document embeddings. Weighted reciprocal-rank fusion combines the ranked lists; ties use ingest recency, and three of the k slots are reserved for the latest-dated matches. The answerer receives original document text preceded by date and author metadata.
Memg++ appends up to five entity-relation triples to Mem++ retrieval, but those triples have no date and are not filtered by the as-of condition. This makes Memg++ a useful comparison for graph augmentation but a less temporally disciplined variant.
equation 1
Meaning: Defines Mem++ as a persistent record store, an optional consolidation operator, and a read-time evidence retrieval operator.
equation 2
Meaning: Defines the immutable document row: full text x, author a, event date t, embedding e, and state or supersession metadata sigma.
equation 3
Meaning: Specifies read-time evidence selection by fusing temporally scoped lexical, tag, and vector rankings and retaining the top k records.
equation 4
Meaning: Defines the candidate set for an as-of query: only active records whose event date is no later than theta are eligible.
Algorithm 1 in Section A gives the pseudocode.
Steps: Algorithm 1 in Section A gives the pseudocode. 3
📊 Experiments
The evaluation covers OrgMemBench, LoCoMo, and LongMemEvalS. OrgMemBench is the primary organizational benchmark: its evaluated medium tier contains 443 dated artifacts across 157 threads over 18 months, 73 questions, six capabilities, and 247 rubric-graded facets. LoCoMo contains 10 multi-session dialogues of about 24K tokens each and 1,540 questions across temporal reasoning, open domain, multi-hop, and single-hop categories. LongMemEvalS contains 500 questions distributed over roughly 48 sessions and 494 turns, with conversations exceeding 100K tokens.
The primary metric is LLM-judge score, scaled from 0 to 100 with higher values better; LoCoMo additionally reports F1 and BLEU-1. On OrgMemBench, Mem++ reaches 57.6 with gpt-4.1-mini, ahead of RAG at 55.0, and reaches 44.2 with gpt-4o-mini, behind Memg++ at 44.4 but ahead of RAG at 43.4. It leads Supersession with both answerers at 67.7 and 57.0 and leads Audit Replay with gpt-4.1-mini at 83.0.
On LoCoMo, Mem++ has the highest average LLM score for both answerers: 81.5 with gpt-4.1-mini and 77.4 with gpt-4o-mini. On LongMemEvalS, Mem++ ranks second overall at 74.7 with gpt-4.1-mini and 72.2 with gpt-4o-mini, behind Memg++ at 75.3 and 72.4. The paper reports that Mem++ trails RAG on some OrgMemBench contradiction and bi-temporal cases, and every method remains below 23 on Justification Chain.
The claude-sonnet-4-6 ablation identifies the vector leg as the central retrieval component. Removing it reduces scores by 35.7 on OrgMemBench and 76.0 on LongMemEvalS. Removing the lexical leg changes scores by at most 0.7, while consolidation and fact indexing reduce performance on the reported benchmarks. The graph variant does not produce a consistent gain and adds two language-model calls per document at write time plus one per question at read time.
tab:orgmem
Caption: Performance on OrgMemBench categorized by question type. Overall is weighted by the number of questions in each category and reported with its standard deviation across runs. Bold indicates the best performance. Underline indicates the second best performance.
| Answerer | Method | Supersession | Decision Provenance | Bi-temporal | Audit Replay | Justification Chain | Contradiction | Overall |
|---|---|---|---|---|---|---|---|---|
| gpt-4.1-mini | Full Context | 0.0 | 27.5 | 0.0 | 33.0 | 22.7 | 8.3 | 17.8±0.1 |
| gpt-4.1-mini | RAG | 66.1 | 79.7 | 66.7 | 40.9 | 19.6 | 75.4 | 55.0±0.3 |
| gpt-4.1-mini | Zep | 31.4 | 42.5 | 23.0 | 76.8 | 20.7 | 43.3 | 41.0±0.2 |
| gpt-4.1-mini | Mem0 | 40.8 | 55.0 | 50.0 | 25.1 | 14.4 | 52.8 | 36.9±0.0 |
| gpt-4.1-mini | A-Mem | 58.3 | 63.3 | 43.6 | 36.5 | 20.7 | 43.5 | 44.5±0.3 |
| gpt-4.1-mini | gbrain | 54.8 | 60.7 | 30.3 | 48.5 | 22.7 | 22.9 | 43.2±0.1 |
| gpt-4.1-mini | Mem++ | 67.7 | 65.9 | 50.0 | 83.0 | 21.5 | 47.2 | 57.6±0.1 |
| gpt-4.1-mini | Memg++ | 67.2 | 63.9 | 50.6 | 79.7 | 19.8 | 44.4 | 55.9±0.1 |
| gpt-4o-mini | Full Context | 2.5 | 21.1 | 0.0 | 35.7 | 17.0 | 13.9 | 16.8±0.1 |
| gpt-4o-mini | RAG | 56.7 | 47.8 | 30.2 | 55.1 | 10.4 | 67.2 | 43.4±0.2 |
| gpt-4o-mini | Zep | 27.5 | 38.3 | 34.9 | 74.9 | 7.4 | 29.6 | 36.2±0.3 |
| gpt-4o-mini | Mem0 | 32.3 | 47.2 | 33.3 | 32.4 | 11.8 | 62.8 | 33.8±0.1 |
| gpt-4o-mini | A-Mem | 42.8 | 60.0 | 28.6 | 21.3 | 4.4 | 41.7 | 32.6±0.3 |
| gpt-4o-mini | gbrain | 47.8 | 46.1 | 25.9 | 46.6 | 7.7 | 21.9 | 34.7±0.2 |
| gpt-4o-mini | Mem++ | 57.0 | 49.2 | 33.3 | 70.7 | 8.9 | 39.8 | 44.2±0.1 |
| gpt-4o-mini | Memg++ | 56.7 | 48.6 | 32.1 | 71.9 | 8.8 | 37.7 | 44.4±0.1 |
Why It Matters: Mem++ leads overall performance with gpt-4.1-mini and leads Supersession with both answerers, while RAG remains stronger on several contradiction and bi-temporal cases.
tab:locomo
Caption: Performance on LoCoMo categorized by type. Bold, underline indicates best and second best performance. The baseline performance comes from Nan et al. (2025).
| Answerer | Method | Average LLM ↑ | Average F1 ↑ | Average BLEU ↑ |
|---|---|---|---|---|
| gpt-4.1-mini | Full Context | 80.6 | 53.3 | 45.0 |
| gpt-4.1-mini | RAG-2048 | 74.5 | 50.9 | 41.3 |
| gpt-4.1-mini | RAG-4096 | 32.9 | 23.5 | 19.2 |
| gpt-4.1-mini | Zep | 61.6 | 36.9 | 30.9 |
| gpt-4.1-mini | Mem0 | 66.3 | 43.5 | 36.5 |
| gpt-4.1-mini | A-Mem | 61.4 | 39.4 | 33.2 |
| gpt-4.1-mini | Nemori | 79.4 | 53.4 | 45.6 |
| gpt-4.1-mini | Mem++ | 81.5 | 52.1 | 44.2 |
| gpt-4.1-mini | Memg++ | 80.4 | 51.9 | 44.3 |
| gpt-4o-mini | Full Context | 72.3 | 46.2 | 37.8 |
| gpt-4o-mini | RAG-2048 | 70.3 | 47.2 | 37.1 |
| gpt-4o-mini | RAG-4096 | 30.7 | 21.4 | 16.8 |
| gpt-4o-mini | Zep | 58.5 | 37.5 | 30.9 |
| gpt-4o-mini | Mem0 | 61.3 | 41.5 | 34.2 |
| gpt-4o-mini | A-Mem | 52.5 | 32.4 | 27.0 |
| gpt-4o-mini | Nemori | 74.4 | 49.5 | 38.5 |
| gpt-4o-mini | Mem++ | 77.4 | 53.1 | 42.0 |
| gpt-4o-mini | Memg++ | 76.7 | 52.1 | 41.0 |
Why It Matters: Mem++ obtains the highest average LLM-judge score for both answerers, although Nemori remains stronger on some F1 and BLEU comparisons.
tab:longmemeval
Caption: Performance on LongMemEvalS by question type. LLM-judge accuracy is reported. Bold indicates the best performance. Underline indicates the second best performance.
| Answerer | Question Type | Full Context | Zep | Nemori | Mem++ | Memg++ |
|---|---|---|---|---|---|---|
| gpt-4o-mini | Single-session preference | 6.7 | 20.0 | 46.7 | 46.7 | 50.0 |
| gpt-4o-mini | Single-session assistant | 89.3 | 80.4 | 83.9 | 94.6 | 96.4 |
| gpt-4o-mini | Temporal reasoning | 42.1 | 62.4 | 61.7 | 56.7 | 56.7 |
| gpt-4o-mini | Multi-session | 38.3 | 57.9 | 51.1 | 63.4 | 62.0 |
| gpt-4o-mini | Knowledge update | 78.2 | 83.3 | 61.5 | 84.3 | 85.2 |
| gpt-4o-mini | Single-session user | 78.6 | 92.9 | 88.6 | 98.4 | 98.4 |
| gpt-4o-mini | Average | 55.0 | 68.0 | 64.2 | 72.2 | 72.4 |
| gpt-4.1-mini | Single-session preference | 16.7 | 22.5 | 86.7 | 47.8 | 53.3 |
| gpt-4.1-mini | Single-session assistant | 98.2 | 83.1 | 92.9 | 94.6 | 94.6 |
| gpt-4.1-mini | Temporal reasoning | 60.2 | 64.5 | 72.2 | 68.5 | 69.0 |
| gpt-4.1-mini | Multi-session | 51.1 | 57.6 | 55.6 | 56.1 | 56.8 |
| gpt-4.1-mini | Knowledge update | 76.9 | 83.1 | 79.5 | 89.9 | 89.9 |
| gpt-4.1-mini | Single-session user | 85.7 | 96.3 | 90.0 | 100.0 | 100.0 |
| gpt-4.1-mini | Average | 65.6 | 69.4 | 74.6 | 74.7 | 75.3 |
Why It Matters: Mem++ ranks second overall on LongMemEvalS behind Memg++, and both variants substantially outperform Full Context on the long-context benchmark.
fig:conv-vs-org

Caption: Conversational vs. organizational memory. Top: a single narrator restates their own facts, and LLM extraction at ingest keeps only the surviving fact. Bottom: different authors in an organization write the same fact, and an earlier decision may still hold.
Why It Matters: It motivates the paper's central distinction: organizational records contain independently authored versions whose earlier states may remain relevant.
fig:architecture

Caption: Overview of framework. (3.2) Each document is stored whole with its author, date and sentence embedding, and it is indexed three ways. (3.3) Records valid at the as-of date are retrieved from the three indexes and fused with weighted RRF, and three of the k slots are reserved for the latest-dated matches. (3.4) The top-k records are passed with their dates and authors to a fixed answering model. graph triples carry no date.
Why It Matters: It summarizes the complete write, retrieval, temporal-scoping, fusion, and evidence-presentation pipeline.
fig:orgmembench

Caption: The six capabilities of OrgMemBench, with questions and answers examples. The tier we evaluated holds 443 artifacts across 157 threads for 18 months, with 73 questions.
Why It Matters: It shows the organizational benchmark scope and the six capability categories used to evaluate temporal and provenance-sensitive memory.
🔮 Conclusion
The paper's central conclusion is that retaining original documents and deferring version selection to read time is a strong baseline for organizational memory. Mem++ outperforms write-time extraction systems on OrgMemBench, leads average LLM-judge performance on LoCoMo, and remains competitive on LongMemEvalS.
The ablations support a narrower engineering conclusion: semantic retrieval over intact documents drives most of the benefit, while consolidation, extracted-fact indexing, and graph augmentation add complexity without a consistent accuracy improvement. The authors explicitly acknowledge that the organizational evaluation relies on one synthetic organization and 73 questions.
🛠️ Future Research Improvements
The paper's stated next step is to collect more diverse organizational records and test Mem++ beyond the current synthetic organizational setting. This should include multiple organizations, document channels, author identities, revision practices, and questions requiring historical reconstruction.
A builder-oriented extension would separately evaluate retrieval quality, version selection, and answer generation, because the current system leaves inference to the answering model. Future studies should also test stronger author metadata, explicit date validation, larger stores, missing-date behavior, and temporally constrained graph retrieval.
The reported tag pathway is weak partly because OrgMemBench supplies no author field, so a meaningful author-aware evaluation remains an open empirical question rather than a demonstrated capability.
🏭 Potential Industry Use Scenarios
Mem++ is a plausible foundation for internal decision assistants that answer what policy, project scope, approval, or operating procedure held at a specified date. Its whole-document retention and date-aware retrieval are also relevant to audit replay, incident review, and reconstructing the sequence of organizational decisions.
Other potential uses include engineering change history, compliance evidence retrieval, cross-team handoff memory, and executive briefing systems that must distinguish current decisions from superseded proposals. These are deployment hypotheses supported by the benchmark capabilities, not demonstrated production results.
For industry adoption, teams would need access controls, source provenance, retention policies, robust author identity, date normalization, and safeguards against presenting a retrieved historical record as the current policy.
💬 Critical Analysis
The strongest aspect of Mem++ is its clean alignment between storage semantics and organizational evidence: immutable full documents preserve context that fact extraction may discard, while temporal filtering makes historical questions operationally explicit. The result is especially compelling on Supersession and Audit Replay, where the architecture directly targets the benchmark structure.
The central weakness is evaluation scope. OrgMemBench is small and synthetic, has no author field on the reported tier, and contains only 73 questions. The paper therefore demonstrates a strong benchmark fit more clearly than broad real-world organizational robustness. The low Justification Chain scores across all methods also show that retaining evidence does not automatically solve multi-step reasoning.
The graph comparison is informative but not decisive: Memg++ sometimes leads LongMemEvalS, yet its triples are undated and its cost is higher. The ablation suggests that the simpler Mem++ design is a better default, but production systems should still measure latency, storage, retrieval recall, model-context limits, and answer reliability under larger and noisier organizational corpora.