LongMemEval results
LongMemEval evaluates long-term memory in chat assistants using multi-session conversation histories where answers depend on facts distributed across past sessions. Jennah is evaluated through its public memory API in two ingestion modes:
- Per-turn: every conversation turn is committed as an individual chunk with
memory:commit. The caller controls chunk boundaries and storage directly. - Formation: each session is sent to
memory:form, which extracts salient facts and relationships, reconciles updates against existing workspace memory, and constructs knowledge graph edges alongside vector embeddings.
Both modes evaluate questions using a single semantic memory:query retrieving the top 30 results.
In formation mode, the answer context also includes the original conversation turns that
the retrieved facts were formed from, read back through each fact's
provenance: whole turns for facts the assistant supplied, and
the user's own turns, clipped, for facts the user stated.
Results
| Per-turn | Formation | |
|---|---|---|
| Answer accuracy | 93.3% | 84.6% |
| Memories stored | 44,842 | 19,687 (2.3x fewer) |
Retrieval coverage for per-turn ingest is 96.7%: for 87 of 90 questions, every session holding part of the answer had at least one chunk in the top 30.
By question type
| Question type | Per-turn | Formation |
|---|---|---|
| Temporal reasoning | 100% | 100% |
| Knowledge update | 98.3% | 98.5% |
| Single-session (user) | 86.7% | 93.3% |
| Single-session (assistant) | 100% | 80.0% |
| Single-session (preference) | 84.4% | 74.1% |
| Multi-session | 90.6% | 61.5% |
Trade-offs and analysis
Storage volume and query latency. Formation reduces stored chunk volume by 2.3x relative to per-turn ingestion. Because exact cosine vector search scales with chunk count, smaller indexes yield lower query latency. Formation also generates structured entities and relationships, maintaining fact revision history when subsequent turns contradict earlier statements.
High-accuracy categories. Formation matches or exceeds per-turn accuracy on temporal reasoning (100%), knowledge updates (98.5% vs. 98.3%), and user-stated facts (93.3% vs. 86.7%), where reconciling new statements against prior facts is required.
Information loss and granular recall. Accuracy is lower on multi-session queries
(61.5% vs. 90.6%) and assistant-generated content (80.0% vs. 100%), such as structured
tables or lists. Formation extracts semantic propositions rather than preserving raw turn
transcripts, and search runs over those propositions rather than over the raw turns. A
fine-grained conversational detail is therefore reachable only through a retrieved fact
it belongs to (the formation figures above already include those facts' source turns).
Workspaces requiring full verbatim recall can commit raw turns directly using
memory:commit, or combine both approaches within the same workspace.
Method
- Dataset: LongMemEval (approximately 115k tokens of conversation history per question).
- Sample: 90 of the dataset's 500 questions, using a stratified sample of 15 questions per question type.
- Retrieval: one semantic query per question, top 30 results, no reranking. For formation, the source turns of the retrieved facts are added to the answer context as described above.
- Scoring: an LLM generates candidate answers from retrieved context, and an independent LLM judge evaluates each against the reference answer. Each scoring run generates and judges every question three times. Reported accuracy is the mean of four scoring runs for per-turn and three for formation, to reduce scoring variance.
- Retrieval coverage: percentage of questions where every session containing answer evidence is represented in the top 30 retrieval results.