Skip to content

LongMemEval results

LongMemEval evaluates long-term memory in chat assistants using multi-session conversation histories where answers depend on facts distributed across past sessions. Jennah is evaluated through its public memory API in two ingestion modes:

  • Per-turn: every conversation turn is committed as an individual chunk with memory:commit. The caller controls chunk boundaries and storage directly.
  • Formation: each session is sent to memory:form, which extracts salient facts and relationships, reconciles updates against existing workspace memory, and constructs knowledge graph edges alongside vector embeddings.

Both modes evaluate questions using a single semantic memory:query retrieving the top 30 results. In formation mode, the answer context also includes the original conversation turns that the retrieved facts were formed from, read back through each fact's provenance: whole turns for facts the assistant supplied, and the user's own turns, clipped, for facts the user stated.

Results

Per-turn Formation
Answer accuracy 93.3% 84.6%
Memories stored 44,842 19,687 (2.3x fewer)

Retrieval coverage for per-turn ingest is 96.7%: for 87 of 90 questions, every session holding part of the answer had at least one chunk in the top 30.

By question type

Question type Per-turn Formation
Temporal reasoning 100% 100%
Knowledge update 98.3% 98.5%
Single-session (user) 86.7% 93.3%
Single-session (assistant) 100% 80.0%
Single-session (preference) 84.4% 74.1%
Multi-session 90.6% 61.5%

Trade-offs and analysis

Storage volume and query latency. Formation reduces stored chunk volume by 2.3x relative to per-turn ingestion. Because exact cosine vector search scales with chunk count, smaller indexes yield lower query latency. Formation also generates structured entities and relationships, maintaining fact revision history when subsequent turns contradict earlier statements.

High-accuracy categories. Formation matches or exceeds per-turn accuracy on temporal reasoning (100%), knowledge updates (98.5% vs. 98.3%), and user-stated facts (93.3% vs. 86.7%), where reconciling new statements against prior facts is required.

Information loss and granular recall. Accuracy is lower on multi-session queries (61.5% vs. 90.6%) and assistant-generated content (80.0% vs. 100%), such as structured tables or lists. Formation extracts semantic propositions rather than preserving raw turn transcripts, and search runs over those propositions rather than over the raw turns. A fine-grained conversational detail is therefore reachable only through a retrieved fact it belongs to (the formation figures above already include those facts' source turns). Workspaces requiring full verbatim recall can commit raw turns directly using memory:commit, or combine both approaches within the same workspace.

Method

  • Dataset: LongMemEval (approximately 115k tokens of conversation history per question).
  • Sample: 90 of the dataset's 500 questions, using a stratified sample of 15 questions per question type.
  • Retrieval: one semantic query per question, top 30 results, no reranking. For formation, the source turns of the retrieved facts are added to the answer context as described above.
  • Scoring: an LLM generates candidate answers from retrieved context, and an independent LLM judge evaluates each against the reference answer. Each scoring run generates and judges every question three times. Reported accuracy is the mean of four scoring runs for per-turn and three for formation, to reduce scoring variance.
  • Retrieval coverage: percentage of questions where every session containing answer evidence is represented in the top 30 retrieval results.