Retrieval channels
Jennah provides two channels for semantic memory retrieval:
- Vector retrieval: Embeds queries and ranks chunks by cosine distance. Vector search identifies content with similar semantic meaning, even when exact terminology differs.
- Lexical retrieval: Matches terms and ranks by whole-word relevance. Lexical search targets exact terms that embedding models may index less distinctively, such as error codes, product SKUs, identifiers, and proper nouns.
Vector retrieval is the default mode. Hybrid retrieval combines vector and lexical channels to query both simultaneously.
The two channels
Vector retrieval
Vector retrieval embeds queryText (or uses a caller-supplied embedding) and ranks candidate chunks within the workspace by cosine distance.
Lexical retrieval
Lexical retrieval performs whole-word token matching against rawContent and ranks matches by relevance.
Lexical matching operates on whole words rather than substrings or prefixes. Searching for quota matches exact token occurrences but does not match prefix variants like quotas.
Using hybrid mode
To query both channels simultaneously, set retrievalMode to RETRIEVAL_MODE_HYBRID in the semantic section of memory:query:
{
"semantic": {
"queryText": "ERR_QUOTA_EXHAUSTED_88213",
"limit": 10,
"retrievalMode": "RETRIEVAL_MODE_HYBRID"
}
}
Omitted or RETRIEVAL_MODE_VECTOR_ONLY executes vector retrieval alone.
queryTextrequired,embeddingnot allowed: Lexical retrieval requires search text. A request that carriesembeddingis served from the vector alone and itsqueryTextis ignored, so hybrid mode with anembedding(with or withoutqueryText) returnsINVALID_ARGUMENT.- Enterprise enablement: Hybrid mode must be enabled for the enterprise before use. Requests made prior to enablement return
FAILED_PRECONDITION. See Enabling the lexical channel.
Reciprocal rank fusion
Hybrid retrieval combines results using Reciprocal Rank Fusion (RRF) rather than merging raw similarity scores:
- Each channel independently ranks candidate chunks.
- A chunk's fused score is computed as
sum(1 / (k + rank))across matching channels, wherekis a fixed ranking constant. - Ranks are deterministic and avoid score normalization discrepancies across different vector and lexical scoring distributions.
- Chunks matched by multiple channels accumulate higher scores, while chunks matched by only a single channel remain eligible for retrieval.
Reading channel provenance
Query results include channel provenance identifying which channel or channels surfaced each chunk:
{
"semantic": {
"matches": [
{
"chunkId": "runbook-quota",
"rawContent": "when a quota is exhausted, first check...",
"distance": 0.12,
"channels": [
"RETRIEVAL_CHANNEL_VECTOR",
"RETRIEVAL_CHANNEL_LEXICAL"
],
"lexicalScore": 2.10,
"rrfScore": 0.0323
},
{
"chunkId": "incident-4471",
"rawContent": "incident ERR_QUOTA_EXHAUSTED_88213 was closed...",
"distance": 0,
"channels": ["RETRIEVAL_CHANNEL_LEXICAL"],
"lexicalScore": 4.82,
"rrfScore": 0.0164
}
]
}
}
channels: Array containingRETRIEVAL_CHANNEL_VECTOR,RETRIEVAL_CHANNEL_LEXICAL, or both. In vector-only queries, this field is[RETRIEVAL_CHANNEL_VECTOR].distance: Cosine distance (lower is closer). Set to0for lexical-only matches.lexicalScore: Whole-word relevance score (higher is more relevant). Set to0for vector-only matches.rrfScore: Fused score determining rank in hybrid mode. Set to0in vector-only queries.validAt: The chunk's valid-time start. Omitted if unset (indicating unbounded past validity, distinct from a zero timestamp). Reported identically across both vector and lexical channels. See Temporal memory.invalidAt: The chunk's valid-time end. Omitted while the chunk is currently valid, so it is always absent in a default query; anasOfValidquery reports it for chunks retired after the requested instant.
Matches are ordered by rrfScore (highest first) in hybrid mode and by distance (lowest first) in vector-only mode.
limit defaults to 10 and is at most 2,000; your plan may set a lower maximum. A larger value is lowered to the maximum rather than refused.
Filter consistency
All query filters apply identically across both channels:
- Scope boundaries: Reads remain clamped to authorized workspaces and subjects.
- Temporal validity: By default, superseded and inactive chunks are excluded from both channels. With
asOfValid(and optionallyasOfTx), both channels return the chunks valid at that instant instead. See Querying historical state. - Metadata filters: Key-value metadata predicates filter chunks before ranking in both vector and lexical evaluations. A query that spans several scopes (
additionalScopes) does not accept metadata filters and returnsINVALID_ARGUMENT; query each scope on its own to filter. - Read timestamp: Both channels evaluate state at the same consistent read timestamp.
Reranking
Setting rerank: true on the semantic section reorders the results with an external
reranking model before they are returned.
Reranking can be enabled for either RETRIEVAL_MODE_VECTOR_ONLY or
RETRIEVAL_MODE_HYBRID. During reranking, the platform retrieves a broader candidate
pool, scores each candidate against queryText, and returns the highest-scoring chunks
up to limit:
When reranking is enabled, matches include the following fields:
rerankScore: Relevance score from the reranking model (higher is more relevant, unlikedistance). The score is on the reranker's own scale and is not comparable todistanceorrrfScore, nor across model versions.rerankTruncated: Boolean indicating whether the chunk was truncated to fit the reranker per-record token limit. When true, the reranker evaluated only an initial prefix ofrawContent. The full chunk remains stored and returned in the match.
Requirements and error conditions
Requests that specify reranking fail explicitly rather than falling back to unreranked results:
queryTextrequired: Reranking requires search text. A request that carriesembeddingignores itsqueryText, so reranking with anembeddingreturnsINVALID_ARGUMENT.- Reranking unavailable or failed: When reranking is not available for the
workspace, or the reranking call fails or does not score every candidate, the
request returns
FAILED_PRECONDITION. Retry, or send the request withoutrerankto get the ordinary ranking.
Regional availability
Reranking availability varies by region. If the workspace's region does not offer
reranking, requests return FAILED_PRECONDITION rather than falling back to another
region. In some regions, reranking inference may be served from outside the
workspace's region.
Evaluation and benchmarks
Reranking is disabled by default because the additional model inference increases query latency. Benchmark evaluations on representative corpora show that reranking can degrade retrieval quality:
| corpus | metric | without reranking | with reranking |
|---|---|---|---|
| LongMemEval, 90 questions, multi-session memory | cover@10 | 88.9% | 84.4% |
| SciFact, 300 queries, 5,183 documents | nDCG@10 | 0.8979 | 0.7793 |
| SciFact | recall@10 | 0.9717 | 0.8972 |
On both benchmarks, reranking degraded recall by demoting relevant documents within the expanded candidate pool. General-purpose reranking models are trained for standard search queries and can perform poorly on conversational recall or scientific claim verification. Evaluate reranking against your specific workload before enabling it in production.
Reranking also makes the retrieval mode largely irrelevant. Measured on the same SciFact
ingest, RETRIEVAL_MODE_HYBRID with reranking scored identically to
RETRIEVAL_MODE_VECTOR_ONLY with reranking (nDCG@10 0.7793 and recall@10 0.8972 for
both):
| mode | nDCG@10 | recall@10 |
|---|---|---|
| vector | 0.8979 | 0.9717 |
| hybrid | 0.8976 | 0.9717 |
| vector + rerank | 0.7793 | 0.8972 |
| hybrid + rerank | 0.7793 | 0.8972 |
This follows from what the stage does: reranking rescores candidates from their content and discards whatever ordering the channels produced, so fusion's contribution does not survive it. If you enable reranking, the choice of retrieval mode affects only which candidates reach the reranker, not how they are ordered afterwards.
Enabling the lexical channel
Lexical indexing is configured per enterprise. Contact support to enable hybrid retrieval for your organization.
When enabled:
- New chunks are indexed for lexical search as they are committed.
- Existing chunks are indexed automatically in the background, with no re-commits required.
- Hybrid queries are served once indexing across all enterprise data regions is complete.