Skip to content

Retrieval channels

Jennah provides two channels for semantic memory retrieval:

  • Vector retrieval: Embeds queries and ranks chunks by cosine distance. Vector search identifies content with similar semantic meaning, even when exact terminology differs.
  • Lexical retrieval: Matches terms and ranks by whole-word relevance. Lexical search targets exact terms that embedding models may index less distinctively, such as error codes, product SKUs, identifiers, and proper nouns.

Vector retrieval is the default mode. Hybrid retrieval combines vector and lexical channels to query both simultaneously.

The two channels

Vector retrieval

Vector retrieval embeds queryText (or uses a caller-supplied embedding) and ranks candidate chunks within the workspace by cosine distance.

Lexical retrieval

Lexical retrieval performs whole-word token matching against rawContent and ranks matches by relevance.

Lexical matching operates on whole words rather than substrings or prefixes. Searching for quota matches exact token occurrences but does not match prefix variants like quotas.

Using hybrid mode

To query both channels simultaneously, set retrievalMode to RETRIEVAL_MODE_HYBRID in the semantic section of memory:query:

{
  "semantic": {
    "queryText": "ERR_QUOTA_EXHAUSTED_88213",
    "limit": 10,
    "retrievalMode": "RETRIEVAL_MODE_HYBRID"
  }
}

Omitted or RETRIEVAL_MODE_VECTOR_ONLY executes vector retrieval alone.

  • queryText required, embedding not allowed: Lexical retrieval requires search text. A request that carries embedding is served from the vector alone and its queryText is ignored, so hybrid mode with an embedding (with or without queryText) returns INVALID_ARGUMENT.
  • Enterprise enablement: Hybrid mode must be enabled for the enterprise before use. Requests made prior to enablement return FAILED_PRECONDITION. See Enabling the lexical channel.

Reciprocal rank fusion

Hybrid retrieval combines results using Reciprocal Rank Fusion (RRF) rather than merging raw similarity scores:

  • Each channel independently ranks candidate chunks.
  • A chunk's fused score is computed as sum(1 / (k + rank)) across matching channels, where k is a fixed ranking constant.
  • Ranks are deterministic and avoid score normalization discrepancies across different vector and lexical scoring distributions.
  • Chunks matched by multiple channels accumulate higher scores, while chunks matched by only a single channel remain eligible for retrieval.

Reading channel provenance

Query results include channel provenance identifying which channel or channels surfaced each chunk:

{
  "semantic": {
    "matches": [
      {
        "chunkId": "runbook-quota",
        "rawContent": "when a quota is exhausted, first check...",
        "distance": 0.12,
        "channels": [
          "RETRIEVAL_CHANNEL_VECTOR",
          "RETRIEVAL_CHANNEL_LEXICAL"
        ],
        "lexicalScore": 2.10,
        "rrfScore": 0.0323
      },
      {
        "chunkId": "incident-4471",
        "rawContent": "incident ERR_QUOTA_EXHAUSTED_88213 was closed...",
        "distance": 0,
        "channels": ["RETRIEVAL_CHANNEL_LEXICAL"],
        "lexicalScore": 4.82,
        "rrfScore": 0.0164
      }
    ]
  }
}
  • channels: Array containing RETRIEVAL_CHANNEL_VECTOR, RETRIEVAL_CHANNEL_LEXICAL, or both. In vector-only queries, this field is [RETRIEVAL_CHANNEL_VECTOR].
  • distance: Cosine distance (lower is closer). Set to 0 for lexical-only matches.
  • lexicalScore: Whole-word relevance score (higher is more relevant). Set to 0 for vector-only matches.
  • rrfScore: Fused score determining rank in hybrid mode. Set to 0 in vector-only queries.
  • validAt: The chunk's valid-time start. Omitted if unset (indicating unbounded past validity, distinct from a zero timestamp). Reported identically across both vector and lexical channels. See Temporal memory.
  • invalidAt: The chunk's valid-time end. Omitted while the chunk is currently valid, so it is always absent in a default query; an asOfValid query reports it for chunks retired after the requested instant.

Matches are ordered by rrfScore (highest first) in hybrid mode and by distance (lowest first) in vector-only mode.

limit defaults to 10 and is at most 2,000; your plan may set a lower maximum. A larger value is lowered to the maximum rather than refused.

Filter consistency

All query filters apply identically across both channels:

  • Scope boundaries: Reads remain clamped to authorized workspaces and subjects.
  • Temporal validity: By default, superseded and inactive chunks are excluded from both channels. With asOfValid (and optionally asOfTx), both channels return the chunks valid at that instant instead. See Querying historical state.
  • Metadata filters: Key-value metadata predicates filter chunks before ranking in both vector and lexical evaluations. A query that spans several scopes (additionalScopes) does not accept metadata filters and returns INVALID_ARGUMENT; query each scope on its own to filter.
  • Read timestamp: Both channels evaluate state at the same consistent read timestamp.

Reranking

Setting rerank: true on the semantic section reorders the results with an external reranking model before they are returned.

Reranking can be enabled for either RETRIEVAL_MODE_VECTOR_ONLY or RETRIEVAL_MODE_HYBRID. During reranking, the platform retrieves a broader candidate pool, scores each candidate against queryText, and returns the highest-scoring chunks up to limit:

{
  "semantic": {
    "queryText": "ERR_QUOTA_EXHAUSTED_88213",
    "limit": 10,
    "rerank": true
  }
}

When reranking is enabled, matches include the following fields:

  • rerankScore: Relevance score from the reranking model (higher is more relevant, unlike distance). The score is on the reranker's own scale and is not comparable to distance or rrfScore, nor across model versions.
  • rerankTruncated: Boolean indicating whether the chunk was truncated to fit the reranker per-record token limit. When true, the reranker evaluated only an initial prefix of rawContent. The full chunk remains stored and returned in the match.

Requirements and error conditions

Requests that specify reranking fail explicitly rather than falling back to unreranked results:

  • queryText required: Reranking requires search text. A request that carries embedding ignores its queryText, so reranking with an embedding returns INVALID_ARGUMENT.
  • Reranking unavailable or failed: When reranking is not available for the workspace, or the reranking call fails or does not score every candidate, the request returns FAILED_PRECONDITION. Retry, or send the request without rerank to get the ordinary ranking.

Regional availability

Reranking availability varies by region. If the workspace's region does not offer reranking, requests return FAILED_PRECONDITION rather than falling back to another region. In some regions, reranking inference may be served from outside the workspace's region.

Evaluation and benchmarks

Reranking is disabled by default because the additional model inference increases query latency. Benchmark evaluations on representative corpora show that reranking can degrade retrieval quality:

corpus metric without reranking with reranking
LongMemEval, 90 questions, multi-session memory cover@10 88.9% 84.4%
SciFact, 300 queries, 5,183 documents nDCG@10 0.8979 0.7793
SciFact recall@10 0.9717 0.8972

On both benchmarks, reranking degraded recall by demoting relevant documents within the expanded candidate pool. General-purpose reranking models are trained for standard search queries and can perform poorly on conversational recall or scientific claim verification. Evaluate reranking against your specific workload before enabling it in production.

Reranking also makes the retrieval mode largely irrelevant. Measured on the same SciFact ingest, RETRIEVAL_MODE_HYBRID with reranking scored identically to RETRIEVAL_MODE_VECTOR_ONLY with reranking (nDCG@10 0.7793 and recall@10 0.8972 for both):

mode nDCG@10 recall@10
vector 0.8979 0.9717
hybrid 0.8976 0.9717
vector + rerank 0.7793 0.8972
hybrid + rerank 0.7793 0.8972

This follows from what the stage does: reranking rescores candidates from their content and discards whatever ordering the channels produced, so fusion's contribution does not survive it. If you enable reranking, the choice of retrieval mode affects only which candidates reach the reranker, not how they are ordered afterwards.

Enabling the lexical channel

Lexical indexing is configured per enterprise. Contact support to enable hybrid retrieval for your organization.

When enabled:

  1. New chunks are indexed for lexical search as they are committed.
  2. Existing chunks are indexed automatically in the background, with no re-commits required.
  3. Hybrid queries are served once indexing across all enterprise data regions is complete.