Forming memory from conversation
memory:form extracts memorable facts and relationships from raw conversation
turns, reconciles candidates against existing workspace memory, and commits the
updates atomically. It automates chunking, entity extraction, ID generation, and
contradiction detection.
POST /v1/agents/{agentInstanceId}/memory:form
{
"turns": [
{
"role": "TURN_ROLE_USER",
"content": "I live in Tokyo and I work at Alphaus."
},
{
"role": "TURN_ROLE_ASSISTANT",
"content": "Noted. Anything else about your setup?"
},
{
"role": "TURN_ROLE_USER",
"content": "I prefer dark mode in every tool."
}
]
}
From the command line, jnh sends the same request from a JSON or YAML
file, or from standard input with --from-file -, and prints the receipt:
Without --formation-key, the command derives a key from the request's content,
so running the same command again after a committed formation replays its receipt
instead of forming twice. --no-formation-key sends no key, to form the same
conversation again on purpose. --observed-at sets
observedAt, overriding the
file's value; a bare date means the start of that day in your local time zone.
memory:form is published for agent workspaces only: there is no
/v1/scopes/{scopeId}/memory:form route, and the CLI form is
jnh agents memory form.
Formation is additive: direct commits (memory:commit) and queries
(memory:query) remain unchanged.
Declaring a memory vocabulary additionally classifies the
entities a formation writes, so a query can ask for every Person a scope knows
about. It steers classification only: it never changes which facts and
relationships are extracted.
When the conversation took place (observedAt)
Use observedAt to specify when the conversation occurred. Leave it out for
real-time formation: formed facts are then stored with no valid-time start (valid
for all past time), formed relationships start at the commit time, and the request
receipt time sets the boundary for revisions.
The observedAt timestamp:
- Sets the valid-time start (
validAt) for newly formed facts and relationships, and serves as the boundary timestamp for revisions. This ensures valid-time queries reflect conversation time rather than ingestion time. Set this field when backfilling or importing historical conversations. - Enforces chronological supersession during reconciliation. Reconciliation only
assigns
revisedwhen the conversation occurred after the contradicted memory became valid.
This prevents historical statements from superseding newer facts. For example, if a
workspace holds "lives in Tokyo" valid from 2026, an imported 2023 conversation
stating "I used to live in Osaka before moving" represents older information.
Treating it as a revision would incorrectly retire current data. The candidate is
instead rejected as MEMORY_DECISION_REJECTED with an explanatory reason, writing
no supersession and leaving current memory intact.
A memory with no recorded valid-time start can be revised by a conversation at any date, as an absent start indicates unbounded past validity.
observedAt is included in the request fingerprint for formationKey. Resending
the same turns under the same key with a different observedAt is treated as key
reuse rather than an idempotent retry.
Omitting observedAt preserves default valid-time behavior for each memory type.
Operational behavior
Latency: Formation performs model inference round-trips and candidate retrieval before writing, taking seconds rather than the milliseconds of a direct commit. Run formation asynchronously or via background workers after a conversation completes.
Per-call bound: Each model call a formation makes is bounded in time. A call
that exceeds the bound fails the formation with DEADLINE_EXCEEDED, whose message
names the bound that was exceeded. Nothing is written and the call is not retried;
resend under the same formationKey to retry. The attempt still
counts as one formation, because the model was called.
The bound applies to one call, not to the formation as a whole. A formation
makes two model calls (extraction, then reconciliation), so it can run well past
its typical duration. No request is allowed more than 30 minutes end to end.
Size client timeouts to that ceiling, not to the typical duration; jnh waits 30
minutes by default (--timeout).
Disconnecting cancels: Closing the connection before the receipt arrives cancels
a formation still in progress. Nothing is written, but a formation that already
reached the model still counts toward your allowance. Resend under the same
formationKey to try again.
Empty commits: Conversations containing no new or memorable facts produce no memory writes. The API returns a successful receipt indicating the evaluation outcome rather than an error. The formation still counts toward your allowance.
External inference: Conversation turns are processed by an external inference model (Vertex AI). See Privacy and data handling below.
Candidate decisions
Each candidate fact or relationship extracted from the conversation receives one of three reconciliation decisions:
| Decision | Meaning | Action |
|---|---|---|
new |
The workspace does not hold this fact | Inserts a chunk, node, or edge |
revised |
The workspace holds a superseded assertion | Closes the prior validity window and inserts a replacement (supersession) |
known |
The workspace already holds this fact | No write |
A fourth value, rejected, indicates a candidate that was not written and
includes a rejectionReason. A revision is rejected when it targets a non-existent
memory, targets a different memory type, is identical to the existing record, or
would retire a memory that became valid after the conversation took place (see
observedAt). Rejected candidates
are not converted into other decisions; the receipt reflects the rejection
explicitly.
Candidates that touch the same memory. One conversation can produce several candidates about the same memory, and formation resolves each one on its own instead of failing the whole formation:
- Two revisions of one memory (for example, "B owns T" replaced by both "A owns
T" and "B reports to A"): the prior memory is retired once, and every
replacement is stored, valid from the same instant. Both candidates are
revisedwith the samematchedId. The receipt'schunkSupersessionsoredgeSupersessionscounts the retired memory once. - A
newcandidate identical to a revision's replacement: the content is stored once, and both candidates keep their decisions. - A
newcandidate for a memory another candidate retires: rejected, because one formation cannot both keep and retire the same memory. - A revision whose replacement another candidate already stores, or which is itself a memory another candidate retires: rejected, and the memory it named stays current.
Non-destructive updates: Formation does not delete or overwrite data. When an assertion is revised, its previous validity window is closed and a replacement is inserted, preserving history for valid-time queries (see Temporal memory).
Example revision in a receipt:
{
"commitTimestamp": "2026-08-25T08:48:01.658Z",
"chunkSupersessions": "1",
"edgeSupersessions": "1",
"candidates": [
{
"candidateId": "fct_1f0c...",
"kind": "CANDIDATE_KIND_FACT",
"text": "The user lives in Osaka.",
"decision": "MEMORY_DECISION_REVISED",
"matchedId": "fct_29e93c79faf7786e5adb8bcb0b49d421"
}
]
}
matchedId identifies the existing memory row against which the candidate was
reconciled (the superseded assertion for revised, or the matching row for
known).
Automatic contradiction detection
While memory:commit requires callers to specify supersessions explicitly,
memory:form infers contradictions and executes supersessions
automatically.
Querying formed memory
Formed memory uses the same storage primitives as direct commits: facts are
stored as vector chunks, relationships as graph nodes and edges, and the
submitted turns of a formation that writes memory as an execution-log step. Retrieve formed memory using standard
memory:query calls.
To inspect workspace state prior to a formation commit, use valid-time queries
(asOfValid) set before the commit timestamp (see
Temporal memory).
candidateId is not a lookup key
candidateId correlates candidates within a receipt. To retrieve a specific
formed chunk, query using semantic search.
Provenance
Every fact formation stores records where it came from, as ordinary metadata on the chunk. The reference survives in the workspace, so a fact remains traceable to its conversation long after the receipt that reported it is gone.
| Key | On | Value |
|---|---|---|
jennah.source_step |
each formed fact chunk | The execution-log step id holding the conversation the fact was extracted from |
jennah.source_turns |
each formed fact chunk | Comma-joined ascending turn numbers within that step, when extraction attributed the fact to turns that exist |
jennah.turn_offsets |
the formation's log step | Comma-joined byte offsets at which each rendered turn begins in the step's toolInput |
These are ordinary metadata, not a separate surface: they are returned wherever the chunk is returned, and they filter like any other key. Retrieving everything one conversation produced is a metadata filter:
{
"semantic": {
"queryText": "engineering team",
"filters": [
{
"key": "jennah.source_step",
"value": "frm_...",
"operator": "OPERATOR_EQUALS"
}
]
}
}
To read the conversation itself, fetch the step and slice its toolInput at the
offsets: turn i runs from offsets[i] to offsets[i+1], and the last runs to
the end. Slice by offset rather than parsing for turn headers, since turn content
can span lines and can contain text shaped like a header.
jennah.source_turns is absent when extraction's attribution was unusable
(it named turns that do not exist, or named none). The fact is still stored and
still carries jennah.source_step: the conversation is known, the position
within it is not. Absent is distinguishable from turn 0.
Provenance names the most recent conversation that stated a fact as new. Writes are idempotent on a content-derived id, so re-extracting an identical fact replaces the chunk's metadata and moves the reference to the later conversation. A fact reconciliation judges already known writes nothing and keeps the reference it had. A revision's replacement chunk carries the revising conversation, and the chunk it retires keeps its own, so each state of a superseded fact names the conversation that produced it.
The jennah. prefix is reserved on extracted attributes. Extraction reads
caller-supplied text, so a conversation could otherwise instruct the model to tag
a fact with a source key naming a different conversation. A candidate whose
extracted attributes use the prefix is rejected, and the receipt names the key.
The restriction applies only to extraction: a caller may write jennah. keys on
their own chunks through memory:commit.
Provenance counts toward the per-item metadata limit. A fact carries one or two platform keys in addition to its extracted attributes. A candidate whose attributes leave no room is rejected with a reason naming the limit, rather than being stored with its provenance silently dropped.
Tables, code, and other structured content
Extraction summarizes conversation turns into prose facts. For structured payloads such as tables, rosters, and configuration blocks, summarization can lose row-level detail. Formation handles structured content through two mechanisms:
Structure preservation:
Turns are scanned for markdown pipe tables and fenced code blocks before
extraction. When a structured block contains memorable content, formation
preserves the verbatim structure within the fact text and flags the candidate
receipt with preservesStructure: true:
{
"candidateId": "fct_8b21...",
"kind": "CANDIDATE_KIND_FACT",
"text": "Shift rotation:\n\n| agent | shift |\n|---|---|\n| Admon | early |",
"decision": "MEMORY_DECISION_NEW",
"preservesStructure": true
}
preservesStructure is asserted only when the input turn contained recognized
structured syntax and the extracted fact retained that syntax verbatim.
Loss reporting (summarizedStructures):
When recognized structured blocks are summarized into prose rather than preserved
verbatim, the receipt reports them in summarizedStructures:
This indicates that individual row entries may not be retrievable via vector
search. For content requiring exact tabular retrieval, use
memory:commit to define explicit chunk boundaries per row or
section.
summarizedCount reports the total number of summarized structures detected
within each turn.
Scope of structure detection
Structure detection recognizes markdown pipe tables and fenced code blocks
in turn content. Indented lists, ASCII tables, and tool execution traces are
summarized as prose without being recorded in summarizedStructures.
What the agent said
Formation extracts from both sides of the conversation. Whether a statement is worth remembering depends on what it says, not on which participant said it, so substantive content the agent supplied to the user is stored on the same terms as statements the user made:
- a recommendation the agent made
- a specific answer it gave to a question
- a name, figure, or reference it introduced
- an artifact it produced for the user to keep, such as a plan, an itinerary, an outline, or a list
This is what makes "remind me of the restaurant you recommended" or "how many subjects were in that study you cited" answerable from memory. A fact derived from an agent turn is written so that it says the agent supplied it, because a recommendation stored as though the user had made it changes whose statement it was:
{
"candidateId": "fct_4c1a...",
"kind": "CANDIDATE_KIND_FACT",
"text": "The assistant recommended Miss Bee Providore for nasi goreng.",
"decision": "MEMORY_DECISION_NEW"
}
What is not stored. The exclusion is about turns that add nothing, not about who was speaking. The agent's acknowledgements, its restatements of what the user just said, its offers of further help, and the mechanics of the conversation are not memory. A conversation in which the agent said nothing worth returning to forms nothing from its turns, which is a normal outcome.
Tool inputs and outputs are content the agent read rather than statements it made to the user. They are memorable only where the agent passed their substance on in its own reply.
Idempotent retries (formationKey)
Because model extraction is non-deterministic, re-sending a timed-out request without a key can create duplicate or divergent memory entries.
Pass a formationKey to ensure idempotent execution:
The formation key record is committed in the same transaction as the extracted
memory. Retrying a request with an existing formationKey returns the original
commit receipt without re-running extraction. That receipt's formationUnits is
what the original formation consumed; the retry itself consumes nothing, so check
whether a receipt is a replay before reading its units as this request's cost.
Reusing a key with different conversation turns returns INVALID_ARGUMENT.
When another write commits first, the losing formation writes nothing and
returns ABORTED (HTTP 409). The error's reason identifies the conflict:
| Reason | Conflict | Action |
|---|---|---|
FORMATION_KEY_CLAIMED |
Another request with the same formationKey committed first |
Resend with the same formationKey; it returns that formation's receipt |
FORMATION_SUPERSEDED_FIRST |
Another formation superseded the same memory assertion first | Resend; the retry reconciles against the current memory |
FORMATION_WRITE_CONFLICT |
Another write to the agent conflicted with this formation | Resend |
Forming the same conversation into the same agent again, with no
formationKey or a new one, is not a conflict. The formation succeeds and its
facts reference the execution-log step the first formation recorded; its receipt
reports executionLogRows 0.
Token usage
Every formation that receives a response from the model records the input and output tokens the model reported, summed across its calls. That includes a formation that stored nothing, and one that failed after a model call returned. Output includes the model's reasoning tokens. Input and output are kept as two separate counts because they are priced differently. The counts are what the model reported, not a bill: a call abandoned before it returned has nothing to report.
Read the totals with jnh usage tokens, or with
GET /v1/billing/usage/formation-tokens, which requires startTime (the 30-day
default below is the CLI's). The answer is always aggregated, never
per formation. Each bucket also carries the formation units
its formations consumed (the UNITS column, units in JSON), summed per
formation:
# The last 30 days, in UTC
jnh usage tokens
# Per Tokyo calendar day, split by model
jnh usage tokens --since 2026-09-01 --per day --by model --tz Asia/Tokyo
# Per ISO week for one API key
jnh usage tokens --per week \
--caller api_key:25678126-312d-4804-abee-f0adf163acbc
--pertakesday,week(ISO weeks, Monday start) ormonth, taken in--tz, which defaults to UTC.--bytakes one ofscope,caller,modelorregion. A caller isuser:<id>orapi_key:<id>, and an API key is recorded as itself, never as the user who created it.--scope,--caller,--modeland--regionnarrow what is counted without grouping by it.- The range can reach back at most 400 days, the retention period for usage records. Usage outlives the scopes it was spent on, so a destroyed scope still appears in a breakdown.
- A formation whose model calls reported no tokens (one abandoned at its time bound) counts 1 unit toward your allowance but has no tokens to report, so it does not appear here.
Reading usage requires billing.usage:read, which the built-in member role does
not carry (see Access control). Grouping or filtering by
scope also requires reach over every agent workspace and every subject scope in
the enterprise. Without it the request is refused rather than answered with only
the scopes you can reach, so a per-scope breakdown always adds up to the
enterprise's totals.
Formation units
Formation allowances are counted in formation units. A formation consumes the largest of:
- 1
- its input tokens divided by 16,384, rounded up
- its output tokens, including reasoning, divided by 8,192, rounded up
Tokens are summed over all model calls in the formation:
| Input tokens | Output tokens | Units |
|---|---|---|
| 2,700 | 1,700 | 1 |
| 2,680 | 16,222 | 2 |
| 16,552 | 70,000 | 9 |
| 0 (abandoned) | 0 (abandoned) | 1 |
How many units a formation consumes depends mostly on how much the model writes,
and its reasoning counts as output. The same short conversation can consume 1
unit on one run and 2 or 3 on another when the model reasons at length, so plan
by the units your receipts report rather than by the number of formations. Every
receipt reports the units consumed in formationUnits. Replaying a completed formation reports the original units and
consumes nothing. A formation whose model calls reported no tokens (such as one
abandoned at its time bound) counts 1 unit toward the allowance, but has no
tokens to report in usage.
Limits and bounds
| Bound | Behavior when exceeded |
|---|---|
| Turns per request (200) | INVALID_ARGUMENT |
| Total input tokens | INVALID_ARGUMENT, naming your plan's cap (8,192 tokens on Free; 16,384 on the trial, Pro, Scale and Team; 100,000 on Enterprise) |
| Candidates per formation (64) | Excess candidates dropped; candidatesDropped and candidateCap reported on receipt |
formationKey length (256 characters) |
INVALID_ARGUMENT |
| Model call time | DEADLINE_EXCEEDED naming the bound; nothing is written, the attempt uses its formation units; resend with the same formationKey |
| Model call output (32,768 tokens per call on Free; bounded on every plan) | ABORTED naming the bound; nothing is written, the attempt uses its formation units; resend with the same formationKey |
| Formation output per day (50,000 tokens on Free) | RESOURCE_EXHAUSTED naming the allowance and when it resets; a formation that reaches what is left stops with ABORTED as above |
| Formations in progress | RESOURCE_EXHAUSTED with a retry delay; nothing is spent; retry once the delay has passed |
Submissions exceeding input bounds are rejected with INVALID_ARGUMENT rather
than truncated. Split large conversations across multiple requests.
candidatesDropped > 0 indicates that the conversation exceeded the candidate
capacity for a single formation commit.
Formation usage is metered separately from direct commit and query quotas.
Exceeding the formation allowance returns RESOURCE_EXHAUSTED, stating the
allowance and the units used, and naming whether it resets daily, monthly, or not
before the trial ends. On a plan with overage, formations continue past
the allowance up to your overage limit instead. A formation started while units remain runs to completion
and consumes its units, so usage can end slightly above the allowance.
Every formation that reaches the model uses its units, whatever it produced:
one that stored memory, one that found nothing worth storing, and one that failed
after the model was called (on its time or output bound, for example). A resend
under the same formationKey that reaches the model uses units again. A resend
answered from the stored receipt, a refusal at the allowance, and a submission
refused for size make no model call and use nothing.
Overage
On the Pro, Scale and Team plans bought through AWS Marketplace, formations continue past the monthly allowance instead of being refused. Each formation unit past the allowance is billed as overage, at the per-unit price on the listing, on your monthly AWS bill.
An enterprise is billable for overage while it holds an active AWS Marketplace subscription on a plan that includes overage. Every other enterprise is refused at its allowance, as described above: Free, the trial, a plan set up any other way, and a subscription awaiting renewal verification. A subscription that ends stops overage at the next formation; the overage already used is still billed.
The overage limit
Overage stops at a hard limit: the allowance plus an overage band. The band is a multiple of the plan's monthly fee, counted in units, and the default multiple is 2:
| Plan | Allowance | Default overage band | Formations refused from |
|---|---|---|---|
| Pro | 550 / month | 992 units | 1,542 units |
| Scale | 1,800 / month | 2,992 units | 4,792 units |
| Team | 6,500 / month | 9,992 units | 16,492 units |
A formation past the limit is refused with RESOURCE_EXHAUSTED and the reason
FORMATION_OVERAGE_LIMIT, naming the limit and the units used. As with the
allowance, a formation started below the limit runs to completion and is billed for
all of its units.
A root or administrator changes the multiple on the console's Billing page, from 0 to 10 in steps of 0.5:
- A multiple of 0 turns overage off, so formations are refused at the allowance.
- A change applies to the next formation, for the current month. Raising the multiple admits formations straight away, without waiting for the next month or a new sign-in.
- Lowering it below what the month has used refuses the next formation. The overage already used is still billed.
- The multiple keeps its meaning across a plan change: 3 is 1,488 units on Pro and 4,488 on Scale.
It cannot be granted by permission, and it cannot be changed with an API key, even one created by an administrator, so an agent can never raise its own spending limit. An enterprise that is not billable can still set a multiple; it applies once the enterprise is billable.
The billing state (GET /v1/billing) reports the month's allowance, units used,
overage units, the multiple, the resulting band, and whether the enterprise is
billable, in units only. Any member who can read the billing state sees them.
Privacy and data handling
Inference routing: Submitted conversation turns are sent to Google Vertex AI
for processing. Jennah does not filter or redact input text; sensitive data (such
as credentials or private identifiers) should be sanitized before calling
memory:form.
Direct commits (memory:commit) and queries (memory:query) stay in the
workspace's region, but formation inference may be processed outside it.
Audit logs: When a formation writes memory, its submitted turns are stored as an execution-log step within the same commit transaction, so extracted memory can be audited back to source turns. A formation that writes nothing (every candidate known or rejected) stores no log step. Turn logs follow the workspace retention lifecycle.
Limitations and security
Heuristic extraction: Model extraction and reconciliation are probabilistic. Formation may occasionally extract extraneous facts or omit relevant details.
Untrusted input: User inputs in submitted conversation turns can influence extracted facts. Audit receipts and execution logs to trace memory origins.
Tenant isolation: Workspace and tenant boundaries are enforced strictly by caller authentication tokens and request routing, not model outputs. Extracted content cannot cross tenant or workspace boundaries.
Deduplication: Identically worded facts generate matching content hashes and merge onto existing rows. Differently phrased statements of the same fact rely on semantic reconciliation and may not always deduplicate.