Skip to content
Build with Mellow

Inside memory

Trace memory processing, storage and the diagnostics behind remembered context.

In this topic

Memory is a separate continuity system beside conversation history. A transcript preserves what was said; memory extracts useful context and decides what to bring into a later request. Understanding the write and read paths helps explain why a saved chat may not immediately appear as a recalled fact.

For everyday controls, use Memory. This chapter follows the services that implement those controls.

Capture turns without pausing the conversation

MemoryService.bufferTurn accepts the user message, optional assistant message, agent identifier, conversation identifier, optional session date, and optional project identifier. It records a pending signal and rearms the conversation's debounce. It does not run the extraction model synchronously on every chat turn.

When the debounce expires, the session changes, or a caller explicitly flushes the session, the service can distill the buffered material. One structured extraction combines episode information, entities, identity changes, and candidate pinned facts. The configured memory model must be available; pending signals are not the same thing as successfully distilled memory.

Project membership is captured when the turn is buffered, while the live conversation still knows its project. Reconstructing that relationship later could lose the correct project association after navigation or a session reset.

Delegated work has separate session bookkeeping

bufferDelegatedTurn records work performed on an agent's behalf without replacing that agent's active direct-chat conversation. A delegated task arriving during a chat therefore does not force the direct session to flush simply by claiming to be the new active conversation.

When adding a new caller to the memory pipeline, choose the appropriate entry point. Reusing the direct-chat call for background delegation can change lifecycle behavior even when the stored text looks correct.

Distinguish the stored material

MaterialPurpose
Pending signalsDurable input awaiting extraction
TranscriptConversation turns available for historical recall
EpisodesDistilled session context such as topics and decisions
Pinned factsUseful individual facts retained across sessions
Identity informationSmall stable context about the person or agent

Distillation need not create a pinned fact from every conversation. A short exchange may warrant only an episode, and irrelevant content should not be promoted just to increase a memory count.

Retrieve only what the turn needs

The read path evaluates the incoming message, selects relevant memory material, and constructs a bounded context block. Identity, episode, pinned-fact, and transcript recall serve different requests. Asking for the exact wording of an earlier conversation is different from asking which project the person works on.

A token budget limits the material inserted into a request. Consequently, a fact can exist in storage without appearing in every answer. Diagnose retrieval selection and the final model request separately from storage. Increasing a budget is not a substitute for identifying why the correct record was not selected.

Explicit memory search is also available through the tool surface where enabled. Inspect the tool's current schema and results rather than assuming that a free-form request searches every agent's memory.

Consolidation is background maintenance

MemoryConsolidator performs maintenance outside the request path: reducing salience over time, combining similar episodes, promoting repeatedly supported facts, and pruning material according to retention settings. The scheduled path respects its configured interval and idle conditions; an on-demand action is a separate trigger.

Retention, salience, and relevance have different roles. Retention determines how long material is kept. Salience influences its importance. Retrieval decides whether it helps this request. A troubleshooting report should name which of these behaviors is unexpected.

Configuration fields for the memory pipeline

The current configuration model defines these defaults and normalization bounds. Read the saved setting for your installation before interpreting a run; defaults are not evidence of the effective value.

FieldDefaultAccepted normalized rangeMeaning
memoryBudgetTokens800100–4000Maximum retrieved-memory context budget
summaryDebounceSeconds6010–3600Quiet interval before session distillation
consolidationIntervalHours241–168Scheduled maintenance interval
salienceFloor0.20–1Salience threshold used by maintenance
episodeRetentionDays3650–3650Retention configuration for historical material

These fields live in MemoryConfiguration.swift. Changing a debounce value is not a way to repair an unavailable extraction model. Check readiness first, then tune timing to the intended workflow.

Search indexes and agent boundaries

MemorySearchService maintains agent-scoped search namespaces and separate entry points for pinned facts, episodes, and transcript. Removing an agent's index, purging its namespace storage, and rebuilding an index are distinct maintenance operations. A rebuild should restore searchable state from the intended underlying records; it should not broaden the caller's agent scope.

When retrieval returns nothing, check both the stored record and indexing state. An extraction failure and an indexing failure can produce the same visible empty search while requiring different recovery. Keep index-failure diagnostics with the affected agent and material type.

Attribute API conversations correctly

Clients can use X-Mellow-Agent-Id with an agent identifier from GET /agents. Keep a stable conversation identifier for turns that belong together. An agent header is attribution, not a way to override authentication or another agent's permissions.

The bulk ingestion endpoint accepts a conversation and turn list:

{
  "agent_id": "AGENT_ID",
  "conversation_id": "research-notes-01",
  "turns": [
    {"user": "The launch review is on Tuesday.", "assistant": "I will use Tuesday in this conversation."}
  ]
}

Send this to POST /memory/ingest only when intentionally adding that material. session_date supplies historical timing; skip_extraction requests transcript insertion without the extraction pass. Successful ingestion is not a guarantee that an unavailable extraction model produced a memory summary.

Trace missing memory in order

  1. Confirm memory is enabled and the turn is nonempty.
  2. Check agent and conversation attribution.
  3. Establish whether the caller reached bufferTurn or the delegated equivalent.
  4. Inspect pending signals and extraction-model readiness.
  5. Confirm a distilled record exists.
  6. Inspect retrieval selection and budget on a later request.
  7. Check whether the model used the supplied evidence accurately.

The service's buffering telemetry separates no caller activity from early exits, including disabled memory and empty messages. Preserve that distinction in diagnostics. Source landmarks include MemoryService.swift, MemorySearchService.swift, the memory database, and the consolidation service. No stage's source implementation alone proves recall quality for a released build.

Continue exploring · Build with MellowInside folder watchers →Understand event handling, batching and dispatch for file-triggered work.