LaserData Cloud
Laser SDKAdvanced

Memory in depth

Scopes, recall modes, vector ranking, lineage, named items, context blocks, consolidation, and custom backends for agent memory

This page is the full reference for Memory. Start there for the short version.

How memory is stored

laser.memory(namespace) returns a MemoryHandle. Every remember, improve, and forget appends a record to the memory topic. forget writes a tombstone and improve writes a feedback weight, so nothing is edited in place and the topic keeps the full history.

A managed deployment folds the memory topic into a versioned key-value read view. Default recall and fetch read that view, so they need Laser Stack or LaserData Cloud. Folded recall rebuilds memory from the topic inside your process instead, so it works with Apache Iggy alone.

When a forget or improve call names a conversation in its scope, it acts only on an item remembered in that conversation. A call without a conversation acts on the item in any conversation. A scoped memory always passes its own conversation, so it cannot forget or reweight another conversation's item.

Namespaces

A client with a default stream scopes the memory namespace to the stream its memory records ride, so laser.memory("notes") reads and writes stream:<stream>/notes. A name that already starts with stream: is sent as written. A client without a default stream sends the bare name. See Stream-scoped resource names for the opt out.

Every record carries its namespace in a header. Folded recall reads only records with the handle's namespace, and the managed view keys rows by namespace, so two namespaces on one topic never mix.

Scope

Each item has a scope. A scope names the stream, the agent, and the conversation the item belongs to, plus an optional user and application.

FieldRemember (Rust, TypeScript)Python keywordMemoryScope field
Conversation.scope(conversation)conversation=conversation
Agent.agent(id)agent=agent
Stream.stream(name)stream=stream
User.user(user)user=user
Application.application(app)application=app in Rust and TypeScript

Recall narrows by every field you set. A field you leave out widens the read, so a recall without a conversation reads across conversations. TypeScript calls recall() with no argument, Rust calls recall(None), and Python omits conversation=.

A stream filter on a log handle returns nothing when it names a stream other than the one the records ride.

.durable() (Python: durable=True) sets the scope lifetime to durable. The default is session. Custom backends receive the lifetime. Built-in backends do not store it, and recall never filters on it.

Memory topic

By default memory records ride agent.memory on the connection's default stream. bootstrap(partitions, retention) creates that topic with the other agent topics.

memory_on_topic(topic, stream) (TypeScript: memoryOnTopic) uses a topic that already exists. The stream is optional and defaults to the connection's default stream: Rust takes an Option<&str>, TypeScript an optional second argument, and Python stream=. memory_topic(topic) (TypeScript: memoryTopic) creates the topic first and sets its partition count and message expiry. Both use the topic name as the namespace.

SettingTypeScriptRustPython
Stream.stream(name).stream(name)stream=
Partitions, default 1.partitions(count).partitions(count)partitions=
Expiry, default 30 days.ttl(ttlMs).ttl(Duration)ttl_ms= in milliseconds
Never expire.noExpiry().no_expiry()ttl_ms=0

TypeScript and Rust finish the builder with .build(). Python awaits laser.memory_topic(topic, ..) directly.

Records of one conversation share a partition key, so each conversation keeps its write order on a multi-partition topic. Topic expiry and view retention are separate. Once a record expires from the topic, folded recall can no longer rebuild it, while the managed view keeps its own retention.

Recall

recall(conversation) starts a recall builder in Rust and TypeScript, and recall(None) in Rust or recall() in TypeScript recalls every conversation. Finish it with .fetch(). Python passes everything as keywords to recall(..). .limit(n) (Python: limit=) caps the result. The default is 50.

Managed and folded recall

Default recall reads the managed key-value view and returns the newest items first. Newest means the latest broker append time, with the source position breaking ties, so items from different partitions come back in arrival order. The view folds the topic in the background, so a fresh write can take a moment to appear.

.folded() (Python: folded=True) rebuilds memory from the topic in your process. Keep one handle for folded reads. A handle remembers what it has read and reads only new records on the next call. A fresh handle reads the whole topic on its first recall.

Feedback from improve reorders folded recall and vector recall for every strategy except recent. Managed recall always returns newest first.

Strategies

StrategyRust and TypeScriptPythonRanks by
Recent.recent()strategy="recent"Newest first. Query text and feedback do not change it
Semantic.semantic(text)semantic=textEmbedding similarity to the text
Keyword.keyword(text)strategy="keyword", semantic=textExact terms, which suits names and identifiers
Hybrid.hybrid(text)strategy="hybrid", semantic=textThe sum of semantic, keyword, and feedback scores

Python also accepts auto, graph, and temporal. Rust and TypeScript set any value with .strategy(..). auto is the default.

Only the vector backend ranks by text. Log memory uses the text as a filter: semantic, keyword, and hybrid recall keep items that share at least one word with the text, then apply the limit. Recent recall ignores the text.

Each item's signals records which strategy surfaced it, its rank within that signal (0 is best), and that signal's score. Feedback appears as an auto signal in every backend.

fuse_reciprocal_rank(signals, limit) (TypeScript: fuseReciprocalRank) merges several ranked lists into one. Each item scores the sum of 1 / (60 + rank) over the lists that contain it, so agreement between lists outranks one list's top hit.

Vector memory

The vector backend embeds each item when you remember it and ranks recall by cosine similarity. It keeps items in process memory and drops them when the handle goes away.

TaskTypeScriptRustPython
Open a governed vector handlelaser.memoryWith(ns, MemoryBackend.Vector, embedder)laser.memory_with(ns, MemoryBackend::Vector).embedder(e)laser.memory_with(ns, "vector", embedder=e)
Same backend, built directlyVectorMemory.governed(laser, embedder)VectorMemory::governed(laser, embedder)VectorMemory.governed(laser, embedder)
Standalone, no connection, no governanceMemoryHandle.vector(embedder) or new VectorMemory(embedder)MemoryHandle::vector(embedder) or VectorMemory::new(embedder)MemoryHandle.vector(embedder) or VectorMemory(embedder)

The embedder comes from your application. It is a Rust Embedder implementation, a TypeScript object with embed(text), or a Python callable or object with embed(text) that returns a list of floats, directly or through an awaitable. A synchronous Python embedder runs on an SDK worker thread, so it must not block.

Python and TypeScript take the embedder when the handle opens, so they check it there: the vector backend without an embedder, or another backend with one, is an invalid error. Rust attaches the embedder after opening with .embedder(..), so a Rust vector handle without one fails with a configuration error at the first remember or semantic recall. .embedder(..) on a vector handle builds a fresh, empty index over the new embedder, so attach it before you write. On a log or custom handle .embedder(..) changes nothing.

Keyword recall on the vector backend scores the share of query words found in the item and never calls the embedder. Without query text, vector recall returns the newest items, ranked by feedback when any exists.

Rerankers

A reranker is a second pass over the recalled candidates. Attach one with .reranker(..) on any handle, or wrap a backend in RerankedMemory (Rust: RerankedMemory::new(inner, reranker), Python: RerankedMemory(inner, reranker), TypeScript: new RerankedMemory(inner, reranker)). It runs only when the query carries text, and it also runs on folded recall. Writes pass through unchanged.

Rust takes a Reranker implementation. TypeScript takes an object with rerank(query, items). Python takes a callable (query, items) -> items or an object with rerank(query, items), synchronous or async.

import { ConversationId, MemoryBackend } from "@laserdata/laser-sdk"

const conversation = ConversationId.new()
const memory = laser.memoryWith("runbooks", MemoryBackend.Vector, embedder).reranker(reranker)

const encoder = new TextEncoder()
await memory.remember(encoder.encode("node-7 sits in the eu-west pool")).scope(conversation).send()
await memory.remember(encoder.encode("node-9 rotates keys monthly")).scope(conversation).send()

const hits = await memory.recall(conversation).semantic("which pool is node-7 in").limit(3).fetch()
use laser_sdk::prelude::full::*;

let conversation = ConversationId::new();
let memory = laser
    .memory_with("runbooks", MemoryBackend::Vector)
    .embedder(embedder)
    .reranker(reranker);

memory
    .remember("node-7 sits in the eu-west pool".as_bytes())
    .scope(conversation)
    .send()
    .await?;
memory
    .remember("node-9 rotates keys monthly".as_bytes())
    .scope(conversation)
    .send()
    .await?;

let hits = memory
    .recall(conversation)
    .semantic("which pool is node-7 in")
    .limit(3)
    .fetch()
    .await?;
import laser_sdk as ls

conversation = ls.new_conversation_id()
memory = laser.memory_with("runbooks", "vector", embedder=embed).reranker(rerank)

await memory.remember("node-7 sits in the eu-west pool", conversation=conversation)
await memory.remember("node-9 rotates keys monthly", conversation=conversation)

hits = await memory.recall(
    conversation=conversation,
    semantic="which pool is node-7 in",
    limit=3,
)

embedder and reranker are your own implementations. The runnable memory example ships a small bag-of-words embedder in each language.

Writing items

Kinds

remember(..) stores a Fact unless you set a kind with .kind(..) (Python: kind="summary" and the other lowercase names).

KindClassHolds
FactSemanticA standalone fact. The default
MessageEpisodicA conversation turn or event
SummarySemanticA summary distilled from other items
EntitySemanticAn extracted entity
FeedbackSemanticA ranking signal
ProcedureProceduralA reusable workflow

MemoryKind::class() in Rust, memory_kind_class(kind) in Python, and memoryClass(kind) in TypeScript report the class. MemoryKind::code(), memory_kind_code(kind), and memoryKindCode(kind) return the stable one-byte code.

Content IDs

remember returns a random ID unless you add .dedup() (Python: dedup=True). With dedup the ID comes from the owner scope (stream, agent, user, and application), the kind, and the body. The conversation is not part of the ID, so remembering the same fact twice stores it once, and two users never share an ID. MemoryId::content(scope, kind, body) in Rust, memory_id_content(kind, body, stream=, agent=, user=, application=) in Python, and MemoryId.content(scope, kind, body) in TypeScript compute the same ID without writing anything. Every SDK produces the same ID for the same input. Graph uses the same approach for node IDs.

Typed append

Typed append writes an ID and a kind that you supply. Built-in backends keep both. Use it when the ID comes from somewhere else, such as a content ID. Rust calls the Memory::append trait method, Python uses append(id, payload, kind=, ..), and TypeScript uses append(scope, id, kind, payload).

import { AgentId, MemoryId, MemoryKind } from "@laserdata/laser-sdk"

const scope = { agent: AgentId.new("planner"), conversation }
const body = new TextEncoder().encode("node-7 rotates keys monthly")
const id = MemoryId.content(scope, MemoryKind.Entity, body)

await memory.append(scope, id, MemoryKind.Entity, body)
let scope = MemoryScope::builder()
    .agent("planner".parse()?)
    .conversation(conversation)
    .build();
let body = b"node-7 rotates keys monthly".to_vec();
let id = MemoryId::content(&scope, MemoryKind::Entity, &body);

Memory::append(&memory, &scope, id, MemoryKind::Entity, body).await?;
body = b"node-7 rotates keys monthly"
memory_id = ls.memory_id_content("entity", body, agent="planner")

await memory.append(
    memory_id,
    body,
    kind="entity",
    agent="planner",
    conversation=conversation,
)

Lineage

An item can record where it came from. origin is the log record that motivated it, such as the message a fact was extracted from. producer names the agent, extraction policy, or consolidator that wrote it, with a name and a version. Both are optional, never filter recall, and never change the item's content ID.

Set them on one write with .origin(source) and .producer(info) on the remember builder (Python: origin= and producer=). with_lineage(origin, producer) on a scoped memory (TypeScript: withLineage) stamps every item it remembers. Inside a session, linked_memory() (TypeScript: linkedMemory) stamps the session's agent as the producer and the record the session acts on as the origin. Call acting_on(source) (TypeScript: actingOn) on the session first to name that record.

const facts = laser
  .context(conversation)
  .memory("facts")
  .withLineage(undefined, { name: "fact-extractor", version: "2" })
await facts.remember(new TextEncoder().encode("the reporter prefers email")).send()

await session.linkedMemory().remember(new TextEncoder().encode("ticket 42 is a login bug")).send()
use laser_sdk::wire::graph::ProducerInfo;

let facts = laser.context(conversation).memory("facts").with_lineage(
    None,
    Some(ProducerInfo {
        name: "fact-extractor".into(),
        version: "2".into(),
    }),
);
facts.remember("the reporter prefers email".as_bytes()).send().await?;

session
    .linked_memory()
    .remember("ticket 42 is a login bug".as_bytes())
    .send()
    .await?;
facts = (
    laser.context(conversation)
    .memory("facts")
    .with_lineage(producer={"name": "fact-extractor", "version": "2"})
)
await facts.remember("the reporter prefers email")

await session.linked_memory().remember("ticket 42 is a login bug")

A recalled item also carries source, the memory record it was folded from (stream, topic, partition, and offset), so a reader can open the original record while the log still holds it. A message source can carry the topic creation time, so a reader can tell a recreated topic from the original. In Python, source is a tuple (stream, topic, partition, offset, generation, conversation). A managed deployment also links an item to the session that wrote it, and record_retrieval(query, items) (TypeScript: recordRetrieval) on a session records which items entered its context.

Item fields

FieldHolds
idThe item's ID, a ULID unless it is a content ID
payloadThe body. Read it as text with text() in Rust and Python or memoryItemText(item) in TypeScript, or as JSON with json() or memoryItemJson(item)
kindThe item kind
provenanceThe conversation and agent it was stored under
scoreThe ranking score, empty for unranked recall
signalsThe per-strategy rank and score behind a ranked result
sourceThe memory record it was folded from
origin, producerLineage, when the item was written with it. Managed recall does not return them

Named items

For facts you address directly, use named items instead of a recall search. All three SDKs call set(key, body), fetch(key), update(key, patch), and remove(key) on the memory handle. Each write is a record on the memory topic, stored under <namespace>/<key>.

update applies a JSON merge patch (RFC 7386): fields in the patch overwrite and null removes. It reads the current value by folding the topic, so it works with Iggy alone. fetch reads the managed key-value view. fetch_folded(key) (TypeScript: fetchFolded) folds the topic in your process instead. remove of an absent key succeeds. Vector and custom handles return an unsupported error for named items.

The LogMemory class carries the same operations as set_named, fetch_named, fetch_named_folded, update_named, and forget_named (TypeScript: camelCase).

Context blocks

context(..) recalls a conversation's items and formats them as one prompt block. Rust takes context(conversation, Some(token_budget)), Python takes context(conversation, token_budget=), and TypeScript takes context(scope, { tokenBudget }). The recall builder renders the same block with .block(Some(budget)) in Rust, .block(budget) in TypeScript, and recall(.., token_budget=n, block=True) in Python.

The token estimate is the byte count divided by four, rounded up, in every SDK. When the budget runs out, the block drops the remaining items and appends a marker such as [... 3 more recalled item(s) omitted ...]. The first item is always kept, even when it alone exceeds the budget.

context and a plain block use default recall, so a log handle needs the managed view. With Iggy alone, use .folded().block(Some(budget)) in Rust, .folded().block(budget) in TypeScript, or recall(.., folded=True, token_budget=budget, block=True) in Python. To format items you already hold, call to_context_block(items, token_budget) (TypeScript: toContextBlock(items, tokenBudget)).

In Python, token_budget= without block=True is passed to the backend as a hint and does not trim the returned list.

Consolidation

consolidate(scope, max_items) runs one pass over a scope. The default pass keeps the newest max_items items in arrival order and forgets the rest. It examines at most 10,000 items per pass, and a limit of zero prunes that whole window. Python takes the scope as keywords: consolidate(max_items, conversation=, agent=, ..).

The pass returns a ConsolidationReport with four counts: summarized, reweighted, pruned, and derived. The default pass fills only pruned, and it counts an item after the backend accepts the forget. reweighted and derived are for custom consolidators.

A summarizer adds a step before pruning. It groups the Message items by conversation and writes one Summary per conversation with a durable lifetime. With prune_summarized, the pass then forgets the messages it folded and counts them in pruned. Your application supplies the summarizer, usually a model call. The SDK ships no model client.

ControlTypeScriptRustPython
Run a passconsolidate(scope, maxItems, options)consolidate(&scope, max_items)consolidate(max_items, ..)
Add a summarizer{ summarizer }consolidate_with(&scope, max_items, summarizer, prune_summarized)summarizer=
Forget folded messages{ pruneSummarized: true }prune_summarized set to true in consolidate_withprune_summarized=True

The TypeScript summarizer is an object with summarize(bodies). Rust implements the Summarizer trait from laser_sdk::memory. Python passes a callable from a list of bytes to a body, or an object with summarize(bodies), synchronous or async. Scoped handles take the same options, and Rust spells it consolidate_with(max_items, summarizer, prune_summarized) on a scoped handle. DefaultConsolidator::new(&memory, max).with_summarizer(..).prune_summarized() builds the same pass as a Rust Consolidator for an agent timer.

Consolidation reads through default recall, so a log handle needs the managed view. A vector handle consolidates in process. A custom backend without a typed append receives the summary through its remember.

const report = await memory.consolidate({ conversation }, 100, {
  summarizer: {
    summarize: async (bodies) => new TextEncoder().encode(`${bodies.length} turns about node-7`)
  },
  pruneSummarized: true
})
console.log(report.summarized, report.pruned)
use laser_sdk::memory::Summarizer;
use laser_sdk::prelude::full::*;

struct Digest;

impl Summarizer for Digest {
    async fn summarize(&self, bodies: Vec<Vec<u8>>) -> Result<Vec<u8>, LaserError> {
        Ok(format!("{} turns about node-7", bodies.len()).into_bytes())
    }
}

let scope = MemoryScope::builder().conversation(conversation).build();
let consolidator = DefaultConsolidator::new(&memory, 100)
    .with_summarizer(Digest)
    .prune_summarized();
let report = consolidator.consolidate(&scope).await?;
println!("{} summarized, {} pruned", report.summarized, report.pruned);
report = await memory.consolidate(
    100,
    conversation=conversation,
    summarizer=lambda bodies: f"{len(bodies)} turns about node-7".encode(),
    prune_summarized=True,
)
print(report.summarized, report.pruned)

Agent turns and periodic consolidation

MemoryHandler wraps an agent handler and remembers each message the handler finishes. After the inner handler succeeds, it remembers the payload under the message's conversation and agent. Remembering stays off until you choose a kind with auto_remember(kind) (TypeScript: autoRemember(kind)). A handler that fails remembers nothing and follows the normal retry path. A failed memory write is dropped, so it never repeats the handler's effects.

An agent can also consolidate on a timer, off the handler loop. Set both an interval and a consolidator, because one alone does nothing. The first pass runs when the agent starts and one more runs each interval. The consolidator owns its memory handle. Its scope names the agent ID, or is empty when a Python or TypeScript agent has no ID. A failed pass does not stop later ones. Shutdown stops the timer. TypeScript and Python reject an interval that is not positive. In Rust, a zero interval makes the agent fail with a configuration error, which ready() returns. Agents covers the agent lifecycle.

ControlTypeScriptRustPython
IntervalconsolidateEvery(intervalMs)consolidate_every(Duration)consolidate_every_ms=
Consolidatorconsolidator(obj), an object with consolidate(scope)consolidator(SharedConsolidator)consolidator=, an object with consolidate(scope)
import { Agent, AgentId, AgentTopic, MemoryHandler, MemoryKind } from "@laserdata/laser-sdk"

const memory = laser.memory("triage")
const handler = new MemoryHandler(inner, memory).autoRemember(MemoryKind.Message)

const agent = Agent.builder()
  .id(AgentId.new("triage"))
  .listenOn(AgentTopic.Sessions)
  .handler(handler)
  .consolidateEvery(60_000)
  .consolidator({ consolidate: (scope) => memory.consolidate(scope, 500) })
  .build()
  .spawn(laser)
use laser_sdk::memory::SharedConsolidator;
use laser_sdk::prelude::full::*;
use std::time::Duration;

let handler = MemoryHandler::new(inner, laser.memory("triage"))
    .auto_remember(MemoryKind::Message);
let consolidator =
    SharedConsolidator::new(DefaultConsolidator::new(laser.memory("triage"), 500));

let agent = Agent::builder()
    .id("triage".parse()?)
    .listen_on(AgentTopic::Sessions)
    .handler(handler)
    .consolidate_every(Duration::from_secs(60))
    .consolidator(consolidator)
    .build()
    .spawn(laser.clone());
memory = laser.memory("triage")
handler = ls.MemoryHandler(inner, memory).auto_remember("message")


class Compact:
    async def consolidate(self, scope):
        await memory.consolidate(500)


agent = laser.spawn_agent(
    "triage",
    ls.AgentTopic.Sessions,
    handler,
    consolidate_every_ms=60_000,
    consolidator=Compact(),
)
await agent.ready()

inner is your own handler.

Custom backends

memory_custom(backend) (TypeScript: memoryCustom) plugs in your own store. The backend owns its scoping and ID minting. It implements remember, recall, improve, and forget, and it can add a typed append. Without append, the SDK falls back to remember and your backend picks the ID and kind. The handle's backend reports log, vector, or custom.

In Rust, implement the Memory trait and pass an Arc to memory_custom. In TypeScript, pass an object that implements the Memory interface. In Python, pass an object with remember(scope, payload), recall(scope, query), improve(scope, feedback), forget(scope, id), and optionally append(scope, id, kind, payload). Each method returns directly or through an awaitable, and scopes, queries, and feedback arrive as dicts. recall returns MemoryItem objects or dicts with a payload and optional id, kind, score, and conversation. Synchronous Python methods run on an SDK worker thread, so they must not block. Python also builds a handle without a connection with MemoryHandle.custom(backend).

Governance and access

Memory writes from a handle built on a Laser pass through the connection's action governor. That covers log and vector handles. The policy sees the proposed item body. A forget or feedback write presents its encoded record. A standalone vector handle has no connection, so no policy applies. See Governance.

The managed read view is shared across principals. Reading needs kv:read on the memory namespace, and writing needs permission to publish to the memory topic. A conversation filter narrows a read but does not grant or limit access. Sharing a namespace between sessions does not merge their histories.

Requirements

Run bootstrap(partitions, retention) once per stream before you write to the default topic. In Rust, memory needs the agent feature, and default recall and fetch also need kv, which managed includes. Without kv, those calls return an unsupported error.

Works with Apache Iggy aloneNeeds Laser Stack or LaserData Cloud
Remember, improve, forget, typed appendDefault recall
Folded recall and fetch_foldedfetch
Named set, update, removecontext and block without folding
Vector and custom handlesConsolidation on a log handle

Language differences

  • TypeScript uses camelCase names and Uint8Array payloads.
  • Python passes scope and options as keywords (agent=, conversation=, user=, application=, kind=, durable=, dedup=, origin=, producer=) instead of builder calls, and accepts str or bytes payloads.
  • Python improve(target, weight, ..) and forget(id, ..) take the scope as keywords. Rust and TypeScript take a scope and a Feedback or ID.

Key operations

CallWhat it does
laser.memory(namespace)Open log-backed memory in a namespace
laser.memory_with(namespace, backend)Choose the log or vector backend. TypeScript adds the embedder as a third argument and Python as embedder=
remember(payload)Store an item, with .scope, .agent, .stream, .user, .application, .kind, .dedup, .durable, .origin, .producer
append(..)Write an item under an ID and kind you supply
recall(..) + .recent(), .semantic(t), .keyword(t), .hybrid(t)Choose a recall strategy
.folded()Read the memory topic in process, with Iggy alone
.limit(n)Cap the result, 50 by default
.embedder(..), .reranker(..)Attach an embedding function or a reranking pass
improve(..)Add a feedback weight to an item
forget(..)Write a tombstone for an item
memory_on_topic(topic), memory_topic(topic)Use your own memory topic
memory_custom(backend)Use your own backend
LogMemory, VectorMemory, RerankedMemoryBuild a backend directly
set, fetch, fetch_folded, update, removeRead and write named items
context(..), block(budget), to_context_block(..)Render items as one prompt block
consolidate(..)Prune a scope, optionally with a summarizer
MemoryHandler + auto_remember(kind)Remember each successful agent turn
consolidate_every + consolidatorConsolidate on a timer from the agent builder
fuse_reciprocal_rank(lists, limit)Merge ranked lists

On this page