Memory in depth
Scopes, recall modes, vector ranking, lineage, named items, context blocks, consolidation, and custom backends for agent memory
This page is the full reference for Memory. Start there for the short version.
How memory is stored
laser.memory(namespace) returns a MemoryHandle. Every remember, improve, and forget appends a record to the memory topic. forget writes a tombstone and improve writes a feedback weight, so nothing is edited in place and the topic keeps the full history.
A managed deployment folds the memory topic into a versioned key-value read view. Default recall and fetch read that view, so they need Laser Stack or LaserData Cloud. Folded recall rebuilds memory from the topic inside your process instead, so it works with Apache Iggy alone.
When a forget or improve call names a conversation in its scope, it acts only on an item remembered in that conversation. A call without a conversation acts on the item in any conversation. A scoped memory always passes its own conversation, so it cannot forget or reweight another conversation's item.
Namespaces
A client with a default stream scopes the memory namespace to the stream its memory records ride, so laser.memory("notes") reads and writes stream:<stream>/notes. A name that already starts with stream: is sent as written. A client without a default stream sends the bare name. See Stream-scoped resource names for the opt out.
Every record carries its namespace in a header. Folded recall reads only records with the handle's namespace, and the managed view keys rows by namespace, so two namespaces on one topic never mix.
Scope
Each item has a scope. A scope names the stream, the agent, and the conversation the item belongs to, plus an optional user and application.
| Field | Remember (Rust, TypeScript) | Python keyword | MemoryScope field |
|---|---|---|---|
| Conversation | .scope(conversation) | conversation= | conversation |
| Agent | .agent(id) | agent= | agent |
| Stream | .stream(name) | stream= | stream |
| User | .user(user) | user= | user |
| Application | .application(app) | application= | app in Rust and TypeScript |
Recall narrows by every field you set. A field you leave out widens the read, so a recall without a conversation reads across conversations. TypeScript calls recall() with no argument, Rust calls recall(None), and Python omits conversation=.
A stream filter on a log handle returns nothing when it names a stream other than the one the records ride.
.durable() (Python: durable=True) sets the scope lifetime to durable. The default is session. Custom backends receive the lifetime. Built-in backends do not store it, and recall never filters on it.
Memory topic
By default memory records ride agent.memory on the connection's default stream. bootstrap(partitions, retention) creates that topic with the other agent topics.
memory_on_topic(topic, stream) (TypeScript: memoryOnTopic) uses a topic that already exists. The stream is optional and defaults to the connection's default stream: Rust takes an Option<&str>, TypeScript an optional second argument, and Python stream=. memory_topic(topic) (TypeScript: memoryTopic) creates the topic first and sets its partition count and message expiry. Both use the topic name as the namespace.
| Setting | TypeScript | Rust | Python |
|---|---|---|---|
| Stream | .stream(name) | .stream(name) | stream= |
| Partitions, default 1 | .partitions(count) | .partitions(count) | partitions= |
| Expiry, default 30 days | .ttl(ttlMs) | .ttl(Duration) | ttl_ms= in milliseconds |
| Never expire | .noExpiry() | .no_expiry() | ttl_ms=0 |
TypeScript and Rust finish the builder with .build(). Python awaits laser.memory_topic(topic, ..) directly.
Records of one conversation share a partition key, so each conversation keeps its write order on a multi-partition topic. Topic expiry and view retention are separate. Once a record expires from the topic, folded recall can no longer rebuild it, while the managed view keeps its own retention.
Recall
recall(conversation) starts a recall builder in Rust and TypeScript, and recall(None) in Rust or recall() in TypeScript recalls every conversation. Finish it with .fetch(). Python passes everything as keywords to recall(..). .limit(n) (Python: limit=) caps the result. The default is 50.
Managed and folded recall
Default recall reads the managed key-value view and returns the newest items first. Newest means the latest broker append time, with the source position breaking ties, so items from different partitions come back in arrival order. The view folds the topic in the background, so a fresh write can take a moment to appear.
.folded() (Python: folded=True) rebuilds memory from the topic in your process. Keep one handle for folded reads. A handle remembers what it has read and reads only new records on the next call. A fresh handle reads the whole topic on its first recall.
Feedback from improve reorders folded recall and vector recall for every strategy except recent. Managed recall always returns newest first.
Strategies
| Strategy | Rust and TypeScript | Python | Ranks by |
|---|---|---|---|
| Recent | .recent() | strategy="recent" | Newest first. Query text and feedback do not change it |
| Semantic | .semantic(text) | semantic=text | Embedding similarity to the text |
| Keyword | .keyword(text) | strategy="keyword", semantic=text | Exact terms, which suits names and identifiers |
| Hybrid | .hybrid(text) | strategy="hybrid", semantic=text | The sum of semantic, keyword, and feedback scores |
Python also accepts auto, graph, and temporal. Rust and TypeScript set any value with .strategy(..). auto is the default.
Only the vector backend ranks by text. Log memory uses the text as a filter: semantic, keyword, and hybrid recall keep items that share at least one word with the text, then apply the limit. Recent recall ignores the text.
Each item's signals records which strategy surfaced it, its rank within that signal (0 is best), and that signal's score. Feedback appears as an auto signal in every backend.
fuse_reciprocal_rank(signals, limit) (TypeScript: fuseReciprocalRank) merges several ranked lists into one. Each item scores the sum of 1 / (60 + rank) over the lists that contain it, so agreement between lists outranks one list's top hit.
Vector memory
The vector backend embeds each item when you remember it and ranks recall by cosine similarity. It keeps items in process memory and drops them when the handle goes away.
| Task | TypeScript | Rust | Python |
|---|---|---|---|
| Open a governed vector handle | laser.memoryWith(ns, MemoryBackend.Vector, embedder) | laser.memory_with(ns, MemoryBackend::Vector).embedder(e) | laser.memory_with(ns, "vector", embedder=e) |
| Same backend, built directly | VectorMemory.governed(laser, embedder) | VectorMemory::governed(laser, embedder) | VectorMemory.governed(laser, embedder) |
| Standalone, no connection, no governance | MemoryHandle.vector(embedder) or new VectorMemory(embedder) | MemoryHandle::vector(embedder) or VectorMemory::new(embedder) | MemoryHandle.vector(embedder) or VectorMemory(embedder) |
The embedder comes from your application. It is a Rust Embedder implementation, a TypeScript object with embed(text), or a Python callable or object with embed(text) that returns a list of floats, directly or through an awaitable. A synchronous Python embedder runs on an SDK worker thread, so it must not block.
Python and TypeScript take the embedder when the handle opens, so they check it there: the vector backend without an embedder, or another backend with one, is an invalid error. Rust attaches the embedder after opening with .embedder(..), so a Rust vector handle without one fails with a configuration error at the first remember or semantic recall. .embedder(..) on a vector handle builds a fresh, empty index over the new embedder, so attach it before you write. On a log or custom handle .embedder(..) changes nothing.
Keyword recall on the vector backend scores the share of query words found in the item and never calls the embedder. Without query text, vector recall returns the newest items, ranked by feedback when any exists.
Rerankers
A reranker is a second pass over the recalled candidates. Attach one with .reranker(..) on any handle, or wrap a backend in RerankedMemory (Rust: RerankedMemory::new(inner, reranker), Python: RerankedMemory(inner, reranker), TypeScript: new RerankedMemory(inner, reranker)). It runs only when the query carries text, and it also runs on folded recall. Writes pass through unchanged.
Rust takes a Reranker implementation. TypeScript takes an object with rerank(query, items). Python takes a callable (query, items) -> items or an object with rerank(query, items), synchronous or async.
import { ConversationId, MemoryBackend } from "@laserdata/laser-sdk"
const conversation = ConversationId.new()
const memory = laser.memoryWith("runbooks", MemoryBackend.Vector, embedder).reranker(reranker)
const encoder = new TextEncoder()
await memory.remember(encoder.encode("node-7 sits in the eu-west pool")).scope(conversation).send()
await memory.remember(encoder.encode("node-9 rotates keys monthly")).scope(conversation).send()
const hits = await memory.recall(conversation).semantic("which pool is node-7 in").limit(3).fetch()use laser_sdk::prelude::full::*;
let conversation = ConversationId::new();
let memory = laser
.memory_with("runbooks", MemoryBackend::Vector)
.embedder(embedder)
.reranker(reranker);
memory
.remember("node-7 sits in the eu-west pool".as_bytes())
.scope(conversation)
.send()
.await?;
memory
.remember("node-9 rotates keys monthly".as_bytes())
.scope(conversation)
.send()
.await?;
let hits = memory
.recall(conversation)
.semantic("which pool is node-7 in")
.limit(3)
.fetch()
.await?;import laser_sdk as ls
conversation = ls.new_conversation_id()
memory = laser.memory_with("runbooks", "vector", embedder=embed).reranker(rerank)
await memory.remember("node-7 sits in the eu-west pool", conversation=conversation)
await memory.remember("node-9 rotates keys monthly", conversation=conversation)
hits = await memory.recall(
conversation=conversation,
semantic="which pool is node-7 in",
limit=3,
)embedder and reranker are your own implementations. The runnable memory example ships a small bag-of-words embedder in each language.
Writing items
Kinds
remember(..) stores a Fact unless you set a kind with .kind(..) (Python: kind="summary" and the other lowercase names).
| Kind | Class | Holds |
|---|---|---|
Fact | Semantic | A standalone fact. The default |
Message | Episodic | A conversation turn or event |
Summary | Semantic | A summary distilled from other items |
Entity | Semantic | An extracted entity |
Feedback | Semantic | A ranking signal |
Procedure | Procedural | A reusable workflow |
MemoryKind::class() in Rust, memory_kind_class(kind) in Python, and memoryClass(kind) in TypeScript report the class. MemoryKind::code(), memory_kind_code(kind), and memoryKindCode(kind) return the stable one-byte code.
Content IDs
remember returns a random ID unless you add .dedup() (Python: dedup=True). With dedup the ID comes from the owner scope (stream, agent, user, and application), the kind, and the body. The conversation is not part of the ID, so remembering the same fact twice stores it once, and two users never share an ID. MemoryId::content(scope, kind, body) in Rust, memory_id_content(kind, body, stream=, agent=, user=, application=) in Python, and MemoryId.content(scope, kind, body) in TypeScript compute the same ID without writing anything. Every SDK produces the same ID for the same input. Graph uses the same approach for node IDs.
Typed append
Typed append writes an ID and a kind that you supply. Built-in backends keep both. Use it when the ID comes from somewhere else, such as a content ID. Rust calls the Memory::append trait method, Python uses append(id, payload, kind=, ..), and TypeScript uses append(scope, id, kind, payload).
import { AgentId, MemoryId, MemoryKind } from "@laserdata/laser-sdk"
const scope = { agent: AgentId.new("planner"), conversation }
const body = new TextEncoder().encode("node-7 rotates keys monthly")
const id = MemoryId.content(scope, MemoryKind.Entity, body)
await memory.append(scope, id, MemoryKind.Entity, body)let scope = MemoryScope::builder()
.agent("planner".parse()?)
.conversation(conversation)
.build();
let body = b"node-7 rotates keys monthly".to_vec();
let id = MemoryId::content(&scope, MemoryKind::Entity, &body);
Memory::append(&memory, &scope, id, MemoryKind::Entity, body).await?;body = b"node-7 rotates keys monthly"
memory_id = ls.memory_id_content("entity", body, agent="planner")
await memory.append(
memory_id,
body,
kind="entity",
agent="planner",
conversation=conversation,
)Lineage
An item can record where it came from. origin is the log record that motivated it, such as the message a fact was extracted from. producer names the agent, extraction policy, or consolidator that wrote it, with a name and a version. Both are optional, never filter recall, and never change the item's content ID.
Set them on one write with .origin(source) and .producer(info) on the remember builder (Python: origin= and producer=). with_lineage(origin, producer) on a scoped memory (TypeScript: withLineage) stamps every item it remembers. Inside a session, linked_memory() (TypeScript: linkedMemory) stamps the session's agent as the producer and the record the session acts on as the origin. Call acting_on(source) (TypeScript: actingOn) on the session first to name that record.
const facts = laser
.context(conversation)
.memory("facts")
.withLineage(undefined, { name: "fact-extractor", version: "2" })
await facts.remember(new TextEncoder().encode("the reporter prefers email")).send()
await session.linkedMemory().remember(new TextEncoder().encode("ticket 42 is a login bug")).send()use laser_sdk::wire::graph::ProducerInfo;
let facts = laser.context(conversation).memory("facts").with_lineage(
None,
Some(ProducerInfo {
name: "fact-extractor".into(),
version: "2".into(),
}),
);
facts.remember("the reporter prefers email".as_bytes()).send().await?;
session
.linked_memory()
.remember("ticket 42 is a login bug".as_bytes())
.send()
.await?;facts = (
laser.context(conversation)
.memory("facts")
.with_lineage(producer={"name": "fact-extractor", "version": "2"})
)
await facts.remember("the reporter prefers email")
await session.linked_memory().remember("ticket 42 is a login bug")A recalled item also carries source, the memory record it was folded from (stream, topic, partition, and offset), so a reader can open the original record while the log still holds it. A message source can carry the topic creation time, so a reader can tell a recreated topic from the original. In Python, source is a tuple (stream, topic, partition, offset, generation, conversation). A managed deployment also links an item to the session that wrote it, and record_retrieval(query, items) (TypeScript: recordRetrieval) on a session records which items entered its context.
Item fields
| Field | Holds |
|---|---|
id | The item's ID, a ULID unless it is a content ID |
payload | The body. Read it as text with text() in Rust and Python or memoryItemText(item) in TypeScript, or as JSON with json() or memoryItemJson(item) |
kind | The item kind |
provenance | The conversation and agent it was stored under |
score | The ranking score, empty for unranked recall |
signals | The per-strategy rank and score behind a ranked result |
source | The memory record it was folded from |
origin, producer | Lineage, when the item was written with it. Managed recall does not return them |
Named items
For facts you address directly, use named items instead of a recall search. All three SDKs call set(key, body), fetch(key), update(key, patch), and remove(key) on the memory handle. Each write is a record on the memory topic, stored under <namespace>/<key>.
update applies a JSON merge patch (RFC 7386): fields in the patch overwrite and null removes. It reads the current value by folding the topic, so it works with Iggy alone. fetch reads the managed key-value view. fetch_folded(key) (TypeScript: fetchFolded) folds the topic in your process instead. remove of an absent key succeeds. Vector and custom handles return an unsupported error for named items.
The LogMemory class carries the same operations as set_named, fetch_named, fetch_named_folded, update_named, and forget_named (TypeScript: camelCase).
Context blocks
context(..) recalls a conversation's items and formats them as one prompt block. Rust takes context(conversation, Some(token_budget)), Python takes context(conversation, token_budget=), and TypeScript takes context(scope, { tokenBudget }). The recall builder renders the same block with .block(Some(budget)) in Rust, .block(budget) in TypeScript, and recall(.., token_budget=n, block=True) in Python.
The token estimate is the byte count divided by four, rounded up, in every SDK. When the budget runs out, the block drops the remaining items and appends a marker such as [... 3 more recalled item(s) omitted ...]. The first item is always kept, even when it alone exceeds the budget.
context and a plain block use default recall, so a log handle needs the managed view. With Iggy alone, use .folded().block(Some(budget)) in Rust, .folded().block(budget) in TypeScript, or recall(.., folded=True, token_budget=budget, block=True) in Python. To format items you already hold, call to_context_block(items, token_budget) (TypeScript: toContextBlock(items, tokenBudget)).
In Python, token_budget= without block=True is passed to the backend as a hint and does not trim the returned list.
Consolidation
consolidate(scope, max_items) runs one pass over a scope. The default pass keeps the newest max_items items in arrival order and forgets the rest. It examines at most 10,000 items per pass, and a limit of zero prunes that whole window. Python takes the scope as keywords: consolidate(max_items, conversation=, agent=, ..).
The pass returns a ConsolidationReport with four counts: summarized, reweighted, pruned, and derived. The default pass fills only pruned, and it counts an item after the backend accepts the forget. reweighted and derived are for custom consolidators.
A summarizer adds a step before pruning. It groups the Message items by conversation and writes one Summary per conversation with a durable lifetime. With prune_summarized, the pass then forgets the messages it folded and counts them in pruned. Your application supplies the summarizer, usually a model call. The SDK ships no model client.
| Control | TypeScript | Rust | Python |
|---|---|---|---|
| Run a pass | consolidate(scope, maxItems, options) | consolidate(&scope, max_items) | consolidate(max_items, ..) |
| Add a summarizer | { summarizer } | consolidate_with(&scope, max_items, summarizer, prune_summarized) | summarizer= |
| Forget folded messages | { pruneSummarized: true } | prune_summarized set to true in consolidate_with | prune_summarized=True |
The TypeScript summarizer is an object with summarize(bodies). Rust implements the Summarizer trait from laser_sdk::memory. Python passes a callable from a list of bytes to a body, or an object with summarize(bodies), synchronous or async. Scoped handles take the same options, and Rust spells it consolidate_with(max_items, summarizer, prune_summarized) on a scoped handle. DefaultConsolidator::new(&memory, max).with_summarizer(..).prune_summarized() builds the same pass as a Rust Consolidator for an agent timer.
Consolidation reads through default recall, so a log handle needs the managed view. A vector handle consolidates in process. A custom backend without a typed append receives the summary through its remember.
const report = await memory.consolidate({ conversation }, 100, {
summarizer: {
summarize: async (bodies) => new TextEncoder().encode(`${bodies.length} turns about node-7`)
},
pruneSummarized: true
})
console.log(report.summarized, report.pruned)use laser_sdk::memory::Summarizer;
use laser_sdk::prelude::full::*;
struct Digest;
impl Summarizer for Digest {
async fn summarize(&self, bodies: Vec<Vec<u8>>) -> Result<Vec<u8>, LaserError> {
Ok(format!("{} turns about node-7", bodies.len()).into_bytes())
}
}
let scope = MemoryScope::builder().conversation(conversation).build();
let consolidator = DefaultConsolidator::new(&memory, 100)
.with_summarizer(Digest)
.prune_summarized();
let report = consolidator.consolidate(&scope).await?;
println!("{} summarized, {} pruned", report.summarized, report.pruned);report = await memory.consolidate(
100,
conversation=conversation,
summarizer=lambda bodies: f"{len(bodies)} turns about node-7".encode(),
prune_summarized=True,
)
print(report.summarized, report.pruned)Agent turns and periodic consolidation
MemoryHandler wraps an agent handler and remembers each message the handler finishes. After the inner handler succeeds, it remembers the payload under the message's conversation and agent. Remembering stays off until you choose a kind with auto_remember(kind) (TypeScript: autoRemember(kind)). A handler that fails remembers nothing and follows the normal retry path. A failed memory write is dropped, so it never repeats the handler's effects.
An agent can also consolidate on a timer, off the handler loop. Set both an interval and a consolidator, because one alone does nothing. The first pass runs when the agent starts and one more runs each interval. The consolidator owns its memory handle. Its scope names the agent ID, or is empty when a Python or TypeScript agent has no ID. A failed pass does not stop later ones. Shutdown stops the timer. TypeScript and Python reject an interval that is not positive. In Rust, a zero interval makes the agent fail with a configuration error, which ready() returns. Agents covers the agent lifecycle.
| Control | TypeScript | Rust | Python |
|---|---|---|---|
| Interval | consolidateEvery(intervalMs) | consolidate_every(Duration) | consolidate_every_ms= |
| Consolidator | consolidator(obj), an object with consolidate(scope) | consolidator(SharedConsolidator) | consolidator=, an object with consolidate(scope) |
import { Agent, AgentId, AgentTopic, MemoryHandler, MemoryKind } from "@laserdata/laser-sdk"
const memory = laser.memory("triage")
const handler = new MemoryHandler(inner, memory).autoRemember(MemoryKind.Message)
const agent = Agent.builder()
.id(AgentId.new("triage"))
.listenOn(AgentTopic.Sessions)
.handler(handler)
.consolidateEvery(60_000)
.consolidator({ consolidate: (scope) => memory.consolidate(scope, 500) })
.build()
.spawn(laser)use laser_sdk::memory::SharedConsolidator;
use laser_sdk::prelude::full::*;
use std::time::Duration;
let handler = MemoryHandler::new(inner, laser.memory("triage"))
.auto_remember(MemoryKind::Message);
let consolidator =
SharedConsolidator::new(DefaultConsolidator::new(laser.memory("triage"), 500));
let agent = Agent::builder()
.id("triage".parse()?)
.listen_on(AgentTopic::Sessions)
.handler(handler)
.consolidate_every(Duration::from_secs(60))
.consolidator(consolidator)
.build()
.spawn(laser.clone());memory = laser.memory("triage")
handler = ls.MemoryHandler(inner, memory).auto_remember("message")
class Compact:
async def consolidate(self, scope):
await memory.consolidate(500)
agent = laser.spawn_agent(
"triage",
ls.AgentTopic.Sessions,
handler,
consolidate_every_ms=60_000,
consolidator=Compact(),
)
await agent.ready()inner is your own handler.
Custom backends
memory_custom(backend) (TypeScript: memoryCustom) plugs in your own store. The backend owns its scoping and ID minting. It implements remember, recall, improve, and forget, and it can add a typed append. Without append, the SDK falls back to remember and your backend picks the ID and kind. The handle's backend reports log, vector, or custom.
In Rust, implement the Memory trait and pass an Arc to memory_custom. In TypeScript, pass an object that implements the Memory interface. In Python, pass an object with remember(scope, payload), recall(scope, query), improve(scope, feedback), forget(scope, id), and optionally append(scope, id, kind, payload). Each method returns directly or through an awaitable, and scopes, queries, and feedback arrive as dicts. recall returns MemoryItem objects or dicts with a payload and optional id, kind, score, and conversation. Synchronous Python methods run on an SDK worker thread, so they must not block. Python also builds a handle without a connection with MemoryHandle.custom(backend).
Governance and access
Memory writes from a handle built on a Laser pass through the connection's action governor. That covers log and vector handles. The policy sees the proposed item body. A forget or feedback write presents its encoded record. A standalone vector handle has no connection, so no policy applies. See Governance.
The managed read view is shared across principals. Reading needs kv:read on the memory namespace, and writing needs permission to publish to the memory topic. A conversation filter narrows a read but does not grant or limit access. Sharing a namespace between sessions does not merge their histories.
Requirements
Run bootstrap(partitions, retention) once per stream before you write to the default topic. In Rust, memory needs the agent feature, and default recall and fetch also need kv, which managed includes. Without kv, those calls return an unsupported error.
| Works with Apache Iggy alone | Needs Laser Stack or LaserData Cloud |
|---|---|
| Remember, improve, forget, typed append | Default recall |
Folded recall and fetch_folded | fetch |
Named set, update, remove | context and block without folding |
| Vector and custom handles | Consolidation on a log handle |
Language differences
- TypeScript uses camelCase names and
Uint8Arraypayloads. - Python passes scope and options as keywords (
agent=,conversation=,user=,application=,kind=,durable=,dedup=,origin=,producer=) instead of builder calls, and acceptsstrorbytespayloads. - Python
improve(target, weight, ..)andforget(id, ..)take the scope as keywords. Rust and TypeScript take a scope and aFeedbackor ID.
Key operations
| Call | What it does |
|---|---|
laser.memory(namespace) | Open log-backed memory in a namespace |
laser.memory_with(namespace, backend) | Choose the log or vector backend. TypeScript adds the embedder as a third argument and Python as embedder= |
remember(payload) | Store an item, with .scope, .agent, .stream, .user, .application, .kind, .dedup, .durable, .origin, .producer |
append(..) | Write an item under an ID and kind you supply |
recall(..) + .recent(), .semantic(t), .keyword(t), .hybrid(t) | Choose a recall strategy |
.folded() | Read the memory topic in process, with Iggy alone |
.limit(n) | Cap the result, 50 by default |
.embedder(..), .reranker(..) | Attach an embedding function or a reranking pass |
improve(..) | Add a feedback weight to an item |
forget(..) | Write a tombstone for an item |
memory_on_topic(topic), memory_topic(topic) | Use your own memory topic |
memory_custom(backend) | Use your own backend |
LogMemory, VectorMemory, RerankedMemory | Build a backend directly |
set, fetch, fetch_folded, update, remove | Read and write named items |
context(..), block(budget), to_context_block(..) | Render items as one prompt block |
consolidate(..) | Prune a scope, optionally with a summarizer |
MemoryHandler + auto_remember(kind) | Remember each successful agent turn |
consolidate_every + consolidator | Consolidate on a timer from the agent builder |
fuse_reciprocal_rank(lists, limit) | Merge ranked lists |