Skip to content

First 10 Minutes

Get MemoryLayer running and store your first memories in three commands.

1. Start the server

Pick the embedding provider that fits your environment. For a no-API-key local smoke test, the mock provider works out of the box:

Terminal window
pip install memorylayer-server memorylayer-client
MEMORYLAYER_EMBEDDING_PROVIDER=mock memorylayer serve

For real semantic search, use OpenAI or Google embeddings:

Terminal window
pip install "memorylayer-server[openai]" memorylayer-client
export MEMORYLAYER_EMBEDDING_PROVIDER=openai
export MEMORYLAYER_EMBEDDING_OPENAI_API_KEY=sk-...
memorylayer serve

Or self-host with a memorylayer-embed-server peer (the default embed_server provider points at it):

Terminal window
pip install "memorylayer-embed-server[gpu]"
memorylayer-embed serve --port 61051 &
pip install memorylayer-server memorylayer-client
memorylayer serve # MEMORYLAYER_EMBED_SERVER_URL defaults to http://localhost:61051

The server starts on port 61001. Verify it’s up:

Terminal window
curl http://localhost:61001/health
{"status": "ok", "version": "0.1.22"}

2. Store memories

Store three memories of different types — a preference, a fact, and a procedure.

curl:

Terminal window
# Preference (semantic)
curl -s -X POST http://localhost:61001/v1/memories \
-H "Content-Type: application/json" \
-d '{
"content": "User prefers Python for backend and TypeScript for frontend",
"type": "semantic",
"subtype": "preference",
"importance": 0.8,
"tags": ["preferences", "languages"]
}'
# Architectural decision (semantic)
curl -s -X POST http://localhost:61001/v1/memories \
-H "Content-Type: application/json" \
-d '{
"content": "We use PostgreSQL for the main database; SQLite for local dev",
"type": "semantic",
"subtype": "decision",
"importance": 0.9,
"tags": ["architecture", "database"]
}'
# Deploy procedure (procedural)
curl -s -X POST http://localhost:61001/v1/memories \
-H "Content-Type: application/json" \
-d '{
"content": "To deploy: run npm run build, then railway up --detach",
"type": "procedural",
"subtype": "workflow",
"importance": 0.7,
"tags": ["deployment"]
}'

Python:

import asyncio
from memorylayer import MemoryLayerClient, MemoryType, MemorySubtype
async def main():
async with MemoryLayerClient(base_url="http://localhost:61001") as client:
m1 = await client.remember(
content="User prefers Python for backend and TypeScript for frontend",
type=MemoryType.SEMANTIC,
subtype=MemorySubtype.PREFERENCE,
importance=0.8,
tags=["preferences", "languages"],
)
m2 = await client.remember(
content="We use PostgreSQL for the main database; SQLite for local dev",
type=MemoryType.SEMANTIC,
subtype=MemorySubtype.DECISION,
importance=0.9,
tags=["architecture", "database"],
)
m3 = await client.remember(
content="To deploy: run npm run build, then railway up --detach",
type=MemoryType.PROCEDURAL,
subtype=MemorySubtype.WORKFLOW,
importance=0.7,
tags=["deployment"],
)
asyncio.run(main())

Each call returns the stored memory:

{
"memory": {
"id": "mem_01j8x4k9pz3qrv2n5w7c",
"workspace_id": "_default",
"content": "User prefers Python for backend and TypeScript for frontend",
"type": "semantic",
"subtype": "preference",
"importance": 0.8,
"tags": ["preferences", "languages"],
"status": "active",
"created_at": "2026-04-02T09:14:22Z"
}
}

3. Recall memories

Query with natural language. The server runs vector similarity search and returns ranked results with relevance scores.

curl:

Terminal window
curl -s -X POST http://localhost:61001/v1/memories/recall \
-H "Content-Type: application/json" \
-d '{
"query": "What programming languages and tools does this project use?",
"limit": 5
}'

Python:

results = await client.recall(
query="What programming languages and tools does this project use?",
limit=5,
)
for mem in results.memories:
print(f"[{mem.relevance_score:.2f}] {mem.content}")

Response:

{
"memories": [
{
"id": "mem_01j8x4k9pz3qrv2n5w7c",
"content": "User prefers Python for backend and TypeScript for frontend",
"type": "semantic",
"subtype": "preference",
"importance": 0.8,
"relevance_score": 0.94
},
{
"id": "mem_01j8x4m2qr7yw3n8v1d",
"content": "We use PostgreSQL for the main database; SQLite for local dev",
"type": "semantic",
"subtype": "decision",
"importance": 0.9,
"relevance_score": 0.81
}
],
"total_count": 2,
"search_latency_ms": 28,
"mode_used": "rag"
}

4. Associate two memories

Link the database decision to the language preference so graph traversal can surface them together.

curl:

Terminal window
# Replace {source_id} and {target_id} with the IDs from step 2
curl -s -X POST http://localhost:61001/v1/memories/{source_id}/associate \
-H "Content-Type: application/json" \
-d '{
"target_id": "{target_id}",
"relationship": "RELATED_TO",
"strength": 0.8
}'

Python:

from memorylayer import RelationshipType
assoc = await client.associate(
source_id=m1.id,
target_id=m2.id,
relationship=RelationshipType.RELATED_TO,
strength=0.8,
)

Response:

{
"association": {
"id": "asc_01j8x5r3pt2mn4v6w8y",
"source_id": "mem_01j8x4k9pz3qrv2n5w7c",
"target_id": "mem_01j8x4m2qr7yw3n8v1d",
"relationship_type": "RELATED_TO",
"strength": 0.8
}
}

5. Recall with graph expansion

Add include_associations: true and the server traverses the knowledge graph — memories linked to your top hits are pulled in automatically.

curl:

Terminal window
curl -s -X POST http://localhost:61001/v1/memories/recall \
-H "Content-Type: application/json" \
-d '{
"query": "language preferences",
"limit": 5,
"include_associations": true,
"traverse_depth": 1
}'

Python:

results = await client.recall(
query="language preferences",
limit=5,
include_associations=True,
traverse_depth=1,
)
for mem in results.memories:
scope = " (graph)" if mem.source_scope == "graph_expansion" else ""
print(f"[{mem.relevance_score:.2f}]{scope} {mem.content}")

Response — the database decision appears via graph expansion even though the query didn’t mention databases:

{
"memories": [
{
"id": "mem_01j8x4k9pz3qrv2n5w7c",
"content": "User prefers Python for backend and TypeScript for frontend",
"relevance_score": 0.94,
"source_scope": "same_context"
},
{
"id": "mem_01j8x4m2qr7yw3n8v1d",
"content": "We use PostgreSQL for the main database; SQLite for local dev",
"relevance_score": 0.76,
"source_scope": "graph_expansion"
}
],
"total_count": 2,
"search_latency_ms": 34,
"mode_used": "rag"
}

What you just did

In five steps you:

  1. Started a local server with embedded vector search (no API key required)
  2. Stored memories with cognitive types (semantic, procedural) and domain subtypes (preference, decision, workflow)
  3. Retrieved them with a natural language query and got back relevance scores
  4. Created a typed association between two memories
  5. Ran a graph-expanded recall that surfaced the linked memory automatically

Next steps

  • Core Concepts — memory types, importance scoring, workspaces, and the knowledge graph
  • Memory Types Reference — all cognitive types, subtypes, and recall modes
  • MCP Integration — connect Claude Code or Claude Desktop to your memory store
  • Installation — Docker setup, embedding providers, and SDK options