First 10 Minutes
Get MemoryLayer running and store your first memories in three commands.
1. Start the server
Pick the embedding provider that fits your environment. For a no-API-key local smoke test, the mock provider works out of the box:
pip install memorylayer-server memorylayer-clientMEMORYLAYER_EMBEDDING_PROVIDER=mock memorylayer serveFor real semantic search, use OpenAI or Google embeddings:
pip install "memorylayer-server[openai]" memorylayer-clientexport MEMORYLAYER_EMBEDDING_PROVIDER=openaiexport MEMORYLAYER_EMBEDDING_OPENAI_API_KEY=sk-...memorylayer serveOr self-host with a memorylayer-embed-server peer (the default embed_server provider points at it):
pip install "memorylayer-embed-server[gpu]"memorylayer-embed serve --port 61051 &
pip install memorylayer-server memorylayer-clientmemorylayer serve # MEMORYLAYER_EMBED_SERVER_URL defaults to http://localhost:61051The server starts on port 61001. Verify it’s up:
curl http://localhost:61001/health{"status": "ok", "version": "0.1.22"}2. Store memories
Store three memories of different types — a preference, a fact, and a procedure.
curl:
# Preference (semantic)curl -s -X POST http://localhost:61001/v1/memories \ -H "Content-Type: application/json" \ -d '{ "content": "User prefers Python for backend and TypeScript for frontend", "type": "semantic", "subtype": "preference", "importance": 0.8, "tags": ["preferences", "languages"] }'
# Architectural decision (semantic)curl -s -X POST http://localhost:61001/v1/memories \ -H "Content-Type: application/json" \ -d '{ "content": "We use PostgreSQL for the main database; SQLite for local dev", "type": "semantic", "subtype": "decision", "importance": 0.9, "tags": ["architecture", "database"] }'
# Deploy procedure (procedural)curl -s -X POST http://localhost:61001/v1/memories \ -H "Content-Type: application/json" \ -d '{ "content": "To deploy: run npm run build, then railway up --detach", "type": "procedural", "subtype": "workflow", "importance": 0.7, "tags": ["deployment"] }'Python:
import asynciofrom memorylayer import MemoryLayerClient, MemoryType, MemorySubtype
async def main(): async with MemoryLayerClient(base_url="http://localhost:61001") as client: m1 = await client.remember( content="User prefers Python for backend and TypeScript for frontend", type=MemoryType.SEMANTIC, subtype=MemorySubtype.PREFERENCE, importance=0.8, tags=["preferences", "languages"], )
m2 = await client.remember( content="We use PostgreSQL for the main database; SQLite for local dev", type=MemoryType.SEMANTIC, subtype=MemorySubtype.DECISION, importance=0.9, tags=["architecture", "database"], )
m3 = await client.remember( content="To deploy: run npm run build, then railway up --detach", type=MemoryType.PROCEDURAL, subtype=MemorySubtype.WORKFLOW, importance=0.7, tags=["deployment"], )
asyncio.run(main())Each call returns the stored memory:
{ "memory": { "id": "mem_01j8x4k9pz3qrv2n5w7c", "workspace_id": "_default", "content": "User prefers Python for backend and TypeScript for frontend", "type": "semantic", "subtype": "preference", "importance": 0.8, "tags": ["preferences", "languages"], "status": "active", "created_at": "2026-04-02T09:14:22Z" }}3. Recall memories
Query with natural language. The server runs vector similarity search and returns ranked results with relevance scores.
curl:
curl -s -X POST http://localhost:61001/v1/memories/recall \ -H "Content-Type: application/json" \ -d '{ "query": "What programming languages and tools does this project use?", "limit": 5 }'Python:
results = await client.recall( query="What programming languages and tools does this project use?", limit=5,)
for mem in results.memories: print(f"[{mem.relevance_score:.2f}] {mem.content}")Response:
{ "memories": [ { "id": "mem_01j8x4k9pz3qrv2n5w7c", "content": "User prefers Python for backend and TypeScript for frontend", "type": "semantic", "subtype": "preference", "importance": 0.8, "relevance_score": 0.94 }, { "id": "mem_01j8x4m2qr7yw3n8v1d", "content": "We use PostgreSQL for the main database; SQLite for local dev", "type": "semantic", "subtype": "decision", "importance": 0.9, "relevance_score": 0.81 } ], "total_count": 2, "search_latency_ms": 28, "mode_used": "rag"}4. Associate two memories
Link the database decision to the language preference so graph traversal can surface them together.
curl:
# Replace {source_id} and {target_id} with the IDs from step 2curl -s -X POST http://localhost:61001/v1/memories/{source_id}/associate \ -H "Content-Type: application/json" \ -d '{ "target_id": "{target_id}", "relationship": "RELATED_TO", "strength": 0.8 }'Python:
from memorylayer import RelationshipType
assoc = await client.associate( source_id=m1.id, target_id=m2.id, relationship=RelationshipType.RELATED_TO, strength=0.8,)Response:
{ "association": { "id": "asc_01j8x5r3pt2mn4v6w8y", "source_id": "mem_01j8x4k9pz3qrv2n5w7c", "target_id": "mem_01j8x4m2qr7yw3n8v1d", "relationship_type": "RELATED_TO", "strength": 0.8 }}5. Recall with graph expansion
Add include_associations: true and the server traverses the knowledge graph — memories linked to your top hits are pulled in automatically.
curl:
curl -s -X POST http://localhost:61001/v1/memories/recall \ -H "Content-Type: application/json" \ -d '{ "query": "language preferences", "limit": 5, "include_associations": true, "traverse_depth": 1 }'Python:
results = await client.recall( query="language preferences", limit=5, include_associations=True, traverse_depth=1,)
for mem in results.memories: scope = " (graph)" if mem.source_scope == "graph_expansion" else "" print(f"[{mem.relevance_score:.2f}]{scope} {mem.content}")Response — the database decision appears via graph expansion even though the query didn’t mention databases:
{ "memories": [ { "id": "mem_01j8x4k9pz3qrv2n5w7c", "content": "User prefers Python for backend and TypeScript for frontend", "relevance_score": 0.94, "source_scope": "same_context" }, { "id": "mem_01j8x4m2qr7yw3n8v1d", "content": "We use PostgreSQL for the main database; SQLite for local dev", "relevance_score": 0.76, "source_scope": "graph_expansion" } ], "total_count": 2, "search_latency_ms": 34, "mode_used": "rag"}What you just did
In five steps you:
- Started a local server with embedded vector search (no API key required)
- Stored memories with cognitive types (
semantic,procedural) and domain subtypes (preference,decision,workflow) - Retrieved them with a natural language query and got back relevance scores
- Created a typed association between two memories
- Ran a graph-expanded recall that surfaced the linked memory automatically
Next steps
- Core Concepts — memory types, importance scoring, workspaces, and the knowledge graph
- Memory Types Reference — all cognitive types, subtypes, and recall modes
- MCP Integration — connect Claude Code or Claude Desktop to your memory store
- Installation — Docker setup, embedding providers, and SDK options