Skip to content

Server Overview

The MemoryLayer server (memorylayer-server) is a FastAPI-based HTTP server that provides the core memory infrastructure. It handles memory storage, vector search, relationship graphs, session management, and embedding generation.

Features

  • REST API — Full-featured HTTP API for all memory operations
  • MCP Compatible — Pairs with the separate MCP server package (@scitrera/memorylayer-mcp-server) for Claude and other LLMs
  • Vector Search — SQLite + sqlite-vec for efficient similarity search (Turso/libSQL also supported)
  • Embedding Providers — embed_server (default; delegates to a memorylayer-embed-server peer for self-hosted text/multi-vector), openai, google, and mock (testing)
  • Cognitive Memory Types — Episodic, semantic, procedural, and working memory
  • Knowledge Graph — 63 typed relationship types across 11 categories for memory associations
  • Context Environment — Server-side Python sandbox with RLM (Recursive Language Model) for memory analysis without consuming the client’s context window
  • Session Management — Working memory with TTL, token-budget-aware extraction, and commit to long-term storage

Installation

Terminal window
# Basic install (you still need an embedding provider — see below)
pip install memorylayer-server
# OpenAI embeddings (cloud)
pip install memorylayer-server[openai]
# Google GenAI embeddings (cloud)
pip install memorylayer-server[google]
# Cloud embedding extras bundled (openai + google)
pip install memorylayer-server[embeddings]
# All optional dependencies (cloud embeddings + LLMs + document parsers)
pip install memorylayer-server[all]

For self-hosted embeddings on GPU, install the separate memorylayer-embed-server peer:

Terminal window
pip install "memorylayer-embed-server[gpu]"
memorylayer-embed serve --port 61051

Then point the core server at it via MEMORYLAYER_EMBED_SERVER_URL=http://embed-host:61051.

Removed in v0.1.x: the in-process local (sentence-transformers), colpali, and qwen3-vl providers were removed from memorylayer-server. All self-hosted embedding now goes through the memorylayer-embed-server peer using the embed_server provider.

Quick Start

Terminal window
# Start the HTTP server
memorylayer serve --port 61001

API Usage

from memorylayer import MemoryLayerClient
client = MemoryLayerClient(base_url="http://localhost:61001")
# Store a memory
memory = await client.remember(
content="User prefers Python for backend development",
type="semantic",
importance=0.8,
tags=["preferences", "programming"]
)
# Recall memories
results = await client.recall(
query="What programming languages does the user like?",
limit=5
)

Architecture

The server is organized into these layers:

┌──────────────────────────────────────────────┐
│ API Layer (FastAPI routes) │
│ - /v1/memories (recall, reflect, associate) │
│ - /v1/sessions, /v1/workspaces │
│ - /v1/context/*, /health │
├──────────────────────────────────────────────┤
│ Service Layer │
│ - MemoryService, SessionService │
│ - WorkspaceService, AssociationService │
│ - ContradictionService, EmbeddingService │
│ - ContextEnvironmentService, LLMService │
├──────────────────────────────────────────────┤
│ Storage Layer │
│ - SQLite + sqlite-vec │
│ - Embedding providers │
└──────────────────────────────────────────────┘

Source Code

The server source code is located at memorylayer-core-python/ and published to PyPI as memorylayer-server.