Architecture
System Overview
MemoryLayer is designed as infrastructure — a memory backend that agent frameworks consume, similar to how applications use PostgreSQL for relational data or Redis for caching.
┌─────────────────────────────────────────────────────────────┐│ Agent Frameworks (LangChain, LlamaIndex, Claude Code) │├─────────────────────────────────────────────────────────────┤│ Client SDKs (Python, TypeScript) │├─────────────────────────────────────────────────────────────┤│ MemoryLayer Server ││ ┌─────────────┬──────────────┬─────────────────────┐ ││ │ REST API │ MCP Server │ CLI │ ││ ├─────────────┴──────────────┴─────────────────────┤ ││ │ Service Layer │ ││ │ Memory · Reflect · Session · Workspace │ ││ │ Association · Context · Decay · Cache │ ││ ├───────────────────────────────────────────────────┤ ││ │ Storage Layer │ ││ │ SQLite + sqlite-vec │ Embedding Providers │ ││ └───────────────────────────────────────────────────┘ │└─────────────────────────────────────────────────────────────┘Component Descriptions
API Layer
The API layer handles all external communication. It is a FastAPI application exposing three interfaces:
- REST API — HTTP endpoints for all memory operations, consumed by the Python and TypeScript SDKs
- MCP Server — Model Context Protocol integration for AI assistants (Claude Code, Claude Desktop)
- CLI — Command-line interface for server management and administration
Key Endpoints
| Endpoint | Description |
|---|---|
POST /v1/memories | Store a new memory |
POST /v1/memories/recall | Semantic search across memories |
POST /v1/memories/reflect | LLM-powered synthesis across memories |
GET /v1/memories/{id} | Retrieve a specific memory |
POST /v1/sessions | Create a working memory session |
POST /v1/sessions/{id}/commit | Commit session to long-term storage |
POST /v1/workspaces | Create a workspace |
POST /v1/memories/{id}/associate | Create a relationship |
POST /v1/memories/{id}/traverse | Graph traversal from a memory |
POST /v1/context/execute | Execute code in a context sandbox |
Service Layer
Business logic is organized into focused, single-responsibility services. Each service is loaded as a plugin, making the architecture extensible.
| Service | Responsibility |
|---|---|
| MemoryService | Core remember/recall/forget operations. Coordinates embedding generation, deduplication, storage, and post-store pipelines (association, contradiction detection, tier generation). |
| ReflectService | LLM-powered synthesis that recalls relevant memories and generates a coherent reflection with source citations. |
| SessionService | Working memory lifecycle: create, touch/extend, commit, expire. Manages key-value working memory entries within sessions. |
| WorkspaceService | Workspace CRUD, tenant isolation, workspace settings, export/import. |
| AssociationService | Relationship graph management. Handles auto-association of similar memories and manual typed associations. |
| EmbeddingService | Pluggable embedding providers for generating vector representations. Supports multiple backend providers. |
| DeduplicationService | Content-hash and semantic duplicate detection. Returns SKIP, UPDATE, MERGE, or CREATE actions. |
| ExtractionService | Fact decomposition (splitting composite memories into atomic facts) and content classification. |
| DecayService | Time-based importance decay. Calculates access-based importance boosts and identifies archival candidates. |
| ContradictionService | Detects contradictory memories using semantic comparison. Supports resolution strategies (keep_a, keep_b, keep_both, merge). |
| SemanticTieringService | Generates abstract and overview tiers for memories, enabling progressive detail retrieval. |
| RerankerService | Cross-encoder reranking of recall candidates for improved relevance ordering. |
| CacheService | Query result caching for recall and association expansion. Invalidated on writes. |
| OntologyService | Manages the relationship type ontology (63 types across 11 categories). |
| ContextEnvironmentService | Server-side Python sandbox for code execution against loaded memory data. |
| TaskService | Background task scheduling for async operations (fact decomposition, auto-enrich, tier generation). |
Storage Layer
The storage layer provides persistent data storage with vector search capabilities.
- SQLite — Primary relational store for memories, sessions, workspaces, associations, and contradictions. Uses WAL mode for concurrent read performance.
- sqlite-vec — Vector extension for hardware-accelerated cosine similarity search. Falls back to Python-computed cosine similarity when unavailable.
- FTS5 — Full-text search index (Porter stemming) for keyword-based memory retrieval.
All data is stored in a single SQLite database file, making deployment and backup straightforward.
Embedding Service
The embedding service is a pluggable provider system for generating vector embeddings:
- Receives text content from the memory service
- Returns a fixed-dimension float vector
- Supports multiple backend providers (OpenAI, local models, etc.)
- Embeddings are stored as binary BLOBs alongside memory records
Session Manager
The session manager handles the working memory lifecycle:
- Create — Allocates a session with TTL, auto-creates workspace if needed
- Touch — Extends session TTL via sliding window
- Commit — Extracts important working memory entries, deduplicates, and stores as long-term memories
- Cleanup — Background task identifies expired sessions, triggers auto-commit, then deletes
Request Flow
Remember (Write Path)
Client → REST API (validate request) → MemoryService.remember() → Generate content hash → Generate embedding (EmbeddingService) → Check for duplicates (DeduplicationService) → SKIP: return existing memory → UPDATE: update existing memory → MERGE: merge content + re-embed → CREATE: store new memory → Store in SQLite (StorageBackend) → Post-store pipeline: → Cache invalidation → Semantic tier generation (abstract/overview) → Contradiction detection → Auto-association with similar memories → Fact decomposition (if content is multi-sentence) ← Return Memory objectRecall (Read Path)
Client → REST API (validate request) → MemoryService.recall() → Check recall cache → Classify query intent (rule-based: entity / temporal / event / general) → selects which retrieval arms may fire for this query → Retrieval arms, fused by Reciprocal Rank Fusion: → A. Dense vector search (sqlite-vec) → B. Keyword search (FTS5 / BM25) ← hybrid, on by default → D. Entity-anchored candidates → E. Fact-restricted candidates → G. Graph traversal (bounded) → cue-anchor similarity → Apply scope boosts (context, workspace, global) → Apply recency + backlink-salience boosts → Expand via association graph (BFS traversal) → Rerank candidates (RerankerService) → Apply detail level filtering (abstract/overview/full) → Increment access counts → Annotate trust scores, freshness, and match signals → Cache result ← Return RecallResultHybrid retrieval is the default, in open source. Recall is not vector
similarity alone: the dense and keyword arms are fused via RRF so that exact
terms, names, and identifiers — which embeddings routinely miss — are retrieved
alongside semantic matches. Additional arms (entity, fact, graph, cue) are gated
per query by intent classification rather than always running, and each result
carries match_signals explaining why it matched.
RecallMode values are RAG (active default), AGENTIC, and the deprecated
LLM / HYBRID.
Reflect (Synthesis Path)
Client → REST API (validate request) → ReflectService.reflect() → Recall relevant memories (via MemoryService) → Build LLM prompt with memory context → LLM synthesis (LLMService) ← Return ReflectResult with reflection + source IDsPlugin Architecture
MemoryLayer uses a plugin system for extensibility. Each service is loaded via a plugin factory that selects the appropriate implementation based on configuration:
Service Plugin Base ├── DefaultMemoryServicePlugin (built-in) ├── SqliteStorageBackendPlugin (built-in) ├── DefaultOntologyServicePlugin (built-in) └── ... (enterprise plugins)Plugins are resolved at startup via environment variables. For example, MEMORYLAYER_STORAGE_BACKEND=sqlite loads the SQLite plugin. This allows enterprise deployments to swap in PostgreSQL, Redis, or other backends without changing application code.
Enterprise Architecture Additions
The retrieval engine described above — hybrid search, the knowledge graph, reflection, consolidation, temporal recall, the context sandbox — is entirely open source. Enterprise does not unlock retrieval features; it replaces the substrate underneath them and adds the surface needed to run memory as shared production infrastructure.
| Component | Description |
|---|---|
| PostgreSQL + pgvector backend | Replaces SQLite for durability, concurrency, and horizontal scale |
| Redis cache | Distributed caching layer for recall results and association expansion |
| Background workers | Dedicated worker processes for fact decomposition, tier generation, and decay |
| Hot / warm / cold tiering | Lifecycle management that archives, restores, and warms memories by access pattern |
| Collections | General-purpose vector-search API for items outside the memory model |
| Datasets | Tabular upload, profiling, slicing, and dataset-derived memories |
| Document chat | Chat over ingested documents with page-level grounding |
| Trajectories | Recall traces for debugging and tuning retrieval behaviour in production |
| Admin & multi-tenancy | Cross-workspace stats, storage usage, jobs, users, applications, RBAC, audit, SSO/OIDC via Aether |
| Enterprise data connectors | Cloud, SaaS, and team-chat sources feeding the ingestion pipeline |
| MemoryLayer Storage | Content-addressed dedup + compression under document and tensor blobs |
See memorylayer.ai for enterprise options.
Client SDKs
Both SDKs wrap the REST API with idiomatic language constructs:
Python SDK (memorylayer-client)
- Async/await with
httpx - Context manager support
- Pydantic models for type safety
- Exception hierarchy mapping HTTP status codes
TypeScript SDK (@scitrera/memorylayer-sdk)
- Native
fetchAPI - Full TypeScript type definitions
- Promise-based async operations
- Typed error hierarchy
Go SDK (github.com/scitrera/memorylayer/memorylayer-sdk-go)
- Idiomatic
context.Context-aware client over the same REST API - Typed request/response structs
MCP Server
The MCP server (@scitrera/memorylayer-mcp-server) wraps the TypeScript SDK to provide MCP-compatible tools:
MCP Client (Claude) ←→ MCP Server ←→ TypeScript SDK ←→ REST API ←→ ServerIt provides 25 tools in the default cc profile (up to 37 in full), across these categories:
- 5 core/utility memory tools: remember, recall, reflect, forget, briefing
- 4 session management tools: session_start, session_end, session_commit, session_status
- 8 context environment tools: exec, inspect, load, inject, query, rlm, status, checkpoint
- 4 chat thread tools: create, append, get, list
- 4 skills tools: list, get, get_file, search
The full profile additionally exposes associate, statistics, graph_query, audit, chat_thread_decompose, chat_thread_delete, skills_save, and 5 MCP server registry tools.
Three-Layer Memory Hierarchy
Memories are organized in a three-layer hierarchy enabling progressive abstraction:
Category Layer (Aggregated Summaries) ↑ aggregatesItem Layer (Discrete Memory Units) ↑ extracted fromResource Layer (Raw Source Material)- Resource Layer — Raw conversations, documents, events
- Item Layer — Individual facts, preferences, decisions (the primary API surface)
- Category Layer — Auto-generated summaries that evolve based on content patterns
Data Flow
Remember (Write Path)
Client → REST API → MemoryService → Generate Embedding → Store in SQLiteRecall (Read Path)
Client → REST API → MemoryService → Vector Search (sqlite-vec) → Rank & Filter → ReturnReflect (Synthesis Path)
Client → REST API → ReflectService → Recall relevant memories → LLM synthesis → ReturnDeployment
Local Development
Single-process, single-file SQLite database. Zero configuration:
pip install memorylayer-servermemorylayer serveProduction (Enterprise)
The enterprise edition adds:
- PostgreSQL backend for scalability
- Redis caching layer
- Horizontal scaling
- Advanced analytics and reflection engine
See memorylayer.ai for enterprise options.