Skip to content

Architecture

System Overview

MemoryLayer is designed as infrastructure — a memory backend that agent frameworks consume, similar to how applications use PostgreSQL for relational data or Redis for caching.

┌─────────────────────────────────────────────────────────────┐
│ Agent Frameworks (LangChain, LlamaIndex, Claude Code) │
├─────────────────────────────────────────────────────────────┤
│ Client SDKs (Python, TypeScript) │
├─────────────────────────────────────────────────────────────┤
│ MemoryLayer Server │
│ ┌─────────────┬──────────────┬─────────────────────┐ │
│ │ REST API │ MCP Server │ CLI │ │
│ ├─────────────┴──────────────┴─────────────────────┤ │
│ │ Service Layer │ │
│ │ Memory · Reflect · Session · Workspace │ │
│ │ Association · Context · Decay · Cache │ │
│ ├───────────────────────────────────────────────────┤ │
│ │ Storage Layer │ │
│ │ SQLite + sqlite-vec │ Embedding Providers │ │
│ └───────────────────────────────────────────────────┘ │
└─────────────────────────────────────────────────────────────┘

Component Descriptions

API Layer

The API layer handles all external communication. It is a FastAPI application exposing three interfaces:

  • REST API — HTTP endpoints for all memory operations, consumed by the Python and TypeScript SDKs
  • MCP Server — Model Context Protocol integration for AI assistants (Claude Code, Claude Desktop)
  • CLI — Command-line interface for server management and administration

Key Endpoints

EndpointDescription
POST /v1/memoriesStore a new memory
POST /v1/memories/recallSemantic search across memories
POST /v1/memories/reflectLLM-powered synthesis across memories
GET /v1/memories/{id}Retrieve a specific memory
POST /v1/sessionsCreate a working memory session
POST /v1/sessions/{id}/commitCommit session to long-term storage
POST /v1/workspacesCreate a workspace
POST /v1/memories/{id}/associateCreate a relationship
POST /v1/memories/{id}/traverseGraph traversal from a memory
POST /v1/context/executeExecute code in a context sandbox

Service Layer

Business logic is organized into focused, single-responsibility services. Each service is loaded as a plugin, making the architecture extensible.

ServiceResponsibility
MemoryServiceCore remember/recall/forget operations. Coordinates embedding generation, deduplication, storage, and post-store pipelines (association, contradiction detection, tier generation).
ReflectServiceLLM-powered synthesis that recalls relevant memories and generates a coherent reflection with source citations.
SessionServiceWorking memory lifecycle: create, touch/extend, commit, expire. Manages key-value working memory entries within sessions.
WorkspaceServiceWorkspace CRUD, tenant isolation, workspace settings, export/import.
AssociationServiceRelationship graph management. Handles auto-association of similar memories and manual typed associations.
EmbeddingServicePluggable embedding providers for generating vector representations. Supports multiple backend providers.
DeduplicationServiceContent-hash and semantic duplicate detection. Returns SKIP, UPDATE, MERGE, or CREATE actions.
ExtractionServiceFact decomposition (splitting composite memories into atomic facts) and content classification.
DecayServiceTime-based importance decay. Calculates access-based importance boosts and identifies archival candidates.
ContradictionServiceDetects contradictory memories using semantic comparison. Supports resolution strategies (keep_a, keep_b, keep_both, merge).
SemanticTieringServiceGenerates abstract and overview tiers for memories, enabling progressive detail retrieval.
RerankerServiceCross-encoder reranking of recall candidates for improved relevance ordering.
CacheServiceQuery result caching for recall and association expansion. Invalidated on writes.
OntologyServiceManages the relationship type ontology (63 types across 11 categories).
ContextEnvironmentServiceServer-side Python sandbox for code execution against loaded memory data.
TaskServiceBackground task scheduling for async operations (fact decomposition, auto-enrich, tier generation).

Storage Layer

The storage layer provides persistent data storage with vector search capabilities.

  • SQLite — Primary relational store for memories, sessions, workspaces, associations, and contradictions. Uses WAL mode for concurrent read performance.
  • sqlite-vec — Vector extension for hardware-accelerated cosine similarity search. Falls back to Python-computed cosine similarity when unavailable.
  • FTS5 — Full-text search index (Porter stemming) for keyword-based memory retrieval.

All data is stored in a single SQLite database file, making deployment and backup straightforward.

Embedding Service

The embedding service is a pluggable provider system for generating vector embeddings:

  • Receives text content from the memory service
  • Returns a fixed-dimension float vector
  • Supports multiple backend providers (OpenAI, local models, etc.)
  • Embeddings are stored as binary BLOBs alongside memory records

Session Manager

The session manager handles the working memory lifecycle:

  1. Create — Allocates a session with TTL, auto-creates workspace if needed
  2. Touch — Extends session TTL via sliding window
  3. Commit — Extracts important working memory entries, deduplicates, and stores as long-term memories
  4. Cleanup — Background task identifies expired sessions, triggers auto-commit, then deletes

Request Flow

Remember (Write Path)

Client
→ REST API (validate request)
→ MemoryService.remember()
→ Generate content hash
→ Generate embedding (EmbeddingService)
→ Check for duplicates (DeduplicationService)
→ SKIP: return existing memory
→ UPDATE: update existing memory
→ MERGE: merge content + re-embed
→ CREATE: store new memory
→ Store in SQLite (StorageBackend)
→ Post-store pipeline:
→ Cache invalidation
→ Semantic tier generation (abstract/overview)
→ Contradiction detection
→ Auto-association with similar memories
→ Fact decomposition (if content is multi-sentence)
← Return Memory object

Recall (Read Path)

Client
→ REST API (validate request)
→ MemoryService.recall()
→ Check recall cache
→ Classify query intent (rule-based: entity / temporal / event / general)
→ selects which retrieval arms may fire for this query
→ Retrieval arms, fused by Reciprocal Rank Fusion:
→ A. Dense vector search (sqlite-vec)
→ B. Keyword search (FTS5 / BM25) ← hybrid, on by default
→ D. Entity-anchored candidates
→ E. Fact-restricted candidates
→ G. Graph traversal (bounded)
→ cue-anchor similarity
→ Apply scope boosts (context, workspace, global)
→ Apply recency + backlink-salience boosts
→ Expand via association graph (BFS traversal)
→ Rerank candidates (RerankerService)
→ Apply detail level filtering (abstract/overview/full)
→ Increment access counts
→ Annotate trust scores, freshness, and match signals
→ Cache result
← Return RecallResult

Hybrid retrieval is the default, in open source. Recall is not vector similarity alone: the dense and keyword arms are fused via RRF so that exact terms, names, and identifiers — which embeddings routinely miss — are retrieved alongside semantic matches. Additional arms (entity, fact, graph, cue) are gated per query by intent classification rather than always running, and each result carries match_signals explaining why it matched.

RecallMode values are RAG (active default), AGENTIC, and the deprecated LLM / HYBRID.

Reflect (Synthesis Path)

Client
→ REST API (validate request)
→ ReflectService.reflect()
→ Recall relevant memories (via MemoryService)
→ Build LLM prompt with memory context
→ LLM synthesis (LLMService)
← Return ReflectResult with reflection + source IDs

Plugin Architecture

MemoryLayer uses a plugin system for extensibility. Each service is loaded via a plugin factory that selects the appropriate implementation based on configuration:

Service Plugin Base
├── DefaultMemoryServicePlugin (built-in)
├── SqliteStorageBackendPlugin (built-in)
├── DefaultOntologyServicePlugin (built-in)
└── ... (enterprise plugins)

Plugins are resolved at startup via environment variables. For example, MEMORYLAYER_STORAGE_BACKEND=sqlite loads the SQLite plugin. This allows enterprise deployments to swap in PostgreSQL, Redis, or other backends without changing application code.

Enterprise Architecture Additions

The retrieval engine described above — hybrid search, the knowledge graph, reflection, consolidation, temporal recall, the context sandbox — is entirely open source. Enterprise does not unlock retrieval features; it replaces the substrate underneath them and adds the surface needed to run memory as shared production infrastructure.

ComponentDescription
PostgreSQL + pgvector backendReplaces SQLite for durability, concurrency, and horizontal scale
Redis cacheDistributed caching layer for recall results and association expansion
Background workersDedicated worker processes for fact decomposition, tier generation, and decay
Hot / warm / cold tieringLifecycle management that archives, restores, and warms memories by access pattern
CollectionsGeneral-purpose vector-search API for items outside the memory model
DatasetsTabular upload, profiling, slicing, and dataset-derived memories
Document chatChat over ingested documents with page-level grounding
TrajectoriesRecall traces for debugging and tuning retrieval behaviour in production
Admin & multi-tenancyCross-workspace stats, storage usage, jobs, users, applications, RBAC, audit, SSO/OIDC via Aether
Enterprise data connectorsCloud, SaaS, and team-chat sources feeding the ingestion pipeline
MemoryLayer StorageContent-addressed dedup + compression under document and tensor blobs

See memorylayer.ai for enterprise options.

Client SDKs

Both SDKs wrap the REST API with idiomatic language constructs:

Python SDK (memorylayer-client)

  • Async/await with httpx
  • Context manager support
  • Pydantic models for type safety
  • Exception hierarchy mapping HTTP status codes

TypeScript SDK (@scitrera/memorylayer-sdk)

  • Native fetch API
  • Full TypeScript type definitions
  • Promise-based async operations
  • Typed error hierarchy

Go SDK (github.com/scitrera/memorylayer/memorylayer-sdk-go)

  • Idiomatic context.Context-aware client over the same REST API
  • Typed request/response structs

MCP Server

The MCP server (@scitrera/memorylayer-mcp-server) wraps the TypeScript SDK to provide MCP-compatible tools:

MCP Client (Claude) ←→ MCP Server ←→ TypeScript SDK ←→ REST API ←→ Server

It provides 25 tools in the default cc profile (up to 37 in full), across these categories:

  • 5 core/utility memory tools: remember, recall, reflect, forget, briefing
  • 4 session management tools: session_start, session_end, session_commit, session_status
  • 8 context environment tools: exec, inspect, load, inject, query, rlm, status, checkpoint
  • 4 chat thread tools: create, append, get, list
  • 4 skills tools: list, get, get_file, search

The full profile additionally exposes associate, statistics, graph_query, audit, chat_thread_decompose, chat_thread_delete, skills_save, and 5 MCP server registry tools.

Three-Layer Memory Hierarchy

Memories are organized in a three-layer hierarchy enabling progressive abstraction:

Category Layer (Aggregated Summaries)
↑ aggregates
Item Layer (Discrete Memory Units)
↑ extracted from
Resource Layer (Raw Source Material)
  • Resource Layer — Raw conversations, documents, events
  • Item Layer — Individual facts, preferences, decisions (the primary API surface)
  • Category Layer — Auto-generated summaries that evolve based on content patterns

Data Flow

Remember (Write Path)

Client → REST API → MemoryService → Generate Embedding → Store in SQLite

Recall (Read Path)

Client → REST API → MemoryService → Vector Search (sqlite-vec) → Rank & Filter → Return

Reflect (Synthesis Path)

Client → REST API → ReflectService → Recall relevant memories → LLM synthesis → Return

Deployment

Local Development

Single-process, single-file SQLite database. Zero configuration:

Terminal window
pip install memorylayer-server
memorylayer serve

Production (Enterprise)

The enterprise edition adds:

  • PostgreSQL backend for scalability
  • Redis caching layer
  • Horizontal scaling
  • Advanced analytics and reflection engine

See memorylayer.ai for enterprise options.