Skip to content

Overview

MemoryLayer is an API-first memory infrastructure for LLM-powered agents. It solves the fundamental challenge of stateless LLMs by providing persistent, queryable, multi-modal memory that agents can read from and write to during execution.

What is MemoryLayer?

MemoryLayer provides cognitive memory capabilities for AI agents, including episodic, semantic, procedural, and working memory with vector-based retrieval and graph-based associations.

Think of it as a memory backend for your AI applications — the same way you’d use PostgreSQL for relational data or Redis for caching, you use MemoryLayer for agent memory.

Key Features

  • Cognitive Memory Architecture — Episodic, semantic, procedural, and working memory types modeled after human cognition
  • Vector Search — SQLite + sqlite-vec for efficient similarity search; cloud embeddings (OpenAI / Google) or self-hosted text & multi-vector via the memorylayer-embed-server peer
  • Knowledge Graph — 63 typed relationships across 11 categories (causal, solution, learning, workflow, hierarchical, etc.) connecting memories into a traversable semantic graph
  • Context Environment — Server-side Python sandbox with RLM (Recursive Language Model) so heavy memory analysis doesn’t consume the agent’s context window
  • Session Management — Working memory with TTL, token-budget-aware extraction triggers, and commit to long-term storage
  • Multi-Platform SDKs — Python and TypeScript client libraries with full type safety
  • Framework Integrations — Drop-in support for LangChain, LlamaIndex, Claude Code (plugin + MCP), and OpenCode (plugin + MCP)
  • MCP Server — 25 tools in the default profile (38 in full) for Claude Code, Claude Desktop, Cursor, and other MCP-compatible LLMs
  • REST API — FastAPI-based HTTP server with OpenAPI documentation

Architecture Overview

┌─────────────────────────────────────────────────────────────┐
│ Agent Frameworks (LangChain, LlamaIndex, Claude Code) │
├─────────────────────────────────────────────────────────────┤
│ Client SDKs (Python, TypeScript) │
├─────────────────────────────────────────────────────────────┤
│ MemoryLayer Server (REST API + MCP) │
├─────────────────────────────────────────────────────────────┤
│ Storage (SQLite + sqlite-vec) │
└─────────────────────────────────────────────────────────────┘

Packages

PackageInstallDescription
Serverpip install memorylayer-serverFastAPI server with SQLite storage
Python SDKpip install memorylayer-clientAsync Python client library
TypeScript SDKnpm install @scitrera/memorylayer-sdkTypeScript/JavaScript client
MCP Servernpm install @scitrera/memorylayer-mcp-serverModel Context Protocol integration
LangChainpip install memorylayer-langchainLangChain memory backend
LlamaIndexpip install memorylayer-llamaindexLlamaIndex chat store

Quick Example

from memorylayer import MemoryLayerClient, MemoryType
async with MemoryLayerClient(base_url="http://localhost:61001") as client:
# Store a memory
memory = await client.remember(
content="User prefers Python for backend development",
type=MemoryType.SEMANTIC,
importance=0.8,
tags=["preferences", "programming"]
)
# Recall memories
results = await client.recall(
query="What programming languages does the user like?",
limit=5
)
# Synthesize insights
reflection = await client.reflect(
query="Summarize user's technology preferences"
)
print(reflection.reflection)

Next Steps