Skip to content

LlamaIndex Integration

The MemoryLayer LlamaIndex integration (memorylayer-llamaindex) provides a persistent BaseChatStore implementation for LlamaIndex applications. Chat history is stored as episodic memories in MemoryLayer, enabling persistence across application restarts.

Installation

Terminal window
pip install memorylayer-llamaindex

This installs the integration package along with its dependencies: memorylayer-client and llama-index-core>=0.10.0.

Prerequisites

A running MemoryLayer server:

Terminal window
pip install "memorylayer-server[openai]" # or [google], [all], or pair with memorylayer-embed-server
memorylayer serve

Features

  • LlamaIndex Native — Implements the BaseChatStore interface
  • Persistent Memory — Chat history survives application restarts
  • Multi-Session Support — Isolated conversations per user/session via chat keys
  • Full Async Support — Every sync method has an async equivalent
  • Agent Compatible — Works with LlamaIndex agents and chat engines
  • Context Manager — Supports both sync and async context managers for resource cleanup

Quick Start

from llama_index.core.llms import ChatMessage, MessageRole
from llama_index.core.memory import ChatMemoryBuffer
from memorylayer_llamaindex import MemoryLayerChatStore
chat_store = MemoryLayerChatStore(
base_url="http://localhost:61001",
workspace_id="ws_123"
)
memory = ChatMemoryBuffer.from_defaults(
chat_store=chat_store,
chat_store_key="user_alice",
token_limit=3000
)
# Messages persist across application restarts
memory.put(ChatMessage(role=MessageRole.USER, content="Hello!"))
memory.put(ChatMessage(role=MessageRole.ASSISTANT, content="Hi there!"))
history = memory.get()

Usage with Chat Engines

SimpleChatEngine

from llama_index.core.chat_engine import SimpleChatEngine
from llama_index.core.memory import ChatMemoryBuffer
from llama_index.llms.openai import OpenAI
from memorylayer_llamaindex import MemoryLayerChatStore
chat_store = MemoryLayerChatStore(
base_url="http://localhost:61001",
workspace_id="ws_demo"
)
memory = ChatMemoryBuffer.from_defaults(
chat_store=chat_store,
chat_store_key="chat_session_1",
token_limit=4000
)
chat_engine = SimpleChatEngine.from_defaults(
memory=memory,
llm=OpenAI(model="gpt-4o-mini"),
system_prompt="You are a helpful assistant with persistent memory."
)
response = chat_engine.chat("Hello! I'm Sarah.")
# Later (even after restart), Sarah's history is still there
response = chat_engine.chat("What's my name?")

FunctionAgent

from llama_index.core.agent.workflow import FunctionAgent
from llama_index.core.memory import ChatMemoryBuffer
from llama_index.core.tools import FunctionTool
from memorylayer_llamaindex import MemoryLayerChatStore
chat_store = MemoryLayerChatStore(
base_url="http://localhost:61001",
workspace_id="ws_agents"
)
memory = ChatMemoryBuffer.from_defaults(
chat_store=chat_store,
chat_store_key="agent_session_1",
token_limit=8000
)
agent = FunctionAgent(tools=[...], llm=llm)
response = await agent.run("What time is it?", ctx=ctx, memory=memory)

Direct ChatStore Operations

The MemoryLayerChatStore class implements the full BaseChatStore interface for direct message management.

Setting and Getting Messages

from llama_index.core.llms import ChatMessage, MessageRole
# Set messages (replaces any existing messages for the key)
chat_store.set_messages("user_123", [
ChatMessage(role=MessageRole.USER, content="Hello!"),
ChatMessage(role=MessageRole.ASSISTANT, content="Hi!")
])
# Get messages (ordered by index)
messages = chat_store.get_messages("user_123")
# Add a single message (appended with correct index)
chat_store.add_message("user_123",
ChatMessage(role=MessageRole.USER, content="New message"))
# List all chat keys in the workspace
keys = chat_store.get_keys()

Deleting Messages

# Delete all messages for a key (returns deleted messages)
deleted = chat_store.delete_messages("user_123")
# Delete a specific message by index
deleted_msg = chat_store.delete_message("user_123", idx=2)
# Delete the last message
deleted_last = chat_store.delete_last_message("user_123")

Async Operations

All sync methods have async equivalents. Use async with for proper resource management:

async with MemoryLayerChatStore(
base_url="http://localhost:61001",
workspace_id="ws_123"
) as chat_store:
await chat_store.aset_messages("user", messages)
messages = await chat_store.aget_messages("user")
await chat_store.async_add_message("user", message)
deleted = await chat_store.adelete_messages("user")
deleted_msg = await chat_store.adelete_message("user", idx=0)
deleted_last = await chat_store.adelete_last_message("user")
keys = await chat_store.aget_keys()

Configuration

ParameterTypeDefaultDescription
base_urlstr"http://localhost:61001"MemoryLayer API URL
api_keystr | NoneNoneAPI key for authentication
workspace_idstr | NoneNoneWorkspace ID for multi-tenant isolation
timeoutfloat30.0Request timeout in seconds

Persistence Across Restarts

# Session 1: Store messages
chat_store = MemoryLayerChatStore(base_url="http://localhost:61001")
memory = ChatMemoryBuffer.from_defaults(
chat_store=chat_store,
chat_store_key="persistent_session"
)
memory.put(ChatMessage(role=MessageRole.USER, content="Remember this."))
# === Application Restart ===
# Session 2: Messages are still there
chat_store2 = MemoryLayerChatStore(base_url="http://localhost:61001")
memory2 = ChatMemoryBuffer.from_defaults(
chat_store=chat_store2,
chat_store_key="persistent_session"
)
history = memory2.get() # Contains previous messages

How It Works

Messages are stored as episodic memories in MemoryLayer with tags for key-based isolation:

  • Each message gets a llamaindex_chat_key:<key> tag for filtering
  • Messages include chat_key, message_index, role, additional_kwargs, and blocks in metadata
  • Message content is stored as the role-prefixed text (e.g., "user: Hello!")
  • Retrieval uses semantic search with the chat key tag and sorts by message_index
  • Block data from LlamaIndex ChatMessage objects is preserved in metadata for lossless reconstruction
  • The add_message method uses per-key threading locks (sync) and asyncio locks (async) to prevent race conditions on index computation

Role Mapping

The integration maps between LlamaIndex MessageRole enum values and string representations:

MessageRoleStringAliases
USER"user""human"
ASSISTANT"assistant""ai"
SYSTEM"system"—
TOOL"tool"—
CHATBOT"chatbot"—
MODEL"model"—
FUNCTION"function"—

Utility Functions

The package exports utility functions for advanced use cases:

from memorylayer_llamaindex import (
chat_message_to_memory_payload, # Convert ChatMessage to MemoryLayer payload
memory_to_chat_message, # Convert MemoryLayer memory to ChatMessage
message_role_to_string, # MessageRole enum to string
string_to_message_role, # String to MessageRole enum
get_message_index, # Extract message index from memory
get_chat_key, # Extract chat key from memory
)

Requirements

  • Python 3.12+
  • memorylayer-client — MemoryLayer Python SDK
  • llama-index-core>=0.10.0 — LlamaIndex core library