Skip to content

Embed Server

The memorylayer-embed-server package is a stateless GPU embedding / transcription server that runs as a peer to memorylayer-server. It serves all heavy ML work (text embeddings, multi-vector / ColPali, OCR, transcription) over HTTP, while keeping the core server lightweight.

In v0.1.x the in-process providers local (sentence-transformers), colpali, and qwen3-vl were removed from memorylayer-server — all self-hosted embedding now routes through this package via the embed_server provider.

When to Use It

Run memorylayer-embed-server when you want any of the following:

  • Self-hosted text embeddings (sentence-transformers, BGE, or any HuggingFace/SBERT model)
  • Multi-vector / ColPali visual embeddings for document pages
  • On-GPU OCR (GLM-OCR etc.)
  • Transcription
  • Air-gapped or compliance-sensitive deployments that cannot call OpenAI or Google

If you are happy with cloud embeddings (openai / google) or just need a smoke test (mock), you do not need to install the embed server.

Installation

Terminal window
# Core install (no GPU dependencies)
pip install memorylayer-embed-server
# Common GPU bundle: OCR + vLLM + ColPali
pip install "memorylayer-embed-server[gpu]"
# Everything: GPU + Google embeddings + observability
pip install "memorylayer-embed-server[all]"

Optional extras:

ExtraPurpose
ocrOCR via Transformers (GLM-OCR, etc.) — transformers, torch, accelerate
vllmHigh-throughput vLLM-served text models
colpaliColPali / late-interaction visual embedding
googleGoogle GenAI embedding/transcription proxy
observabilityPrometheus /metrics + OpenTelemetry tracing
gpuocr + vllm + colpali
allgpu + google + observability
devpytest + ruff

The visual-tokenizer (Qwen3.5) lives in the proprietary memorylayer-embed-server-enterprise package; install that separately if you need it.

Quick Start

Terminal window
# Start on the default port (61051)
memorylayer-embed serve
# Custom host/port
memorylayer-embed serve --host 0.0.0.0 --port 61051
# Verbose logging
memorylayer-embed -v serve

Verify the server is up:

Terminal window
curl http://localhost:61051/health
curl http://localhost:61051/health/ready

Wiring It to the Core Server

Once the embed server is running, point memorylayer-server at it:

Terminal window
export MEMORYLAYER_EMBEDDING_PROVIDER=embed_server # default
export MEMORYLAYER_EMBED_SERVER_URL=http://embed-host:61051
memorylayer serve

For cross-datacenter or mTLS deployments, use Aether transport instead of plain HTTP:

Terminal window
export MEMORYLAYER_EMBED_TRANSPORT=aether
export MEMORYLAYER_EMBED_AETHER_TARGET=sv::memorylayer-embed::default

When EMBED_TRANSPORT=aether, the embed server runs behind the Aether proxy-sidecar (PID 1 in the official Docker image, supervisor mode) which terminates inbound ProxyHttpRequest envelopes and forwards to the embed FastAPI on 127.0.0.1:61051. See Aether Transport for the broader picture of what the mesh brings (mTLS, identity headers, OBO delegation, service discovery).

CLI

CommandDescription
memorylayer-embed serveStart the HTTP server (--host, --port)
memorylayer-embed versionPrint the package version

Global flag -v / --verbose enables debug logging.

Configuration

VariableDefaultDescription
MEMORYLAYER_EMBED_SERVER_HOST127.0.0.1Bind address
MEMORYLAYER_EMBED_SERVER_PORT61051Listening port
MEMORYLAYER_EMBED_MODEL_TEXT(provider default)Override the default text-embedding model
MEMORYLAYER_EMBED_MODEL_COLPALI(provider default)Override the default ColPali model
EMBED_SERVER_RUN_SIDECARtrue (Docker)When true, the Aether sidecar is PID 1 and supervises the embed server; when false, the embed server runs directly as plain HTTP (used by integration tests and dev shells).

See the provider modules under src/memorylayer_embed_server/ for full per-model environment variables.

Docker

The repository ships a CUDA-runtime Dockerfile at oss/memorylayer-embed-server/Dockerfile. It builds three stages:

  1. A Go stage that builds the Aether proxy-sidecar binary.
  2. A CUDA-devel builder stage that compiles native Python extensions (torch, vllm, colpali-engine, OCR transformers) into a venv via uv with --torch-backend=auto.
  3. A CUDA-runtime stage that copies the venv + sidecar, exposes port 61051, and runs /usr/local/bin/embed-server-entrypoint.sh as a non-root user (uid 65532).
Terminal window
# Build (from the repository root)
docker build -f oss/memorylayer-embed-server/Dockerfile -t memorylayer-embed-server .
# Run as a plain HTTP peer
docker run -d \
--name memorylayer-embed \
--gpus all \
-p 61051:61051 \
-e EMBED_SERVER_RUN_SIDECAR=false \
memorylayer-embed-server

Test variants (Dockerfile.test, Dockerfile.real-test, Dockerfile.real-test-full) are used by the integration test harness; Dockerfile.test builds a lighter CPU image suitable for CI.

Health Checks

  • GET /health — process is up
  • GET /health/ready — model(s) loaded and ready to serve
  • GET /health/load — per-provider load snapshot (in-flight requests, utilization) for upstream LB routing

The Docker image’s HEALTHCHECK targets /health (30s interval, 60s start period).

LLM Inference Hosting (optional)

When MEMORYLAYER_EMBED_LLM_ENABLED=true, the embed server hosts chat / completions LLM workloads alongside embeddings. One or more vllm serve child processes run inside the same pod, each serving a distinct LLM, and the FastAPI app exposes OpenAI-compatible chat endpoints:

  • POST /v1/chat/completions — OpenAI chat (streaming + non-streaming)
  • POST /v1/completions — legacy text completions
  • GET /v1/models — model list (profile names + aliases + underlying model IDs)

Routing is by the OpenAI-standard model request field, so any OpenAI client (including the core server’s embed_server LLM provider) works out of the box. The payload is forwarded verbatim to the underlying vLLM — tool calls, structured output, multimodal images, and reasoning fields all pass through.

Enabling

Terminal window
export MEMORYLAYER_EMBED_LLM_ENABLED=true
export MEMORYLAYER_EMBED_LLM_PROFILES=qwen # comma-list of profile names
export MEMORYLAYER_EMBED_LLM_DEFAULT_PROFILE=qwen # used when request `model` doesn't match
export MEMORYLAYER_EMBED_LLM_PRELOAD=false # default lazy; flip to eager-start at boot
export MEMORYLAYER_EMBED_LLM_PORT_RANGE=18100-18199 # private port pool for internal vllm subprocesses
# Per-profile (substitute the uppercased profile name)
export MEMORYLAYER_EMBED_LLM_PROFILE_QWEN_MODEL=Qwen/Qwen2.5-7B-Instruct
export MEMORYLAYER_EMBED_LLM_PROFILE_QWEN_ALIASES=qwen-7b,qwen2.5
export MEMORYLAYER_EMBED_LLM_PROFILE_QWEN_GPU_MEM_UTIL=0.4

Internal vLLM ports are auto-assigned from the configured port pool; operators never set ports manually.

Pointing the core server at a hosted LLM

The core server has an explicit embed_server LLM provider type that piggybacks on the same dual HTTP / Aether transport already wired for embeddings:

Terminal window
export MEMORYLAYER_LLM_PROFILE_INFERENCE_PROVIDER=embed_server
export MEMORYLAYER_LLM_PROFILE_INFERENCE_MODEL=qwen # matches a profile/alias on the embed server
export MEMORYLAYER_LLM_ASSIGN_REFLECT=inference # route reflection through it

Each core-side LLM profile can independently override _EMBED_SERVER_URL, _EMBED_SERVER_TRANSPORT, _EMBED_SERVER_AETHER_TARGET, and _EMBED_SERVER_TIMEOUT — so two profiles can fan out to two different embed-server peers (one HTTP, one Aether, mixed across regions). See Server Configuration for the full schema.

Full reference

The complete env-var schema, routing rules, health-check shapes, and recipes live in the in-repo doc: oss/memorylayer-embed-server/docs/embedding-providers.md → “LLM Inference Profiles”.

Versioning

memorylayer-embed-server is released in lockstep with memorylayer-server (currently 0.1.22). The version pin in its dependencies keeps client and server protocol versions aligned.