Embed Server
The memorylayer-embed-server package is a stateless GPU embedding / transcription server that runs as a peer to memorylayer-server. It serves all heavy ML work (text embeddings, multi-vector / ColPali, OCR, transcription) over HTTP, while keeping the core server lightweight.
In v0.1.x the in-process providers local (sentence-transformers), colpali, and qwen3-vl were removed from memorylayer-server — all self-hosted embedding now routes through this package via the embed_server provider.
When to Use It
Run memorylayer-embed-server when you want any of the following:
- Self-hosted text embeddings (sentence-transformers, BGE, or any HuggingFace/SBERT model)
- Multi-vector / ColPali visual embeddings for document pages
- On-GPU OCR (GLM-OCR etc.)
- Transcription
- Air-gapped or compliance-sensitive deployments that cannot call OpenAI or Google
If you are happy with cloud embeddings (openai / google) or just need a smoke test (mock), you do not need to install the embed server.
Installation
# Core install (no GPU dependencies)pip install memorylayer-embed-server
# Common GPU bundle: OCR + vLLM + ColPalipip install "memorylayer-embed-server[gpu]"
# Everything: GPU + Google embeddings + observabilitypip install "memorylayer-embed-server[all]"Optional extras:
| Extra | Purpose |
|---|---|
ocr | OCR via Transformers (GLM-OCR, etc.) — transformers, torch, accelerate |
vllm | High-throughput vLLM-served text models |
colpali | ColPali / late-interaction visual embedding |
google | Google GenAI embedding/transcription proxy |
observability | Prometheus /metrics + OpenTelemetry tracing |
gpu | ocr + vllm + colpali |
all | gpu + google + observability |
dev | pytest + ruff |
The visual-tokenizer (Qwen3.5) lives in the proprietary memorylayer-embed-server-enterprise package; install that separately if you need it.
Quick Start
# Start on the default port (61051)memorylayer-embed serve
# Custom host/portmemorylayer-embed serve --host 0.0.0.0 --port 61051
# Verbose loggingmemorylayer-embed -v serveVerify the server is up:
curl http://localhost:61051/healthcurl http://localhost:61051/health/readyWiring It to the Core Server
Once the embed server is running, point memorylayer-server at it:
export MEMORYLAYER_EMBEDDING_PROVIDER=embed_server # defaultexport MEMORYLAYER_EMBED_SERVER_URL=http://embed-host:61051memorylayer serveFor cross-datacenter or mTLS deployments, use Aether transport instead of plain HTTP:
export MEMORYLAYER_EMBED_TRANSPORT=aetherexport MEMORYLAYER_EMBED_AETHER_TARGET=sv::memorylayer-embed::defaultWhen EMBED_TRANSPORT=aether, the embed server runs behind the Aether proxy-sidecar (PID 1 in the official Docker image, supervisor mode) which terminates inbound ProxyHttpRequest envelopes and forwards to the embed FastAPI on 127.0.0.1:61051. See Aether Transport for the broader picture of what the mesh brings (mTLS, identity headers, OBO delegation, service discovery).
CLI
| Command | Description |
|---|---|
memorylayer-embed serve | Start the HTTP server (--host, --port) |
memorylayer-embed version | Print the package version |
Global flag -v / --verbose enables debug logging.
Configuration
| Variable | Default | Description |
|---|---|---|
MEMORYLAYER_EMBED_SERVER_HOST | 127.0.0.1 | Bind address |
MEMORYLAYER_EMBED_SERVER_PORT | 61051 | Listening port |
MEMORYLAYER_EMBED_MODEL_TEXT | (provider default) | Override the default text-embedding model |
MEMORYLAYER_EMBED_MODEL_COLPALI | (provider default) | Override the default ColPali model |
EMBED_SERVER_RUN_SIDECAR | true (Docker) | When true, the Aether sidecar is PID 1 and supervises the embed server; when false, the embed server runs directly as plain HTTP (used by integration tests and dev shells). |
See the provider modules under src/memorylayer_embed_server/ for full per-model environment variables.
Docker
The repository ships a CUDA-runtime Dockerfile at oss/memorylayer-embed-server/Dockerfile. It builds three stages:
- A Go stage that builds the Aether
proxy-sidecarbinary. - A CUDA-devel builder stage that compiles native Python extensions (
torch,vllm,colpali-engine, OCR transformers) into a venv viauvwith--torch-backend=auto. - A CUDA-runtime stage that copies the venv + sidecar, exposes port 61051, and runs
/usr/local/bin/embed-server-entrypoint.shas a non-root user (uid 65532).
# Build (from the repository root)docker build -f oss/memorylayer-embed-server/Dockerfile -t memorylayer-embed-server .
# Run as a plain HTTP peerdocker run -d \ --name memorylayer-embed \ --gpus all \ -p 61051:61051 \ -e EMBED_SERVER_RUN_SIDECAR=false \ memorylayer-embed-serverTest variants (Dockerfile.test, Dockerfile.real-test, Dockerfile.real-test-full) are used by the integration test harness; Dockerfile.test builds a lighter CPU image suitable for CI.
Health Checks
GET /health— process is upGET /health/ready— model(s) loaded and ready to serveGET /health/load— per-provider load snapshot (in-flight requests, utilization) for upstream LB routing
The Docker image’s HEALTHCHECK targets /health (30s interval, 60s start period).
LLM Inference Hosting (optional)
When MEMORYLAYER_EMBED_LLM_ENABLED=true, the embed server hosts chat / completions LLM workloads alongside embeddings. One or more vllm serve child processes run inside the same pod, each serving a distinct LLM, and the FastAPI app exposes OpenAI-compatible chat endpoints:
POST /v1/chat/completions— OpenAI chat (streaming + non-streaming)POST /v1/completions— legacy text completionsGET /v1/models— model list (profile names + aliases + underlying model IDs)
Routing is by the OpenAI-standard model request field, so any OpenAI client (including the core server’s embed_server LLM provider) works out of the box. The payload is forwarded verbatim to the underlying vLLM — tool calls, structured output, multimodal images, and reasoning fields all pass through.
Enabling
export MEMORYLAYER_EMBED_LLM_ENABLED=trueexport MEMORYLAYER_EMBED_LLM_PROFILES=qwen # comma-list of profile namesexport MEMORYLAYER_EMBED_LLM_DEFAULT_PROFILE=qwen # used when request `model` doesn't matchexport MEMORYLAYER_EMBED_LLM_PRELOAD=false # default lazy; flip to eager-start at bootexport MEMORYLAYER_EMBED_LLM_PORT_RANGE=18100-18199 # private port pool for internal vllm subprocesses
# Per-profile (substitute the uppercased profile name)export MEMORYLAYER_EMBED_LLM_PROFILE_QWEN_MODEL=Qwen/Qwen2.5-7B-Instructexport MEMORYLAYER_EMBED_LLM_PROFILE_QWEN_ALIASES=qwen-7b,qwen2.5export MEMORYLAYER_EMBED_LLM_PROFILE_QWEN_GPU_MEM_UTIL=0.4Internal vLLM ports are auto-assigned from the configured port pool; operators never set ports manually.
Pointing the core server at a hosted LLM
The core server has an explicit embed_server LLM provider type that piggybacks on the same dual HTTP / Aether transport already wired for embeddings:
export MEMORYLAYER_LLM_PROFILE_INFERENCE_PROVIDER=embed_serverexport MEMORYLAYER_LLM_PROFILE_INFERENCE_MODEL=qwen # matches a profile/alias on the embed serverexport MEMORYLAYER_LLM_ASSIGN_REFLECT=inference # route reflection through itEach core-side LLM profile can independently override _EMBED_SERVER_URL, _EMBED_SERVER_TRANSPORT, _EMBED_SERVER_AETHER_TARGET, and _EMBED_SERVER_TIMEOUT — so two profiles can fan out to two different embed-server peers (one HTTP, one Aether, mixed across regions). See Server Configuration for the full schema.
Full reference
The complete env-var schema, routing rules, health-check shapes, and recipes live in the in-repo doc: oss/memorylayer-embed-server/docs/embedding-providers.md → “LLM Inference Profiles”.
Versioning
memorylayer-embed-server is released in lockstep with memorylayer-server (currently 0.1.22). The version pin in its dependencies keeps client and server protocol versions aligned.
Related Pages
- Aether Transport — what running the embed peer behind Aether buys you
- Server Configuration — core server env vars for
embed_serverprovider wiring - Installation — end-to-end install flow for both packages
- Architecture — where the embed server fits in the system diagram
- aetherlayer.ai — Aether product site