Skip to content

Aether Transport

MemoryLayer can run as a standalone HTTP service or behind Aether, a service-mesh transport that provides mTLS-secured connectivity, identity-aware request routing, on-behalf-of (OBO) delegation, and cross-datacenter orchestration. Aether is completely optional — everything in memorylayer-server works over plain HTTP without it — but turning it on collapses a lot of authentication, secret-management, and multi-DC plumbing into a single layer that the application code does not have to reason about.

Aether is a separate product. See aetherlayer.ai for the full overview, deployment guides, and operator documentation.


Why You Might Want It

Without Aether, running MemoryLayer at scale means assembling these pieces yourself:

  • TLS termination and certificate rotation between every pair of services that talk to each other
  • An identity provider that issues JWTs or OIDC tokens, plus the verification logic on every server
  • An on-behalf-of (OBO) mechanism so an application can call MemoryLayer as a specific end user without sharing the user’s credentials
  • A secret-distribution channel for API keys, OAuth refresh tokens, and DB credentials across machines / datacenters
  • A way to colocate GPU-hungry services (embeddings, OCR) in one DC while keeping the API surface in another
  • A durable task queue + scheduler so long-running, retryable work (ingestion, enrichment, decay, summarization) survives process restarts
  • A fine-grained agent permission model so the same act_as / on_behalf_of token only grants the scopes the user actually authorized

With Aether, the transport layer carries all of that:

ConcernAether brings
mTLS everywhereEach service holds its own short-lived certificate; the sidecar negotiates the channel
Authenticated identityX-Aether-Subject-* headers are signed by the mesh and trusted by the server
OBO delegationX-Aether-Grant-ID + X-Aether-Authority-Mode let an upstream agent service call MemoryLayer as an end user with audit trail
Service discoveryServices register as sv::memorylayer::default, sv::memorylayer-embed::default, etc.; callers target the topic, not an IP
Cross-DC routingGPU embed peers can live anywhere on the mesh; the core server reaches them with the same call shape it uses locally
Token issuance/v1/tokens proxies to the Aether gRPC token service so issuance / revocation are platform-managed
Durable task execution + schedulingAether provides a durable task queue and scheduler that the rest of the platform (MemoryLayer included) plugs into for retryable, restart-safe work — ingestion jobs, enrichment passes, decay sweeps, RPG sync follow-ups, etc.
Agent permissionsMesh-issued grants carry scoped permissions that downstream services consume directly, so “what is this agent allowed to do on this user’s behalf” is one verifiable decision, not a per-service ACL audit

If you only ever run one MemoryLayer process on localhost, none of this matters and plain HTTP is the right call. If you are deploying across hosts, datacenters, or stitching MemoryLayer into a larger agent platform, Aether dramatically shrinks the surface you have to operate.

Distributed Agent Systems

Aether is designed for the multi-agent case from the ground up: each agent (planner, coder, reviewer, retriever, …) runs as a separately-deployable service that registers a topic, holds a mesh certificate, and is reachable by every other agent in the system through the same proxy_http_async / proxy_grpc_async primitives. That gives you:

  • Inter-agent calls without bespoke RPC plumbing — one agent invoking another looks the same as one agent calling MemoryLayer.
  • OBO across hops — a user’s grant flows from the top-level agent down through every agent it delegates to, with each hop verifiable.
  • Coordinated durable work — the task scheduler lets a fan-out / fan-in workflow span agents and survive any single agent restarting.
  • Pluggable agent placement — run light agents on a small box, GPU-bound agents on a beefier one, and the wire shape never changes.

MemoryLayer is what makes those agents share memory. Once every agent in the system reaches MemoryLayer over Aether using the same OBO grant chain, you get a single coherent, identity-aware memory store across the whole agent fleet — recall in agent B sees what agent A remembered, scoped to the right user, without any of the individual agents having to think about identity, transport, or persistence. That combination — Aether for orchestration + identity, MemoryLayer for shared distributed memory — is the most powerful way to run multi-agent systems we are aware of, and is the primary deployment pattern we recommend for production agent platforms.


How MemoryLayer Uses Aether

MemoryLayer integrates with Aether at four points; any subset of them can be enabled independently.

1. REST-over-Aether front door

When an AetherServiceConnection is present at startup, the FastAPI app attaches itself to the connection and serves the same REST surface (/v1/memories, /v1/sessions, etc.) over Aether’s ProxyHttpRequest envelopes. Clients that already speak Aether reach the server through sv::memorylayer::default instead of opening a TCP socket.

This is additive: the plain-HTTP front door on MEMORYLAYER_SERVER_PORT continues to work for local development and direct curl access.

2. Aether-aware authentication

Setting MEMORYLAYER_AUTHENTICATION_SERVICE=aether swaps in AetherAuthenticationService, which reads the mesh-signed identity headers:

HeaderMeaning
X-Aether-Subject-TypeCaller principal type (User, Application, etc.)
X-Aether-Subject-IDCaller principal ID
X-Aether-Grant-IDOBO grant token (when an upstream service is delegating)
X-Aether-Authority-ModeOBO authority mode (e.g. act_as, on_behalf_of)

Because the headers are mesh-signed, the server does not need its own JWT verifier or IdP integration. The OSS default authentication service (open auth) remains the option for local development.

3. Embed-server transport switch

The core server’s connection to memorylayer-embed-server has two transports:

Terminal window
# Default: plain HTTP
export MEMORYLAYER_EMBED_TRANSPORT=http
export MEMORYLAYER_EMBED_SERVER_URL=http://embed-host:61051
# Aether mesh
export MEMORYLAYER_EMBED_TRANSPORT=aether
export MEMORYLAYER_EMBED_AETHER_TARGET=sv::memorylayer-embed::default

In aether mode, the core server issues proxy_http_async calls through the shared AetherServiceConnection against the configured topic. This is what enables cross-DC GPU placement under mTLS — the embed peer can be anywhere on the mesh that publishes sv::memorylayer-embed::*, and the application code does not change.

See Embed Server for the embed-side configuration.

4. Token issuance via Aether gRPC

When Aether is wired up, the OSS /v1/tokens endpoint proxies create / list / revoke calls to the Aether token gRPC service instead of issuing local tokens. The token shape exposed to clients is unchanged; what differs is who owns the underlying KMS-backed secret material.

This is what makes the OSS /v1/tokens plugin file is_enabled=False by default in OSS-only deployments — it requires an Aether-issued connection to be useful.


Configuration Reference

VariableDefaultDescription
MEMORYLAYER_AETHER_SERVICE_CONNECTIONdefaultPlugin name for the Aether service-connection extension. Set to default to enable when the Aether client library is installed; leave at default (the no-op resolution) to run without Aether.
MEMORYLAYER_AUTHENTICATION_SERVICEdefaultSet to aether to use the mesh-signed identity headers for authentication. Default default is open-auth (OSS-friendly local dev).
MEMORYLAYER_EMBED_TRANSPORThttphttp for direct calls to the embed peer; aether to route through the mesh.
MEMORYLAYER_EMBED_AETHER_TARGETsv::memorylayer-embed::defaultAether service topic to target when EMBED_TRANSPORT=aether.
MEMORYLAYER_EMBED_SERVER_URLhttp://localhost:61051Used only when EMBED_TRANSPORT=http.

The Aether client side ships separately as part of your Aether install. The MemoryLayer server gracefully no-ops every Aether integration point if the connection is not registered, so a build that includes the integration code can be deployed both with and without Aether without conditional flags.


Deployment Topology

┌────────────────────────────────────────────────────────────────────────┐
│ Aether mesh │
│ │
│ ┌──────────────────┐ ┌──────────────────────┐ │
│ │ memorylayer │ │ memorylayer-embed │ │
│ │ -server │ ──────► │ -server (GPU peer) │ │
│ │ sv::memorylayer │ proxy_ │ sv::memorylayer- │ │
│ │ ::default │ http_ │ embed::default │ │
│ └──────────────────┘ async └──────────────────────┘ │
│ ▲ │
│ │ ProxyHttpRequest envelopes (REST-over-Aether) │
│ │ │
│ ┌──────────────────┐ ┌──────────────────────┐ │
│ │ Agent / SDK │ │ Aether gRPC token / │ │
│ │ client │ │ identity service │ │
│ └──────────────────┘ └──────────────────────┘ │
│ │
└────────────────────────────────────────────────────────────────────────┘

Each box is a separately-deployed process; they all hold short-lived mesh certificates and reach each other through service topics rather than IPs. The same memorylayer-server binary that runs without Aether on a developer’s laptop runs inside this topology unchanged — only environment variables change.

The official memorylayer-embed-server Docker image runs the Aether proxy-sidecar as PID 1 in supervisor mode by default; the embed FastAPI on 127.0.0.1:61051 has no Aether knowledge of its own. Set EMBED_SERVER_RUN_SIDECAR=false to run the embed server as plain HTTP for local development.


When Not to Use Aether

Aether is a serious operational commitment. Skip it when:

  • You are running a single-machine local development setup
  • You don’t need cross-datacenter routing or OBO delegation
  • You are happy issuing your own bearer tokens (or running OSS with open auth) and exposing the HTTP port directly
  • You don’t have an existing Aether deployment to plug into

Plain HTTP + MEMORYLAYER_EMBEDDING_PROVIDER=openai / google / a local embed peer covers most self-hosted production deployments without the mesh.