Skip to content

Dataset Management

MemoryLayer Enterprise can ingest tabular datasets, automatically profile column statistics and distributions, and extract structured memories from the data. Supported formats include CSV, TSV, Parquet, JSONL, and Excel (.xlsx).


Upload a Dataset

Upload a dataset using multipart/form-data. The server detects the format, stores the file as Parquet, and starts a background profiling job.

Terminal window
curl -X POST "$MEMORYLAYER_URL/v1/datasets" \
-H "Authorization: Bearer $API_KEY" \
-F "file=@sales_data.csv" \
-F "name=Q1 Sales Data" \
-F "target_context_id=analytics" \
-F "importance=0.6" \
-F "sample_rows=1000" \
-F "detect_time_series=true" \
-F "generate_summaries=true"

Parameters:

ParameterDefaultDescription
file(required)The dataset file
name(from filename)Human-readable dataset name
target_context_id_defaultMemory context for extracted memories
importance0.5Default importance for extracted memories (0.0—1.0)
sample_rows1000Max rows to include in LLM summary sample
detect_time_seriestrueAttempt to detect temporal columns
generate_summariestrueGenerate LLM natural-language summaries

The response includes the dataset record and the profiling job:

{
"dataset": {
"id": "ds_abc123",
"workspace_id": "ws_001",
"name": "Q1 Sales Data",
"filename": "sales_data.csv",
"format": "csv",
"status": "pending",
"size_bytes": 512000,
"row_count": 0,
"column_count": 0,
"created_at": "2026-04-01T12:00:00Z"
},
"job": {
"id": "dsjob_xyz789",
"status": "queued",
"progress_percent": 0,
"created_at": "2026-04-01T12:00:00Z"
}
}

Automatic Profiling

When profiling completes, the dataset record is enriched with:

  • Row and column counts
  • Per-column statistics: null counts/percentages, unique values, min/max/mean/median/std for numeric columns, length stats for strings, top values for categoricals
  • Histograms for numeric columns (configurable bin count)
  • Time series detection: identifies temporal columns with resolution (second, minute, hour, day, week, month, year) and date ranges
  • LLM-generated summary: a natural-language description of the dataset

Retrieve the full profile by fetching the dataset:

Terminal window
curl "$MEMORYLAYER_URL/v1/datasets/$DATASET_ID" \
-H "Authorization: Bearer $API_KEY"

The columns array contains per-column profiling results, and profile_summary contains the natural-language summary.


Dataset Statuses

StatusMeaning
pendingUploaded, waiting for profiling
profilingColumn statistics being computed
summarizingLLM generating natural-language summary
completedProfiling and memory extraction done
failedProcessing encountered an error

Query with DuckDB SQL

Slice into any dataset using DuckDB SQL or structured filters. The dataset is queryable as a table named data.

Raw SQL

Terminal window
curl -X POST "$MEMORYLAYER_URL/v1/datasets/$DATASET_ID/slice" \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"sql": "SELECT region, SUM(revenue) as total FROM data GROUP BY region ORDER BY total DESC LIMIT 10"
}'

Structured filters

Terminal window
curl -X POST "$MEMORYLAYER_URL/v1/datasets/$DATASET_ID/slice" \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"columns": ["product", "revenue", "date"],
"filters": [
{"column": "revenue", "op": ">", "value": 10000},
{"column": "region", "op": "=", "value": "EMEA"}
],
"order_by": "revenue",
"descending": true,
"limit": 50,
"offset": 0
}'

Supported filter operators: =, !=, <, >, <=, >=, in, like.

The response includes column names, data types, rows, total matching count, and the SQL that was executed.


Retrieve Extracted Memories

Terminal window
curl "$MEMORYLAYER_URL/v1/datasets/$DATASET_ID/memories" \
-H "Authorization: Bearer $API_KEY"

Job Tracking

Terminal window
# List dataset jobs
curl "$MEMORYLAYER_URL/v1/datasets/jobs?status=running" \
-H "Authorization: Bearer $API_KEY"
# Get job status
curl "$MEMORYLAYER_URL/v1/datasets/jobs/$JOB_ID" \
-H "Authorization: Bearer $API_KEY"
# Cancel a job
curl -X POST "$MEMORYLAYER_URL/v1/datasets/jobs/$JOB_ID/cancel" \
-H "Authorization: Bearer $API_KEY"

List and Delete Datasets

Terminal window
# List datasets
curl "$MEMORYLAYER_URL/v1/datasets?status=completed&limit=50" \
-H "Authorization: Bearer $API_KEY"
# Delete dataset only
curl -X DELETE "$MEMORYLAYER_URL/v1/datasets/$DATASET_ID" \
-H "Authorization: Bearer $API_KEY"
# Delete dataset and its extracted memories
curl -X DELETE "$MEMORYLAYER_URL/v1/datasets/$DATASET_ID?delete_memories=true" \
-H "Authorization: Bearer $API_KEY"