Dataset Management
MemoryLayer Enterprise can ingest tabular datasets, automatically profile column statistics and distributions, and extract structured memories from the data. Supported formats include CSV, TSV, Parquet, JSONL, and Excel (.xlsx).
Upload a Dataset
Upload a dataset using multipart/form-data. The server detects the format, stores the file as Parquet, and starts a background profiling job.
curl -X POST "$MEMORYLAYER_URL/v1/datasets" \ -H "Authorization: Bearer $API_KEY" \ -F "file=@sales_data.csv" \ -F "name=Q1 Sales Data" \ -F "target_context_id=analytics" \ -F "importance=0.6" \ -F "sample_rows=1000" \ -F "detect_time_series=true" \ -F "generate_summaries=true"Parameters:
| Parameter | Default | Description |
|---|---|---|
file | (required) | The dataset file |
name | (from filename) | Human-readable dataset name |
target_context_id | _default | Memory context for extracted memories |
importance | 0.5 | Default importance for extracted memories (0.0—1.0) |
sample_rows | 1000 | Max rows to include in LLM summary sample |
detect_time_series | true | Attempt to detect temporal columns |
generate_summaries | true | Generate LLM natural-language summaries |
The response includes the dataset record and the profiling job:
{ "dataset": { "id": "ds_abc123", "workspace_id": "ws_001", "name": "Q1 Sales Data", "filename": "sales_data.csv", "format": "csv", "status": "pending", "size_bytes": 512000, "row_count": 0, "column_count": 0, "created_at": "2026-04-01T12:00:00Z" }, "job": { "id": "dsjob_xyz789", "status": "queued", "progress_percent": 0, "created_at": "2026-04-01T12:00:00Z" }}Automatic Profiling
When profiling completes, the dataset record is enriched with:
- Row and column counts
- Per-column statistics: null counts/percentages, unique values, min/max/mean/median/std for numeric columns, length stats for strings, top values for categoricals
- Histograms for numeric columns (configurable bin count)
- Time series detection: identifies temporal columns with resolution (second, minute, hour, day, week, month, year) and date ranges
- LLM-generated summary: a natural-language description of the dataset
Retrieve the full profile by fetching the dataset:
curl "$MEMORYLAYER_URL/v1/datasets/$DATASET_ID" \ -H "Authorization: Bearer $API_KEY"The columns array contains per-column profiling results, and profile_summary contains the natural-language summary.
Dataset Statuses
| Status | Meaning |
|---|---|
pending | Uploaded, waiting for profiling |
profiling | Column statistics being computed |
summarizing | LLM generating natural-language summary |
completed | Profiling and memory extraction done |
failed | Processing encountered an error |
Query with DuckDB SQL
Slice into any dataset using DuckDB SQL or structured filters. The dataset is queryable as a table named data.
Raw SQL
curl -X POST "$MEMORYLAYER_URL/v1/datasets/$DATASET_ID/slice" \ -H "Authorization: Bearer $API_KEY" \ -H "Content-Type: application/json" \ -d '{ "sql": "SELECT region, SUM(revenue) as total FROM data GROUP BY region ORDER BY total DESC LIMIT 10" }'Structured filters
curl -X POST "$MEMORYLAYER_URL/v1/datasets/$DATASET_ID/slice" \ -H "Authorization: Bearer $API_KEY" \ -H "Content-Type: application/json" \ -d '{ "columns": ["product", "revenue", "date"], "filters": [ {"column": "revenue", "op": ">", "value": 10000}, {"column": "region", "op": "=", "value": "EMEA"} ], "order_by": "revenue", "descending": true, "limit": 50, "offset": 0 }'Supported filter operators: =, !=, <, >, <=, >=, in, like.
The response includes column names, data types, rows, total matching count, and the SQL that was executed.
Retrieve Extracted Memories
curl "$MEMORYLAYER_URL/v1/datasets/$DATASET_ID/memories" \ -H "Authorization: Bearer $API_KEY"Job Tracking
# List dataset jobscurl "$MEMORYLAYER_URL/v1/datasets/jobs?status=running" \ -H "Authorization: Bearer $API_KEY"
# Get job statuscurl "$MEMORYLAYER_URL/v1/datasets/jobs/$JOB_ID" \ -H "Authorization: Bearer $API_KEY"
# Cancel a jobcurl -X POST "$MEMORYLAYER_URL/v1/datasets/jobs/$JOB_ID/cancel" \ -H "Authorization: Bearer $API_KEY"List and Delete Datasets
# List datasetscurl "$MEMORYLAYER_URL/v1/datasets?status=completed&limit=50" \ -H "Authorization: Bearer $API_KEY"
# Delete dataset onlycurl -X DELETE "$MEMORYLAYER_URL/v1/datasets/$DATASET_ID" \ -H "Authorization: Bearer $API_KEY"
# Delete dataset and its extracted memoriescurl -X DELETE "$MEMORYLAYER_URL/v1/datasets/$DATASET_ID?delete_memories=true" \ -H "Authorization: Bearer $API_KEY"