System status: all systems operational

Data Agents for your Business.

The agentic data platform that automates data quality, governance, and observability — end to end. Command your data lake with the precision of a technical blueprint: data agents do the heavy lifting, your team stays in the loop.

500+ connectors // powered by dlt, DuckDB & Apache Arrow

The architecture

Multi-layered metadata orchestration designed for infinite scale. DuckLake bridges the gap between raw storage and intelligence — an Iceberg catalog and ACID SQL over your own Parquet.

Metadata layer
Iceberg catalog
Parquet storage
Ref: metadata & storageStatus: live
datajoi architecture: a DuckDB in-process client connected to a DuckLake catalog (ACID SQL, Parquet storage), an Iceberg catalog with metadata, manifest, and data files, and the metadata layer resolving metadata files to manifest lists to manifest files.
The architecture — under the hood

Built on DuckDB + DuckLake

DuckLake is a lakehouse architecture that is, in essence, just a database and some Parquet files. DuckDB runs in-process right next to your data; the DuckLake catalog is served by an ACID-compliant SQL database; and the data itself lives as open Parquet files on local disk or object storage.

Diagram 002 — DuckLake catalogMetadata in the database
DuckLake architecture: the catalog database encloses the entire metadata layer — the current metadata pointer for db1.table1, metadata snapshots s0 and s1, manifest lists, and manifests — while Parquet data files sit below in the data layer.

DuckLake's move: the entire metadata layer — pointers, snapshots, manifest lists, manifests — lives inside one ACID-compliant SQL database. Only the Parquet data files sit in storage, so every metadata change is a single transaction.

Diagram 003 — catalog schemaSQL all the way down
DuckLake catalog schema: SQL tables for snapshots, schemas, tables, columns, table and column statistics, data files, delete files, partition columns, and file-level column statistics, all resolving to Parquet files in the data layer.

The whole catalog is ordinary SQL tables — snapshots, schemas, tables, columns, data files, delete files, and statistics. Every change is a transaction; the data layer stays open Parquet.

Diagram 004 — metadata resolutionIceberg catalog
Iceberg catalog resolution: the database layer holds the current metadata pointer for db1.table1; it resolves through metadata files and manifest lists in the metadata layer down to manifest files and data files in the data layer.

The catalog keeps one current metadata pointer per table. A query resolves pointer → metadata file → manifest list → manifest → data files, so every read sees one consistent view.

Diagram 005 — snapshotsTime travel
Snapshot lineage: metadata files carrying snapshots s0 and s1 point at manifest lists, which point at manifest files shared between snapshots, which point at the underlying data files.

Snapshots (s0, s1, …) are immutable. Each commit writes a new metadata file that shares unchanged manifests with its predecessors — point-in-time travel is just reading an older snapshot.

The architecture — command surface

Command & query

Data agents run the command loop end to end — they write, execute, and validate SQL against your lake, and manage the catalog itself. On DuckLake + DuckDB that means agents register sources, evolve schemas, compact snapshots, and keep statistics fresh: every change an ACID transaction in the catalog, every result a DataStory™ your whole team can read.

0.024s
Avg query latency
99.9%
Schema integrity
datajoi workspace: a SQL query for top customers by revenue running against warehouse.orders, a results table (Northwind Ltd $4.21M, Globex Corp $3.08M, Initech $1.94M), live ORDERS → CUST_ORDERS → SUMMARY lineage, and a DataStory noting the top 5 customers drive 61% of revenue.
The architecture — agentic pipeline

Automated at every step.

An agentic pipeline with human-verified, quality-controlled lineage — end to end.

01_connect.sql

Connect your data

Lake schema and lineage read automatically.

02_agent.sql

Agents write SQL

Plain English becomes a query you can check.

03_run.sql

Run & validate

Executes with automated data-quality checks.

04_review.sql

Quality review

Agents flag anomalies; a human confirms edge cases.

05_story.md

DataStory™ delivered

A verified narrative, ready to share.

Verification checks Schema check passed Row-count drift within 0.2% Human review confirmed Lineage verified 12 sources All checks passing
Pricing

A plan for every data platform

Run your own, run your whole company's, or right-size something custom — datajoi's agents do the heavy lifting.

datajoi platformpro

Pro

$99/mo

For the hands-on data pro

  • Manage your own data platform
  • Or manage platforms for your clients
  • Agents, DataStory™ & verification checks
Start with Pro
datajoi platformMost popular

SMB

$999/mo · up to 10 users

Fits most businesses out of the box

  • Everything in Pro
  • 500+ existing connectors
  • Shared workspace for up to 10 users
Start with SMB
datajoi platformenterprise

Enterprise

Contact us

We'll right-size a plan to your stack

  • Custom connectors
  • Data migrations & data modeling
  • Anything custom, built with you
Talk to us [email protected]

Looking for the Chrome extension? Free trial, then paid →

Reference

Frequently asked questions

What is datajoi?

datajoi is an agentic data platform that automates data quality, governance, and observability end to end. Purpose-built data agents connect sources, write SQL and data models, run quality tests, and document everything — while your team reviews and approves every change before it reaches production.

How do datajoi's data agents work?

datajoi ships purpose-built agents for ingestion, modeling, quality, and cataloging. They work your data stack through MCP, API, and SDK integrations: an agent proposes a change — a new transformation, a schema edit, a quality rule — as a clear, reviewable request in plain language. A human approves, edits, or sends it back. Nothing touches production without a human yes.

What is the DuckLake architecture datajoi is built on?

DuckLake is a lakehouse architecture that is, in essence, just a database and some Parquet files. datajoi runs DuckDB in-process against a DuckLake catalog: an Iceberg-compatible metadata layer — metadata files, manifest lists, and manifest files — points at columnar Parquet data files in your own storage. You get ACID SQL, snapshots, and time travel without a proprietary warehouse or vendor lock-in.

Does datajoi replace my data team?

No. datajoi is built human-in-the-loop by design. Agents handle the repetitive heavy lifting — writing tests, mapping lineage, monitoring pipelines — and your team stays in control with plain-language reviews, one-click approvals, and a full audit trail of who approved what and when. No SQL is required to review changes.

Which data sources does datajoi connect to?

datajoi offers 500+ connectors powered by dlt, with DuckDB and Apache Arrow for compute. It works with warehouses and tools like Snowflake, BigQuery, Databricks, Postgres, dbt, Kafka, and S3 — plus SaaS apps, files, and streams. Data lands from any source in minutes.

What is a DataStory?

A DataStory is datajoi's answer format: ask a question in plain language and get back the generated SQL, the results, and a narrative explanation with live charts — backed by lineage, quality tests, and observability. DataStories refresh on their own, so there is no dashboard to build or maintain.

Early access

Put your data on autopilot.

Join the waitlist and be first to command an agentic, end-to-end data platform — built on DuckLake, with humans in the loop.

No spam. We'll only email you about early access.