01Open Source · MCP

OmniData MCP

Give any LLM a data engine it can actually trust.

OmniData MCP project cover

Protocol

MCP

Engines

Spark + DuckDB

Licence

Open source

Metric profile

01 / Signal

4 verbs

0

Query surface covered

inspect · query · aggregate · export

02 / Signal

Spark ↔ DuckDB

0

Engine routing accuracy

size + locality aware

03 / Signal

1 config

0

Setup friction

single MCP server entry

Overview

OmniData MCP exposes distributed data processing to language models through the Model Context Protocol — so an agent can profile, query and transform real datasets instead of hallucinating about them.

PySpark handles the heavy distributed workloads while DuckDB serves fast, in-process analytical queries on local and columnar files. The MCP layer normalises both behind one typed tool surface.

What makes it work

01

One tool surface, two engines

Agents call the same verbs — inspect, query, aggregate, export — and the server routes to Spark or DuckDB based on data size and locality.

02

Schema-aware responses

Every result carries schema, row counts and sampling metadata so models reason over structure rather than raw text blobs.

03

Built for pipelines

Designed to slot into existing lakehouse workflows: parquet, CSV and SQL sources without a bespoke ingestion step.

Architecture

  • MCP server exposing typed data tools to any compatible client
  • PySpark executor for distributed transformations
  • DuckDB executor for sub-second analytical queries
  • Schema + profiling layer returned with every tool call

Stack

  • PySpark
  • DuckDB
  • MCP
  • Python

Open-source Model Context Protocol project for big data processing using PySpark and DuckDB.

Interested?

Explore the code, or talk about the ideas behind it.