OmniData MCP
Give any LLM a data engine it can actually trust.

Protocol
MCP
Engines
Spark + DuckDB
Licence
Open source
Metric profile
01 / Signal
4 verbs
Query surface covered
inspect · query · aggregate · export
02 / Signal
Spark ↔ DuckDB
Engine routing accuracy
size + locality aware
03 / Signal
1 config
Setup friction
single MCP server entry
Overview
OmniData MCP exposes distributed data processing to language models through the Model Context Protocol — so an agent can profile, query and transform real datasets instead of hallucinating about them.
PySpark handles the heavy distributed workloads while DuckDB serves fast, in-process analytical queries on local and columnar files. The MCP layer normalises both behind one typed tool surface.
What makes it work
One tool surface, two engines
Agents call the same verbs — inspect, query, aggregate, export — and the server routes to Spark or DuckDB based on data size and locality.
Schema-aware responses
Every result carries schema, row counts and sampling metadata so models reason over structure rather than raw text blobs.
Built for pipelines
Designed to slot into existing lakehouse workflows: parquet, CSV and SQL sources without a bespoke ingestion step.
Architecture
- MCP server exposing typed data tools to any compatible client
- PySpark executor for distributed transformations
- DuckDB executor for sub-second analytical queries
- Schema + profiling layer returned with every tool call
Stack
- PySpark
- DuckDB
- MCP
- Python
Open-source Model Context Protocol project for big data processing using PySpark and DuckDB.
Interested?