Query the way data lives.
Nested objects, arrays, missing fields, mixed types. Explore real-world data with JSONiq, a language made for it.
NESTED & HETEROGENEOUS DATABig data. Nested data. Beautifully messy data.
Query, transform, and make sense of it all with one expressive language.
Query your data directly from S3 and HDFS.
Built for the data you actually have.
ONE ENGINE. SO MANY POSSIBILITIES.
From a quick question to a complete data pipeline. RumbleDB brings the tools together, so you can stay in the flow.
Nested objects, arrays, missing fields, mixed types. Explore real-world data with JSONiq, a language made for it.
NESTED & HETEROGENEOUS DATAClean, filter, join, and transform. Turn sprawling datasets into exactly the structures your next step needs.
CLEAN & TRANSFORMValidate with JSound schemas and convert between formats. Take your JSON all the way to efficient Parquet.
VALIDATE & CONVERTWork locally, then scale across a cluster with Apache Spark—including Amazon EMR with data on S3 and HDFS. The same language goes with you.
LOCAL & DISTRIBUTED EXECUTIONRead from your local drive, HDFS, S3, or Azure blob storage. Query JSON, XML, CSV, Parquet, Avro, and more. Work with data lakehouse tables in Delta Lake and Apache Iceberg.
FILES, FORMATS & DATA LAKEHOUSERumbleML brings Spark ML estimators and transformers into JSONiq as function items. Prepare data, train models, and make predictions in one language.
EXPLORE RUMBLEMLLESS CEREMONY. MORE DISCOVERY.
Express what you want in a few readable lines. JSONiq combines familiar query concepts with the flexibility of nested data.
let $forest := [
{ "name": "Coast Redwood",
"trees": [{ "height": 115 }, { "height": 96 }] }
]
for $grove in $forest[]
for $tree in $grove.trees[]
where $tree.height > 100
return { "species": $grove.name,
"height": $tree.height }{
"species": "Coast Redwood",
"height": 115
}Walk through nested arrays naturally. No flattening required.
W3C compliance
RumbleDB 3.0 brings standards-based querying to your data toolkit, with JSONiq and support for W3C XQuery. Expressive languages. A foundation you can build on.
Explore language support ↗THE WHOLE TOOLKIT. INCLUDED.
From expressive queries to distributed execution. Explore everything that comes with RumbleDB.
$0Apache License 2.0
The complete RumbleDB toolkit, free and open source.
We do not even need your data or any registration or contact details.
Every feature below is included.
A NEW CHAPTER
RumbleDB 3.0. Named for a tree that reaches extraordinary heights. Made for data that does, too.
One versatile engine, from your first JSON query to your next data pipeline. Open source, built on Apache Spark, and ready to explore.
YOUR NEXT DISCOVERY STARTS HERE
Open source. Made for the curious.