MEET RUMBLEDB 3.0 / Coast Redwood

The Swiss army
knife for data.

Big data. Nested data. Beautifully messy data.
Query, transform, and make sense of it all with one expressive language.

✳   Open source, Apache 2.0↗   From laptop to cluster95% W3C compliance
Runs on Amazon EMR

Query your data directly from S3 and HDFS.

↗
GROW BEYOND THE FRAME3.0
An illustrated coast redwood with branching layers, representing nested data and growth from laptop to cluster. nested JSON XML Parquet CSV Text limitless possibilities
SEQUOIA SEMPERVIRENSDeep roots. New heights.

Built for the data you actually have.

JSON & JSON LinesParquetCSVXMLApache Spark

ONE ENGINE. SO MANY POSSIBILITIES.

Your data toolkit.
All unfolded.

From a quick question to a complete data pipeline. RumbleDB brings the tools together, so you can stay in the flow.

01

Query the way data lives.

Nested objects, arrays, missing fields, mixed types. Explore real-world data with JSONiq, a language made for it.

NESTED & HETEROGENEOUS DATA
02

Give your data a new shape.

Clean, filter, join, and transform. Turn sprawling datasets into exactly the structures your next step needs.

CLEAN & TRANSFORM
03

Bring structure to the wild.

Validate with JSound schemas and convert between formats. Take your JSON all the way to efficient Parquet.

VALIDATE & CONVERT
04

Start small. Think enormous.

Work locally, then scale across a cluster with Apache Spark—including Amazon EMR with data on S3 and HDFS. The same language goes with you.

LOCAL & DISTRIBUTED EXECUTION
05

Meet data where it is.

Read from your local drive, HDFS, S3, or Azure blob storage. Query JSON, XML, CSV, Parquet, Avro, and more. Work with data lakehouse tables in Delta Lake and Apache Iceberg.

FILES, FORMATS & DATA LAKEHOUSE
06

Machine learning, as functions.

RumbleML brings Spark ML estimators and transformers into JSONiq as function items. Prepare data, train models, and make predictions in one language.

EXPLORE RUMBLEML

LESS CEREMONY. MORE DISCOVERY.

A little query.
A lot of possibility.

Express what you want in a few readable lines. JSONiq combines familiar query concepts with the flexibility of nested data.

THE JSONIQ FIELD GUIDEIllustrative examples
forests.jq
let $forest := [
  { "name": "Coast Redwood",
    "trees": [{ "height": 115 }, { "height": 96 }] }
]
for $grove in $forest[]
for $tree in $grove.trees[]
where $tree.height > 100
return { "species": $grove.name,
         "height": $tree.height }
RESULT● 1 object
{
  "species": "Coast Redwood",
  "height": 115
}

Walk through nested arrays naturally. No flattening required.

BUILT ON SOLID GROUND
95%

W3C compliance

Deep roots in standards.
Room to grow.

RumbleDB 3.0 brings standards-based querying to your data toolkit, with JSONiq and support for W3C XQuery. Expressive languages. A foundation you can build on.

Explore language support ↗

THE WHOLE TOOLKIT. INCLUDED.

Every tool.
One open invitation.

From expressive queries to distributed execution. Explore everything that comes with RumbleDB.

ONE TIER. EVERYTHING INCLUDED.

Free & open source

$0Apache License 2.0

The complete RumbleDB toolkit, free and open source.

We do not even need your data or any registration or contact details.

Every feature below is included.

Languages & reusable building blocks

  • JSONiq language
  • XQuery language
  • Library modules
  • Very large set of builtin functions
  • Importing library modules over HTTP(S)

Queries, transformations & windows

  • FLWOR expressions (more powerful than SQL’s SELECT FROM WHERE)
  • For clause
  • Let clause
  • Where clause
  • Group by clause
  • Order by clause
  • Count clause
  • Tumbling and sliding window clauses
  • Selection
  • Projection
  • Aggregation
  • Sorting
  • Zipping with position
  • Joins between collections

Types, schemas & validation

  • Rich W3C-compliant type system
  • Typed data
  • Static typing
  • XML Schema validation
  • XML Schema annotation
  • JSound validation

Functions, mathematics & logic

  • Higher-order functions
  • Arithmetic
  • W3C-compliant mathematical builtin functions
  • First-order logic
  • Universal and existential quantification on large sequences
  • General comparison
  • Value comparison
  • String concatenation

Control flow & error handling

  • Conditional expressions
  • Switch expressions
  • Typeswitch expressions
  • Try-catch expressions
  • While loops
  • Exit expressions
  • Break expressions
  • Continue expressions

Updates & scripting

  • JSONiq Update Facility
  • Variable assignment
  • Scripting and side effects (based on snapshot semantics)

Input formats

  • Reading JSON
  • Reading JSON Lines
  • Reading XML
  • Reading CSV
  • Reading Parquet
  • Reading Text
  • Reading Avro
  • Reading libSVM
  • Reading ROOT

Output formats

  • Output to JSON
  • Output to JSON Lines
  • Output to XML
  • Output to HTML
  • Output to XHTML
  • Output to CSV
  • Output to Parquet
  • Output to Avro
  • Output to YAML
  • Output to Text

Storage & data lakehouses

  • Read from local file system
  • Write to local file system
  • Read from HDFS
  • Write to HDFS
  • Read from Amazon S3
  • Write to Amazon S3
  • Read from Azure blob storage
  • Write to Azure blob storage
  • Read from HTTP(S)
  • Delta Lake support
  • Apache Iceberg support

Database connectivity

  • Reading PostgreSQL tables (Python edition)
  • Reading MongoDB collections (Python edition)
  • Accessing Hive metastore tables

Scale, execution & optimization

  • kBs of data
  • MBs of data
  • GBs of data
  • TBs of data
  • PBs of data
  • Execution on a laptop
  • Execution on Amazon EMR (8.x)
  • Compatible with Spark 4.0
  • Compatible with Spark 4.1
  • Compatible with Spark 4.2
  • Automatic parallelization of FLWOR clauses on cores or a cluster
  • Automatic parallelization of JSON navigation on cores or a cluster
  • Automatic parallelization of downward XML navigation on cores or a cluster
  • Automatic pushdowns (count, sum, …)
  • Automatic optimizations

Machine learning with RumbleML

  • Spark ML estimators and transformers as JSONiq function items
  • Model training
  • Prediction with trained models
  • Feature preprocessing
  • Machine learning pipelines

Python interoperability & result retrieval

  • Binding Python values to query variables
  • Binding PySpark DataFrames to query variables
  • Retrieving results as PySpark DataFrames
  • Interoperability with Spark SQL
  • Streaming results through an iterator
  • Access to native typed items through the RumbleDB Item API
  • Retrieving heterogeneous results as RDDs (experimental)

APIs, integrations & installation

  • Compatible with Python
  • Compatible with Pandas
  • Java API
  • Available as a command line utility
  • Standalone jar
  • Thin jar to use with a Spark installation
  • Available via Homebrew
  • Available via Docker
  • Interactive query shell
  • Jupyter notebook integration
  • Installation via pip install jsoniq
  • VS Code support with syntax highlighting

A NEW CHAPTER

Meet Coast Redwood.

RumbleDB 3.0. Named for a tree that reaches extraordinary heights. Made for data that does, too.

One versatile engine, from your first JSON query to your next data pipeline. Open source, built on Apache Spark, and ready to explore.

THE COAST REDWOOD EDITION
3.0
RUMBLEDBROOTED IN POSSIBILITY.

YOUR NEXT DISCOVERY STARTS HERE

Bring your data.
We’ll bring the tools.

Open source. Made for the curious.