SQL editor
Worksheets with a schema browser, parameters, explain, formatting, run history and result charts. A command palette gets you anywhere from the keyboard.
Cloud-neutral data and AI platform
Query one governed Iceberg catalog with DataFusion, Ballista, Apache Spark, Sail or SparkHouse. Run it in the cloud and region you choose, or in your own Kubernetes cluster.
The problem
Take a US hyperscaler's SaaS, and the data leaves your jurisdiction. Build a lakehouse from parts, and you spend a year on platform work before the first analyst runs a query.
Either way, one engine rarely fits every workload, and every vendor catalog turns into its own silo.
SparkHouse is the control plane in the middle. It decides where compute runs, which engine runs each statement, and who can see what, while your tables stay in open Iceberg.
How it works
Pick a region and cloud, or point SparkHouse at your own Kubernetes cluster. Attach a residency policy by country, cloud and region.
Use the Iceberg catalog every organization gets, or federate AWS Glue, Databricks Unity Catalog, Snowflake or any Iceberg REST catalog.
Send SQL to one endpoint. The router picks the engine for each statement, and every run shows which one it chose.
Sovereignty
Workspace
SQL editor, notebooks, jobs and a catalog browser share one workspace. Every statement goes through the same router, policies, audit log and metering.
Worksheets with a schema browser, parameters, explain, formatting, run history and result charts. A command palette gets you anywhere from the keyboard.
SQL, Python and Markdown cells in one document, with results, tables and charts under each cell. Import and export Jupyter .ipynb.
Ordered SQL tasks with cron schedules and time zones, retries, timeouts, logs and cancel. Notebooks run as job tasks, with parameters you can override per run.
Schemas, tables, views, columns, sample data and history, with comments, tags and permissions beside them.
Every statement is a run: who ran it, on which engine, where it was placed and how long it took.
Compute and notebook uptime metered per user, endpoint, engine and placement.
spark in a kernel is a Spark Connect session.%%sql SELECT region, sum(amount) AS revenue FROM sales.orders WHERE day = '{{day}}' GROUP BY region
df = sh.sql("SELECT * FROM sales.daily")
display(df)# Notes for the finance reviewEngines
AUTO routes each statement by lane (write, heavy, interactive) and by lowest effective cost. The run view shows the engine and why it was chosen.
Fast interactive queries, with warm pools on dedicated data planes.
The same Arrow-native stack, spread across workers for heavy statements.
Familiar engine for heavy ETL, reached over Spark Connect, standalone or on native Kubernetes.
Spark-compatible engine written in Rust on DataFusion, with no JVM. Runs as a single server or as a driver that launches its own workers on Kubernetes.
Our own Go engine. It is in the product and still maturing, so we label it as a preview.
Open catalog
Every organization gets its own Apache Polaris catalog on highly available PostgreSQL. Browse schemas, tables, views, columns, sample data and history in the Catalog Explorer, and manage comments, tags and permissions beside them.
Federate what you already run
Federated tables use per-table vended credentials, and you choose read-only or writable.
Serverless and cost
Governance and operations
For developers
shctl covers endpoints, clusters, SQL, jobs, runs, catalog, audit and benchmarks.
# sign in to your SparkHouse host shctl login --host https://sparkhouse.example --token shp_... # create a Ballista cluster with four workers shctl clusters create --name etl --engine ballista --workers 4 # run a statement on it shctl sql --endpoint etl -e "SELECT count(*) FROM lineitem"
Roadmap
Planned, order not committed. If one of these blocks you, tell us in the pilot request.
FAQ
In the data plane you choose: a region on GCP or AWS, or your own Kubernetes cluster. The control plane holds metadata, not table data.
Yes. The data-plane agent runs engines in your cluster and connects outbound to the control plane.
GCP, AWS, and any Kubernetes cluster.
The engines and formats underneath are open source (Iceberg, Polaris, DataFusion, Ballista, Apache Spark, Sail). The SparkHouse platform itself is proprietary.
Inside a platform notebook, yes: on an Apache Spark endpoint, spark is a Spark Connect session. Pointing your own PySpark at SparkHouse from outside is not available yet, and a Spark Connect gateway is planned.
Through pilot terms. There is no self-serve signup or public price list yet.
Pilot
We are working with a small number of design partners who have a residency or sovereignty constraint and several teams on one lakehouse. Tell us about your setup and we will reply with next steps.