Back to feed

CPU vs GPU SQL: How Velox and cuDF Make Presto Up to 4x Faster

An IBM Developer demo shows open-source Presto running on GPU via Velox (C++ engine) and NVIDIA cuDF, cutting the same query on the 179-million-row IBM AML dataset from 6 seconds to 1.6 seconds — the only change is switching the Docker image to gpu-nightly.

Imported to Nodesdaily: (UTC+03:00)
Watch on YouTube — KRmln6u9QNk
Reading options

Device speech is unavailable in this browser.

Concept lens

Choose a technical term in this view to read its general definition, teaching example and use in the article.

No terms from our glossary were found in this view. The glossary does not cover every term yet.

We start from a bottleneck every analytics team knows: Presto was born as a distributed SQL engine for big data on CPU, and as data grows query time balloons. In the video, Roy and William from NVIDIA show how to break that ceiling with GPU. The work comes from IBM, NVIDIA and the Presto community; the goal is to add a GPU-accelerated engine on top of an existing warehouse without moving data. The promise is simple: instead of Java-based Presto, run Presto with Velox — a C++ execution layer — so the same SQL fans out on parallel hardware. Think of it as widening a single-lane road to eight lanes — same car, more lanes.

A quick grounding on Presto and Velox helps. Presto started at Facebook in 2012 as an open-source distributed SQL engine that separates storage and compute, scaling via coordinator and workers. In 2019 the community forked into PrestoDB and Trino; PrestoDB today lives under prestodb/presto and the Presto Foundation, hosted by the Linux Foundation. Classic Presto runs on Java. Velox, built by Meta, is a vectorized, columnar C++ execution library — the heart of Presto C++ / Presto Native. It sheds JVM overhead and executes the same query plan in vectors close to hardware. Paired with cuDF, that plan scatters across GPU cores. Analogy: replacing a Java orchestra with a small, fast C++ crew that steps directly onto the stage.

The GPU magic comes via NVIDIA cuDF and its newer extension cuDF-X. Part of RAPIDS (now NVIDIA CUDA-X), cuDF is a GPU DataFrame library — backed by libcudf and pylibcudf — that runs joins, aggregates, filters and sorts in parallel across thousands of CUDA cores. It mirrors the video: take a data frame, join, aggregate, compute — all accelerated on GPU. The bottleneck is moving data into GPU memory (VRAM). William's point is key: pushing a 100-row query to GPU makes no sense; transfer cost eats the gain. But on 179 million rows, grouped and sorted in bulk, parallelism wins clearly. 2026 GB200 NVL72 and DGX B200 runs echo the same insight — if data can move between operators without leaving GPU memory via NVLink, latency drops.

The hands-on setup in the video is three notebooks, all in the Presto Foundation's Prestorials repository. The first notebook brings up a Presto Native cluster with Docker Compose. Steps are crisp: create metastore and warehouse directories on the host, write Presto config files, then start coordinator and worker via docker compose. A verification cell shows both components Running — a CPU coordinator and a CPU worker. Then a few simple queries confirm health. This is a minimal Presto lab that operates on host data instead of building a warehouse from scratch — a one-command playground infra teams love.

The second notebook adds large data. The dataset is IBM's synthetic transactions for anti-money laundering (AML). You find it on Kaggle as "IBM Transactions for Anti Money Laundering" and also on Hugging Face; row count varies by scale, about 176 million rows at large scale. It mimics financial transactions — bank transfers, purchases, credit-card events — and is widely used for GNN and foundation-model research. The notebook loads it into the warehouse and again verifies the cluster is up with simple queries. This step sets the benchmark's realistic volume; synthetic yet prod-like in shape and scale for an analytical workload.

The third notebook delivers the new part: copy the existing warehouse and add the GPU engine on top. In Explorer you see the config and metadata duplicated verbatim, then a few extra GPU properties are added and the GPU cluster is started. Startup looks fast, but the first verification still says Starting; a second try shows Running. Then a smoke test: select star limit 1. William notes the first query takes about a second because it brings up binaries in VRAM — that warm-up cost is expected. Then a deliberate design choice: the GPU side is read-only. The reason is to avoid two writers corrupting the same files while reports and jobs still run on the CPU cluster; the GPU engine just reads. If you plan to run only the GPU engine, you can drop the restriction.

And the head-to-head. The same query — not monstrous, but grouping and sorting over all rows and returning ten rows — is sent to both engines in sequence. The code opens separate connections and cursors for CPU and GPU in a loop and times each run. The result on screen is crisp: CPU about 6 seconds, GPU 1.6 seconds. That is a 3.75× speedup, rounded to four to five times in the narration. The Docker delta is equally minimal: coordinator image prestodb/presto:latest versus gpu-nightly, worker prestodb/presto-native:latest versus gpu-nightly; a side-by-side comparison cell in Jupyter's light theme highlights the two-line difference. So copy the warehouse, flip the image tag, add a few GPU properties — the speed win is that close.

Zooming out, the picture comes together. The three notebooks in Prestorials show how to place a second, faster tier next to a production warehouse without moving data and while keeping the same SQL. The IBM + NVIDIA Velox + cuDF integration, as described in the March 2026 arXiv VLDB workshop paper, focuses on two hard problems: moving data from storage to GPU operators and keeping operator exchange inside GPU memory. Early evaluations on TPC-H-derived queries report up to 6× cost/performance gains. The video gives a concrete lab proof on 179 million rows: 6 versus 1.6 seconds. The closing is practical: if you want to try GPU-accelerated Presto, clone Prestorials, run the same notebooks, and share your own speedup in the comments.

Visualization: nodesdaily AI
PlatformTimeSpeedup
CPU (Presto Native)6.0 s1×
GPU (Presto + Velox/cuDF)1.6 s3.8×

AI commentary

"What struck me is how this demo reframes speed: you gain by swapping the execution engine, not the data. A 3.75× win on 176 million rows for a two-line Docker config delta is generous for a nightly build — yet the first-query VRAM warm-up and the fact that tiny queries still belong on CPU keep the story honest."

AI assessment

Steel-manning the opposite view, CPU remains the rational default for many jobs. On a 100-row filter or point lookup, the cost of moving data into VRAM erases the gain; for small, frequent, interactive queries and write-heavy workloads, JVM-based Presto and even Postgres-like systems stay simpler and cheaper. Moving existing pipelines, monitoring and security certifications to a gpu-nightly image also carries operational risk — a nightly tag promises no stability.

Limits and methodology notes are clear. The video compares two clusters on the same Docker host, on one node; network shuffle and multi-node distributed exchange are not measured. The first-query VRAM warm-up is not broken out as a line item, so second and later queries look better. VRAM capacity — even 13.5 TB HBM3e / 17 TB LPDDR5X on GB200 — is finite; tables larger than memory need spilling or partitioning. And the IBM AML data is synthetic; real prod skew, null ratios and cardinality can push the planner elsewhere.

Whose claim and what needs independent checking? The speed claim rests on the 6 versus 1.6 seconds in-video measurement and the up to 6× cost/performance note in the IBM-NVIDIA arXiv 2606.24647 VLDB workshop paper. For independent validation, the Prestorials notebooks are open — rerun the same data volume and instance type with TPC-H-derived queries; the GB200 NVL72 blog's DGX B200 numbers are a different hardware class, so read them as a scaling reference, not a direct comparison. The NVIDIA CUDA-X / cuDF docs and Prestodb documentation are the primary sources for what versions and nightly tags actually contain.

Practical takeaway: a GPU tier is attractive for repeated large scans — nightly reports and feature extraction with 100M+ rows and heavy aggregates/sorts — especially if you already run Velox-based Presto Native, where the switch is a two-line config delta. For interactive small queries, frequent writes and tight budgets, staying on CPU or evaluating alternatives' own acceleration paths (e.g., Spark RAPIDS) is healthier. Do not decide before measuring your query-size distribution and VRAM budget.

Sources

11 links; no other published story cites them. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.

presto · velox · gpu · nvidia cudf · sql · docker · ibm

Follow the topic

Before this story

A short reading order from earlier stories linked to this event by an editor.

Evidence and sources

Review permitted source passages, versions and origins.

KAYNAKLARLA OKU

Bu haberi açalım.

Hesap kontrol ediliyor…