Skip to main content
100%

Comparing the performances of various DBMS using dplyr for large-ish data

✓ Published0🌍 Public
RRCura
Last edited Nov 20, 2017
Created on Nov 20, 2017

This example compares the execution times of data manipulation operations across different database management systems using the `dplyr` and `dbplyr` R packages. It benchmarks a `group_by` with `summarise` and a `left_join` with `mutate` on a 50-million-row tibble, running the same queries against in-memory data frames, SQLite databases (in-memory and file-based), and PostgreSQL. The code uses `system.time` to measure the elapsed user and system times for each operation, recording results in a benchmark summary table. The visualization highlights performance differences across database backends for typical data wrangling tasks.

AI-generated description

Similar vizzes