1 Guide
This guide starts with a dataframe and then shows how to select columns, choose another representation, and interpret missing values and metadata. If you already know which procedure you need, turn to the Reference.
1.1 Installation and a first dataset
Use Racket 9.3 or later. Install from the Racket catalog:
raco pkg install --auto datasets
This also installs datasets-core and the adapter dependencies. From a checkout of the repository, install both packages from its root:
raco pkg install ./datasets-core |
raco pkg install --auto ./datasets |
The two top-level directories are independently installable multi-collection packages. datasets-core provides datasets/core; datasets adds the convenient loaders and adapters to the same collection. The Polars package supplies its native library. With Nix, nix develop prepares an isolated environment containing both packages and that library.
> (require datasets (prefix-in pl: polars)) ffi-lib: could not load foreign library
path: libcairo.so.2
system error: libcairo.so.2: cannot open shared object
file: No such file or directory
> (define iris (load-iris)) load-iris: undefined;
cannot reference an identifier before its definition
in module: top-level
> (pl:dataframe-height iris) pl:dataframe-height: undefined;
cannot reference an identifier before its definition
in module: top-level
> (pl:dataframe-column-names iris) pl:dataframe-column-names: undefined;
cannot reference an identifier before its definition
in module: top-level
The prefix keeps Polars operations distinct from Racket’s arithmetic and other similarly named procedures. load-iris returns 150 rows: four flower measurements and a species label. Each call creates an independent dataframe. Data is bundled with the package, so loading does not download anything.
1.2 Discovering datasets and selecting columns
dataset-names lists the available identifiers. A named loader, such as load-mtcars, is shorthand for (load-dataset 'mtcars). Both accept #:columns and #:format.
> (dataset-names) dataset-names: undefined;
cannot reference an identifier before its definition
in module: top-level
> (define cars (load-dataset 'mtcars #:columns '("model" "mpg"))) load-dataset: undefined;
cannot reference an identifier before its definition
in module: top-level
> (pl:dataframe-column-names cars) pl:dataframe-column-names: undefined;
cannot reference an identifier before its definition
in module: top-level
> (hash-ref (dataset-info 'mtcars) "identifiers") dataset-info: undefined;
cannot reference an identifier before its definition
in module: top-level
Column names are strings. A selection preserves the order you give and retains all source rows. Identifiers such as "model" are real columns, so they can be kept for display or left out of a numeric matrix. Duplicate names, empty selections, and unknown columns are reported as errors.
1.3 Grouping categorical data
Category labels remain strings in Polars. For example, Iris species labels can be used directly in a group operation:
> (define by-species (pl:~> iris (pl:group-by "species") (pl:agg (pl:mean (pl:col "sepal-length"))))) pl:~>: undefined;
cannot reference an identifier before its definition
in module: top-level
> (pl:dataframe-height by-species) pl:dataframe-height: undefined;
cannot reference an identifier before its definition
in module: top-level
> (pl:dataframe-column-names by-species) pl:dataframe-column-names: undefined;
cannot reference an identifier before its definition
in module: top-level
The result has one row for each species. Grouped output need not follow category order; the intended level order is recorded separately in metadata:
> (define species-info (list-ref (hash-ref (dataset-info 'iris) "columns") 4)) dataset-info: undefined;
cannot reference an identifier before its definition
in module: top-level
> (hash-ref species-info "levels") species-info: undefined;
cannot reference an identifier before its definition
in module: top-level
Keeping labels avoids introducing arbitrary numeric codes. See Value representations for the representation of each value type.
1.4 Handling missing observations
Airquality has missing measurements. The table format uses dataset-missing; the Polars adapter translates that singleton to pl:polars-null.
> (define air (load-airquality)) load-airquality: undefined;
cannot reference an identifier before its definition
in module: top-level
> (pl:polars-null? (pl:series-ref (pl:dataframe-column air "ozone") 4)) pl:polars-null?: undefined;
cannot reference an identifier before its definition
in module: top-level
> (define air-table (load-airquality #:format 'table)) load-airquality: undefined;
cannot reference an identifier before its definition
in module: top-level
> (define ozone (dataset-table-column air-table "ozone")) dataset-table-column: undefined;
cannot reference an identifier before its definition
in module: top-level
> (for/sum ([value (in-vector ozone)]) (if (dataset-missing? value) 1 0)) ozone: undefined;
cannot reference an identifier before its definition
in module: top-level
Missing values are distinct from zero and #f. Decide whether to exclude, replace, or otherwise account for them in your analysis. The loaders preserve them; matrix conversion reports a missing value instead of silently removing it.
1.5 Preparing a numeric matrix
Choose 'matrix for a math/matrix matrix. Rows are observations, and the requested column order determines the matrix’s columns. For Iris, select the four measurements explicitly to leave out the species label:
> (require math/matrix)
> (define measurements (load-iris #:format 'matrix #:columns '("sepal-length" "sepal-width" "petal-length" "petal-width"))) load-iris: undefined;
cannot reference an identifier before its definition
in module: top-level
> (matrix-num-rows measurements) measurements: undefined;
cannot reference an identifier before its definition
in module: top-level
> (matrix-num-cols measurements) measurements: undefined;
cannot reference an identifier before its definition
in module: top-level
> (matrix-ref measurements 0 0) measurements: undefined;
cannot reference an identifier before its definition
in module: top-level
Trying to include a label or a missing observation produces an error naming the column and the zero-based row:
> (load-iris #:format 'matrix #:columns '("species")) load-iris: undefined;
cannot reference an identifier before its definition
in module: top-level
> (load-airquality #:format 'matrix #:columns '("ozone")) load-airquality: undefined;
cannot reference an identifier before its definition
in module: top-level
Diabetes already has ten standardized predictors in the lars representation. See its entry in the Dataset Catalog and Provenance before comparing it with another library’s dataset; raw measurements and exact scikit-learn parity are not promised.
1.6 Using mutable data-frame objects
The 'data-frame format works with the Racket data-frame library. Its NA value is explicitly set to dataset-missing.
> (require (prefix-in df: data-frame) datasets/data-frame) instantiate-linklet: mismatch;
reference to a variable that is uninitialized
name: cairo-lib
exporting instance: "/home/root/racket/share/pkgs/draw-lib
/racket/draw/unsafe/cairo-lib.rkt"
importing instance: "/home/root/racket/share/pkgs/draw-lib
/racket/draw/unsafe/cairo.rkt"
possible reason: modules need to be recompiled because
dependencies changed
possible solution: running `racket -y`, `raco make`, or
`raco setup`
> (define frame (load-airquality #:format 'data-frame #:columns '("ozone" "temp"))) load-airquality: undefined;
cannot reference an identifier before its definition
in module: top-level
> (data-frame-column-names frame) data-frame-column-names: undefined;
cannot reference an identifier before its definition
in module: top-level
> (df:df-is-na? frame "ozone" (df:df-ref frame 4 "ozone")) df:df-is-na?: undefined;
cannot reference an identifier before its definition
in module: top-level
The underlying library stores columns by name and does not promise an order from df:df-series-names. data-frame-column-names retrieves the initial selection order recorded by this adapter.
Changes to one frame do not affect a later load:
> (define first (load-iris #:format 'data-frame)) load-iris: undefined;
cannot reference an identifier before its definition
in module: top-level
> (df:df-set! first 0 -100 "sepal-length") df:df-set!: undefined;
cannot reference an identifier before its definition
in module: top-level
> (df:df-ref first 0 "sepal-length") df:df-ref: undefined;
cannot reference an identifier before its definition
in module: top-level
> (df:df-ref (load-iris #:format 'data-frame) 0 "sepal-length") df:df-ref: undefined;
cannot reference an identifier before its definition
in module: top-level
1.7 Interpreting counts and time series
Titanic’s rows are contingency cells, not individual passengers. Sum the "frequency" column to count people, and retain zero-frequency rows when you need the complete table of combinations:
> (define counts (load-titanic #:format 'table)) load-titanic: undefined;
cannot reference an identifier before its definition
in module: top-level
> (dataset-table-row-count counts) dataset-table-row-count: undefined;
cannot reference an identifier before its definition
in module: top-level
> (for/sum ([count (in-vector (dataset-table-column counts "frequency"))]) count) dataset-table-column: undefined;
cannot reference an identifier before its definition
in module: top-level
AirPassengers contains monthly international airline passenger counts, measured in thousands. Its year and month columns make the chronology explicit:
> (define monthly (load-air-passengers #:format 'matrix)) load-air-passengers: undefined;
cannot reference an identifier before its definition
in module: top-level
> (for/list ([column (in-range 3)]) (matrix-ref monthly 0 column)) monthly: undefined;
cannot reference an identifier before its definition
in module: top-level
> (for/list ([column (in-range 3)]) (matrix-ref monthly 143 column)) monthly: undefined;
cannot reference an identifier before its definition
in module: top-level
> (hash-ref (dataset-info 'air-passengers) "time-series") dataset-info: undefined;
cannot reference an identifier before its definition
in module: top-level
The same metadata interface records source citations, licenses, units, identifiers, and file checksums for every dataset. The catalog is generated from that registry rather than maintained as a second independent list.
1.8 Using core data in another package
A package that constructs its own dataframes can depend on datasets-core and use datasets/core directly. This is particularly useful for Polars documentation: depending only on the core avoids a dependency cycle through datasets.
The following example uses core data with Polars’ constructors:
> (require datasets/core) > (define table (load-dataset-table 'iris #:columns '("sepal-length")))
> (define own-frame (pl:dataframe-new (list (pl:series-new-f64 "sepal-length" (vector->list (dataset-table-column table "sepal-length")))))) pl:dataframe-new: undefined;
cannot reference an identifier before its definition
in module: top-level
> (pl:dataframe-height own-frame) pl:dataframe-height: undefined;
cannot reference an identifier before its definition
in module: top-level
The repository’s examples/core-with-polars.rkt demonstrates the same pattern and is tested before the adapter package is installed. For downstream Nix, fetch this repository with flake = false and install datasets-core/ from that source. This keeps the complete flake and its native dependencies out of the downstream dependency graph.