The files
Every answer on the dashboard is DuckDB reading Parquet out of a public S3 bucket it does not own. Nothing is downloaded: it fetches the footer first, then only the byte ranges of the columns a query touches. This is one file of 8 in the TMAX 2025 partition, out of 14,031 in the dataset.
Read over HTTPS from noaa-ghcn-pds.s3.amazonaws.com/parquet/by_year/YEAR=2025/ELEMENT=TMAX/a88c3528a93343af9532898d6ef5b484_0.snappy.parquet
Rows
500K
1 row groups
File size
1.7 MB
footer 3.6 kB
Columns
7
2 withheld from this role
Compression
1.1×
written by parquet-cpp-arrow 7.0.0
What a query fetches
Pick columns as if writing a SELECT — the map shows the byte ranges DuckDB requests from the CDN, drawn to scale from the file's real chunk offsets.
634.1 kB
of 1.7 MB (36.4%) · ~2 range requests + footer
Fetched by this query
Skipped
Withheld by policy
Footer (always read)
Columns on disk
| Column | Parquet type | Codec | Compressed | Share of data | Ratio |
|---|---|---|---|---|---|
| Value | INT64 | SNAPPY | 630.1 kB | 1.0× | |
| Source flag | BYTE_ARRAY | SNAPPY | 57.7 kB | 2.0× | |
| Measurement flag | BYTE_ARRAY | SNAPPY | 21.7 kB | 1.4× | |
| Quality flag | BYTE_ARRAY | SNAPPY | 4.6 kB | 1.2× | |
| Observed on | BYTE_ARRAY | SNAPPY | 454 B | 1.7× | |
2 columns withheld from this role | 1.0 MB | ||||
Read 13.4 MB from 8 of 14,031 files · 9.20 GB dataset · fetched once, then read from memory
0.1456% of the dataset