Working with Quebec teaching snapshots

UlavalSSD provides fixed snapshots for practising data cleaning and exploratory analysis. These examples run offline and use base R. The same objects can be used with the tidyverse when those packages are installed.

library(UlavalSSD)

Dates and missing measurements

MeteoQuebec retains the raw column names and types expected by the course. Month and day are character strings. The first column, ...1, is an export identifier, not a measurement. Explicitly converting to a base data frame makes subsetting behaviour independent of whether tibble is installed.

weather <- as.data.frame(MeteoQuebec)
weather$date <- as.Date(with(weather, paste(year, month, day, sep = "-")))
range(weather$date)
#> [1] "1970-01-01" "2025-01-22"
head(weather[c("date", "min_temp", "max_temp")])
#>         date min_temp max_temp
#> 1 1970-01-01    -19.4    -12.8
#> 2 1970-01-02    -19.4    -12.8
#> 3 1970-01-03    -18.9    -13.3
#> 4 1970-01-04    -20.6    -13.3
#> 5 1970-01-05    -21.1    -12.8
#> 6 1970-01-06    -18.3    -10.0

A complete sequence of dates does not mean every measurement is observed. Count the missing values before choosing a summary or a model.

measurements <- c(
  "max_temp", "mean_temp", "min_temp", "total_precip",
  "total_rain", "total_snow", "snow_grnd"
)
data.frame(
  variable = measurements,
  missing = vapply(weather[measurements], function(x) sum(is.na(x)), integer(1)),
  missing_proportion = vapply(
    weather[measurements], function(x) mean(is.na(x)),
    numeric(1)
  ),
  row.names = NULL
)
#>       variable missing missing_proportion
#> 1     max_temp      53        0.002635374
#> 2    mean_temp      54        0.002685098
#> 3     min_temp      40        0.001988961
#> 4 total_precip      97        0.004823231
#> 5   total_rain   10490        0.521605092
#> 6   total_snow   10587        0.526428323
#> 7    snow_grnd    8912        0.443140570

For a descriptive plot, retain dates alongside observations. Missing values remain visible as gaps rather than being replaced by zero.

one_year <- weather[weather$year == 2024, , drop = FALSE]
plot(one_year$date, one_year$min_temp,
  type = "l",
  xlab = "Date", ylab = "Daily minimum temperature (degrees C)"
)

Daily minimum temperatures during 2024 in the historical teaching snapshot.

An independent reconstruction from the official ECCC service matched every stored column on 2026-09-22. Station 5251 is used through 1995 and station 26892 from 1996. These data support cleaning exercises; a climate-trend study would require checking station history, measurement flags and consistency first.

Administrative records and formatted amounts

listecondamnation contains historical records. It is not a representative sample of food establishments. One establishment may appear more than once. Counts of these records cannot estimate the probability of an offence or an establishment’s present-day operating conditions.

records <- as.data.frame(listecondamnation)
range(as.Date(records$Date_publication))
#> [1] "2023-02-13" "2025-02-10"
head(sort(table(records$Type_etablissement), decreasing = TRUE))
#> 
#>                 RESTAURANT       REST. SERVICE RAPIDE 
#>                       1353                        187 
#>  RESTAURANT SERVICE RAPIDE RESTAURANT METS A EMPORTER 
#>                        137                         35

Amende is text. Examine its formats before parsing currency. The following example validates a deliberately small set of known formats, rather than silently discarding every non-numeric character in an arbitrary string.

amount_text <- c("1 500 $", "250,50 $", NA_character_)
amount_clean <- gsub("[ $]", "", amount_text)
amount_clean <- sub(",", ".", amount_clean, fixed = TRUE)
valid <- is.na(amount_clean) | grepl("^[0-9]+([.][0-9]{1,2})?$", amount_clean)
stopifnot(all(valid))
amount <- as.numeric(amount_clean)
data.frame(amount_text, amount)
#>   amount_text amount
#> 1     1 500 $ 1500.0
#> 2    250,50 $  250.5
#> 3        <NA>     NA

This example uses synthetic amounts. Applying a conversion to the full dataset requires checking all observed formats and preserving the original text for verification.

Exercise prompts and feedback

French output remains the default. The row feedback is a fixed answer key for the original penguin exercise file distributed on the course website. It is valid only before rows are filtered or reordered.

consulter_taches("histogramme", lang = "en")
#> [1] "Show the distribution of penguin flipper lengths. There seems to be an error in the data. Can you find it?"
verifier_valeur_aberrante(11, lang = "en")
#> [1] "Exercise correction: body mass should be 3300 g (3.3 kg) and bill length should be 37.8 mm."

Static style feedback

Install the optional package lintr to use eval_tidyverse_style(). It examines code without executing it. In a Quarto document it inspects fenced R chunks and reports line numbers in the original file.

if (requireNamespace("lintr", quietly = TRUE)) {
  path <- tempfile(fileext = ".R")
  writeLines(c("daily_mean <- mean(c(2, 4, 6))", "print(daily_mean)"), path)
  feedback <- eval_tidyverse_style(path)
  feedback[c("total", "status", "diagnostics")]
  unlink(path)
}

The score is a transparent formative indicator based on seven observable criteria. It is not a validated assessment of programming competence. Empty or invalid code has no score, and the clarity of reasoning and usefulness of comments remain for human review. See ?eval_tidyverse_style for the criteria and Quarto fence support.

Sources and snapshot limits

Weather observations are attributed to Environment and Climate Change Canada; administrative records are attributed to MAPAQ via Donnees Quebec. See ?MeteoQuebec, ?listecondamnation, and the installed attribution file:

system.file("COPYRIGHTS", package = "UlavalSSD")
#> [1] "/tmp/Rtmp3kXTFK/Rinst16ba4426980a/UlavalSSD/COPYRIGHTS"

The datasets do not refresh during installation, loading, examples or tests.