The exported datasets now use a consistent
<grain>_<measure>_<frequency>
convention:
| 1.x name | 2.0 name |
|---|---|
passengers_entrance |
line_entries_monthly |
passengers_transported |
line_transported_monthly |
station_averages |
station_transported_monthly |
station_daily |
station_entries_daily |
lines |
rail_lines |
stations |
rail_stations |
station_inauguration |
Unshipped |
calendar_spo |
Unchanged |
metro_colors |
Unchanged |
Removed the published line_number = 99 system rows
from line_entries_monthly and
line_transported_monthly. They did not provide a consistent
whole-network measure. No clean network total exists across operators:
METRO line entries include transfers arriving from Lines 4 and 5, while
Lines 4 and 5 count turnstiles only. Do not sum transported counts
because interchange journeys are counted on every line used. Corrected
the operator for Line 5 in rail_lines and
rail_stations to ViaMobilidade (#22).
Renamed the station monthly dataset to
station_transported_monthly: METRO’s file measures
transported passengers (boardings plus transfers), not turnstile
entries. Summed over a line’s stations its mdu usually
comes within 2% of the line’s mdu in
line_transported_monthly.
Dropped Line 5 rows from August 2018 onward from
station_transported_monthly. The Dataverse feed records
turnstiles only after the ViaMobilidade handover, so those rows are
neither measure; station data for that era stays in
station_entries_daily. METRO-era Line 5 (January 2016 to
July 2018) is unchanged.
Added Line 4 to line_transported_monthly from
January 2012 with all five metrics, summing turnstile entries and
transfers from the Dataverse feed. Line 5 stays absent after August
2018: its feed records no transfers.
line_transported_monthly now counts individual
passengers, like the other demand datasets. METRO publishes transported
counts in thousands, and 1.x kept that unit; value is now
multiplied by 1000. read_metro_demand() rescales 1.x
release assets the same way.
Replaced metrosp_cache_dir(),
metrosp_cache_enable(), and
metrosp_cache_list() with metrosp_cache(),
which returns the cache listing and prints its location. Downloads now
use the platform-specific tools::R_user_dir() cache by
default; cache = FALSE,
options(metrosp.cache_dir = "/path"), and
METROSP_CACHE_DIR continue to opt out or override the
location. metrosp_cache_clear() removes only
package-managed vintage directories, preserving unrelated files in a
custom cache root. A cached vintage left unused for 90 days is deleted
on the next read (#26).
read_metro_demand() accepts the new dataset names
and detects 1.x asset names from each release manifest, including the
rolling data-latest release. A 1.x asset is translated to
the 2.0 contract on read — columns renamed, station_id
filled in, and the line_number = 99 rows dropped — so a
read returns the same shape and the same totals whichever vintage it
came from. An unrecognized vintage is now an error rather
than a silent fall back to the bundled snapshot (#25).
Unshipped the incomplete station_inauguration
dataset without replacement. Its source and builder remain under
data-raw/ for future curation (#25).
Added a stable, opaque station_id to both
station-demand datasets and the station geometry table. Named station
members retain their official station_name, while
officially documented physical interchange complexes share an ID across
lines and modes. Complex membership comes from a committed,
source-backed crosswalk rather than name or distance inference
(#24).
Standardized the four demand datasets on value,
renamed the metric columns to metric,
metric_name, and metric_name_pt, made
station_transported_monthly explicitly use the
mdu metric, and made year and
line_number integer columns. calendar_spo now
uses is_optional_holiday and is_long_weekend
(#23).
The metric rename changes meaning: in 1.x it contained
the English label, while in 2.0 it contains the stable abbreviated code.
Code that groups, filters, or pivots on metric columns should use this
mapping:
| 1.x column | 2.0 column | Meaning |
|---|---|---|
metric_abb |
metric |
Stable code such as total or mdu |
metric |
metric_name |
English display label |
metric_pt |
metric_name_pt |
Portuguese display label |
The measure columns passengers and
avg_passenger are now consistently named
value. The calendar columns
is_ponto_facultativo and is_feriadao are now
is_optional_holiday and is_long_weekend.
line_transported_monthly$value is 1000 times its 1.x
value. Remove any * 1000 your code applied before comparing
it with the entry counts.
Code that previously used line_number == 99 has no
replacement network total: METRO line entries include transfers from
Lines 4 and 5 while those lines count turnstiles only, and transported
counts double-count interchange journeys. Do not sum transported counts:
a passenger is counted on every line used, so interchange journeys would
be counted more than once.
Only vintages that were actually published can be pinned. The first
dated archive is data-2026-09; earlier year-month tags do
not exist.
Fixed read_metro_demand() keeping a dated vintage’s
cached manifest forever. Dated manifests now expire on
metrosp.cache_ttl like data-latest, so a month
republished after the first read is picked up. Monthly vintages are
revisable within their month, not immutable.
Fixed archived station-name lookup failing when sf
is attached.
Fixed the package description and the data article describing
station demand as entries. station_transported_monthly
reports average weekday passengers transported, and
station_entries_daily counts transfers from other operators
but not transfers between METRO lines.
Fixed metrosp_cache_clear() and
read_metro_demand() accepting vintage values
such as "data-latest/.." that resolve outside the cache’s
vintage directories. vintage now accepts only
"latest" or a year-month, with or without the
data- prefix.
Corrected the station_entries_daily docs: monthly
station sums usually match line totals in
line_entries_monthly but can differ by a fraction of a
percent (#43).
Excluded .posit/ from source package builds
(#42).
read_metro_demand(), which reads the four demand
datasets from the published GitHub releases instead of the bundled
snapshot. vintage pins a dated monthly batch, and
source chooses between the cache, a download, and the
bundled data.metrosp_cache_dir(),
metrosp_cache_enable(), metrosp_cache_list(),
and metrosp_cache_clear(). Downloads land in a
session-temporary directory until you allow a persistent cache.The bundled snapshot is unchanged in this release.
station_daily
includes observations through July 31, 2026.passengers_entrance, passengers_transported,
and station_averages.passengers_entrance,
passengers_transported, and station_averages.
METRO published those months only as PDFs without a text layer; they
were transcribed from the rendered pages and reconciled against the
totals printed beside them.passengers_entrance for
Lines 1, 2, 3, 5, and 15, and for the network total. The file METRO
published under that name repeats the transported table, so no entrance
figures exist for that month. Line 4 comes from the Dataverse and is
unaffected.line_number = 99) in
passengers_transported. The report reprinted May’s network
column; the per-line values for June are unaffected.mdu, msa, mdo, and
max for Lines 4 and 5 in passengers_entrance,
across the whole series. The Dataverse source is station-level, and the
averages and daily peak were taken over station-days rather than over
line-day totals, so each was divided by the number of stations reporting
that month. Values rise by roughly 5 to 17 times depending on the line
and the year. total is unchanged. Anyone comparing Line 4
or 5 averages against an earlier release should expect a break.NA for 2016–2019 in
passengers_entrance and
passengers_transported. The lookup keyed a named vector on
accented Portuguese text, which stops matching outside a UTF-8 locale,
so a build could drop every label without erroring. Station names and
line numbers in station_averages were affected the same
way. The pipeline now refuses to run outside a UTF-8 locale.No exported value changes. The snapshot was refrozen for two
structural differences: metric_abb in
passengers_entrance and passengers_transported
no longer carries a stray names attribute, and
station_averages drops one all-NA row (Jardim
Colonial, January 2022).
LINHA … / DEMANDA … header and take the line
roster from that header, so a source that adds a title row or drops a
line no longer needs a hand-edited offset. One reader now serves 2016
and 2020 onward, and another serves the 2017–2019 monthly files..import_stn_avg_2016(),
get_skip_offset(),
read_csv_stations_average(), and
clean_stations_average(), all superseded by the shared
readers.dim_line$line_name_full, and derived
metro_lines from dim_line instead of restating
it.psg_line, stn_avg, stn_daily,
suffixed by era) and dropped the . prefix.filter_out() and
replace_values().data-publish.yaml now writes each batch to a dated
data-YYYY-MM release tag as well as the rolling
data-latest tag, so a batch stays retrievable after
data-latest moves on.passengers_transported reports
thousands of passengers, while the other demand datasets count
individual passengers.passengers_transported to August 2018, the month of the
ViaMobilidade handover.data-latest
release to the package landing page and both vignettes.passengers_entrance,
passengers_transported, and station_daily
rebuilt through April 2026.passengers_entrance had one).calendar_spo: a São Paulo holiday and
business-day calendar (2012–2030) covering national, state, and
municipal holidays, for use in demand seasonality and business-day
adjustments.forecasts and
forecast_accuracy exports. These were exported in 1.0.0 by
mistake: metrosp is an observed-demand data package and
forecasting belongs downstream, so they should never have shipped.metro_lines export. Its line-name columns
(line_name, line_name_pt) are already
denormalized onto every passenger/station dataset, and the full line
list (including planned and CPTM lines) is available in
lines. It remains an internal join dimension in the ETL
pipeline.station_daily$line_number /
station_daily$year are now integer (previously
double), matching the other datasets. Values are
unchanged.station_averages, station_daily) and
the geometry datasets (stations), so a station joins
cleanly across sources.NA rows are now trimmed per line
during assembly; interior NAs (e.g. station outages) are
preserved. All datasets rebuilt.stations (Vila Mariana, Line
1) caused by the GeoSampa import not deduplicating after the name/join
cleanup step. Added a regression test asserting stations
has no duplicate rows.passengers_entrance, passengers_transported):
a hardcoded per-year header offset was one row short of where the
source’s “Jan” row actually landed, which shifted every month’s figures
up one slot (February’s numbers were recorded as January’s, and so on)
and mistook the annual total row for December. The offset is now
detected dynamically per file instead of hardcoded, since the header
length has drifted release to release.targets caching gap where the METRO CSV
download target returned a constant directory path, so a fresh download
never invalidated the downstream parsers (entrance_current,
transported_current, averages_current,
daily_current) – they kept serving stale cached data even
right after a real re-download. The target is now content-hashed against
the downloaded files themselves.data-raw/ ETL now runs as a targets pipeline
(targets::tar_make()), replacing the flag-driven
run_pipeline.R orchestrator. Pipeline functions live in
data-raw/R/; the legacy scripts remain and produce
identical output. See CLAUDE.md for the workflow.tarchetypes::tar_force() and skip cleanly when the flags
are off.