---
title: "Build a Mini Reference Library"
author: "Win Cowger"
date: "`r Sys.Date()`"
output: rmarkdown::html_vignette
vignette: >
  %\VignetteIndexEntry{Build a Mini Reference Library}
  %\VignetteEncoding{UTF-8}
  %\VignetteEngine{knitr::rmarkdown}
---

```{r, include = FALSE}
knitr::opts_chunk$set(
  collapse = TRUE,
  comment = "#>",
  warning = FALSE
)

data.table::setDTthreads(2)
```

```{r setup}
library(OpenSpecy)
```

This example combines a few files bundled with OpenSpecy into a small reference
library. Real library builds usually need larger lookup tables and more
curation, but the same helper functions apply.

## Read And Combine Spectra

`c_spec()` and `build_lib()` default to the widest range represented by the
source spectra at resolution 6. Values outside each source's original range are
kept as `NA`, so useful spectral regions are not discarded.
This compact vignette example uses the shared overlapping range to keep the
rendered output small.

```{r mini-library-read}
mini_files <- c(
  read_extdata("raman_hdpe.csv"),
  read_extdata("ftir_ldpe_soil.asp"),
  read_extdata("raman_atacamit.spc")
)

mini_sources <- lapply(mini_files, read_any)
mini_sources <- lapply(mini_sources, function(x) {
  x$metadata$intensity_units <- "absorbance"
  attr(x, "intensity_unit") <- "absorbance"
  x
})
mini_raw <- c_spec(mini_sources, range = "common", res = 6)

check_OpenSpecy(mini_raw)
dim(mini_raw$spectra)
mini_raw$metadata[, "file_name", with = FALSE]
```

## Create And Fill A Lookup

Lookup templates help users see which metadata values need curation. When
`path` is not supplied, the template is returned as a `data.table`; users can
write it to CSV by supplying `path`. Lookup joins use exact values, so edit the
template until the key column matches the object metadata.

```{r mini-library-template}
template <- make_lib_lookup_template(
  mini_raw,
  columns = "file_name",
  add = c("material", "material_type")
)

template
```

For this small example, create the lookup table directly in R. The same shape
could come from a CSV edited outside R.

```{r mini-library-join}
lookup <- data.table::data.table(
  file_name = basename(mini_files),
  library_type = c("example", "example", "example"),
  spectrum_type = c("raman", "ftir", "raman"),
  material = c("hdpe", "ldpe in soil", "atacamite")
)

hierarchy <- data.table::data.table(
  material = c("hdpe", "ldpe in soil", "atacamite"),
  material_class = c("polyethylene", "polyethylene", "copper mineral"),
  material_type = c("plastic", "plastic", "mineral")
)

join_lib_metadata(mini_raw, lookup, by = "file_name",
                  require_complete = TRUE)$metadata[
  , c("file_name", "library_type", "spectrum_type", "material"),
  with = FALSE
]
```

## Build A Mini Library

`build_lib()` can run the ordinary lookup and material hierarchy joins before
applying named recipes. Recipe names become names in the returned list. Empty
recipes keep the merged spectra unchanged; other recipe lists are passed to
`process_spec()`. Missing values and processing attributes are handled
automatically. Metadata column names are also cleaned to lowercase underscore
names. Known aliases are coalesced using an editable lookup table. Variants
that differ only by underscores or one terminal plural `s` match
automatically.

Before sources are merged, `build_lib()` also converts declared reflectance and
transmittance spectra to absorbance by default. A nonempty
`attr(x, "intensity_unit")` is the primary truth for the whole object;
otherwise, `metadata$intensity_units` is evaluated spectrum by spectrum.
Unknown or missing units are left unchanged with a warning. Use
`convert_intensity = FALSE` when unit handling has already been completed
outside the builder.

The normal input is one or more file paths; one `OpenSpecy` or a list of
`OpenSpecy` objects supports sources already loaded in memory. An RDS path may
store either one object or a list of objects. Other formats are read with
`read_any()`. Large same-axis source lists are prepared in bulk, which is much
faster for legacy RDS files that store many one-spectrum objects. Named stages
and elapsed time are reported while the library is built; use
`progress = FALSE` when quiet output is preferred. Supplying
`restrict_range_args` triggers the existing `restrict_range()` operation before
deduplication and recipes; multiple retained ranges can exclude a known silent
region without custom workflow code.

```{r metadata-name-lookup}
name_lookup <- lib_metadata_name_lookup(
  project_code = c("campaign id", "study code"),
  regex = list(instrument_mode = "^method_[0-9]+$")
)
name_lookup[
  canonical_name %in% c("material_color", "number_of_accumulations")
]
lib_clean_name(c("User Name", "Laser (%)", "Method...3"))
```

Named arguments add exact aliases to the defaults, while `regex` adds patterns
evaluated against cleaned names. Overlapping regex patterns produce an error
that identifies the source column and matching rules. Pass the result as
`metadata_name_lookup`. Set `clean_metadata_values = TRUE` in `build_lib()`
or `clean_values = TRUE` in `lib_clean_metadata()` to lowercase, trim, and
ASCII-normalize character metadata values before joins. Ordinary and
hierarchical joins run whenever their corresponding lookup input is non-`NULL`.
`spectrum_identity` receives one additional deterministic cleanup: recognizable
paths are reduced to their basename and trailing file extensions supported by
`read_any()` are removed. Numeric OPUS extensions include `.10` and any other
terminal period followed only by digits. Exact lookup keys receive the same
cleanup. Each built library records changed values and counts in its
`spectrum_identity_cleanup_report` attribute. Keep flexible class patterns in a
separate table and apply them afterward with `predict_class_reference()`.
Automatic ordinary lookups use the single shared column with overlapping values
and unique lookup keys. Lookups with no usable shared key are skipped with a
message; lookups with multiple usable shared keys are treated as ambiguous, so
wrap a lookup as `list(lookup = table, by = "key")` when an explicit key is
needed. Named `by` vectors map metadata names to different lookup names.
Canonical source metadata is standardized before these external joins. Every
spectrum receives `library_name` from a populated `organization`, otherwise
from `user_name`; this is the canonical source-library grouping used by build
assessments. A blank `organization` is still filled from the reviewed internal
`user_name` alias for the existing type lookup. The older lookup-level
`fallback_by` field is deprecated. Set `fill_only = TRUE` to fill blank values
such as `library_type` or `spectrum_type` without replacing populated source
metadata.

```{r mini-library-build}
mini_libs <- build_lib(
  mini_files,
  recipes = list(
    raw = list(),
    derivative = list(
      conform_spec = FALSE,
      smooth_intens = TRUE,
      smooth_intens_args = list(window = 15, derivative = 1),
      make_rel = TRUE
    ),
    nobaseline = list(
      conform_spec = FALSE,
      smooth_intens = FALSE,
      subtr_baseline = TRUE,
      make_rel = TRUE
    )
  ),
  metadata_lookups = lookup,
  material_hierarchy = hierarchy,
  clean_metadata_values = TRUE,
  convert_intensity = FALSE,
  assess = TRUE,
  dedupe = FALSE
)

names(mini_libs)
check_OpenSpecy(mini_libs$raw)
check_OpenSpecy(mini_libs$derivative)
attr(mini_libs$derivative, "derivative_order")
attr(mini_libs$nobaseline, "baseline")
mini_libs$raw$metadata[
  , .(file_name, material, material_class, material_type, sn,
      assessment_flag, assessment_checks)
]
```

`prune_lib()` performs the spectrum-supported cleanup used before medoid or
model creation. With `cross_class = TRUE`, it first flags correlations strictly
above `cross_class_threshold` between different reviewed classes within each
source library. The spectrum with the most active wrong-class neighbors is
removed first and scores are recalculated; adjacent maximum-score ties remove
both endpoints. A second pass then uses the number of independent opposing
libraries, removing the higher-evidence endpoint or both on a tie. Generic and
unclassified labels are excluded. The audit records phase, view, component,
identities/classes/libraries, active degree or independent evidence, decision
round, threshold, and reason. Within the same
spectral-technique pool, `other` may then be
reassigned to the nearest established class, `other plastic` only to a
candidate whose `material_type` is `plastic`, and `other material` only to
`organic matter` or `mineral`. The matched `material_type` is copied with the
class. It then evaluates material classes from largest to smallest using
bounded correlation blocks. FTIR and NIR share a candidate pool; Raman is
separate; 2200--2420 cm^-1^ is excluded by default. After generic reassignment,
each `spectrum_type` and material-class group must contain at least `min_n`
spectra across the complete input database; `min_n` is not applied separately
to each source library. Smaller resolved groups are reassigned as a whole when an eligible
destination exists and removed only when no destination exists; groups exactly
at the threshold are retained whole and larger groups are never reduced below
it. Constant or unmatched spectra in eligible classes are retained, and
deterministic IDs resolve correlation ties. Use `return = "report"` for the
retained IDs, frozen class schedule, reassignment correlations, spectrum-level
removals, and `excluded_classes`, which reports observed support and the number
of additional or reassigned spectra needed. End-to-end builds combine these
class rows in `assessments$cleanup$summary` with the recipe name.
`build_lib(prune = ...)` maps recipe names to `prune_lib()` argument lists, so
each selected spectral representation is pruned independently. Resolved
class/type groups below `min_n` are reassigned as a whole to the
most-correlated established class in the same technique pool and material type;
they are dropped only when no eligible correlated destination exists. The
pruning assessment records the destination, mean class correlation, action,
and reason. In the official
workflow, an otherwise unresolved identity is temporarily labeled `other`.
The default `remove_other = TRUE` removes blank identities and the unresolved
literal `other` class before quality control. Reviewed broad `other plastic`
and `other material` categories stay in the reference libraries and remain
eligible for constrained reassignment during derivative/no-baseline pruning.
The builder keeps reviewed identifiers, source metadata, prior labels, reasons,
actions, and typed before/after counts in `assessments$cleanup$summary`. The
separate `assessments$cleanup$dropped_spectrum_identities` table contains only
the sorted distinct source identities removed anywhere in cleanup. Set
`remove_other = FALSE` to retain generic rows for the
constrained semisupervised `prune_lib()` reassignment described above.

```{r mini-library-prune}
pruned <- prune_lib(
  mini_libs$derivative,
  min_n = 1,
  return = "report",
  progress = FALSE
)
pruned$summary
pruned$schedule
pruned$excluded_classes
```

The composable default leaves the new first pass off. Enable it explicitly for
a library carrying canonical `library_name` metadata; the official derivative
and no-baseline recipes do this by default:

```{r cross-class-prune-policy, eval=FALSE}
prune_lib(
  reference_library,
  cross_class = TRUE,
  cross_class_threshold = 0.9
)
```

## Official Reference-Library Workflow

The version-controlled
[`workflows/OpenSpecy_reference_library.R`](https://github.com/wincowgerDEV/OpenSpecy-package/blob/main/workflows/OpenSpecy_reference_library.R)
script retains the source and output path finders and passes those paths to one
`build_lib()` call. The complete visual flow,
including checkpoint reuse and old/new assessments, is maintained in
[`.specify/memory/build-lib-diagram.html`](https://github.com/wincowgerDEV/OpenSpecy-package/blob/main/.specify/memory/build-lib-diagram.html).
Canonically named exact-class, regex-class, library-type, material-hierarchy,
known-bad-ID, material-form-regex, and common-use tables live under
`workflows/data/`. Raw-data corrections are completed externally before this
workflow runs. The source-library paths and output directory are explicit
inputs. When `workflow_data` is not supplied, `build_lib()` looks for the
helper tables under `data/` beside the calling script, then under `data/` or
`workflows/data/` in the current working directory. Calling `build_lib()` with
no `x` now stops with an actionable error instead of guessing external paths.
The official build derives `library_name` from organization first and user name
second, fills a missing organization from `user_name` before one source-type
lookup, and asserts complete library/spectrum-type coverage. Per-recipe
retention rows in `assessments$cleanup$summary` report counts at preparation,
exclusion, deduplication, broad-category review, quality control, pruning,
post-transform, and final partition stages. A completely dropped source names
the first empty stage and its reason. The
exact class lookup runs first;
`predict_class_reference()` then evaluates the separate regex table only for
blank materials. Exact/regex overlaps are audited without overwriting exact
values, and conflicting regex predictions stop for review. `material_class`
describes chemistry rather than physical form. Every reviewed standard plastic
class begins with `poly` for discoverability (`other plastic` is the explicit
catch-all exception); polyethylene and polypropylene remain separate, and
`polyhydroxy(meth)acrylates` retains its chemically meaningful optional group
notation. Poly(styrene-butadiene) rubber and
poly(ethylene-propylene-diene) rubber have distinct classes. Paint binder labels
such as acrylic, alkyd, urethane, nitrocellulose, and vinyl combinations remain
in `spectrum_identity`, `paint` remains in `material_form`, and the hierarchy
maps each reviewed binder to its relevant chemical polymer family. Remaining
uncertain identities follow the selected generic-row review policy.

The official workflow also standardizes two enrichment fields. It normalizes
and concatenates all atomic metadata values once per spectrum, then evaluates
the reviewable `material_form_regex.csv`. A uniquely matching existing
`material_form` takes precedence; otherwise the full metadata row is searched.
Cross-category matches remain `NA` and are retained in the clash audit. The
controlled vocabulary currently covers paint, rubber, hard plastic, fiber,
pellet, film plastic, fragment, foam, and sphere/bead.
`common_use_reference.csv` is a material-class lookup for the proposed
predominant global ultimate end-market. Quantitative rows retain consumer,
industrial, and unresolved mass shares: `consumer` and `industrial` require
more than 50%, while `mixed` requires at least 25% on each side, at least 75%
classified, and neither side above 50%. Where comparable mass shares do not
exist, a qualitative proposal is allowed only with a cited application source,
retrieval date, and explicit review note; share cells remain blank. Consumer
endpoints include products used directly by individuals, such as passenger
tires, household goods, clothing, and patient-facing healthcare products.
Classes with substantial consumer and industrial endpoints use `mixed`.
Ambiguous catch-alls and poorly supported specialty groups remain `NA`.
Commodity shares use the [OECD 2019 polymer-by-application mass
dataset](https://data-explorer.oecd.org/vis?df%5Bag%5D=OECD.ENV.EEI&df%5Bid%5D=DSD_PU_2019%40DF_PU_2019);
PTFE uses a separate material-specific end-use source recorded in the CSV.
Coverage, provenance, matched evidence, and clashes are retained with the
upstream build assessments. Derivative and
nobaseline spectra pass three quality gates before pruning: FTIR CO2 is
flattened when the 2200--2420 maximum divided by the 2420--2550 silent-region
maximum is greater than two; high-tail detection ignores NA padding, trims each
spectrum on its finite support, and drops failed corrections; and finite
running SNR below two (or unavailable SNR) is removed. Same-library majority
closure then precedes independent-library evidence, generic reassignment, and
conventional top-match pruning. After rounding and type partitioning, the full
and model-range views are closed again. Raw remains unpruned. Every spectrum
excluded by cross-class closure is preserved in the review quarantine.
Each completed library, medoid, model, and assessment component is written with
a SHA-256 manifest under `output_dir/checkpoints`. `reuse = TRUE` loads it only when
the source/lookup signatures, arguments, package version, and builder contract
still match. Validated output is promoted into a versioned release directory,
which contains the seven legacy RDS names, explicit algorithm-named model RDS
files, one aggregate build RDS, and a release manifest. Existing versioned
payloads are immutable: a different hash stops promotion. The release also
contains `quarantined_spectra.rds`, a versioned bundle of valid `OpenSpecy`
objects by recipe/type, row-level removal metadata, the long conflict audit, and
build provenance. Its review copy and CSV mirrors are checkpointed under
`output_dir/review` before downstream model and comparison stages.

The returned object has four top-level elements: `libraries`, `medoids`,
`models`, and `assessments`. Libraries are recipe-first and then keyed by
`ftir`, `raman`, or `nir`: full Raman uses 200--4000, FTIR uses 400--4000,
and NIR uses 4000--12000 cm^-1^. FTIR/Raman medoids and models use 800--3200;
the NIR identification interval is derived from the longest region with at
least 90% finite coverage (falling back to available coverage) and recorded on
the object. All-blank metadata columns are dropped independently by type.
The enrichment columns are retained even when every value for one type is
missing; other all-blank metadata columns are dropped independently by type.
Each spectrum must observe at least 10% of its type-specific identification
axis to enter medoid selection or model fitting. PAM operates on a temporary
relative matrix where each spectrum's missing values are filled by that
spectrum's finite mean. Selected medoid IDs are then pulled from the original
range-restricted object, so published medoids retain their original `NA`
positions. Groups with more than 3,000 spectra use five deterministic
1,000-spectrum PAM samples; each candidate set is scored against the complete
group with the same correlation distance, avoiding a full oversized
dissimilarity matrix. `models` is algorithm-first (`logistic_regression` or
`random_forest`), then recipe and type. Logistic regression retains
derivative/nobaseline medoid training and fills restored gaps with the finite
mean at each wavenumber across training medoids. Random forests use the complete
eligible raw, derivative, and nobaseline type libraries, relative normalization,
wavenumber-mean filling, inverse-frequency balanced case sampling, 500
probability trees, permutation importance, and all available ranger threads.
Balanced sampling improves minority representation in each bootstrap without
also applying a class-vote correction. Because this resampling can make
out-of-bag accuracy optimistic, treat OOB results as fit diagnostics. The
release assessment separately applies each deployed model to its complete
corresponding source dataset.
`train_spec_model()` exposes both engines; `build_model_lib()` remains a
compatible wrapper. Each model stores its filler, support audit, training
diagnostics, and one tidy `tests` table. Logistic lambda selection remains
maximum out-of-fold macro class accuracy with its calibrated alpha,
no-intercept, grouped multinomial, and weighting settings unchanged.

Full assessment uses every candidate and downloaded legacy artifact. Candidate
and legacy artifacts independently use a seeded approximately ten-percent
class/type-stratified sample. Rows sharing a physical identifier or exact
spectral content are one group and cannot cross training and test. This
source-local policy tolerates taxonomy evolution
without fuzzy cross-version class matching; its results describe each
artifact on its own source population and are not paired causal deltas.
Held-out groups are removed from full reference libraries before their
identification assessment. Medoid artifacts are instead tested as deployed:
each medoid library identifies every spectrum in its complete corresponding
processed library (for example, `medoid_derivative.rds` searches
`derivative.rds`). Existing logistic and random-forest artifacts likewise
identify their complete corresponding source libraries once; assessment does
not select new medoids or call `train_spec_model()`. `match_spec()` uses stored
fillers for partial spectra. Macro class
accuracy is the primary metric, followed by evaluated coverage, overall
accuracy, and clear old/new
`assess_spec(report = "all")` shifts. The review surface contains at most ten
tables nested under `cleanup`, `ref_lib`, `medoid`, `model`, and
`functionality`. Old and new values occupy adjacent columns; accuracy,
misidentification counts, absolute diagnostic correlations, and warning/error
rate shifts are sorted from largest to smallest. Pass rows are omitted from
quality shifts. Machine-level split and row-test evidence is available through
`attr(reference_library_build$assessments, "evidence")`. Published accuracy
tables contain aggregate overall and macro metrics, not per-class accuracy
rows. A versioned release writes all global cleanup, quality, pruning, model,
and comparison evidence to `assessments.rds`; its library, medoid, and model
files contain only runtime data and prediction state. The accompanying
`reference_library_build.rds` is a lightweight index that can be supplied to
`rebuild_lib_artifacts()`. The smaller 1,000-spectrum run in the manual
benchmark is development-only and is not part of `build_lib()` or final
acceptance evidence.

The standard call therefore has this shape:

```{r official-call-shape, eval=FALSE}
reference_library_build <- build_lib(
  files,
  output_dir = output_dir,
  previous_library_dir = "system",
  remove_other = TRUE
)
```

After upstream libraries have already completed, rebuild only the downstream
artifacts into a new explicit output location:

```{r downstream-rebuild-shape, eval=FALSE}
updated_reference_library_build <- rebuild_lib_artifacts(
  completed_build_dir,
  output_dir = downstream_output_dir,
  previous_library_dir = "system"
)
```

`completed_build_dir` may instead be a `reference_library_build.rds`, a
`libraries.rds` checkpoint, or the corresponding in-memory build object.

The script is excluded from package builds but remains available
in the GitHub repository so library releases can be reviewed and reproduced.
