Under the hood

charr backend

charr contains three complete implementations of the stringr API.

Backend Implementation Returns
reference Calls stringi / stringr ordinary character vectors
base charr’s own optimized C++ with base strings ordinary character vectors
altrep (default) charr’s own optimized C++ with ALTREP strings ALTREP vectors

The reference backend is the original stringr implementation, kept unchanged so there is always a reference to compare against.

base and altrep are separate implementations. Each uses its own R and C++ functions, because ALTREP storage has a different shape which changes the best way to process the data.

All three backends are semantically equivalent and produce the same behavior except for a small nuance noted below.

ICU and C++17

ICU (International Components for Unicode) is the low-level C++ library that processes Unicode string data.

charr uses system ICU for 78.2 and 78.3. If neither of those versions are installed, the configure script selects the bundled ICU 78.3 statically linked with every ICU C symbol suffixed _78_charr so it doesn’t collide with the ICU inside stringi or any other loaded package.

CHARR_SYSTEM_ICU=yes or no forces either side; unset, configure auto-detects and falls back to the bundle. Windows always uses the bundle, and macOS uses it unless CHARR_SYSTEM_ICU=yes is set.

ICU 78 requires C++17, and charr compiles as C++17 so it can always use a current ICU: the system ICU 78.2 or 78.3 when present, and its bundled ICU 78.3 otherwise. At the time of writing, stringi (version 1.8.9) uses the system ICU when one is installed and otherwise its bundled ICU 74.1, which is always the case on Windows. When the two packages end up on different ICU versions, results can differ slightly.

Emoji are the easiest place to see it. 🫩 (U+1FAE9 face with bags under eyes) was added in Unicode 16.0, so ICU 78 knows what it is and ICU 74.1 treats it as an unassigned code point. With the reference on its bundled ICU 74.1:

str_width("🫩")
#> reference (ICU 74.1):     1
#> base and altrep (ICU 78): 2

str_detect("🫩", regex("\\p{Emoji}"))
#> reference (ICU 74.1):     FALSE
#> base and altrep (ICU 78): TRUE

str_sort(c("🫩", "apple", "zebra"))
#> reference (ICU 74.1):     "apple" "zebra" "🫩"
#> base and altrep (ICU 78): "🫩" "apple" "zebra"

Two is the correct width. str_width() reports how many columns a string takes in a monospaced terminal, and Unicode classifies emoji as double width.

More generally, the main change from ICU 74.1 to 78 is Unicode itself. ICU 74.1 implements Unicode 15.1 and ICU 78 implements Unicode 17.0, so the two optimized backends support newer Unicode changes. For example, Unicode 16.0 adds seven scripts on its own, among them Garay, an alphabet devised in Senegal in the 1960s.

Extended benchmarks

The corpus is multilingual sentences from Tatoeba, sampled in 1,000,000, 100,000 and 10,000 sizes. Each sample is used twice: once as an ordinary character vector and once as a charvec, an optimized ALTREP class (see charport).

To comprehensively evaluate charr, the benchmark exercises 67 operations across the entire stringr API. Each measurement is one operation in a fresh R process. Each condition gets five repetitions and the reported figure is the median.

Three of the query patterns are first derived from the corpus before benchmarks are run:

Each panel below is a family of operations, a grouping of operations that do similar work, either producing the same kind of output or using the same patterns.

Bars are how many times faster charr is than the reference on the same input, and the dashed line is the reference.

References

charr incorporates and builds off of three impressive external packages:

charr’s own code and the parts derived from stringr are MIT licensed. Code adapted from stringi is under its BSD 3-Clause License, and the bundled ICU source and data under the Unicode License v3. LICENSE.note in the source package explains which parts fall under which license, and the installed COPYRIGHTS file carries the full notices.