| Type: | Package |
| Title: | Modelling Compositional Data with Zero Values |
| Version: | 1.1 |
| Date: | 2026-10-06 |
| Author: | Michail Tsagris [aut, cre] |
| Maintainer: | Michail Tsagris <mtsagris@uoc.gr> |
| Depends: | R (≥ 4.0) |
| Imports: | cluster, graphics, grDevices, MASS, rangen, Rfast, stats |
| Suggests: | Compositional, CompositionalZADR, mziln, Rfast2 |
| Description: | Modelling structural zeros in compositional data using a conditional logistic normal model as described by Aitchison (1986), where MLE (Maximum Likelihood Estimation) is performed via the EM (Expectation-Maximization) algorithm. The relevant paper is Alzeley and Tsagris (2026) <doi:10.48550/arXiv.2608.29954>. |
| License: | GPL-2 | GPL-3 [expanded from: GPL (≥ 2)] |
| NeedsCompilation: | no |
| Packaged: | 2026-10-06 20:14:19 UTC; mtsag |
| Repository: | CRAN |
| Date/Publication: | 2026-10-06 22:00:31 UTC |
Modelling Compositional Data with Zero Values
Description
Modelling Compositional Data with Zero Values.
Details
| Package: | Compositionalcln |
| Type: | Package |
| Version: | 1.1 |
| Date: | 2026-10-06 |
Maintainers
Michail Tsagris <mtsagris@uoc.gr>.
Author(s)
Michail Tsagris mtsagris@uoc.gr
References
Alzeley O. and Tsagris M. (2026). Modelling compositional data with structural zero values https://arxiv.org/pdf/2608.29954
Aitchison J. (1986). The statistical analysis of compositional data.
Maximum likelihood estimation of the conditional logistic normal model
Description
Maximum likelihood estimation of the conditional logistic normal model.
Usage
boot.clnmle(y, tol = 1e-6, maxit = 500, R = 1000)
Arguments
y |
A numerical matrix with compositional data with zero values. |
tol |
The tolerance value to terminate the EM algorithm. |
maxit |
The maximum number of iterations allowed for the EM algorithm. |
R |
The number of bootstrap resamples to generate. |
Details
The function performs non-parametric bootstrap for the conditional multivariate logistic normal distribution (Alzeley and Tsagris, 2026).
Value
A list including:
m |
The bootsrap mean vectors in the Euclidean space. |
mesi |
The bootstrap mean vectors in the simplex space. |
S |
The bootstrap estimated covariance matrices. |
Author(s)
Michail Tsagris.
R implementation and documentation: Michail Tsagris mtsagris@uoc.gr.
References
Alzeley O. and Tsagris M. (2026). Modelling compositional data with structural zero values https://arxiv.org/pdf/2608.29954
Aitchison J. (1986). The statistical analysis of compositional data, pages 272–274.
See Also
Examples
y <- as.matrix(iris[, 1:3])
y <- y / rowSums(y)
ind <- sample(150, 10)
for ( k in ind ) y[k, sample(3, 1)] <- 0
y <- y / rowSums(y)
mod <- boot.clnmle(y, R = 300)
Bootstrap for the conditional logistic normal regression
Description
Bootstrap for the conditional logistic normal regression.
Usage
boot.clnreg(y, x, tol = 1e-6, maxit = 500, R = 1000)
Arguments
y |
A matrix with the compositional data (dependent variable). The number of observations (vectors) with no zero values should be more than the columns of the predictor variables. Otherwise, the initial values will not be calculated. |
x |
The predictor variable(s), they can be either continnuous or categorical or both. |
tol |
The tolerance value to terminate the EM algorithm. |
maxit |
The maximum number of iterations allowed for the EM algorithm. |
R |
The number of bootstrap resamples to generate. |
Details
Bootstrap to estimate the covariance matrix of the coefficients of the conditional logistic normal regression is performed.
Value
A list including:
beta |
The bootstrapped beta coefficients. |
sigma |
The bootstrap estimate of the covariance matrix of the regression parameters. |
Author(s)
Michail Tsagris.
R implementation and documentation: Michail Tsagris mtsagris@uoc.gr.
References
Alzeley O. and Tsagris M. (2026). Modelling compositional data with structural zero values https://arxiv.org/pdf/2608.29954
Aitchison J. (1986). The statistical analysis of compositional data, pages 272–274.
See Also
Examples
x <- as.vector(iris[, 4])
y <- as.matrix(iris[, 1:3])
y <- y / rowSums(y)
ind <- sample(150, 10)
for ( k in ind ) y[k, sample(3, 1)] <- 0
y <- y / rowSums(y)
mod <- cln.reg(y, x)
Non-parametric confidence region of the simplicial mean on a ternary diagram
Description
Non-parametric confidence region of the simplicial mean on a ternary diagram.
Usage
cln.bcr(y, tol = 1e-6, maxit = 500, alpha = 0.05, R = 1000,
dg = TRUE, hg = TRUE, ploty = TRUE)
Arguments
y |
A matrix with the compositional data (with zero values). |
tol |
The tolerance value to terminate the EM algorithm. |
maxit |
The maximum number of iterations allowed for the EM algorithm. |
alpha |
The significance level to construct the (1-alpha)% confidence region. |
R |
The number of bootstrap resamples to generate. |
dg |
Do you want diagonal grid lines to appear? If yes, set this TRUE. |
hg |
Do you want horizontal grid lines to appear? If yes, set this TRUE. |
ploty |
Do you want to plot the data as well? |
Details
A non-parametric confidence region of the mean in S^2 is constructed based
on the method of Fiksel, Zeger and Datta (2022).
Value
A ternary diagram with the simplicial mean and the non-parametric confidence region.
Author(s)
Michail Tsagris.
R implementation and documentation: Michail Tsagris mtsagris@uoc.gr.
References
Fiksel J., Zeger S. and Datta A. (2022). A transformation-free linear regression for compositional outcomes and predictors. Biometrics, 78(3): 974–987.
Alzeley O. and Tsagris M. (2026). Modelling compositional data with structural zero values https://arxiv.org/pdf/2608.29954
Aitchison J. (1986). The statistical analysis of compositional data, pages 272–274.
See Also
Examples
y <- as.matrix(iris[1:50, 1:3])
y <- y / rowSums(y)
ind <- sample(50, 10)
for ( k in ind ) y[k, sample(3, 1)] <- 0
y <- y / rowSums(y)
cln.bcr(y, R = 300, ploty = FALSE)
Biplot based on the conditional logistic normal model
Description
Biplot based on the conditional logistic normal model.
Usage
cln.biplot(y, m, S)
Arguments
y |
A vector or a matrix with compositional data with zero values. |
m |
The mean vector in |
S |
The covariance matrix in |
Details
The function first transforms the covariance matrix that is based on the alr transformation to the covariance matrix based on the clr transformation. Then, for the vectors with zero values it uses the vectors with the estimated zis on the alr-space (see Alzeley and Tsagris, 2026) from the EM algorithm. Then it produces the classical biplot. The observations that contain zero values will be plotted with a red colour and the observations with no zero values are plotted with a grey colour.
Value
The biplot of the alr transformed data.
Author(s)
Michail Tsagris.
R implementation and documentation: Michail Tsagris mtsagris@uoc.gr.
References
Alzeley O. and Tsagris M. (2026). Modelling compositional data with structural zero values https://arxiv.org/pdf/2608.29954
Aitchison J. (1986). The statistical analysis of compositional data.
See Also
cln.biplot, cln.contour, cln.mle
Examples
y <- as.matrix(iris[, 1:4])
y <- y / rowSums(y)
ind <- sample(150, 10)
for ( k in ind ) y[k, sample(4, 1)] <- 0
y <- y / rowSums(y)
mod <- cln.mle(y)
cln.biplot(y, mod$m, mod$S)
Many conditional logistic normal regressions given some covariates
Description
Many conditional logistic normal regressions given some covariates.
Usage
cln.condregs(y, X, prior, tol = 1e-6, maxit = 500)
Arguments
y |
A matrix with the compositional data (dependent variable). The number of observations (vectors) with no zero values should be more than the columns of the predictor variables. Otherwise, the initial values will not be calculated. |
X |
A matrix with the predictor variable(s). A regression model will be fit to each of them, separately. |
prior |
A matrix or a vector with some predictors that will be in the model. |
tol |
The tolerance value to terminate the EM algorithm. |
maxit |
The maximum number of iterations allowed for the EM algorithm. |
Details
A conditional logistic normal regression is being fitted for each of the columns
of X. The difference from cln.regs, is that this time, each of the
the regression models fit contain the prior variables.
Value
A list including:
stat |
A vector with the log-likelihood ratio test statistics. |
pvalue |
A vector with the corresponding logged p-values. |
Author(s)
Michail Tsagris.
R implementation and documentation: Michail Tsagris mtsagris@uoc.gr.
References
Alzeley O. and Tsagris M. (2026). Modelling compositional data with structural zero values https://arxiv.org/pdf/2608.29954
See Also
Examples
x <- matrix( rnorm(150 * 10), ncol = 10 )
prior <- iris[, 4]
y <- as.matrix(iris[, 1:3])
y <- y / rowSums(y)
ind <- sample(150, 10)
for ( k in ind ) y[k, sample(3, 1)] <- 0
y <- y / rowSums(y)
mod <- cln.condregs(y, x, prior = prior)
Contour plot of the conditional logistic normal distribution in S^2
Description
Contour plot of the conditional logistic normal distribution in S^2.
Usage
cln.contour(m, S, n = 100, y = NULL, cont.line = FALSE)
Arguments
m |
A value with the mean vector in |
S |
The covariance matrix in |
n |
The number of grid points to consider over which the density is calculated. |
y |
This is either NULL (no data) or contains a 3 column matrix with compositional data. |
cont.line |
Do you want the contour lines to appear? If yes, set this TRUE. |
Details
The user can plot only the contour lines of a zero adjusted Dirichlet distribution with som given parameters, or can also add the relevant data should he/she wish to.
Value
A ternary diagram with the points and the conditional multivariate normal contour lines.
Author(s)
Michail Tsagris.
R implementation and documentation: Michail Tsagris mtsagris@uoc.gr.
References
Alzeley O. and Tsagris M. (2026). Modelling compositional data with structural zero values https://arxiv.org/pdf/2608.29954
Aitchison J. (1986). The statistical analysis of compositional data, pages 272–274.
See Also
Examples
y <- as.matrix(iris[, 1:3])
y <- y / rowSums(y)
ind <- sample(150, 10)
for ( k in ind ) y[k, sample(3, 1)] <- 0
y <- y / rowSums(y)
mod <- cln.mle(y)
cln.contour( m = mod$m, S = mod$S )
Confidence region of the simplicial mean on a ternary diagram
Description
Confidence region of the simplicial mean on a ternary diagram.
Usage
cln.cr(y, m, S, alpha = 0.05, dg = TRUE, hg = TRUE, ploty = TRUE)
Arguments
y |
A matrix with the compositional data (with zero values). |
m |
A value with the mean vector in |
S |
The covariance matrix in |
alpha |
The significance level to construct the (1-alpha)% confidence region. |
dg |
Do you want diagonal grid lines to appear? If yes, set this TRUE. |
hg |
Do you want horizontal grid lines to appear? If yes, set this TRUE. |
ploty |
Do you want to plot the data as well? |
Details
The confidence region of the mean in S^2 is constructed and then mapped back onto the simplex.
Value
A ternary diagram with the simplicial mean and the confidence region.
Author(s)
Michail Tsagris.
R implementation and documentation: Michail Tsagris mtsagris@uoc.gr.
References
Alzeley O. and Tsagris M. (2026). Modelling compositional data with structural zero values https://arxiv.org/pdf/2608.29954
Aitchison J. (1986). The statistical analysis of compositional data, pages 272–274.
See Also
Examples
y <- as.matrix(iris[1:50, 1:3])
y <- y / rowSums(y)
ind <- sample(50, 10)
for ( k in ind ) y[k, sample(3, 1)] <- 0
y <- y / rowSums(y)
mod <- cln.mle(y)
m <- mod$m
S <- mod$S
cln.cr(y, m, S, ploty = FALSE)
James test for equality of two mean vectors based on the conditional logistic normal model
Description
James test for equality of two mean vectors based on the conditional logistic normal model.
Usage
cln.james(y1, y2, a = 0.05, tol = 1e-6, maxit = 500)
Arguments
y1 |
A numerical matrix with compositional data with zero values. |
y2 |
A numerical matrix with compositional data with zero values. |
a |
The significance level, set to 0.05 by default. |
tol |
The tolerance value to terminate the EM algorithm. |
maxit |
The maximum number of iterations allowed for the EM algorithm. |
Details
The function fits the conditional multivariate logistic normal distribution (Alzeley and Tsagris, 2026) to each group, estimates the mean vectors and covariance matrices and uses them to perform the James test for equality of two mean vectors without assuming equality of the covariance matrices.
Value
A vector with 4 numbers, the value of the test statistic, the p-value, the correction and the corrected critical value.
Author(s)
Michail Tsagris.
R implementation and documentation: Michail Tsagris mtsagris@uoc.gr.
References
Alzeley O. and Tsagris M. (2026). Modelling compositional data with structural zero values https://arxiv.org/pdf/2608.29954
Aitchison J. (1986). The statistical analysis of compositional data, pages 272–274.
See Also
Examples
y1 <- as.matrix(iris[1:75, 1:4])
ind <- sample(75, 10)
for ( k in ind ) y1[k, sample(3, 1)] <- 0
y1 <- y1 / rowSums(y1)
y2 <- as.matrix(iris[76:150, 1:4])
ind <- sample(75, 10)
for ( k in ind ) y2[k, sample(3, 1)] <- 0
y2 <- y2 / rowSums(y2)
cln.james(y1, y2)
Maximum likelihood estimation of the conditional logistic normal model
Description
Maximum likelihood estimation of the conditional logistic normal model.
Usage
cln.mle(y, tol = 1e-6, maxit = 500)
Arguments
y |
A numerical matrix with compositional data with zero values. |
tol |
The tolerance value to terminate the EM algorithm. |
maxit |
The maximum number of iterations allowed for the EM algorithm. |
Details
The function fits the conditional multivariate logistic normal distribution (Alzeley and Tsagris, 2026).
Value
A list including:
m |
The mean vector in the Euclidean space. |
mesi |
The mean vector in the simplex space. |
S |
The estimated covariance matrix. |
loglik |
The log-likelhiood value. |
patterns |
A matrix with the patterns of zeros, where the value of 0 indicates the presence of a zero, and the last column contains the percentage of occurrence each pattern. This is useful for random values simulation. |
iters |
The number of iterations required by the EM algorithm. |
Author(s)
Michail Tsagris.
R implementation and documentation: Michail Tsagris mtsagris@uoc.gr.
References
Alzeley O. and Tsagris M. (2026). Modelling compositional data with structural zero values https://arxiv.org/pdf/2608.29954
Aitchison J. (1986). The statistical analysis of compositional data, pages 272–274.
See Also
Examples
y <- as.matrix(iris[, 1:3])
y <- y / rowSums(y)
ind <- sample(150, 10)
for ( k in ind ) y[k, sample(3, 1)] <- 0
y <- y / rowSums(y)
cln.mle(y)
Principal component analysis based on the conditional logistic normal model
Description
Principal component analysis based on the conditional logistic normal model.
Usage
cln.pca(y, m, S)
Arguments
y |
A vector or a matrix with compositional data with zero values. |
m |
The mean vector in |
S |
The covariance matrix in |
Details
The function first transforms the covariance matrix that is based on the alr transformation to the covariance matrix based on the clr transformation. Then, for the vectors with zero values it uses the vectors with the estimated zis on the alr-space (see Alzeley and Tsagris, 2026) from the EM algorithm. Then it produces the classical principal component analysis.
Value
A list including:
values |
A vector with the eigenvalues. |
prop |
A vector with the proportion of the eigenvalues |
loadings |
A matrix with the estimated loadings |
scores |
A vector with the estimated scores. |
Author(s)
Michail Tsagris.
R implementation and documentation: Michail Tsagris mtsagris@uoc.gr.
References
Alzeley O. and Tsagris M. (2026). Modelling compositional data with structural zero values https://arxiv.org/pdf/2608.29954
Aitchison J. (1986). The statistical analysis of compositional data.
See Also
cln.biplot, cln.contour, cln.mle
Examples
y <- as.matrix(iris[, 1:4])
y <- y / rowSums(y)
ind <- sample(150, 10)
for ( k in ind ) y[k, sample(3, 1)] <- 0
y <- y / rowSums(y)
mod <- cln.mle(y)
pc <- cln.pca(y, mod$m, mod$S)
sc <- pc$scores[, 1:2] # centered, but unstandardised
# orthonormal basis of the clr plane of 3 parts (ilr basis)
Psi <- cbind( c(1, -1, 0) / sqrt(2), c(1, 1, -2) / sqrt(6) )
est <- tcrossprod( sc, Psi)
phat <- exp(est) ; phat <- est / rowSums(est)
ternary(phat)
Conditional logistic normal regression
Description
Conditional logistic normal regression.
Usage
cln.reg(y, x, tol = 1e-6, maxit = 500, seb = FALSE, xnew = NULL)
Arguments
y |
A matrix with the compositional data (dependent variable). The number of observations (vectors) with no zero values should be more than the columns of the predictor variables. Otherwise, the initial values will not be calculated. |
x |
The predictor variable(s), they can be either continnuous or categorical or both. |
tol |
The tolerance value to terminate the EM algorithm. |
maxit |
The maximum number of iterations allowed for the EM algorithm. |
seb |
If you want the standard errors of the regression coefficients set this to TRUE. |
xnew |
If you have new data use it, otherwise leave it NULL. |
Details
The conditional logistic normal regression is being fitted. The likelihood conists of two components. The contributions of the non zero compositional values and the contributions of the compositional vectors with at least one zero value. The second component may have many different sub-categories, one for each pattern of zeros.
Value
A list including:
patterns |
A matrix with the patterns of zeros, where the value of 0 indicates the presence of a zero, and the last column contains the percentage of occurrence each pattern. This is useful for random values simulation. |
iters |
The number of iterations required by the EM algorithm. |
S |
The covariance of the latent residuals. |
beta |
The beta coefficients. |
loglik |
The value of the log-likelihood. |
est |
The fitted or the predicted values (if xnew is not NULL). |
Author(s)
Michail Tsagris.
R implementation and documentation: Michail Tsagris mtsagris@uoc.gr.
References
Alzeley O. and Tsagris M. (2026). Modelling compositional data with structural zero values https://arxiv.org/pdf/2608.29954
Aitchison J. (1986). The statistical analysis of compositional data, pages 272–274.
See Also
Examples
x <- as.vector(iris[, 4])
y <- as.matrix(iris[, 1:3])
ind <- sample(150, 10)
for ( k in ind ) y[k, sample(3, 1)] <- 0
y <- y / rowSums(y)
mod <- cln.reg(y, x)
Many conditional logistic normal regressions
Description
Many conditional logistic normal regressions.
Usage
cln.regs(y, X, tol = 1e-6, maxit = 500)
Arguments
y |
A matrix with the compositional data (dependent variable). The number of observations (vectors) with no zero values should be more than the columns of the predictor variables. Otherwise, the initial values will not be calculated. |
X |
A matrix with the predictor variable(s). A regression model will be fit to each of them, separately. |
tol |
The tolerance value to terminate the EM algorithm. |
maxit |
The maximum number of iterations allowed for the EM algorithm. |
Details
A conditional logistic normal regression is being fitted for each of the columns of X.
Value
A list including:
stat |
A vector with the log-likelihood ratio test statistics. |
pvalue |
A vector with the corresponding logged p-values. |
Author(s)
Michail Tsagris.
R implementation and documentation: Michail Tsagris mtsagris@uoc.gr.
References
Alzeley O. and Tsagris M. (2026). Modelling compositional data with structural zero values https://arxiv.org/pdf/2608.29954
See Also
Examples
x <- matrix( rnorm(150 * 10), ncol = 10 )
y <- as.matrix(iris[, 1:3])
y <- y / rowSums(y)
ind <- sample(150, 10)
for ( k in ind ) y[k, sample(3, 1)] <- 0
y <- y / rowSums(y)
mod <- cln.regs(y, x)
Residual diagnostic plot for the CLN and CLNR models
Description
Residual diagnostic plot for the CLN and CLNR models.
Usage
cln.residplot(y, m = NULL, S, x = NULL, beta = NULL)
Arguments
y |
A numerical matrix with the compositional data, possibly containing zero values. The last column is the reference component of the additive log-ratio transformation. |
m |
The estimated mean vector of the alr transformed data, of length |
S |
The estimated |
x |
The predictor variable(s), the same as those given to |
beta |
The |
Details
For the i-th composition let b_i be the vector of the log-ratios among
its C_i nonzero components, each taken against the last nonzero one. Under the
model, b_i follows a multivariate normal distribution with mean
Q_i \mu_i and covariance matrix Q_i \Sigma Q_i^\top, where
Q_i depends only on the zero pattern and \mu_i is either the common mean
or B^\top x_i. The squared Mahalanobis distance
m_i = (b_i - Q_i \mu_i)^\top (Q_i \Sigma Q_i^\top)^{-1} (b_i - Q_i \mu_i)
is compared to a chi-square distribution with c_i = C_i - 1 degrees of
freedom. The chi-square probabilities are transformed to normal scores,
z_i = \Phi^{-1}(G_{c_i}(m_i)), which have a common scale for
all rows, and are plotted against the standard normal quantiles. Only the observed part of each
vector is used, so no zero value is replaced.
The reference distribution ignores that the parameters were estimated from the same data, so
for small samples or many predictors the plot may look slightly better than the fit deserves.
Compositions with only one nonzero component contain no log-ratio and are omitted (NA).
Value
The QQ-plot is produced and a list including:
m2 |
The squared Mahalanobis distance |
dof |
The degrees of freedom of each distance, that is, the number of nonzero components minus one. |
cdf |
The value of the chi-square cumulative distribution function with |
z |
The normal scores of |
Author(s)
Michail Tsagris.
R implementation and documentation: Michail Tsagris mtsagris@uoc.gr.
References
Alzeley O. and Tsagris M. (2026). Modelling compositional data with structural zero values https://arxiv.org/pdf/2608.29954
Aitchison J. (2003). The statistical analysis of compositional data. New Jersey: Reprinted by The Blackburn Press.
See Also
Examples
x <- as.vector(iris[, 4])
y <- as.matrix(iris[, 1:3])
y <- y / rowSums(y)
ind <- sample(150, 10)
for ( k in ind ) y[k, sample(3, 1)] <- 0
y <- y / rowSums(y)
## without covariates
mod <- cln.mle(y)
cln.residplot(y, m = mod$m, S = mod$S)
## with covariates
mod <- cln.reg(y, x)
cln.residplot(y, S = mod$S, x = x, beta = mod$beta)
Density of the conditional logistic normal model
Description
Density of the conditional logistic normal model.
Usage
dcln(y, m, S, logged = FALSE)
Arguments
y |
A vector or a matrix with compositional data with zero values. |
m |
The mean vector in |
S |
The covariance matrix in |
logged |
If you want the log of the density set this TRUE, otherwise leave it FALSE. |
Details
The function computes the density values of the conditional multivariate normal model.
Value
The density values at the given compositional data x.
Author(s)
Michail Tsagris.
R implementation and documentation: Michail Tsagris mtsagris@uoc.gr.
References
Alzeley O. and Tsagris M. (2026). Modelling compositional data with structural zero values https://arxiv.org/pdf/2608.29954
Aitchison J. (1986). The statistical analysis of compositional data, pages 272–274.
See Also
Examples
x <- as.vector(iris[, 4])
y <- as.matrix(iris[, 1:3])
y <- y / rowSums(y)
ind <- sample(150, 10)
for ( k in ind ) y[k, sample(3, 1)] <- 0
y <- y / rowSums(y)
mod <- cln.mle(y)
f <- dcln(y, mod$m, mod$S)
Forward-backward with early dropping variable selection
Description
Forward-backward with early dropping variable selection.
Usage
fbed.cln(y, X, prior = NULL, alpha = 0.05, univ = NULL, K = 0,
tol = 1e-6, maxit = 500)
Arguments
y |
A matrix with the compositional data (dependent variable). The number of observations (vectors) with no zero values should be more than the columns of the predictor variables. Otherwise, the initial values will not be calculated. |
X |
A matrix with the available predictor variables. |
prior |
A matrix or a vector with some predictors that will be in the model. |
alpha |
The significance level. |
univ |
If you have the test statistics and the p-values (the output of |
K |
How many times should the FBED be repeated? |
tol |
The tolerance value to terminate the EM algorithm. |
maxit |
The maximum number of iterations allowed for the EM algorithm. |
Details
The function performs the FBED variable selection algorithm of Borboudakis and Tsamardinos (2019).
Value
A list including:
univ |
The result of the simple regressions, the same result as |
res |
A matrix with the selected variables, their test statistic and their logged p-value. |
info |
A matrix with two columns, the number of variables selected at each step and the number of tests performed at each step. Each step means at each value of K. |
runtime |
The running time of the algorithm. |
Author(s)
Michail Tsagris.
R implementation and documentation: Michail Tsagris mtsagris@uoc.gr.
References
Alzeley O. and Tsagris M. (2026). Modelling compositional data with structural zero values https://arxiv.org/pdf/2608.29954
Borboudakis G. and Tsamardinos I. (2019). Forward-backward selection with early dropping. Journal of Machine Learning Research, 20(8): 1-39.
See Also
Examples
y <- as.matrix(iris[, 1:3])
y <- y / rowSums(y)
ind <- sample(150, 10)
for ( k in ind ) y[k, sample(3, 1)] <- 0
y <- y / rowSums(y)
X <- matrix(rnorm(150 * 30), ncol = 30)
mod <- fbed.cln(y, X)
Ternary diagram
Description
Ternary diagram.
Usage
ternary(y, dg = FALSE, hg = FALSE, m = NULL, colour = NULL)
Arguments
y |
A matrix with the compositional data (with zero values present). |
dg |
Do you want diagonal grid lines to appear? If yes, set this TRUE. |
hg |
Do you want horizontal grid lines to appear? If yes, set this TRUE. |
m |
If you know the estimated simplicial mean pass it here, otherwise leave it NULL. It will appear with the symbol +. |
colour |
If you want the points to appear in different colour put a vector with the colour numbers or colours. |
Details
There are two ways to create a ternary graph. We used here that one where each edge is equal to 1 and it is what Aitchison (1986) uses. For every given point, the sum of the distances from the edges is equal to 1. Horizontal and or diagonal grid lines, and the mean can appear. Zeros in the data appear with red x.
Value
The ternary plot. Additionally, horizontal or diagonal grid lines can appear as well.
Author(s)
Michail Tsagris.
R implementation and documentation: Michail Tsagris mtsagris@uoc.gr.
References
Aitchison, J. (1983). Principal component analysis of compositional data. Biometrika 70(1): 57–65.
Aitchison J. (1986). The statistical analysis of compositional data. Chapman & Hall.
See Also
Examples
y <- as.matrix(iris[, 1:3])
y <- y / rowSums(y)
ind <- sample(150, 10)
for ( k in ind ) y[k, sample(3, 1)] <- 0
y <- y / rowSums(y)
mod <- cln.mle(y)
ternary(y, hg = TRUE, dg = TRUE, m = mod$mesi)