dissimilarity              package:clue              R Documentation

_D_i_s_s_i_m_i_l_a_r_i_t_y _B_e_t_w_e_e_n _P_a_r_t_i_t_i_o_n_s _o_r _H_i_e_r_a_r_c_h_i_e_s

_D_e_s_c_r_i_p_t_i_o_n:

     Compute the dissimilarity between (ensembles) of partitions or
     hierarchies.

_U_s_a_g_e:

     cl_dissimilarity(x, y = NULL, method = "euclidean")

_A_r_g_u_m_e_n_t_s:

       x: an ensemble of partitions or hierarchies, or something
          coercible to that (see 'cl_ensemble').

       y: 'NULL' (default), or as for 'x'.

  method: a character string specifying one of the built-in methods for
          computing dissimilarity, or a function to be taken as a
          user-defined method.  If a character string, its lower-cased
          version is matched against the lower-cased names of the
          available built-in methods using 'pmatch'.  See *Details* for
          available built-in methods.

_D_e_t_a_i_l_s:

     If 'y' is given, its components must be of the same kind as those
     of 'x' (i.e., components must either all be partitions, or all be
     hierarchies).

     If all components are partitions, the following built-in methods
     for measuring dissimilarity between two partitions with respective
     membership matrices u and v (brought to a common number of
     columns) are available:

     '"_e_u_c_l_i_d_e_a_n"' the Euclidean dissimilarity of the memberships,
          i.e., the square root of the minimal sum of the squared
          differences of u and all column permutations of v.  See
          Dimitriadou, Weingessel and Hornik (2002).

     '"_m_a_n_h_a_t_t_a_n"' the Manhattan dissimilarity of the memberships,
          i.e., the minimal sum of the absolute differences of u and
          all column permutations of v.

     '"_c_o_m_e_m_b_e_r_s_h_i_p_s"' the Euclidean dissimilarity of the elements of
          the co-membership matrices C(u) = u u' and C(v), i.e., the
          square root of the sum of the squared differences of C(u) and
          C(v).

     If all components are hierarchies, available built-in methods for
     measuring agreement between two hierarchies with respective
     ultrametrics u and v are as follows.

     '"_e_u_c_l_i_d_e_a_n"' the Euclidean dissimilarity of the ultrametrics
          (i.e., the square root of the sum of the squared differences
          of u and v).

     '"_e_u_c_l_i_d_e_a_n"' the Manhattan dissimilarity of the ultrametrics
          (i.e., the sum of the absolute differences of u and v).

     '"_c_o_p_h_e_n_e_t_i_c"' 1 - c^2, where c is the cophenetic correlation
          coefficient (i.e., the product-moment correlation of the
          ultrametrics).

     '"_g_a_m_m_a"' the rate of inversions between the ultrametrics (i.e.,
          the rate of pairs (i,j) and (k,l) for which u_{ij} < u_{kl}
          and v_{ij} > v_{kl}).

     If a user-defined agreement method is to be employed, it must be a
     function taking two clusterings as its arguments.

     Symmetric dissimilarity objects of class '"cl_dissimilarity"' are
     implemented as symmetric proximity objects with self-proximities
     identical to zero, and inherit from class '"cl_proximity"'.  They
     can be coerced to dense square matrices using 'as.matrix'.  It is
     possible to use 2-index matrix-style subscripting for such
     objects; unless this uses identical row and column indices, this
     results in a (non-symmetric dissimilarity) object of class
     '"cl_cross_dissimilarity"'.

     Symmetric dissimilarity objects also inherit from class '"dist"'
     (although they currently do not "strictly" extend this class),
     thus making it possible to use them directly for clustering
     algorithms based on dissimilarity matrices of this class, see the
     examples.

_V_a_l_u_e:

     If 'y' is 'NULL', an object of class '"cl_dissimilarity"'
     containing the dissimilarities between all pairs of components of
     'x'.  Otherwise, an object of class '"cl_cross_dissimilarity"'
     with the dissimilarities between the components of 'x' and the
     components of 'y'.

_R_e_f_e_r_e_n_c_e_s:

     E. Dimitriadou and A. Weingessel and K. Hornik (2002). A
     combination scheme for fuzzy clustering. _International Journal of
     Pattern Recognition and Artificial Intelligence_, *16*, 901-912.

_S_e_e _A_l_s_o:

     'cl_agreement'

_E_x_a_m_p_l_e_s:

     ## An ensemble of partitions.
     data("CKME")
     pens <- CKME[1 : 30]
     diss <- cl_dissimilarity(pens)
     summary(c(diss))
     cl_dissimilarity(pens[1:5], pens[6:7])
     ## Equivalently, using subscripting.
     diss[1:5, 6:7]
     ## Can use the dissimilarities for "secondary" clustering
     ## (e.g. obtaining hierarchies of partitions):
     hc <- hclust(diss)
     plot(hc)

