dissimilarity              package:clue              R Documentation

_D_i_s_s_i_m_i_l_a_r_i_t_y _B_e_t_w_e_e_n _P_a_r_t_i_t_i_o_n_s _o_r _H_i_e_r_a_r_c_h_i_e_s

_D_e_s_c_r_i_p_t_i_o_n:

     Compute the dissimilarity between (ensembles) of partitions or
     hierarchies.

_U_s_a_g_e:

     cl_dissimilarity(x, y = NULL, method = "euclidean")

_A_r_g_u_m_e_n_t_s:

       x: an ensemble of partitions or hierarchies, or something
          coercible to that (see 'cl_ensemble').

       y: 'NULL' (default), or as for 'x'.

  method: a character string specifying one of the built-in methods for
          computing dissimilarity, or a function to be taken as a
          user-defined method.  If a character string, its lower-cased
          version is matched against the lower-cased names of the
          available built-in methods using 'pmatch'.  See *Details* for
          available built-in methods.

_D_e_t_a_i_l_s:

     If 'y' is given, its components must be of the same kind as those
     of 'x' (i.e., components must either all be partitions, or all be
     hierarchies).

     If all components are partitions, the following built-in methods
     for measuring dissimilarity between two partitions with respective
     membership matrices u and v (brought to a common number of
     columns) are available:

     '"_e_u_c_l_i_d_e_a_n"' the Euclidean dissimilarity of the memberships,
          i.e., the minimal sum of the squared differences of u and all
          column permutations of v, divided by twice the number of
          objects in the partitions.  See Dimitriadou, Weingessel and
          Hornik (2002).

     '"_c_o_m_e_m_b_e_r_s_h_i_p_s"' the Euclidean dissimilarity of the elements of
          the co-membership matrices C(u) = u u' and C(v), i.e., the
          sum of the squared differences of C(u) and C(v), divided by
          the squared number of objects in the partitions.

     If all components are hierarchies, available built-in methods for
     measuring agreement between two hierarchies with respective
     ultrametrics u and v are as follows.

     '"_e_u_c_l_i_d_e_a_n"' the Euclidean dissimilarity of the ultrametrics
          (i.e., the sum of the squared differences of u and v).

     '"_c_o_p_h_e_n_e_t_i_c"' 1 - c^2, where c is the cophenetic correlation
          coefficient (i.e., the product-moment correlation of the
          ultrametrics).

     '"_g_a_m_m_a"' the rate of inversions between the ultrametrics (i.e.,
          the rate of pairs (i,j) and (k,l) for which u_{ij} < u_{kl}
          and v_{ij} > v_{kl}).

     If a user-defined agreement method is to be employed, it must be a
     function taking two clusterings as its arguments.

     Symmetric dissimilarity objects of class '"cl_dissimilarity"' are
     implemented as symmetric proximity objects with self-proximities
     identical to zero, and inherit from class '"cl_proximity"'.  They
     can be coerced to dense square matrices using 'as.matrix'.  It is
     possible to use 2-index matrix-style subscripting for such
     objects; unless this uses identical row and column indices, this
     results in a (non-symmetric dissimilarity) object of class
     '"cl_cross_dissimilarity"'.

     Symmetric dissimilarity objects also inherit from class '"dist"'
     (although they currently do not "strictly" extend this class),
     thus making it possible to use them directly for clustering
     algorithms based on dissimilarity matrices of this class, see the
     examples.

_V_a_l_u_e:

     If 'y' is 'NULL', an object of class '"cl_dissimilarity"'
     containing the dissimilarities between all pairs of components of
     'x'.  Otherwise, an object of class '"cl_cross_dissimilarity"'
     with the dissimilarities between the components of 'x' and the
     components of 'y'.

_R_e_f_e_r_e_n_c_e_s:

     E. Dimitriadou and A. Weingessel and K. Hornik (2002). A
     combination scheme for fuzzy clustering. _International Journal of
     Pattern Recognition and Artificial Intelligence_, *16*, 901-912.

_S_e_e _A_l_s_o:

     'cl_agreement'

_E_x_a_m_p_l_e_s:

     ## An ensemble of partitions.
     data("CKME")
     pens <- CKME[1 : 30]
     diss <- cl_dissimilarity(pens)
     summary(c(diss))
     cl_dissimilarity(pens[1:5], pens[6:7])
     ## Equivalently, using subscripting.
     diss[1:5, 6:7]
     ## Can use the dissimilarities for "secondary" clustering
     ## (e.g. obtaining hierarchies of partitions):
     hc <- hclust(diss)
     plot(hc)

