---
title: "Using dbscan with tidyverse"
author: "Michael Hahsler"
output: rmarkdown::html_vignette
vignette: >
  %\VignetteIndexEntry{Using dbscan with tidyverse}
  %\VignetteEngine{knitr::rmarkdown}
  %\VignetteEncoding{UTF-8}
---

```{r setup, include=FALSE}
knitr::opts_chunk$set(collapse = TRUE, comment = "#>")
```

The **dbscan** package provides `tidy()`, `augment()`, and `glance()` methods
for its clustering algorithms, making them easy to use with tidyverse,
ggplot2, and [tidymodels](https://www.tidymodels.org/learn/statistics/k-means/).

Load the packages and prepare the numeric variables from the iris data:

```{r data, message=FALSE, warning=FALSE}
library(dbscan)
library(tidyverse)

x <- iris[, 1:4]
db <- x %>% dbscan(eps = .42, minPts = 5)
```

Get cluster statistics as a tibble:

```{r tidy}
tidy(db)
```

Visualize the clustering with ggplot2, using an x for noise points:

```{r plot, fig.alt="DBSCAN clusters in the iris data"}
augment(db, x) %>%
  ggplot(aes(x = Petal.Length, y = Petal.Width)) +
  geom_point(aes(color = .cluster, shape = noise)) +
  scale_shape_manual(values = c(19, 4))
```
