Cluster Analysis

An interactive tool for finding an optimal splitting of paths by behavioral metrics. Allows you to inspect clusters in a Segment Overview-style heatmap and offers the best possible splitting from the silhouette score perspective. When a grid is searched, the Silhouette tab charts every candidate and you can click any bar to interpret that partition instead of the top-scoring one — useful when several candidates score alike and the winner doesn't read well. Everything downstream follows the pick, including the code shown and the segment saved by "Save Clusters". Once the splitting looks right, you can label the clusters and save them as a new segment column of the eventstream right from the UI by clicking "Save Clusters".

How it works

Every representation in Path Analysis summarizes paths by their shape — which events, in which order. Clustering goes the other way: it describes each path by a handful of numbers, then groups paths whose numbers look alike. Two users who never visited the same screen can still land in the same cluster if they were equally brief, equally repetitive, or equally unlikely to buy.

The pipeline is three steps, and each maps to one argument:

  1. Describe. features turns every path into a row of numbers, using the shared path metrics registry — length, duration, active_days, a per-event count, whether a pattern matched. This choice is the analysis: clusters can only differ along dimensions you measured.
  2. Scale. scaler puts those numbers on a comparable footing. Without it a duration in seconds (six digits) drowns out a length in events (two digits), and the clustering silently becomes "group by duration". Hence the "minmax" default.
  3. Split. n_clusters fixes the number of groups, or names a range to search. Given a range, the widget clusters at every size in it and keeps the one with the best silhouette score — a measure of how much tighter paths sit within their own cluster than to the nearest other one. It is a hint about structure, not a verdict: a mediocre score across the whole range usually means the paths form a continuum rather than distinct groups.

An optional fourth knob, NMF, sits between steps 2 and 3.

What comes back is not the clusters themselves but a way to read them: the overview heatmap puts overview_metrics in rows and clusters in columns, exactly like Segment Overview, so you can see what actually distinguishes cluster 2 from cluster 3 and give it a name. Note that features and overview_metrics are independent on purpose — profiling clusters on metrics you did not cluster by is how you find out whether the split means anything beyond its own inputs.

Clusters are not a special kind of object. Once a split looks right, "Save Clusters" (or add_clusters) writes the labels into the eventstream as an ordinary segment column, and from there every segment-aware tool — diff mode, Segment Overview, the in_segment metric — works on it unchanged.

Choosing a cluster count is a judgment call, not a computation. Prefer the smallest number of clusters you can still describe in words; a split you cannot name is a split you cannot act on.

Reading past the winner

Because it is a judgment call, the top silhouette score is a starting point rather than an answer. Scores across a range are often close enough that the winner is decided by noise, while the runner-up splits into groups you can actually name.

The Silhouette tab charts every candidate the search tried. Click any bar to interpret that partition instead: the overview heatmap, the H-matrix, the copy-pasteable add_clusters(...) code and "Save Clusters" all follow the pick, so what you read is what you save. The top-scoring point keeps a ★ so you never lose sight of what you traded away, and the shaded band holds every candidate within 5% of it — those are the ones worth reading before settling.

To pin that choice programmatically — in a notebook you re-run, or to open the widget on it — pass the select argument instead of taking the winner.

Usage

stream.cluster_analysis(
    features=[{"metric": "length"}, {"metric": "duration"}, {"metric": "event_count_bulk"}],
    method_args={"n_clusters": "3-6"},
)

stream.cluster_analysis(
    features=[{"metric": "length"}],
    method="hdbscan",
    method_args={"min_cluster_size": 50},
)

Examples

Basic

stream.cluster_analysis(
  features=[
      {"metric": "event_count_bulk", "metric_args": {"events": ["catalog", "product_view", "add_to_cart", "purchase"]}},
  ],
  method_args={"n_clusters": 3},
)

Parameters

Data

Data parameters change the computed result. They are the arguments of the widget's headless twin stream.cluster_analysis_data() — see headless mode below — with one exception, which the twin accepts and this constructor does not:

  • nmf_components — see NMF below. It is not missing from the widget: the sidebar has an NMF toggle and a component-count field, and the result arrives as the H-matrix / W Cluster Means tabs. It is only unavailable as a constructor argument, so set it in the sidebar or use the headless twin.

method and its method_args are accepted here, but the sidebar only edits n_clusters: there is no control to switch the algorithm, and none for hdbscan's min_cluster_size / cluster_selection_epsilon. Passing them to the constructor is the way to run hdbscan in the widget.

ParameterTypeDescription
featureslist of dict, optionalMetric configurations used as clustering features (see the Path Metrics). If omitted, the sidebar starts pre-filled with a wildcard event_count_bulk metric (one column per event in the eventstream) — that pre-fill is a starting point to edit, not something that runs on its own: passing features explicitly (or clicking "Apply" in the sidebar) is what actually triggers clustering.
method{"kmeans", "hdbscan"}, default "kmeans"Clustering algorithm.
method_argsdict, optionalParameters of the chosen method, same shape and schema as cluster_analysis_data: n_clusters for "kmeans" (also editable in the sidebar, so a value passed here is a starting point rather than a lock), min_cluster_size and cluster_selection_epsilon for "hdbscan" — which the sidebar has no fields for, making this the only way to set them.
scaler{"minmax", "std"}, optionalFeature scaler applied before clustering; default "minmax".
overview_metricslist of dict, optionalMetrics shown in the overview heatmap after clustering (independent of features). If omitted, the sidebar starts pre-filled with a wildcard event_count_bulk metric here too (mean count per event); same as features, it only takes effect once you click "Apply" or pass the argument explicitly. Both features and overview_metrics accept metric configs from the same Path Metrics registry.
path_colstr, optionalPath ID column override; defaults to schema.path_col.
selectdict, optionalWhich grid point the widget opens on, e.g. {"n_clusters": 5}; by default the top-scoring one. Its keys are bare parameter names, as they appear in silhouette["params"] — a grid point is a coordinate, not a call. Only meaningful when a method_args value (or nmf_components) is a range, and equivalent to clicking that bar in the Silhouette tab. Persisted with the rest of the widget state.

Display

Display parameters only affect how the widget is rendered.

ParameterTypeDescription
heightint, default 520Widget height in pixels.
sidebar_openbool, default TrueWhether the sidebar starts open.
state_filestr, optionalJSON file the widget state is bound to; see Saving widget state.

NMF

Turning on NMF in the sidebar (or passing nmf_components to the headless twin) inserts a dimensionality-reduction step between scaling and clustering: the feature matrix is factorized into the requested number of non-negative components, and paths are clustered on those components instead of on the raw metrics. It is independent of the clustering algorithm — the usual pairing is NMF with k-means.

Reach for it when features is wide and correlated — a wildcard event_count_bulk over a few dozen events is the standard case. Twenty near-duplicate count columns let the loudest events dominate the distance calculation; a handful of components spreads the signal out. Like n_clusters, nmf_components accepts a single number or a range to grid-search. It sits outside method_args because it is a pipeline step, not a parameter of the clustering algorithm — the same is true of scaler.

The payoff is interpretability, and it arrives as two extra tabs:

  • H-matrix — how each component is built out of your features. This is what names a component: one loading heavily on search and filter_results is "browsing intensity", whatever the algorithm called it.
  • W Cluster Means — how strongly each cluster expresses each component. Read it together with the H-matrix and each cluster gets a description in terms of behavior rather than a number.

The cost is that clusters no longer sit in the space of your original metrics, so overview_metrics (which always report raw per-path values) stays the honest check on whether the split means anything.

From the widget to a segment column

The widget explores; add_clusters persists. The two differ in one place: cluster_analysis searches a range of cluster counts ("3-8" by default) and picks a winner, while add_clusters materializes one specific clustering and therefore needs an exact method_args={"n_clusters": ...}.

You don't have to transcribe it. "Save Clusters" in the sidebar writes the matching add_clusters(...) call — optionally with your renamed cluster labels — and headlessly the same value comes back as best_params:

features = [{"metric": "length"}, {"metric": "active_days"}]

result = stream.cluster_analysis_data(features=features)
result["best_params"]        # {"method": "kmeans", "method_args": {"n_clusters": 3}, "scaler": "minmax"}

stream = stream.add_clusters("behavior", features=features, **result["best_params"])

best_params is shaped as add_clusters keyword arguments, but it carries only how the clustering was configured — pass the same features alongside it.

Headless mode

stream.cluster_analysis_data()

Run cluster analysis headlessly and return a dict of results.

Pass lists inside method_args (or for nmf_components) to trigger grid search with silhouette scoring. For the kmeans method (the default), n_clusters defaults to "3-8" if omitted — including for nmf_components-only searches.

best_params holds the concrete parameter values actually used to produce overview_df (the winning combination when searching, the point named by select if you named one, or just the fixed values passed in otherwise), already shaped as add_clusters keyword arguments (method, method_args, scaler, and nmf_components if one was used) — splat it into add_clusters to materialize the same clustering as a segment column.