pyppur.utils package

Utility functions for pyppur.

pyppur.utils.compute_silhouette(X_embedded, labels)[source]

Compute the silhouette score for the embedding.

The silhouette score measures how well clusters are separated.

Parameters:
  • X_embedded (ndarray) – Low-dimensional embedding.

  • labels (ndarray) – Cluster or class labels.

Returns:

Silhouette score in range [-1, 1].

Return type:

float

pyppur.utils.compute_trustworthiness(X_original, X_embedded, n_neighbors=5)[source]

Compute the trustworthiness score for dimensionality reduction.

Trustworthiness measures how well local neighborhoods are preserved.

Parameters:
  • X_original (ndarray) – Original high-dimensional data.

  • X_embedded (ndarray) – Low-dimensional embedding.

  • n_neighbors (int) – Number of neighbors to consider.

Returns:

Trustworthiness score in range [0, 1].

Return type:

float

pyppur.utils.standardize_data(X, center=True, scale=True, scaler=None)[source]

Standardize data for projection pursuit.

Parameters:
  • X (ndarray) – Input data, shape (n_samples, n_features).

  • center (bool) – Whether to center the data.

  • scale (bool) – Whether to scale the data to unit variance.

  • scaler (StandardScaler | None) – Optional pre-fitted scaler for transform-only operation.

Returns:

Standardized data and the scaler.

Return type:

tuple[ndarray, StandardScaler]

Submodules

pyppur.utils.metrics module

Evaluation metrics for dimensionality reduction.

pyppur.utils.metrics.compute_trustworthiness(X_original, X_embedded, n_neighbors=5)[source]

Compute the trustworthiness score for dimensionality reduction.

Trustworthiness measures how well local neighborhoods are preserved.

Parameters:
  • X_original (ndarray) – Original high-dimensional data.

  • X_embedded (ndarray) – Low-dimensional embedding.

  • n_neighbors (int) – Number of neighbors to consider.

Returns:

Trustworthiness score in range [0, 1].

Return type:

float

pyppur.utils.metrics.compute_silhouette(X_embedded, labels)[source]

Compute the silhouette score for the embedding.

The silhouette score measures how well clusters are separated.

Parameters:
  • X_embedded (ndarray) – Low-dimensional embedding.

  • labels (ndarray) – Cluster or class labels.

Returns:

Silhouette score in range [-1, 1].

Return type:

float

pyppur.utils.metrics.compute_distance_distortion(X_original, X_embedded)[source]

Compute the distance distortion between original and embedded spaces.

Distance distortion measures how well pairwise distances are preserved.

Parameters:
  • X_original (ndarray) – Original high-dimensional data.

  • X_embedded (ndarray) – Low-dimensional embedding.

Returns:

Mean squared distance distortion.

Return type:

float

pyppur.utils.metrics.evaluate_embedding(X_original, X_embedded, labels=None, n_neighbors=5)[source]

Evaluate the quality of an embedding using multiple metrics.

Parameters:
  • X_original (ndarray) – Original high-dimensional data.

  • X_embedded (ndarray) – Low-dimensional embedding.

  • labels (ndarray | None) – Optional cluster or class labels.

  • n_neighbors (int) – Number of neighbors for trustworthiness.

Returns:

Dictionary with evaluation metrics.

Return type:

dict[str, float]

pyppur.utils.preprocessing module

Preprocessing utilities for projection pursuit.

pyppur.utils.preprocessing.standardize_data(X, center=True, scale=True, scaler=None)[source]

Standardize data for projection pursuit.

Parameters:
  • X (ndarray) – Input data, shape (n_samples, n_features).

  • center (bool) – Whether to center the data.

  • scale (bool) – Whether to scale the data to unit variance.

  • scaler (StandardScaler | None) – Optional pre-fitted scaler for transform-only operation.

Returns:

Standardized data and the scaler.

Return type:

tuple[ndarray, StandardScaler]

pyppur.utils.visualization module

Visualization utilities for projection pursuit results.

pyppur.utils.visualization.plot_embedding(X_embedded, labels=None, title='Projection Pursuit Embedding', metrics=None, figsize=(10, 8), cmap='tab10', alpha=0.7, s=30.0, ax=None)[source]

Plot the results of a projection pursuit embedding.

Parameters:
  • X_embedded (ndarray) – Embedded data, shape (n_samples, 2) or (n_samples, 3).

  • labels (ndarray | None) – Optional labels for coloring points.

  • title (str) – Plot title.

  • metrics (dict[str, float] | None) – Optional dictionary of metrics to include in title.

  • figsize (tuple[float, float]) – Figure size (width, height) in inches.

  • cmap (str) – Colormap name.

  • alpha (float) – Transparency of points.

  • s (float) – Point size.

  • ax (Axes | Axes3D | None) – Optional axes to plot on.

Returns:

Figure and Axes objects.

Return type:

tuple[Figure, Axes | Axes3D]

pyppur.utils.visualization.plot_reconstruction(X, X_recon, n_samples=3)[source]

Plot reconstructed samples alongside original samples.

Parameters:
  • X (ndarray) – Original data.

  • X_recon (ndarray) – Reconstructed data.

  • n_samples (int) – Number of samples to plot.

Returns:

matplotlib Figure.

Return type:

Figure

pyppur.utils.visualization.plot_comparison(embeddings, labels=None, metrics=None, title=None, figsize=(15, 5), cmap='tab10', alpha=0.7, s=30.0)[source]

Plot a comparison of multiple embeddings.

Parameters:
  • embeddings (dict[str, ndarray]) – Dictionary of embeddings {name: embedded_data}.

  • labels (ndarray | None) – Optional labels for coloring points.

  • metrics (dict[str, dict[str, float]] | None) – Optional dictionary of metrics for each embedding.

  • title (str | None) – Optional overall figure title.

  • figsize (tuple[float, float]) – Figure size (width, height) in inches.

  • cmap (str) – Colormap name.

  • alpha (float) – Transparency of points.

  • s (float) – Point size.

Returns:

matplotlib Figure object.

Return type:

Figure