API Reference

This page provides detailed documentation for all pyppur classes and functions.

Main Classes

ProjectionPursuit

class pyppur.ProjectionPursuit(n_components=2, objective='distance_distortion', alpha=0.1, max_iter=500, tol=1e-06, random_state=None, optimizer='L-BFGS-B', n_init=3, verbose=False, center=True, scale=True, weight_by_distance=False, tied_weights=True, l2_reg=0.0, use_nonlinearity_in_distance=True, distance_metric='correlation')[source]

Bases: TransformerMixin, BaseEstimator

Implementation of Projection Pursuit for dimensionality reduction.

This class provides methods to find optimal projections by minimizing either reconstruction loss or distance distortion. It supports both initialization strategies and different optimizers.

Parameters:
__init__(n_components=2, objective='distance_distortion', alpha=0.1, max_iter=500, tol=1e-06, random_state=None, optimizer='L-BFGS-B', n_init=3, verbose=False, center=True, scale=True, weight_by_distance=False, tied_weights=True, l2_reg=0.0, use_nonlinearity_in_distance=True, distance_metric='correlation')[source]

Initialize a ProjectionPursuit model.

Parameters:
  • n_components (int) – Number of projection dimensions to use.

  • objective (Objective | str) – Optimization objective enum value.

  • alpha (float) – Steepness parameter for the ridge function g(z) = tanh(alpha * z). Lower values (0.01-0.1) work better for distance preservation.

  • max_iter (int) – Maximum number of iterations for optimization.

  • tol (float) – Tolerance for optimization convergence.

  • random_state (int | None) – Random seed for reproducibility.

  • optimizer (str) – Optimization method (‘L-BFGS-B’ recommended).

  • n_init (int) – Number of random initializations to try.

  • verbose (bool) – Whether to log progress information.

  • center (bool) – Whether to center the data.

  • scale (bool) – Whether to scale the data.

  • weight_by_distance (bool) – Whether to weight distance distortion by inverse of original distances.

  • tied_weights (bool) – Whether to use tied weights (encoder=decoder) for reconstruction.

  • l2_reg (float) – L2 regularization strength for decoder weights (when tied_weights=False).

  • use_nonlinearity_in_distance (bool) – Whether to apply ridge function before computing distances.

  • distance_metric (Literal['mse', 'correlation', 'spearman']) – How to measure distance preservation. Options: ‘mse’ (scale-sensitive), ‘correlation’ (scale-invariant), ‘spearman’ (rank-based). Default is ‘correlation’.

Return type:

None

fit(X, y=None)[source]

Fit the ProjectionPursuit model to the data.

Parameters:
  • X (ndarray) – Input data, shape (n_samples, n_features).

  • y (Any) – Ignored. Accepted for scikit-learn estimator API compatibility.

Returns:

The fitted model.

Return type:

ProjectionPursuit

transform(X)[source]

Apply dimensionality reduction to X.

Parameters:

X (ndarray) – Input data, shape (n_samples, n_features).

Returns:

Transformed data, shape (n_samples, n_components).

Return type:

ndarray

fit_transform(X, y=None, **fit_params)[source]

Fit the model with X and apply dimensionality reduction on X.

Parameters:
  • X (ndarray) – Input data, shape (n_samples, n_features).

  • y (Any) – Ignored. Accepted for scikit-learn estimator API compatibility.

  • **fit_params (Any) – Ignored. Accepted for scikit-learn API compatibility.

Returns:

Transformed data, shape (n_samples, n_components).

Return type:

ndarray

reconstruct(X)[source]

Reconstruct X from the projected data.

Parameters:

X (ndarray) – Input data, shape (n_samples, n_features).

Returns:

Reconstructed data, shape (n_samples, n_features).

Return type:

ndarray

reconstruction_error(X)[source]

Compute the reconstruction error for X.

Parameters:

X (ndarray) – Input data, shape (n_samples, n_features).

Returns:

Mean squared reconstruction error.

Return type:

float

distance_distortion(X)[source]

Compute the distance distortion for X.

Parameters:

X (ndarray) – Input data, shape (n_samples, n_features).

Returns:

Mean squared distance distortion.

Return type:

float

compute_trustworthiness(X, n_neighbors=5)[source]

Compute the trustworthiness score for the dimensionality reduction.

Trustworthiness measures how well the local structure is preserved. A score of 1.0 indicates perfect trustworthiness, while a score of 0.0 indicates that the local structure is not preserved at all.

Parameters:
  • X (ndarray) – Input data, shape (n_samples, n_features).

  • n_neighbors (int) – Number of neighbors to consider for trustworthiness.

Returns:

Trustworthiness score between 0.0 and 1.0.

Return type:

float

compute_silhouette(X, labels)[source]

Compute the silhouette score for the dimensionality reduction.

Silhouette score measures how well clusters are separated. A score close to 1.0 indicates that clusters are well separated, while a score close to -1.0 indicates poor separation.

Parameters:
  • X (ndarray) – Input data, shape (n_samples, n_features).

  • labels (ndarray) – Cluster labels for each sample.

Returns:

Silhouette score between -1.0 and 1.0.

Return type:

float

evaluate(X, labels=None, n_neighbors=5)[source]

Evaluate the dimensionality reduction with multiple metrics.

Parameters:
  • X (ndarray) – Input data, shape (n_samples, n_features).

  • labels (ndarray | None) – Optional cluster labels for silhouette score.

  • n_neighbors (int) – Number of neighbors for trustworthiness.

Returns:

Dictionary with evaluation metrics.

Return type:

dict[str, float]

property n_components_: int

Number of components actually used by the fit.

This equals the n_components constructor argument, except when that exceeded the number of features, in which case it is the number of features. The constructor argument itself is never modified.

Returns:

Number of projection directions found.

property x_loadings_: ndarray

Projection directions (encoder).

Returns:

Projection directions, shape (n_components, n_features).

property decoder_weights_: ndarray | None

Decoder weights (for untied weights only).

Returns:

Decoder weights, shape (n_components, n_features), or None if using tied weights.

property loss_curve_: list[float]

Loss curve during optimization.

Returns:

Loss values during optimization.

property best_loss_: float

Best loss value achieved.

Returns:

Best loss value.

property fit_time_: float

Time taken to fit the model.

Returns:

Time in seconds.

property optimizer_info_: dict[str, Any]

Additional information from the optimizer.

Returns:

Optimizer information.

Objective Types

class pyppur.Objective(*values)[source]

Bases: StrEnum

Objective types for projection pursuit.

DISTANCE_DISTORTION = 'distance_distortion'
RECONSTRUCTION = 'reconstruction'

Objective Functions

Base Objective

class pyppur.objectives.BaseObjective(alpha=1.0, **kwargs)[source]

Bases: ABC

Abstract base class for projection pursuit objective functions.

Parameters:
__init__(alpha=1.0, **kwargs)[source]

Initialize the objective function.

Parameters:
  • alpha (float) – Steepness parameter for ridge functions.

  • **kwargs (Any) – Additional keyword arguments.

Return type:

None

abstractmethod __call__(a_flat, X, k, **kwargs)[source]

Compute the objective function value.

Parameters:
  • a_flat (ndarray) – Flattened projection directions.

  • X (ndarray) – Input data.

  • k (int) – Number of projections.

  • **kwargs (Any) – Additional arguments.

Returns:

Objective function value.

Return type:

float

static g(z, alpha=1.0)[source]

Apply the ridge function (non-linearity) to projected data.

Parameters:
  • z (ndarray) – Input data, shape (n_samples, n_components).

  • alpha (float) – Steepness parameter for the ridge function.

Returns:

Transformed data with the same shape as z.

Return type:

ndarray

static grad_g(z, alpha=1.0)[source]

Compute the gradient of the ridge function.

Parameters:
  • z (ndarray) – Input data, shape (n_samples, n_components).

  • alpha (float) – Steepness parameter for the ridge function.

Returns:

Gradient values with the same shape as z.

Return type:

ndarray

Distance Objective

class pyppur.objectives.DistanceObjective(alpha=0.1, weight_by_distance=False, use_nonlinearity=True, distance_metric='correlation', **kwargs)[source]

Bases: BaseObjective

Distance distortion objective function for projection pursuit.

This objective minimizes the difference between pairwise distances in the original space and the projected space. Can optionally apply ridge function nonlinearity before distance computation.

The distance_metric parameter controls how distance preservation is measured: - ‘mse’: Mean squared error between distance matrices (default, scale-sensitive) - ‘correlation’: Negative Pearson correlation (scale-invariant, recommended) - ‘spearman’: Negative Spearman rank correlation (scale and monotonic-invariant)

Parameters:
  • alpha (float)

  • weight_by_distance (bool)

  • use_nonlinearity (bool)

  • distance_metric (Literal['mse', 'correlation', 'spearman'])

  • kwargs (Any)

__init__(alpha=0.1, weight_by_distance=False, use_nonlinearity=True, distance_metric='correlation', **kwargs)[source]

Initialize the distance distortion objective.

Parameters:
  • alpha (float) – Steepness parameter for the ridge function.

  • weight_by_distance (bool) – Whether to weight distortion by inverse of original distances.

  • use_nonlinearity (bool) – Whether to apply ridge function before computing distances.

  • distance_metric (Literal['mse', 'correlation', 'spearman']) – How to measure distance preservation. Options: - ‘mse’: Mean squared error (scale-sensitive, original behavior) - ‘correlation’: Negative Pearson correlation (scale-invariant) - ‘spearman’: Negative Spearman rank correlation

  • **kwargs (Any) – Additional keyword arguments.

Return type:

None

__call__(a_flat, X, k, dist_X=None, weight_matrix=None, **kwargs)[source]

Compute the distance distortion objective.

Parameters:
  • a_flat (ndarray) – Flattened projection directions.

  • X (ndarray) – Input data.

  • k (int) – Number of projections.

  • dist_X (ndarray | None) – Pairwise distances in original space (optional).

  • weight_matrix (ndarray | None) – Optional weight matrix for distances.

  • **kwargs (Any) – Additional arguments.

Returns:

Distance distortion value (to be minimized).

Return type:

float

Reconstruction Objective

class pyppur.objectives.ReconstructionObjective(alpha=1.0, tied_weights=True, l2_reg=0.0, **kwargs)[source]

Bases: BaseObjective

Reconstruction loss objective function for projection pursuit.

This objective minimizes the reconstruction error when projecting and reconstructing data. Supports both tied-weights (encoder=decoder) and free decoder configurations.

Parameters:
__init__(alpha=1.0, tied_weights=True, l2_reg=0.0, **kwargs)[source]

Initialize the reconstruction objective.

Parameters:
  • alpha (float) – Steepness parameter for the ridge function.

  • tied_weights (bool) – If True, use tied weights (B=A). If False, learn separate decoder B.

  • l2_reg (float) – L2 regularization strength for decoder weights (when tied_weights=False).

  • **kwargs (Any) – Additional keyword arguments.

Return type:

None

__call__(a_flat, X, k, **kwargs)[source]

Compute the reconstruction objective.

Parameters:
  • a_flat (ndarray) – Flattened parameters (encoder A, and decoder B if untied).

  • X (ndarray) – Input data.

  • k (int) – Number of projections.

  • **kwargs (Any) – Additional arguments.

Returns:

Reconstruction loss value (to be minimized).

Return type:

float

reconstruct(X, a_matrix, b_matrix=None)[source]

Reconstruct data from projections.

Parameters:
  • X (ndarray) – Input data.

  • a_matrix (ndarray) – Encoder projection matrix.

  • b_matrix (ndarray | None) – Decoder matrix (if None, uses tied weights with a_matrix).

Returns:

Reconstructed data.

Return type:

ndarray

Optimizers

SciPy Optimizer

class pyppur.optimizers.ScipyOptimizer(objective_func, n_components, method='L-BFGS-B', max_iter=1000, tol=1e-06, random_state=None, verbose=False, **kwargs)[source]

Bases: BaseOptimizer

Optimizer using SciPy’s optimization methods.

This optimizer leverages SciPy’s optimization functionality, particularly the L-BFGS-B method which is well-suited for projection pursuit problems.

Parameters:
__init__(objective_func, n_components, method='L-BFGS-B', max_iter=1000, tol=1e-06, random_state=None, verbose=False, **kwargs)[source]

Initialize the SciPy optimizer.

Parameters:
  • objective_func (Callable[[...], float]) – Objective function to minimize.

  • n_components (int) – Number of projection components.

  • method (str) – SciPy optimization method (default: “L-BFGS-B”).

  • max_iter (int) – Maximum number of iterations.

  • tol (float) – Tolerance for convergence.

  • random_state (int | None) – Random seed for reproducibility.

  • verbose (bool) – Whether to log progress information.

  • **kwargs (Any) – Additional keyword arguments for the optimizer.

Return type:

None

optimize(X, initial_guess=None, **kwargs)[source]

Optimize the projection directions using SciPy’s optimization methods.

Parameters:
  • X (ndarray) – Input data, shape (n_samples, n_features).

  • initial_guess (ndarray | None) – Optional initial guess for projection directions.

  • **kwargs (Any) – Additional arguments for the objective function.

Returns:

  • Optimized projection directions, shape (n_components, n_features)

  • Final objective value

  • Additional optimizer information

Return type:

Tuple containing

Grid Optimizer

class pyppur.optimizers.GridOptimizer(objective_func, n_components, n_directions=250, n_iterations=10, max_iter=1000, tol=1e-06, random_state=None, verbose=False, **kwargs)[source]

Bases: BaseOptimizer

Optimizer using a grid-based search approach.

This optimizer is particularly useful for projection indices that are not differentiable or have many local minima. It systematically explores the space of projection directions using a grid-based approach.

Parameters:
__init__(objective_func, n_components, n_directions=250, n_iterations=10, max_iter=1000, tol=1e-06, random_state=None, verbose=False, **kwargs)[source]

Initialize the grid optimizer.

Parameters:
  • objective_func (Callable[[...], float]) – Objective function to minimize.

  • n_components (int) – Number of projection components.

  • n_directions (int) – Number of random directions to generate per iteration.

  • n_iterations (int) – Number of refinement iterations.

  • max_iter (int) – Maximum number of iterations.

  • tol (float) – Tolerance for convergence.

  • random_state (int | None) – Random seed for reproducibility.

  • verbose (bool) – Whether to print progress information.

  • **kwargs (Any) – Additional keyword arguments for the optimizer.

Return type:

None

optimize(X, initial_guess=None, **kwargs)[source]

Optimize the projection directions using a grid-based approach.

Parameters:
  • X (ndarray) – Input data, shape (n_samples, n_features).

  • initial_guess (ndarray | None) – Optional initial guess for projection directions.

  • **kwargs (Any) – Additional arguments for the objective function.

Returns:

  • Optimized projection directions, shape (n_components, n_features)

  • Final objective value

  • Additional optimizer information

Return type:

Tuple containing

Utility Functions

Metrics

Evaluation metrics for dimensionality reduction.

pyppur.utils.metrics.compute_trustworthiness(X_original, X_embedded, n_neighbors=5)[source]

Compute the trustworthiness score for dimensionality reduction.

Trustworthiness measures how well local neighborhoods are preserved.

Parameters:
  • X_original (ndarray) – Original high-dimensional data.

  • X_embedded (ndarray) – Low-dimensional embedding.

  • n_neighbors (int) – Number of neighbors to consider.

Returns:

Trustworthiness score in range [0, 1].

Return type:

float

pyppur.utils.metrics.compute_silhouette(X_embedded, labels)[source]

Compute the silhouette score for the embedding.

The silhouette score measures how well clusters are separated.

Parameters:
  • X_embedded (ndarray) – Low-dimensional embedding.

  • labels (ndarray) – Cluster or class labels.

Returns:

Silhouette score in range [-1, 1].

Return type:

float

pyppur.utils.metrics.compute_distance_distortion(X_original, X_embedded)[source]

Compute the distance distortion between original and embedded spaces.

Distance distortion measures how well pairwise distances are preserved.

Parameters:
  • X_original (ndarray) – Original high-dimensional data.

  • X_embedded (ndarray) – Low-dimensional embedding.

Returns:

Mean squared distance distortion.

Return type:

float

pyppur.utils.metrics.evaluate_embedding(X_original, X_embedded, labels=None, n_neighbors=5)[source]

Evaluate the quality of an embedding using multiple metrics.

Parameters:
  • X_original (ndarray) – Original high-dimensional data.

  • X_embedded (ndarray) – Low-dimensional embedding.

  • labels (ndarray | None) – Optional cluster or class labels.

  • n_neighbors (int) – Number of neighbors for trustworthiness.

Returns:

Dictionary with evaluation metrics.

Return type:

dict[str, float]

Preprocessing

Preprocessing utilities for projection pursuit.

pyppur.utils.preprocessing.standardize_data(X, center=True, scale=True, scaler=None)[source]

Standardize data for projection pursuit.

Parameters:
  • X (ndarray) – Input data, shape (n_samples, n_features).

  • center (bool) – Whether to center the data.

  • scale (bool) – Whether to scale the data to unit variance.

  • scaler (StandardScaler | None) – Optional pre-fitted scaler for transform-only operation.

Returns:

Standardized data and the scaler.

Return type:

tuple[ndarray, StandardScaler]

Visualization

Visualization utilities for projection pursuit results.

pyppur.utils.visualization.plot_embedding(X_embedded, labels=None, title='Projection Pursuit Embedding', metrics=None, figsize=(10, 8), cmap='tab10', alpha=0.7, s=30.0, ax=None)[source]

Plot the results of a projection pursuit embedding.

Parameters:
  • X_embedded (ndarray) – Embedded data, shape (n_samples, 2) or (n_samples, 3).

  • labels (ndarray | None) – Optional labels for coloring points.

  • title (str) – Plot title.

  • metrics (dict[str, float] | None) – Optional dictionary of metrics to include in title.

  • figsize (tuple[float, float]) – Figure size (width, height) in inches.

  • cmap (str) – Colormap name.

  • alpha (float) – Transparency of points.

  • s (float) – Point size.

  • ax (Axes | Axes3D | None) – Optional axes to plot on.

Returns:

Figure and Axes objects.

Return type:

tuple[Figure, Axes | Axes3D]

pyppur.utils.visualization.plot_reconstruction(X, X_recon, n_samples=3)[source]

Plot reconstructed samples alongside original samples.

Parameters:
  • X (ndarray) – Original data.

  • X_recon (ndarray) – Reconstructed data.

  • n_samples (int) – Number of samples to plot.

Returns:

matplotlib Figure.

Return type:

Figure

pyppur.utils.visualization.plot_comparison(embeddings, labels=None, metrics=None, title=None, figsize=(15, 5), cmap='tab10', alpha=0.7, s=30.0)[source]

Plot a comparison of multiple embeddings.

Parameters:
  • embeddings (dict[str, ndarray]) – Dictionary of embeddings {name: embedded_data}.

  • labels (ndarray | None) – Optional labels for coloring points.

  • metrics (dict[str, dict[str, float]] | None) – Optional dictionary of metrics for each embedding.

  • title (str | None) – Optional overall figure title.

  • figsize (tuple[float, float]) – Figure size (width, height) in inches.

  • cmap (str) – Colormap name.

  • alpha (float) – Transparency of points.

  • s (float) – Point size.

Returns:

matplotlib Figure object.

Return type:

Figure

Helper Functions

Normalization

pyppur.optimizers.scipy_optimizer.normalize_projection_directions(a_flat, n_components, n_features)[source]

Normalize the encoder projection directions to unit norm.

Parameters:
  • a_flat (ndarray) – Flattened parameter vector.

  • n_components (int) – Number of projection components.

  • n_features (int) – Number of features.

Returns:

Normalized parameter vector.

Return type:

ndarray