API Reference¶
This page provides detailed documentation for all pyppur classes and functions.
Main Classes¶
ProjectionPursuit¶
- class pyppur.ProjectionPursuit(n_components=2, objective='distance_distortion', alpha=0.1, max_iter=500, tol=1e-06, random_state=None, optimizer='L-BFGS-B', n_init=3, verbose=False, center=True, scale=True, weight_by_distance=False, tied_weights=True, l2_reg=0.0, use_nonlinearity_in_distance=True, distance_metric='correlation')[source]¶
Bases:
TransformerMixin,BaseEstimatorImplementation of Projection Pursuit for dimensionality reduction.
This class provides methods to find optimal projections by minimizing either reconstruction loss or distance distortion. It supports both initialization strategies and different optimizers.
- Parameters:
n_components (int)
alpha (float)
max_iter (int)
tol (float)
random_state (int | None)
optimizer (str)
n_init (int)
verbose (bool)
center (bool)
scale (bool)
weight_by_distance (bool)
tied_weights (bool)
l2_reg (float)
use_nonlinearity_in_distance (bool)
distance_metric (Literal['mse', 'correlation', 'spearman'])
- __init__(n_components=2, objective='distance_distortion', alpha=0.1, max_iter=500, tol=1e-06, random_state=None, optimizer='L-BFGS-B', n_init=3, verbose=False, center=True, scale=True, weight_by_distance=False, tied_weights=True, l2_reg=0.0, use_nonlinearity_in_distance=True, distance_metric='correlation')[source]¶
Initialize a ProjectionPursuit model.
- Parameters:
n_components (int) – Number of projection dimensions to use.
objective (Objective | str) – Optimization objective enum value.
alpha (float) – Steepness parameter for the ridge function g(z) = tanh(alpha * z). Lower values (0.01-0.1) work better for distance preservation.
max_iter (int) – Maximum number of iterations for optimization.
tol (float) – Tolerance for optimization convergence.
random_state (int | None) – Random seed for reproducibility.
optimizer (str) – Optimization method (‘L-BFGS-B’ recommended).
n_init (int) – Number of random initializations to try.
verbose (bool) – Whether to log progress information.
center (bool) – Whether to center the data.
scale (bool) – Whether to scale the data.
weight_by_distance (bool) – Whether to weight distance distortion by inverse of original distances.
tied_weights (bool) – Whether to use tied weights (encoder=decoder) for reconstruction.
l2_reg (float) – L2 regularization strength for decoder weights (when tied_weights=False).
use_nonlinearity_in_distance (bool) – Whether to apply ridge function before computing distances.
distance_metric (Literal['mse', 'correlation', 'spearman']) – How to measure distance preservation. Options: ‘mse’ (scale-sensitive), ‘correlation’ (scale-invariant), ‘spearman’ (rank-based). Default is ‘correlation’.
- Return type:
None
- fit(X, y=None)[source]¶
Fit the ProjectionPursuit model to the data.
- Parameters:
X (ndarray) – Input data, shape (n_samples, n_features).
y (Any) – Ignored. Accepted for scikit-learn estimator API compatibility.
- Returns:
The fitted model.
- Return type:
- transform(X)[source]¶
Apply dimensionality reduction to X.
- Parameters:
X (ndarray) – Input data, shape (n_samples, n_features).
- Returns:
Transformed data, shape (n_samples, n_components).
- Return type:
ndarray
- fit_transform(X, y=None, **fit_params)[source]¶
Fit the model with X and apply dimensionality reduction on X.
- reconstruct(X)[source]¶
Reconstruct X from the projected data.
- Parameters:
X (ndarray) – Input data, shape (n_samples, n_features).
- Returns:
Reconstructed data, shape (n_samples, n_features).
- Return type:
ndarray
- reconstruction_error(X)[source]¶
Compute the reconstruction error for X.
- Parameters:
X (ndarray) – Input data, shape (n_samples, n_features).
- Returns:
Mean squared reconstruction error.
- Return type:
- distance_distortion(X)[source]¶
Compute the distance distortion for X.
- Parameters:
X (ndarray) – Input data, shape (n_samples, n_features).
- Returns:
Mean squared distance distortion.
- Return type:
- compute_trustworthiness(X, n_neighbors=5)[source]¶
Compute the trustworthiness score for the dimensionality reduction.
Trustworthiness measures how well the local structure is preserved. A score of 1.0 indicates perfect trustworthiness, while a score of 0.0 indicates that the local structure is not preserved at all.
- compute_silhouette(X, labels)[source]¶
Compute the silhouette score for the dimensionality reduction.
Silhouette score measures how well clusters are separated. A score close to 1.0 indicates that clusters are well separated, while a score close to -1.0 indicates poor separation.
- Parameters:
X (ndarray) – Input data, shape (n_samples, n_features).
labels (ndarray) – Cluster labels for each sample.
- Returns:
Silhouette score between -1.0 and 1.0.
- Return type:
- evaluate(X, labels=None, n_neighbors=5)[source]¶
Evaluate the dimensionality reduction with multiple metrics.
- property n_components_: int¶
Number of components actually used by the fit.
This equals the
n_componentsconstructor argument, except when that exceeded the number of features, in which case it is the number of features. The constructor argument itself is never modified.- Returns:
Number of projection directions found.
- property x_loadings_: ndarray¶
Projection directions (encoder).
- Returns:
Projection directions, shape (n_components, n_features).
- property decoder_weights_: ndarray | None¶
Decoder weights (for untied weights only).
- Returns:
Decoder weights, shape (n_components, n_features), or None if using tied weights.
Objective Types¶
Objective Functions¶
Base Objective¶
- class pyppur.objectives.BaseObjective(alpha=1.0, **kwargs)[source]¶
Bases:
ABCAbstract base class for projection pursuit objective functions.
Distance Objective¶
- class pyppur.objectives.DistanceObjective(alpha=0.1, weight_by_distance=False, use_nonlinearity=True, distance_metric='correlation', **kwargs)[source]¶
Bases:
BaseObjectiveDistance distortion objective function for projection pursuit.
This objective minimizes the difference between pairwise distances in the original space and the projected space. Can optionally apply ridge function nonlinearity before distance computation.
The distance_metric parameter controls how distance preservation is measured: - ‘mse’: Mean squared error between distance matrices (default, scale-sensitive) - ‘correlation’: Negative Pearson correlation (scale-invariant, recommended) - ‘spearman’: Negative Spearman rank correlation (scale and monotonic-invariant)
- Parameters:
- __init__(alpha=0.1, weight_by_distance=False, use_nonlinearity=True, distance_metric='correlation', **kwargs)[source]¶
Initialize the distance distortion objective.
- Parameters:
alpha (float) – Steepness parameter for the ridge function.
weight_by_distance (bool) – Whether to weight distortion by inverse of original distances.
use_nonlinearity (bool) – Whether to apply ridge function before computing distances.
distance_metric (Literal['mse', 'correlation', 'spearman']) – How to measure distance preservation. Options: - ‘mse’: Mean squared error (scale-sensitive, original behavior) - ‘correlation’: Negative Pearson correlation (scale-invariant) - ‘spearman’: Negative Spearman rank correlation
**kwargs (Any) – Additional keyword arguments.
- Return type:
None
Reconstruction Objective¶
- class pyppur.objectives.ReconstructionObjective(alpha=1.0, tied_weights=True, l2_reg=0.0, **kwargs)[source]¶
Bases:
BaseObjectiveReconstruction loss objective function for projection pursuit.
This objective minimizes the reconstruction error when projecting and reconstructing data. Supports both tied-weights (encoder=decoder) and free decoder configurations.
- __init__(alpha=1.0, tied_weights=True, l2_reg=0.0, **kwargs)[source]¶
Initialize the reconstruction objective.
- Parameters:
- Return type:
None
- reconstruct(X, a_matrix, b_matrix=None)[source]¶
Reconstruct data from projections.
- Parameters:
X (ndarray) – Input data.
a_matrix (ndarray) – Encoder projection matrix.
b_matrix (ndarray | None) – Decoder matrix (if None, uses tied weights with a_matrix).
- Returns:
Reconstructed data.
- Return type:
ndarray
Optimizers¶
SciPy Optimizer¶
- class pyppur.optimizers.ScipyOptimizer(objective_func, n_components, method='L-BFGS-B', max_iter=1000, tol=1e-06, random_state=None, verbose=False, **kwargs)[source]¶
Bases:
BaseOptimizerOptimizer using SciPy’s optimization methods.
This optimizer leverages SciPy’s optimization functionality, particularly the L-BFGS-B method which is well-suited for projection pursuit problems.
- Parameters:
- __init__(objective_func, n_components, method='L-BFGS-B', max_iter=1000, tol=1e-06, random_state=None, verbose=False, **kwargs)[source]¶
Initialize the SciPy optimizer.
- Parameters:
objective_func (Callable[[...], float]) – Objective function to minimize.
n_components (int) – Number of projection components.
method (str) – SciPy optimization method (default: “L-BFGS-B”).
max_iter (int) – Maximum number of iterations.
tol (float) – Tolerance for convergence.
random_state (int | None) – Random seed for reproducibility.
verbose (bool) – Whether to log progress information.
**kwargs (Any) – Additional keyword arguments for the optimizer.
- Return type:
None
- optimize(X, initial_guess=None, **kwargs)[source]¶
Optimize the projection directions using SciPy’s optimization methods.
- Parameters:
X (ndarray) – Input data, shape (n_samples, n_features).
initial_guess (ndarray | None) – Optional initial guess for projection directions.
**kwargs (Any) – Additional arguments for the objective function.
- Returns:
Optimized projection directions, shape (n_components, n_features)
Final objective value
Additional optimizer information
- Return type:
Tuple containing
Grid Optimizer¶
- class pyppur.optimizers.GridOptimizer(objective_func, n_components, n_directions=250, n_iterations=10, max_iter=1000, tol=1e-06, random_state=None, verbose=False, **kwargs)[source]¶
Bases:
BaseOptimizerOptimizer using a grid-based search approach.
This optimizer is particularly useful for projection indices that are not differentiable or have many local minima. It systematically explores the space of projection directions using a grid-based approach.
- Parameters:
- __init__(objective_func, n_components, n_directions=250, n_iterations=10, max_iter=1000, tol=1e-06, random_state=None, verbose=False, **kwargs)[source]¶
Initialize the grid optimizer.
- Parameters:
objective_func (Callable[[...], float]) – Objective function to minimize.
n_components (int) – Number of projection components.
n_directions (int) – Number of random directions to generate per iteration.
n_iterations (int) – Number of refinement iterations.
max_iter (int) – Maximum number of iterations.
tol (float) – Tolerance for convergence.
random_state (int | None) – Random seed for reproducibility.
verbose (bool) – Whether to print progress information.
**kwargs (Any) – Additional keyword arguments for the optimizer.
- Return type:
None
- optimize(X, initial_guess=None, **kwargs)[source]¶
Optimize the projection directions using a grid-based approach.
- Parameters:
X (ndarray) – Input data, shape (n_samples, n_features).
initial_guess (ndarray | None) – Optional initial guess for projection directions.
**kwargs (Any) – Additional arguments for the objective function.
- Returns:
Optimized projection directions, shape (n_components, n_features)
Final objective value
Additional optimizer information
- Return type:
Tuple containing
Utility Functions¶
Metrics¶
Evaluation metrics for dimensionality reduction.
- pyppur.utils.metrics.compute_trustworthiness(X_original, X_embedded, n_neighbors=5)[source]¶
Compute the trustworthiness score for dimensionality reduction.
Trustworthiness measures how well local neighborhoods are preserved.
- pyppur.utils.metrics.compute_silhouette(X_embedded, labels)[source]¶
Compute the silhouette score for the embedding.
The silhouette score measures how well clusters are separated.
- Parameters:
X_embedded (ndarray) – Low-dimensional embedding.
labels (ndarray) – Cluster or class labels.
- Returns:
Silhouette score in range [-1, 1].
- Return type:
- pyppur.utils.metrics.compute_distance_distortion(X_original, X_embedded)[source]¶
Compute the distance distortion between original and embedded spaces.
Distance distortion measures how well pairwise distances are preserved.
- Parameters:
X_original (ndarray) – Original high-dimensional data.
X_embedded (ndarray) – Low-dimensional embedding.
- Returns:
Mean squared distance distortion.
- Return type:
- pyppur.utils.metrics.evaluate_embedding(X_original, X_embedded, labels=None, n_neighbors=5)[source]¶
Evaluate the quality of an embedding using multiple metrics.
- Parameters:
X_original (ndarray) – Original high-dimensional data.
X_embedded (ndarray) – Low-dimensional embedding.
labels (ndarray | None) – Optional cluster or class labels.
n_neighbors (int) – Number of neighbors for trustworthiness.
- Returns:
Dictionary with evaluation metrics.
- Return type:
Preprocessing¶
Preprocessing utilities for projection pursuit.
Visualization¶
Visualization utilities for projection pursuit results.
- pyppur.utils.visualization.plot_embedding(X_embedded, labels=None, title='Projection Pursuit Embedding', metrics=None, figsize=(10, 8), cmap='tab10', alpha=0.7, s=30.0, ax=None)[source]¶
Plot the results of a projection pursuit embedding.
- Parameters:
X_embedded (ndarray) – Embedded data, shape (n_samples, 2) or (n_samples, 3).
labels (ndarray | None) – Optional labels for coloring points.
title (str) – Plot title.
metrics (dict[str, float] | None) – Optional dictionary of metrics to include in title.
figsize (tuple[float, float]) – Figure size (width, height) in inches.
cmap (str) – Colormap name.
alpha (float) – Transparency of points.
s (float) – Point size.
ax (Axes | Axes3D | None) – Optional axes to plot on.
- Returns:
Figure and Axes objects.
- Return type:
tuple[Figure, Axes | Axes3D]
- pyppur.utils.visualization.plot_reconstruction(X, X_recon, n_samples=3)[source]¶
Plot reconstructed samples alongside original samples.
- Parameters:
X (ndarray) – Original data.
X_recon (ndarray) – Reconstructed data.
n_samples (int) – Number of samples to plot.
- Returns:
matplotlib Figure.
- Return type:
Figure
- pyppur.utils.visualization.plot_comparison(embeddings, labels=None, metrics=None, title=None, figsize=(15, 5), cmap='tab10', alpha=0.7, s=30.0)[source]¶
Plot a comparison of multiple embeddings.
- Parameters:
embeddings (dict[str, ndarray]) – Dictionary of embeddings {name: embedded_data}.
labels (ndarray | None) – Optional labels for coloring points.
metrics (dict[str, dict[str, float]] | None) – Optional dictionary of metrics for each embedding.
title (str | None) – Optional overall figure title.
figsize (tuple[float, float]) – Figure size (width, height) in inches.
cmap (str) – Colormap name.
alpha (float) – Transparency of points.
s (float) – Point size.
- Returns:
matplotlib Figure object.
- Return type:
Figure