Authors
Bombina, P., Coombes, K. R.
Abstract
Motivation: Comparative studies of trajectory inference (TI) methods evaluate complete computational pipelines, making it impossible to isolate how much distortion is introduced specifically by the dimensionality reduction (DR) step. To our knowledge, no study has directly and systematically evaluated how well DR methods alone preserve a known reference path when projecting high-dimensional single-cell data to two dimensions, and no current study has introduced a dedicated set of metrics to quantify the degree of path-preservation quality after dimensionality reduction. This gap matters because DR is a universal preprocessing choice that shapes all downstream trajectory analysis, yet its independent geometric effect on path structure remains uncharacterized, and practitioners have no principled way to quantify it. Methods: We tested a panel of candidate path-preservation metrics on two single-cell datasets with known reference trajectories, one linear and one cyclic, to determine whether the resulting metric values, DR-method rankings, and overall conclusions are sensitive to the number of points used to construct and display the path, and whether they remain stable once that choice is fixed. The primary linear dataset is a CD4+ T-cell surface-protein dataset (3,096 cells, 51 proteins); a ground-truth reference path was constructed from cells lying close to the first principal component (PC1) of a single cluster, providing a known linear trajectory in the high-dimensional space. Sixteen DR methods were applied and twelve geometric path-preservation metrics were computed, spanning log-ratio distortions of length, curvature, and spatial similarity; Spearman rank correlations of pairwise distances and segment lengths; and structural complexity measures including self-intersection frequency and coiling. To test the sensitivity of this evaluation framework to path density, we varied the fraction of cells used to define the reference path from 1% to 10% (31-310 path points) and tracked how method rankings responded. The same analysis was repeated on a topologically distinct reference, a closed B-cell cell-cycle loop detected by persistent homology in a separate CyTOF dataset, to test whether these conclusions about metric and method stability hold for cyclic as well as linear trajectories. Results: The central sensitivity question, whether the number of points used to construct the reference path changes the evaluation's conclusions, was answered negatively on both datasets. On the linear PC1 trajectory, absolute values of all twelve metrics shifted smoothly as the path-density threshold was varied from 1% to 10% (31-310 points), reflecting the broadening of the reference band, but each method's composite rank remained stable across every threshold: no method changed performance tier as the hyperparameter varied. A composite rank aggregating all twelve metrics identified the same consistently high-performing methods (UMAP, MDS, CNPE, TSNE, SPE, LPMIP) and consistently low-performing methods (SPMDS, LPP, DVE, LAPEIG, PHATE) at every density level tested. Considered on its own, `SpatDistSpear`, the single most discriminating metric, separated a high-fidelity group (LPMIP, DM, MDS, SPMDS, DVE, CISOMAP, CNPE; all r > 0.80) from a mid-range group (LAPEIG, SPE, PHATE, UMAP, TSNE, FOSMOD) and a low-fidelity group (PFA, NNP, LPP); global distance preservation and overall composite performance therefore do not always agree on the same "top tier" of methods, but this disagreement in which metric identifies the best methods is itself density-independent rather than an artifact of the specific threshold chosen. The cyclic loop reproduced the same density-independence: absolute metric values drifted with the per-segment band width, yet each method's composite rank again held constant across all eleven density levels. The identity of the best and worst performers was largely, though not entirely, conserved between the two topologies, with CNPE, LPMIP, SPE, and MDS as top performers and LAPEIG, DVE, and SPMDS as poor performers on both the linear path and the closed loop. UMAP and TSNE were exceptions, dropping from top performers on the linear path to the middle of the sixteen-method panel, rather than the worst tier, on the closed loop. This topology-dependence is a property of the reference geometry rather than of path density: it holds consistently regardless of how many points are used to define the path. Significance: This work introduces a direct, pipeline-independent evaluation of how DR methods distort trajectory geometry, a benchmarking dimension absent from existing TI comparisons. The within-dataset rank stability result, demonstrated on both a linear and a cyclic reference trajectory, validates the use of a fixed reference-path threshold as a robust operating point for large-scale DR benchmarking; however, the partial reordering of top performers between topologies shows that a method's DR benchmark ranking is trajectory-shape-dependent and should not be assumed to transfer from a linear to a cyclic reference.
Preprint server:
bioRxiv
The authors list and abstract were imported from bioRxiv on 06 Sep 2026.
Advertisement
Stats
- Recommendations n/a n/a positive of 0 vote(s)
- Views 14
- Comments 0