Informatic21.07.2026
A More Trustworthy Way to Visualize Complex Data
A new method «EmbedOR» uses geometry to keep clusters in data intact where popular tools quietly fail.
Fribourg, July 6, 2026: A joint team of researchers from Columbia University, Yale University, and the University of Fribourg have developed a novel algorithm that produces more faithful and trustworthy two-dimensional representations of complex, high-dimensional data. Such data often occurs in single-cell genomics and neuroscience, for example. The method, called «EmbedOR» (Embedding via Ollivier-Ricci Curvature-based Metric Learning), comes with the mathematical guarantee that it preserves structure that today’s most widely used tools tend to distort. The work appears this week in the Proceedings of the National Academy of Sciences (PNAS).
To make sense of data with hundreds or thousands of dimensions, scientists routinely compress it into a 2D picture so that patterns like clusters become directly visible to the eye. The standard tools for this, most commonly t-SNE and UMAP, appear in a plethora of studies every year. Yet, they all share a known and somewhat consequential flaw in that their pictures can mislead. For instance, they can split apart regions that are genuinely connected, and can miss real clusters even entirely. These kinds of errors are hardest to spot precisely when the data are noisy and high-dimensional.
A tried and tested method
«EmbedOR» corrects for this by borrowing an idea from geometry, namely curvature. Intuitively, curvature measures whether data is bundled together tightly or stretched thinly between groups. Equipped with this lens through which to study data reshaping, «EmbedOR» manages to pull the true clusters together while making the gaps between them stand out. Critically, the team did not just demonstrate that «EmbedOR» works in practice, they also provided a mathematical guarantee for the method.
«We have known for a long time that such visualizations can be misleading, but in many settings, we had little options to figure out when that is the case. With ‹EmbedOR›, we finally get pictures we can believe in, and a proof that we can trust them.» said Bastian Grossenbacher-Rieck, Professor of Machine Learning at the University of Fribourg and one author of the study.
In experiments on both synthetic and real-world data, «EmbedOR» was thus indeed shown to be substantially less likely than t-SNE, UMAP, or related methods to break apart regions that should belong together. The team also showed that EmbedOR’s underlying measurement can be layered onto existing visualizations, including those made with other tools, to highlight exactly where a picture is fragmented and reveal more about the true shape of the data underneath.
The study, «EmbedOR: Provable Cluster-Preserving Visualizations with Curvature-Based Stochastic Neighbor Embeddings,» was authored by Tristan Luca Saidi, Abigail Hickok, Bastian Rieck, and Andrew J. Blumberg.
