Searching. Please wait…
1582
37
170
29108
4409
2600
347
389
Abstract: A robust clustering method for probabilities in Wasserstein space is introduced. This new ?trimmed k-barycenters? approach relies on recent results on barycenters in Wasserstein space that allow intensive computation, as required by clustering algorithms to be feasible. The possibility of trimming the most discrepant distributions results in a gain in stability and robustness, highly convenient in this setting. As a remarkable application, we consider a parallelized clustering setup in which each of m units processes a portion of the data, producing a clustering report, encoded as k probabilities. We prove that the trimmed k-barycenter of the m×k reports produces a consistent aggregation which we consider the result of a ?wide consensus?. We also prove that a weighted version of trimmed k-means algorithms based on k-barycenters in the space of Wasserstein keeps the descending character of the concentration step, guaranteeing convergence to local minima. We illustrate the methodology with simulated and real data examples. These include clustering populations by age distributions and analysis of cytometric data.
Fuente: Statistics and Computing (2019) 29:139-160
Publisher: Springer
Year of publication: 2019
No. of pages: 22
Publication type: Article
DOI: 10.1007/s11222-018-9800-z
ISSN: 0960-3174,1573-1375
Spanish project: MTM2014-56235-C2-1-P ; MTM2014-56235-C2-2
Publication Url: https://doi.org/10.1007/s11222-018-9800-z
Read publication
BARRIO, E. DEL
JUAN ANTONIO CUESTA ALBERTOS
MAYO-ISCAR, A.
Back