Abstract
A method for clustering and re-clustering in a storage system, the method includes (a) randomly selecting a first number of sets of clusters of a vast number of clusters stored in the storage system; wherein the vast number of clusters comprises first clusters and second clusters, wherein each first cluster is generated regardless of similarities between members of the first cluster, wherein each second cluster is generated based on at least similarities between members of the second cluster; (b) for each set of clusters, calculating by a processing circuit, a set re-clustering score that is based on (i) a sparseness measure of each clusters of the set of clusters, and (ii) one or more inter-clusters relationship measure indicative of spatial relationships between the clusters of the set of clusters; (c) identifying, by the processing circuit and based on set re-clustering scores of the first number of sets of clusters, one or more sets of clusters to be re-clustered; (d) re-clustering each one or more sets of clusters to provide updated second clusters; and (e) storing the updated second clusters in the storage system.