首页 / 资料库 / 文献详情

Squeezer:An Efficient Algorithm for Clustering Categorical Data

何增有徐晓飞邓胜春

2002Acta Scientiarum Naturalium Universitatis SunyatseniComputer Science被引 108

出版方页面 →

摘要

This paper presents a new efficient algorithm for clustering categorical data,Squeezer, which can produce high quality clustering results and at the same time deservegood scalability. The Squeezer algorithm reads each tuple t in sequence, either assigning tto an existing cluster (initially none), or creating t as a new cluster, which is determined bythe similarities between t and clusters. Due to its characteristics, the proposed algorithm isextremely suitable for clustering data streams, where given a sequence of points, the objective isto maintain consistently good clustering of the sequence so far, using a small amount of memoryand time. Outliers can also be handled efficiently and directly in Squeezer. Experimental resultson real-life and synthetic datasets verify the superiority of Squeezer.

引用本文(GB/T 7714)

何增有, 徐晓飞, 邓胜春. Squeezer:An Efficient Algorithm for Clustering Categorical Data[J]. Acta Scientiarum Naturalium Universitatis Sunyatseni, 2002.

引文网络

参考文献与被引分析加载中…

本站仅收录题录与摘要供学习参考,全文版权归属出版方;如有侵权请联系我们删除。