维普中文期刊产品整合服务

CLS-Miner: efficient and effective closed high-utility itemset mining

查看全文 作  者:Thu-Lan [1,2]DAM;Kenli [1,3,4]LI;Philippe FOURNIER-[5]VIGER;Quang-Huy [6]DUONG 高影响力作者 机构地区:[1]College of Computer Science and Electronic Engineering, Hunan University, Changsha 410082, China;[2]Faculty of Information Technology, Hanoi University of Industry, Hanoi, Vietnam;[3]CIC of HPC, National University of Defense Technology, Changsha 410073, China;[4]National Supercomputing Center in Changsha, Changsha 410082, China;[5]School of Natural Sciences and Humanities, Harbin Institute of Technology Shenzhen Graduate School, Shenzhen 518055, China;[6]Department of Computer Science, Norwegian University of Science and Technology, Trondheim, Norway高影响力机构 出  处:《Frontiers of Computer Science》索引2019年第13卷第2期,共25页高影响力期刊 基  金:the National Natural Science Foundation of China (Grant Nos. 61133005, 61432005, 61370095, 61472124, 61202109, and 61472126);the International Science and Technology Cooperation Program of China (2015DFA11240 and 2014DFBS0010). 摘  要:High-utility itemset mining (HUIM) is a popular data mining task with applications in numerous domains. However, traditional HUIM algorithms often produce a very large set of high-utility itemsets (HUIs). As a result, analyzing HUIs can be very time consuming for users. Moreover, a large set of HUIs also makes HUIM algorithms less efficient in terms of execution time and memory consumption. To address this problem, closed high-utility itemsets (CHUIs), concise and lossless representations of all HUIs, were proposed recently. Although mining CHUIs is useful and desirable, it remains a computationally expensive task. This is because current algorithms often generate a huge number of candidate itemsets and are unable to prune the search space effectively. In this paper, we address these issues by proposing a novel algorithm called CLS-Miner. The proposed algorithm utilizes the utility-list structure to directly compute the utilities of itemsets without producing candidates. It also introduces three novel strategies to reduce the search space, namely chain-estimated utility co-occurrence pruning, lower branch pruning, and pruning by coverage. Moreover, an effective method for checking whether an itemset is a subset of another itemset is introduced to further reduce the time required for discovering CHUIs. To evaluate the performance of the proposed algorithm and its novel strategies, extensive experiments have been conducted on six benchmark datasets having various characteristics. Results show that the proposed strategies are highly efficient and effective, that the proposed CLS-Miner algorithm outperforms the current state-ofthe- art CHUD and CHUI-Miner algorithms, and that CLSMiner scales linearly. 关 键 词:UTILITY MINING high-utility ITEMSET MINING CLOSED ITEMSET MINING CLOSED high-utility ITEMSET MINING
相关文献

参考文献(36)

引证文献(10)

网站首页 | 关于我们 | 联系我们 | 产品服务 | 客服中心 | 广告服务 | 版权声明 | 网站联盟 | 友情链接 | 售卡网点

版权所有© 渝B2-20050021-1 渝公网安备 50019002500403号 违法和不良信息举报中心

互联网出版许可证 新出网证(渝)字10号 全国400电话 - 免长途话费