TY - GEN
T1 - MB3-Miner: efficiently mining eMBedded subTREEs using Tree Model Guided candidate generation
AU - Tan, H.
AU - Dillon, T.
AU - Hadzic, F.
AU - Chang, E.
AU - Feng, L.
N1 - Imported from EWI/DB PMS [db-utwente:inpr:0000003700], http://eric.univ-lyon2.fr/~mcd05/
PY - 2005/11
Y1 - 2005/11
N2 - Tree mining has many useful applications in areas such as Bioinformatics, XML mining, Web mining, etc. In general, most of the formally represented information in these domains is a tree structured form. In this paper we focus on mining frequent embedded subtrees from databases of rooted labeled ordered subtrees. We propose a novel and unique embedding list representation that is suitable for describing embedded subtrees. This representation is completely different from the string-like or conventional adjacency list representation previously utilized for trees. We present the mathematical model of a breadth-first-search Tree Model Guided (TMG) candidate generation approach previously introduced in [8]. The key characteristic of the TMG approach is that it enumerates fewer candidates by ensuring that only valid candidates that conform to the structural aspects of the data are generated as opposed to the join approach. Our experiments with both synthetic and real-life datasets provide comparisons against one of the state-of-the-art algorithms, TreeMiner [15], and they demonstrate the effectiveness and the efficiency of the technique.
AB - Tree mining has many useful applications in areas such as Bioinformatics, XML mining, Web mining, etc. In general, most of the formally represented information in these domains is a tree structured form. In this paper we focus on mining frequent embedded subtrees from databases of rooted labeled ordered subtrees. We propose a novel and unique embedding list representation that is suitable for describing embedded subtrees. This representation is completely different from the string-like or conventional adjacency list representation previously utilized for trees. We present the mathematical model of a breadth-first-search Tree Model Guided (TMG) candidate generation approach previously introduced in [8]. The key characteristic of the TMG approach is that it enumerates fewer candidates by ensuring that only valid candidates that conform to the structural aspects of the data are generated as opposed to the join approach. Our experiments with both synthetic and real-life datasets provide comparisons against one of the state-of-the-art algorithms, TreeMiner [15], and they demonstrate the effectiveness and the efficiency of the technique.
KW - DB-DM: DATA MINING
KW - EWI-7338
KW - IR-63537
KW - METIS-229585
M3 - Conference contribution
SN - 0-9738918-8-2
SP - 103
EP - 110
BT - Proceedings of the 1st International Workshop on Mining Complex Data 2005 (MCD 2005)
PB - IEEE
CY - Halivax, Nova Scotia, Canada
T2 - 1st International Workshop on Mining Complex Data 2005 (MCD 2005)
Y2 - 27 November 2005 through 27 November 2005
ER -