Tag Archives: data science

Subgraph mining datasets

In this post, I will provide two standard benchmark datasets that can be used for frequent subgraph mining. Moreover, I will provide a set of small graph datasets that I have created for debugging subgraph mining algorithms. The format of … Continue reading

Posted in Big data, Data Mining, Data science, Graph mining | Tagged , , , | Leave a comment

How to discover interesting patterns in data?

Discovering interesting patterns in data is often referred as data mining, data science or big data.  In the last few years, I have written several blog posts providing introduction to data mining and key topics in data mining: An Introduction to … Continue reading

Posted in Big data, Data Mining, Data science, Research | Tagged , , , , | Leave a comment

This is why you should visualize your data!

In the data science and data mining communities, several practitioners are applying various algorithms on data, without attempting to visualize the data.  This is a big mistake because sometimes, visualizing the data greatly helps to understand the data. Some phenomena are obvious … Continue reading

Posted in Big data, Data Mining, Data science | Tagged , , , , | 3 Comments

An Introduction to Sequential Pattern Mining

In this blog post, I will give an introduction to sequential pattern mining, an important data mining task with a wide range of applications from text analysis to market basket analysis.  This blog post is aimed to be a short … Continue reading

Posted in Big data, Data Mining, Data science | Tagged , , , , , , , | 22 Comments

An introduction to frequent subgraph mining

In this blog post, I will give an introduction to an interesting data mining task called frequent subgraph mining, which consists of discovering interesting patterns in graphs. This task is important since data is naturally represented as graph in many domains (e.g. … Continue reading

Posted in Big data, Data Mining, Data science, Graph mining | Tagged , , , , , , | 15 Comments

We are launching a new data mining journal

In this blog post, I will discuss one of my recent and current project. I have been recently working with my colleague Chun-Wei Lin on launching a new journal, titled “Data Science and Pattern Recognition“. This is a new open-access journal, … Continue reading

Posted in Big data, Data Mining, Data science, Research | Tagged , , , | 2 Comments

Introduction to clustering: the K-Means algorithm (with Java code)

In this blog post, I will introduce the popular data mining task of clustering (also called cluster analysis).  I will explain what is the goal of clustering, and then introduce the popular K-Means algorithm with an example. Moreover, I will briefly explain how an open-source Java implementation of … Continue reading

Posted in Big data, Data Mining, Data science, Open-source | Tagged , , , , , , | 13 Comments

Introduction to time series mining with SPMF

This blog post briefly explain how time series data mining can be performed with the Java open-source data mining library SPMF (v.2.06).  It first explain what is a time series and then discuss how data mining can be performed on time series. What is … Continue reading

Posted in Big data, Data Mining, Open-source, Time series | Tagged , , , , , , , , | 2 Comments

Report of the PAKDD 2014 conference (part 3)

This post continue my report of the PAKDD 2014 in Tainan (Taiwan). The panel about big data Friday, there was a great panel about big data with 7 top researchers from the field of data mining.  I will try to faithfully report some … Continue reading

Posted in Academia, Big data, Conference, Data Mining, Data science | Tagged , , , , | 1 Comment

Report of the PAKDD 2014 conference (part 2)

This post will continue my report of the PAKDD 2014 in Tainan (Taiwan). About big data Another interesting talk at this conference was given by Jian Pei. The topic was Big Data. Some key ideas in this talk was that to make … Continue reading

Posted in Academia, Big data, Conference, Data Mining, Data science | Tagged , , , , | 2 Comments