Statistical DataMining

Course Outline

This is an introductory class of statistical datamining. In this course, we will solve the real scientific/financial problems using datamining techniques. The methodologies covered in this course are:

l Data Visualization

l       Missing data handling

l       Multiple regression

l       CART (Classification and Regression Tree)

l       Model Comparison using CV (Cross-Validation)

l       Neural Networks

l       Bagging (RandomForest), Boosting(Gradient boosting)

l       Clusterings (K-means, PAM, Hierarchical Clustering, GMM)

l       Classifications(LDA, QDA, Logistic Regression, CART, SVM)

Note : A statistical language R is an essential tool for this class. We will use R extensively in entire class.

전산실 규모로 정원은 50명만 받습니다. 더 이상 받으면 수업, 시험 진행이 많이 어렵습니다.

중간고사 10월 16일(목) 오후 6시 - 9시.

기말고사 12월 4일(목) 오후 6시 - 9시.

Textbook

R을 이용한 데이터마이닝 (2nd edition).

Referencebook : Datamining with R, Introduction to Statistical Learning

Evaluation

Homework & Quiz  10%, Midterm 45%, Final exam 45%.