WEKA and Machine Learning Algorithms. Algorithm Types Classification (supervised) Given -> A set of classified examples “instances” Produce -> A way of.

Slides:

Advertisements

Similar presentations

Florida International University COP 4770 Introduction of Weka.

Advertisements

Weka & Rapid Miner Tutorial By Chibuike Muoh. WEKA:: Introduction A collection of open source ML algorithms – pre-processing – classifiers – clustering.

Department of Computer Science, University of Waikato, New Zealand Eibe Frank WEKA: A Machine Learning Toolkit The Explorer Classification and Regression.

UNIVERSITY OF JYVÄSKYLÄ DEPARTMENT OF MATHEMATICAL INFORMATION TECHNOLOGY Tutorial 1: Introduction to WEKA and YALETIES443: Introduction to DM 1 Tutorial.

WEKA Evaluation of WEKA Waikato Environment for Knowledge Analysis Presented By: Manoj Wartikar & Sameer Sagade.

Department of Computer Science, University of Waikato, New Zealand Eibe Frank WEKA: A Machine Learning Toolkit The Explorer Classification and Regression.

March 25, 2004Columbia University1 Machine Learning with Weka Lokesh S. Shrestha.

An Extended Introduction to WEKA. Data Mining Process.

1 Statistical Learning Introduction to Weka Michel Galley Artificial Intelligence class November 2, 2006.

ML ALGORITHMS. Algorithm Types Classification (supervised) Given -> A set of classified examples “instances” Produce -> A way of classifying new examples.

Machine Learning with WEKA. WEKA: the bird Copyright: Martin Kramer

1 SIMS 290-2: Applied Natural Language Processing Marti Hearst October 30, 2006 Some slides by Preslav Nakov and Eibe Frank.

1 How to use Weka How to use Weka. 2 WEKA: the software Waikato Environment for Knowledge Analysis Collection of state-of-the-art machine learning algorithms.

CSCI 347 / CS 4206: Data Mining Module 05: WEKA Topic 04: Data Preparation Tools.

CSCI 347 / CS 4206: Data Mining Module 06: Evaluation Topic 07: Cost-Sensitive Measures.

1 © Goharian & Grossman 2003 Introduction to Data Mining (CS 422) Fall 2010.

An Exercise in Machine Learning

CSCI 347 / CS 4206: Data Mining Module 05: WEKA Topic 01: WEKA Navigation.

 The Weka The Weka is an well known bird of New Zealand..  W(aikato) E(nvironment) for K(nowlegde) A(nalysis)  Developed by the University of Waikato.

Contributed by Yizhou Sun 2008 An Introduction to WEKA.

Department of Computer Science, University of Waikato, New Zealand Geoff Holmes WEKA project and team Data Mining process Data format Preprocessing Classification.

Evaluation – next steps

WEKA - Explorer (sumber: WEKA Explorer user Guide for Version 3-5-5)

Appendix: The WEKA Data Mining Software

In part from: Yizhou Sun 2008 An Introduction to WEKA Explorer.

Department of Computer Science, University of Waikato, New Zealand Bernhard Pfahringer (based on material by Eibe Frank, Mark Hall, and Peter Reutemann)

Data Mining: Classification & Predication Hosam Al-Samarraie, PhD. Centre for Instructional Technology & Multimedia Universiti Sains Malaysia.

Department of Computer Science, University of Waikato, New Zealand Eibe Frank WEKA: A Machine Learning Toolkit The Explorer Classification and Regression.

Machine Learning with Weka Cornelia Caragea Thanks to Eibe Frank for some of the slides.

Weka: Experimenter and Knowledge Flow interfaces Neil Mac Parthaláin

Department of Computer Science, University of Waikato, New Zealand Eibe Frank WEKA: A Machine Learning Toolkit The Explorer Classification and Regression.

Weka – A Machine Learning Toolkit October 2, 2008 Keum-Sung Hwang.

Data Mining Practical Machine Learning Tools and Techniques By I. H. Witten, E. Frank and M. A. Hall Chapter 5: Credibility: Evaluating What’s Been Learned.

Introduction to Weka Xingquan (Hill) Zhu Slides copied from Jeffrey Junfeng Pan (UST)

 A collection of open source ML algorithms ◦ pre-processing ◦ classifiers ◦ clustering ◦ association rule  Created by researchers at the University.

Machine Learning Tutorial-2. Recall, Precision, F-measure, Accuracy Ch. 5.

W E K A Waikato Environment for Knowledge Aquisition.

Data Mining Practical Machine Learning Tools and Techniques By I. H. Witten, E. Frank and M. A. Hall DM Finals Study Guide Rodney Nielsen.

An Exercise in Machine Learning

***Classification Model*** Hosam Al-Samarraie, PhD. CITM-USM.

Weka Tutorial. WEKA:: Introduction A collection of open source ML algorithms – pre-processing – classifiers – clustering – association rule Created by.

Weka. Weka A Java-based machine vlearning tool Implements numerous classifiers and other ML algorithms Uses a common.

Machine Learning with WEKA - Yohan Chin. WEKA ? Waikato Environment for Knowledge Analysis A Collection of Machine Learning algorithms for data tasks.

WEKA's Knowledge Flow Interface Data Mining Knowledge Discovery in Databases ELIE TCHEIMEGNI Department of Computer Science Bowie State University, MD.

In part from: Yizhou Sun 2008 An Introduction to WEKA Explorer.

@relation age sex { female, chest_pain_type { typ_angina, asympt, non_anginal,

WEKA: A Practical Machine Learning Tool WEKA ： A Practical Machine Learning Tool.

Department of Computer Science, University of Waikato, New Zealand Eibe Frank WEKA: A Machine Learning Toolkit The Explorer Classification and Regression.

Department of Computer Science, University of Waikato, New Zealand Geoff Holmes WEKA project and team Data Mining process Data format Preprocessing Classification.

An Introduction to WEKA

Machine Learning with WEKA

Waikato Environment for Knowledge Analysis

Sampath Jayarathna Cal Poly Pomona

An Introduction to WEKA

Machine Learning with WEKA

Machine Learning with WEKA

Weka Package Weka package is open source data mining software written in Java. Weka can be applied to your dataset from the GUI, the command line or called.

Machine Learning with Weka

An Introduction to WEKA

CSCI N317 Computation for Scientific Applications Unit Weka

Machine Learning with Weka

Machine Learning with WEKA

Lecture 10 – Introduction to Weka

Statistical Learning Introduction to Weka

Copyright: Martin Kramer

Machine Learning: Decision Trees in AIMA and WEKA

Data Mining CSCI 307, Spring 2019 Lecture 7

Data Mining CSCI 307, Spring 2019 Lecture 8

Presentation transcript:

WEKA and Machine Learning Algorithms

Algorithm Types Classification (supervised) Given -> A set of classified examples “instances” Produce -> A way of classifying new examples Instances: described by fixed set of features “attributes” Classes: discrete or continuous “classification” “regression” Interested in: – Results? (classifying new instances) – Model? (how the decision is made) Clustering (unsupervised) There are no classes Association rules Look for rules that relate features to other features

Classification

Clustering

It is expected that similarity among members of a cluster should be high and similarity among objects of different clusters should be low. The objectives of clustering – knowing which data objects belong to which cluster – understanding common characteristics of the members of a specific cluster

Clustering vs Classification There is some similarity between clustering and classification. Both classification and clustering are about assigning appropriate class or cluster labels to data records. However, clustering differs from classification in two aspects. – First, in clustering, there are no pre-defined classes. This means that the number of classes or clusters and the class or cluster label of each data record are not known before the operation. – Second, clustering is about grouping data rather than developing a classification model. Therefore, there is no distinction between data records and examples. The entire data population is used as input to the clustering process.

Association Mining

Overfitting Memorization vs generalization To fix, use – Training data — to form rules – Validation data — to decide on best rule – Test data — to determine system performance Cross-validation

Baseline Experiments In order to evaluate the efficiency of the classifiers used in experiments, we use baselines: – Majority based random classification (Kappa=0) – Class distribution based random classification (Kappa=0) Kappa statistics, is used as a measure to assess the improvement of a classifier’s accuracy over a predictor employing chance as its guide. P 0 is the accuracy of the classifier and P c is the expected accuracy that can be achieved by a randomly guessing classifier on the same data set. Kappa statistics has a range between 1 and 1, where 1 is total disagreement (i.e., total misclassification) and 1 is perfect agreement (i.e., a 100% accurate classification). Kappa score over 0.4 indicates a reasonable agreement beyond chance. 9

Data Mining Process

WEKA: the software Machine learning/data mining software written in Java (distributed under the GNU Public License) Used for research, education, and applications Complements “Data Mining” by Witten & Frank Main features: – Comprehensive set of data pre-processing tools, learning algorithms and evaluation methods – Graphical user interfaces (incl. data visualization) – Environment for comparing learning algorithms

Weka’s Role in the Big Picture Input Raw data Input Raw data Data Mining by Weka Pre-processing Classification Regression Clustering Association Rules Visualization Data Mining by Weka Pre-processing Classification Regression Clustering Association Rules Visualization Output Result Output Result

13 WEKA: Terminology Some synonyms/explanations for the terms used by WEKA: Attribute: feature Relation: collection of examples Instance: collection in use Class: category

@relation age sex { female, chest_pain_type { typ_angina, asympt, non_anginal, cholesterol exercise_induced_angina { no, class { present, 63,male,typ_angina,233,no,not_present 67,male,asympt,286,yes,present 67,male,asympt,229,yes,present 38,female,non_anginal,?,no,not_present... WEKA only deals with “flat” files

Explorer: pre-processing the data Data can be imported from a file in various formats: ARFF, CSV, C4.5, binary Data can also be read from a URL or from an SQL database (using JDBC) Pre-processing tools in WEKA are called “filters” WEKA contains filters for: – Discretization, normalization, resampling, attribute selection, transforming and combining attributes, …

10/2/ Explorer: building “classifiers” Classifiers in WEKA are models for predicting nominal or numeric quantities Implemented learning schemes include: – Decision trees and lists, instance-based classifiers, support vector machines, multi-layer perceptrons, logistic regression, Bayes’ nets, … “Meta”-classifiers include: – Bagging, boosting, stacking, error-correcting output codes, locally weighted learning, …

Classifiers - Workflow Labeled Data Learning Algorithm Classifier Unlabeled Data Predictions

Evaluation Accuracy – Percentage of Predictions that are correct – Problematic for some disproportional Data Sets Precision – Percent of positive predictions correct Recall (Sensitivity) – Percent of positive labeled samples predicted as positive Specificity – The percentage of negative labeled samples predicted as negative.

20 Contains information about the actual and the predicted classification All measures can be derived from it: accuracy: (a+d)/(a+b+c+d) recall: d/(c+d) => R precision: d/(b+d) => P F-measure: 2PR/(P+R) false positive (FP) rate: b /(a+b) true negative (TN) rate: a /(a+b) false negative (FN) rate: c /(c+d) Confusion matrix predicted –+ true –ab +cd

10/2/ Explorer: clustering data WEKA contains “clusterers” for finding groups of similar instances in a dataset Implemented schemes are: – k-Means, EM, Cobweb, X-means, FarthestFirst Clusters can be visualized and compared to “true” clusters (if given) Evaluation based on loglikelihood if clustering scheme produces a probability distribution

10/2/ Explorer: finding associations WEKA contains an implementation of the Apriori algorithm for learning association rules – Works only with discrete data Can identify statistical dependencies between groups of attributes: – milk, butter  bread, eggs (with confidence 0.9 and support 2000) Apriori can compute all rules that have a given minimum support and exceed a given confidence

10/2/ Explorer: attribute selection Panel that can be used to investigate which (subsets of) attributes are the most predictive ones Attribute selection methods contain two parts: – A search method: best-first, forward selection, random, exhaustive, genetic algorithm, ranking – An evaluation method: correlation-based, wrapper, information gain, chi-squared, … Very flexible: WEKA allows (almost) arbitrary combinations of these two

10/2/ Explorer: data visualization Visualization very useful in practice: e.g. helps to determine difficulty of the learning problem WEKA can visualize single attributes (1-d) and pairs of attributes (2-d) – To do: rotating 3-d visualizations (Xgobi-style) Color-coded class values “Jitter” option to deal with nominal attributes (and to detect “hidden” data points) “Zoom-in” function

10/2/ Performing experiments Experimenter makes it easy to compare the performance of different learning schemes For classification and regression problems Results can be written into file or database Evaluation options: cross-validation, learning curve, hold-out Can also iterate over different parameter settings Significance-testing built in!

10/2/ The Knowledge Flow GUI New graphical user interface for WEKA Java-Beans-based interface for setting up and running machine learning experiments Data sources, classifiers, etc. are beans and can be connected graphically Data “flows” through components: e.g., “data source” -> “filter” -> “classifier” -> “evaluator” Layouts can be saved and loaded again later

27 Beyond the GUI How to reproduce experiments with the command-line/API – GUI, API, and command-line all rely on the same set of Java classes – Generally easy to determine what classes and parameters were used in the GUI. – Tree displays in Weka reflect its Java class hierarchy. > java -cp ~galley/weka/weka.jar weka.classifiers.trees.J48 –C 0.25 –M 2 -t -T

28 Important command-line parameters where options are: Create/load/save a classification model: -t : training set -l : load model file -d : save model file Testing: -x : N-fold cross validation -T : test set -p : print predictions + attribute selection S > java -cp ~galley/weka/weka.jar weka.classifiers. [classifier_options] [options]

Problem with Running Weka Solution : java -Xmx1000m -jar weka.jar Problem : Out of memory for large data set