Computing and Statistical Data Analysis / Stat 4

Slides:



Advertisements
Similar presentations
Fundamentals of Probability
Advertisements

Introductory Mathematics & Statistics for Business
STATISTICS Random Variables and Distribution Functions
Lecture 2 ANALYSIS OF VARIANCE: AN INTRODUCTION
Experimental Particle Physics PHYS6011 Joel Goldstein, RAL 1.Introduction & Accelerators 2.Particle Interactions and Detectors (2) 3.Collider Experiments.
Statistical Issues at LHCb Yuehong Xie University of Edinburgh (on behalf of the LHCb Collaboration) PHYSTAT-LHC Workshop CERN, Geneva June 27-29, 2007.
On Comparing Classifiers : Pitfalls to Avoid and Recommended Approach
Detection Chia-Hsin Cheng. Wireless Access Tech. Lab. CCU Wireless Access Tech. Lab. 2 Outlines Detection Theory Simple Binary Hypothesis Tests Bayes.
Chapter 4 Inference About Process Quality
G. Cowan Statistical Methods in Particle Physics1 Statistical Methods in Particle Physics Day 2: Multivariate Methods (I) 清华大学高能物理研究中心 2010 年 4 月 12—16.
Topics in statistical data analysis for particle physics
Nov 7, 2013 Lirong Xia Hypothesis testing and statistical decision theory.
Lecture XXIII.  In general there are two kinds of hypotheses: one concerns the form of the probability distribution (i.e. is the random variable normally.
Statistical Data Analysis / Stat 2
8. Statistical tests 8.1 Hypotheses K. Desch – Statistical methods of data analysis SS10 Frequent problem: Decision making based on statistical information.
G. Cowan 2007 CERN Summer Student Lectures on Statistics1 Introduction to Statistics − Day 3 Lecture 1 Probability Random variables, probability densities,
Statistics for HEP / LAL Orsay, 3-5 January 2012 / Lecture 1
G. Cowan Lectures on Statistical Data Analysis 1 Statistical Data Analysis: Lecture 8 1Probability, Bayes’ theorem, random variables, pdfs 2Functions of.
G. Cowan 2011 CERN Summer Student Lectures on Statistics / Lecture 41 Introduction to Statistics − Day 4 Lecture 1 Probability Random variables, probability.
G. Cowan Discovery and limits / DESY, 4-7 October 2011 / Lecture 1 1 Statistical Methods for Discovery and Limits Lecture 1: Introduction, Statistical.
G. Cowan RHUL Physics Statistical Methods for Particle Physics / 2007 CERN-FNAL HCP School page 1 Statistical Methods for Particle Physics CERN-FNAL Hadron.
G. Cowan Lectures on Statistical Data Analysis Lecture 10 page 1 Statistical Data Analysis: Lecture 10 1Probability, Bayes’ theorem 2Random variables and.
G. Cowan Statistics for HEP / NIKHEF, December 2011 / Lecture 1 1 Statistical Methods for Particle Physics Lecture 1: Introduction, Statistical Tests,
Introduction to Statistics − Day 2
G. Cowan Lectures on Statistical Data Analysis 1 Statistical Data Analysis: Lecture 7 1Probability, Bayes’ theorem, random variables, pdfs 2Functions of.
G. Cowan RHUL Physics Bayesian Higgs combination page 1 Bayesian Higgs combination using shapes ATLAS Statistics Meeting CERN, 19 December, 2007 Glen Cowan.
G. Cowan SUSSP65, St Andrews, August 2009 / Statistical Methods 3 page 1 Statistical Methods in Particle Physics Lecture 3: Multivariate Methods.
Hypothesis Testing – Introduction
G. Cowan Lectures on Statistical Data Analysis Lecture 7 page 1 Statistical Data Analysis: Lecture 7 1Probability, Bayes’ theorem 2Random variables and.
880.P20 Winter 2006 Richard Kass 1 Confidence Intervals and Upper Limits Confidence intervals (CI) are related to confidence limits (CL). To calculate.
Uniovi1 Some distributions Distribution/pdfExample use in HEP BinomialBranching ratio MultinomialHistogram with fixed N PoissonNumber of events found UniformMonte.
G. Cowan 2009 CERN Summer Student Lectures on Statistics1 Introduction to Statistics − Day 4 Lecture 1 Probability Random variables, probability densities,
G. Cowan Lectures on Statistical Data Analysis Lecture 3 page 1 Lecture 3 1 Probability (90 min.) Definition, Bayes’ theorem, probability densities and.
Irakli Chakaberia Final Examination April 28, 2014.
G. Cowan Statistical Methods in Particle Physics1 Statistical Methods in Particle Physics Day 3: Multivariate Methods (II) 清华大学高能物理研究中心 2010 年 4 月 12—16.
G. Cowan Lectures on Statistical Data Analysis Lecture 1 page 1 Lectures on Statistical Data Analysis London Postgraduate Lectures on Particle Physics;
G. Cowan CERN Academic Training 2010 / Statistics for the LHC / Lecture 11 Statistics for the LHC Lecture 1: Introduction Academic Training Lectures CERN,
Statistics In HEP Helge VossHadron Collider Physics Summer School June 8-17, 2011― Statistics in HEP 1 How do we understand/interpret our measurements.
G. Cowan CLASHEP 2011 / Topics in Statistical Data Analysis / Lecture 21 Topics in Statistical Data Analysis for HEP Lecture 2: Statistical Tests CERN.
G. Cowan CLASHEP 2011 / Topics in Statistical Data Analysis / Lecture 1 1 Topics in Statistical Data Analysis for HEP Lecture 1: Parameter Estimation CERN.
G. Cowan Lectures on Statistical Data Analysis Lecture 2 page 1 Lecture 2 1 Probability Definition, Bayes’ theorem, probability densities and their properties,
Statistical Decision Theory Bayes’ theorem: For discrete events For probability density functions.
G. Cowan 2011 CERN Summer Student Lectures on Statistics / Lecture 31 Introduction to Statistics − Day 3 Lecture 1 Probability Random variables, probability.
G. Cowan S0S 2010 / Statistical Tests and Limits 1 Statistical Tests and Limits Lecture 1: general formalism IN2P3 School of Statistics Autrans, France.
1 Introduction to Statistics − Day 4 Glen Cowan Lecture 1 Probability Random variables, probability densities, etc. Lecture 2 Brief catalogue of probability.
G. Cowan Statistical Data Analysis / Stat 2 1 Statistical Data Analysis Stat 2: Monte Carlo Method, Statistical Tests London Postgraduate Lectures on Particle.
G. Cowan Lectures on Statistical Data Analysis Lecture 8 page 1 Statistical Data Analysis: Lecture 8 1Probability, Bayes’ theorem 2Random variables and.
1 Introduction to Statistics − Day 3 Glen Cowan Lecture 1 Probability Random variables, probability densities, etc. Brief catalogue of probability densities.
G. Cowan TAE 2010 / Statistics for HEP / Lecture 11 Statistical Methods for HEP Lecture 1: Introduction Taller de Altas Energías 2010 Universitat de Barcelona.
G. Cowan Computing and Statistical Data Analysis / Stat 9 1 Computing and Statistical Data Analysis Stat 9: Parameter Estimation, Limits London Postgraduate.
G. Cowan IDPASC School of Flavour Physics, Valencia, 2-7 May 2013 / Statistical Analysis Tools 1 Statistical Analysis Tools for Particle Physics IDPASC.
1 Introduction to Statistics − Day 2 Glen Cowan Lecture 1 Probability Random variables, probability densities, etc. Brief catalogue of probability densities.
G. Cowan Lectures on Statistical Data Analysis Lecture 9 page 1 Statistical Data Analysis: Lecture 9 1Probability, Bayes’ theorem 2Random variables and.
G. Cowan Lectures on Statistical Data Analysis Lecture 6 page 1 Statistical Data Analysis: Lecture 6 1Probability, Bayes’ theorem 2Random variables and.
G. Cowan Lectures on Statistical Data Analysis Lecture 5 page 1 Statistical Data Analysis: Lecture 5 1Probability, Bayes’ theorem 2Random variables and.
G. Cowan Lectures on Statistical Data Analysis Lecture 10 page 1 Statistical Data Analysis: Lecture 10 1Probability, Bayes’ theorem 2Random variables and.
G. Cowan Aachen 2014 / Statistics for Particle Physics, Lecture 21 Statistical Methods for Particle Physics Lecture 2: statistical tests, multivariate.
G. Cowan Statistical methods for HEP / Freiburg June 2011 / Lecture 1 1 Statistical Methods for Discovery and Limits in HEP Experiments Day 1: Introduction,
Introduction to Statistics − Day 2
Comment on Event Quality Variables for Multivariate Analyses
Multi-dimensional likelihood
TAE 2017 / Statistics Lecture 1
Statistics for Particle Physics Lecture 2: Statistical Tests
Computing and Statistical Data Analysis Stat 5: Multivariate Methods
TAE 2018 Benasque, Spain 3-15 Sept 2018 Glen Cowan Physics Department
Computing and Statistical Data Analysis / Stat 6
Computing and Statistical Data Analysis / Stat 7
TRISEP 2016 / Statistics Lecture 2
SUSSP65, St Andrews, August 2009 / Statistical Methods 3
Presentation transcript:

Computing and Statistical Data Analysis / Stat 4 Computing and Statistical Data Analysis Stat 4: MC, Intro to Statistical Tests London Postgraduate Lectures on Particle Physics; University of London MSci course PH4515 Glen Cowan Physics Department Royal Holloway, University of London g.cowan@rhul.ac.uk www.pp.rhul.ac.uk/~cowan Course web page: www.pp.rhul.ac.uk/~cowan/stat_course.html G. Cowan Computing and Statistical Data Analysis / Stat 4

The transformation method Given r1, r2,..., rn uniform in [0, 1], find x1, x2,..., xn that follow f (x) by finding a suitable transformation x (r). Require: i.e. That is, set and solve for x (r). G. Cowan Computing and Statistical Data Analysis / Stat 4

Example of the transformation method Exponential pdf: Set and solve for x (r). → works too.) G. Cowan Computing and Statistical Data Analysis / Stat 4

The acceptance-rejection method Enclose the pdf in a box: (1) Generate a random number x, uniform in [xmin, xmax], i.e. r1 is uniform in [0,1]. (2) Generate a 2nd independent random number u uniformly distributed between 0 and fmax, i.e. (3) If u < f (x), then accept x. If not, reject x and repeat. G. Cowan Computing and Statistical Data Analysis / Stat 4

Example with acceptance-rejection method If dot below curve, use x value in histogram. G. Cowan Computing and Statistical Data Analysis / Stat 4

Improving efficiency of the acceptance-rejection method The fraction of accepted points is equal to the fraction of the box’s area under the curve. For very peaked distributions, this may be very low and thus the algorithm may be slow. Improve by enclosing the pdf f(x) in a curve C h(x) that conforms to f(x) more closely, where h(x) is a pdf from which we can generate random values and C is a constant. Generate points uniformly over C h(x). If point is below f(x), accept x. G. Cowan Computing and Statistical Data Analysis / Stat 4

Monte Carlo event generators Simple example: e+e- → m+m- Generate cosq and f: Less simple: ‘event generators’ for a variety of reactions: e+e- → m+m-, hadrons, ... pp → hadrons, D-Y, SUSY,... e.g. PYTHIA, HERWIG, ISAJET... Output = ‘events’, i.e., for each event we get a list of generated particles and their momentum vectors, types, etc. G. Cowan Computing and Statistical Data Analysis / Stat 4

Computing and Statistical Data Analysis / Stat 4 A simulated event PYTHIA Monte Carlo pp → gluino-gluino G. Cowan Computing and Statistical Data Analysis / Stat 4

Monte Carlo detector simulation Takes as input the particle list and momenta from generator. Simulates detector response: multiple Coulomb scattering (generate scattering angle), particle decays (generate lifetime), ionization energy loss (generate D), electromagnetic, hadronic showers, production of signals, electronics response, ... Output = simulated raw data → input to reconstruction software: track finding, fitting, etc. Predict what you should see at ‘detector level’ given a certain hypothesis for ‘generator level’. Compare with the real data. Estimate ‘efficiencies’ = #events found / # events generated. Programming package: GEANT G. Cowan Computing and Statistical Data Analysis / Stat 4

Computing and Statistical Data Analysis / Stat 4 Hypotheses A hypothesis H specifies the probability for the data, i.e., the outcome of the observation, here symbolically: x. x could be uni-/multivariate, continuous or discrete. E.g. write x ~ f (x|H). x could represent e.g. observation of a single particle, a single event, or an entire “experiment”. Possible values of x form the sample space S (or “data space”). Simple (or “point”) hypothesis: f (x|H) completely specified. Composite hypothesis: H contains unspecified parameter(s). The probability for x given H is also called the likelihood of the hypothesis, written L(x|H). G. Cowan Computing and Statistical Data Analysis / Stat 4

Computing and Statistical Data Analysis / Stat 4 Definition of a test Goal is to make some statement based on the observed data x as to the validity of the possible hypotheses. Consider e.g. a simple hypothesis H0 and alternative H1. A test of H0 is defined by specifying a critical region W of the data space such that there is no more than some (small) probability a, assuming H0 is correct, to observe the data there, i.e., P(x  W | H0 ) ≤ a If x is observed in the critical region, reject H0. a is called the size or significance level of the test. Critical region also called “rejection” region; complement is acceptance region. G. Cowan Computing and Statistical Data Analysis / Stat 4

Computing and Statistical Data Analysis / Stat 4 Definition of a test (2) But in general there are an infinite number of possible critical regions that give the same significance level a. So the choice of the critical region for a test of H0 needs to take into account the alternative hypothesis H1. Roughly speaking, place the critical region where there is a low probability to be found if H0 is true, but high if H1 is true: G. Cowan Computing and Statistical Data Analysis / Stat 4

Rejecting a hypothesis Note that rejecting H0 is not necessarily equivalent to the statement that we believe it is false and H1 true. In frequentist statistics only associate probability with outcomes of repeatable observations (the data). In Bayesian statistics, probability of the hypothesis (degree of belief) would be found using Bayes’ theorem: which depends on the prior probability p(H). What makes a frequentist test useful is that we can compute the probability to accept/reject a hypothesis assuming that it is true, or assuming some alternative is true. G. Cowan Computing and Statistical Data Analysis / Stat 4

Computing and Statistical Data Analysis / Stat 4 Type-I, Type-II errors Rejecting the hypothesis H0 when it is true is a Type-I error. The maximum probability for this is the size of the test: P(x  W | H0 ) ≤ a But we might also accept H0 when it is false, and an alternative H1 is true. This is called a Type-II error, and occurs with probability P(x  S - W | H1 ) = b One minus this is called the power of the test with respect to the alternative H1: Power = 1 - b G. Cowan Computing and Statistical Data Analysis / Stat 4

Example setting for statistical tests: the Large Hadron Collider Counter-rotating proton beams in 27 km circumference ring pp centre-of-mass energy 14 TeV Detectors at 4 pp collision points: ATLAS CMS LHCb (b physics) ALICE (heavy ion physics) general purpose G. Cowan Computing and Statistical Data Analysis / Stat 4

Computing and Statistical Data Analysis / Stat 4 The ATLAS detector 2100 physicists 37 countries 167 universities/labs 25 m diameter 46 m length 7000 tonnes ~108 electronic channels G. Cowan Computing and Statistical Data Analysis / Stat 4

Computing and Statistical Data Analysis / Stat 4 A simulated SUSY event high pT muons high pT jets of hadrons p p missing transverse energy G. Cowan Computing and Statistical Data Analysis / Stat 4

Computing and Statistical Data Analysis / Stat 4 Background events This event from Standard Model ttbar production also has high pT jets and muons, and some missing transverse energy. → can easily mimic a SUSY event. G. Cowan Computing and Statistical Data Analysis / Stat 4

Statistical tests (in a particle physics context) Suppose the result of a measurement for an individual event is a collection of numbers x1 = number of muons, x2 = mean pT of jets, x3 = missing energy, ... follows some n-dimensional joint pdf, which depends on the type of event produced, i.e., was it For each reaction we consider we will have a hypothesis for the pdf of , e.g., etc. E.g. call H0 the background hypothesis (the event type we want to reject); H1 is signal hypothesis (the type we want). G. Cowan Computing and Statistical Data Analysis / Stat 4

Computing and Statistical Data Analysis / Stat 4 Selecting events Suppose we have a data sample with two kinds of events, corresponding to hypotheses H0 and H1 and we want to select those of type H1. Each event is a point in space. What ‘decision boundary’ should we use to accept/reject events as belonging to event types H0 or H1? H0 Perhaps select events with ‘cuts’: H1 accept G. Cowan Computing and Statistical Data Analysis / Stat 4

Other ways to select events Or maybe use some other sort of decision boundary: linear or nonlinear H0 H0 H1 H1 accept accept How can we do this in an ‘optimal’ way? G. Cowan Computing and Statistical Data Analysis / Stat 4

Computing and Statistical Data Analysis / Stat 4 Test statistics The decision boundary can be defined by an equation of the form where t(x1,…, xn) is a scalar test statistic. We can work out the pdfs Decision boundary is now a single ‘cut’ on t, which divides the space into the critical (rejection) region and acceptance region. This defines a test. If the data fall in the critical region, we reject H0. G. Cowan Computing and Statistical Data Analysis / Stat 4

Signal/background efficiency Probability to reject background hypothesis for background event (background efficiency): accept b reject b g(t|b) g(t|s) Probability to accept a signal event as signal (signal efficiency): G. Cowan Computing and Statistical Data Analysis / Stat 4

Purity of event selection Suppose only one background type b; overall fractions of signal and background events are ps and pb (prior probabilities). Suppose we select signal events with t > tcut. What is the ‘purity’ of our selected sample? Here purity means the probability to be signal given that the event was accepted. Using Bayes’ theorem we find: So the purity depends on the prior probabilities as well as on the signal and background efficiencies. G. Cowan Computing and Statistical Data Analysis / Stat 4

Constructing a test statistic How can we choose a test’s critical region in an ‘optimal way’? Neyman-Pearson lemma states: To get the highest power for a given significance level in a test of H0, (background) versus H1, (signal) the critical region should have inside the region, and ≤ c outside, where c is a constant which determines the power. Equivalently, optimal scalar test statistic is N.B. any monotonic function of this is leads to the same test. G. Cowan Computing and Statistical Data Analysis / Stat 4

Why Neyman-Pearson doesn’t always help The problem is that we usually don’t have explicit formulae for the pdfs P(x|H0), P(x|H1). Instead we may have Monte Carlo models for signal and background processes, so we can produce simulated data, and enter each event into an n-dimensional histogram. Use e.g. M bins for each of the n dimensions, total of Mn cells. But n is potentially large, → prohibitively large number of cells to populate with Monte Carlo data. Compromise: make Ansatz for form of test statistic with fewer parameters; determine them (e.g. using MC) to give best discrimination between signal and background. G. Cowan Computing and Statistical Data Analysis / Stat 4

Computing and Statistical Data Analysis / Stat 4 Multivariate methods Many new (and some old) methods: Fisher discriminant Neural networks Kernel density methods Support Vector Machines Decision trees Boosting Bagging New software for HEP, e.g., TMVA , Höcker, Stelzer, Tegenfeldt, Voss, Voss, physics/0703039 StatPatternRecognition, I. Narsky, physics/0507143 G. Cowan Computing and Statistical Data Analysis / Stat 4