CS 258 Parallel Computer Architecture Lecture 1 Introduction to Parallel Architecture January 23, 2002 Prof John D. Kubiatowicz.

Slides:



Advertisements
Similar presentations
Multiprocessors— Large vs. Small Scale Multiprocessors— Large vs. Small Scale.
Advertisements

CA 714CA Midterm Review. C5 Cache Optimization Reduce miss penalty –Hardware and software Reduce miss rate –Hardware and software Reduce hit time –Hardware.
POLITECNICO DI MILANO Parallelism in wonderland: are you ready to see how deep the rabbit hole goes? ILP: VLIW Architectures Marco D. Santambrogio:
COE 502 / CSE 661 Parallel and Vector Architectures Prof. Muhamed Mudawar Computer Engineering Department King Fahd University of Petroleum and Minerals.
Chapter1 Fundamental of Computer Design Dr. Bernard Chen Ph.D. University of Central Arkansas.
CpE442 Intro. To Computer Architecture CpE 442 Introduction To Computer Architecture Lecture 1 Instructor: H. H. Ammar These slides are based on the lecture.
1 Lecture 5: Part 1 Performance Laws: Speedup and Scalability.
Ch1. Fundamentals of Computer Design 3. Principles (5) ECE562/468 Advanced Computer Architecture Prof. Honggang Wang ECE Department University of Massachusetts.
ECE669 L15: Mid-term Review March 25, 2004 ECE 669 Parallel Computer Architecture Lecture 15 Mid-term Review.
ECE669 L1: Course Introduction January 29, 2004 ECE 669 Parallel Computer Architecture Lecture 1 Course Introduction Prof. Russell Tessier Department of.
1 COMP 206: Computer Architecture and Implementation Montek Singh Mon, Dec 5, 2005 Topic: Intro to Multiprocessors and Thread-Level Parallelism.
Introduction What is Parallel Algorithms? Why Parallel Algorithms? Evolution and Convergence of Parallel Algorithms Fundamental Design Issues.
Chapter 17 Parallel Processing.
Csci4203/ece43631 Review Quiz. 1)It is less expensive 2)It is usually faster 3)Its average CPI is smaller 4)It allows a faster clock rate 5)It has a simpler.
CIS 629 Parallel Arch. Intro Parallel Computer Architecture Slides blended from those of David Patterson, CS 252 and David Culler, CS 258 UC Berkeley.
EET 4250: Chapter 1 Performance Measurement, Instruction Count & CPI Acknowledgements: Some slides and lecture notes for this course adapted from Prof.
Why Parallel Architecture? Todd C. Mowry CS 495 January 15, 2002.
Parallel Processing Architectures Laxmi Narayan Bhuyan
Implications for Programming Models Todd C. Mowry CS 495 September 12, 2002.
ECE 232 L1 Intro.1 Adapted from Patterson 97 ©UCBCopyright 1998 Morgan Kaufmann Publishers ECE 232 Hardware Organization and Design Lecture 1 Introduction.
CMSC 611: Advanced Computer Architecture Parallel Computation Most slides adapted from David Patterson. Some from Mohomed Younis.
0 Introduction to parallel processing Chapter 1 from Culler & Singh Winter 2007.
CPU Performance Assessment As-Bahiya Abu-Samra *Moore’s Law *Clock Speed *Instruction Execution Rate - MIPS - MFLOPS *SPEC Speed Metric *Amdahl’s.
Computer performance.
Parallel Computing Laxmikant Kale
Introduction CSE 410, Spring 2008 Computer Systems
Multi-core architectures. Single-core computer Single-core CPU chip.
Multi-Core Architectures
ECE 568: Modern Comp. Architectures and Intro to Parallel Processing Fall 2006 Ahmed Louri ECE Department.
Sogang University Advanced Computing System Chap 1. Computer Architecture Hyuk-Jun Lee, PhD Dept. of Computer Science and Engineering Sogang University.
CS 320 Spring 2003 Introduction Laxmikant Kale
SJSU SPRING 2011 PARALLEL COMPUTING Parallel Computing CS 147: Computer Architecture Instructor: Professor Sin-Min Lee Spring 2011 By: Alice Cotti.
Frank Casilio Computer Engineering May 15, 1997 Multithreaded Processors.
Problem is to compute: f(latitude, longitude, elevation, time)  temperature, pressure, humidity, wind velocity Approach: –Discretize the.
Advanced Computer Architecture Fundamental of Computer Design Instruction Set Principles and Examples Pipelining:Basic and Intermediate Concepts Memory.
Super computers Parallel Processing By Lecturer: Aisha Dawood.
1 Today About the class Introductions Any new people? Start of first module: Parallel Computing.
ECE 569: High-Performance Computing: Architectures, Algorithms and Technologies Spring 2006 Ahmed Louri ECE Department.
Spring 2003CSE P5481 Issues in Multiprocessors Which programming model for interprocessor communication shared memory regular loads & stores message passing.
CS252/Patterson Lec 1.1 1/17/01 CMPUT429/CMPE382 Winter 2001 Topic2: Technology Trend and Cost/Performance (Adapted from David A. Patterson’s CS252 lecture.
Ted Pedersen – CS 3011 – Chapter 10 1 A brief history of computer architectures CISC – complex instruction set computing –Intel x86, VAX –Evolved from.
CS- 492 : Distributed system & Parallel Processing Lecture 7: Sun: 15/5/1435 Foundations of designing parallel algorithms and shared memory models Lecturer/
INEL6067 Technology ---> Limitations & Opportunities Wires -Area -Propagation speed Clock Power VLSI -I/O pin limitations -Chip area -Chip crossing delay.
Dept. of Computer Science - CS6461 Computer Architecture CS6461 – Computer Architecture Fall 2015 Lecture 1 – Introduction Adopted from Professor Stephen.
Performance Performance
Outline Why this subject? What is High Performance Computing?
CS433 Spring 2001 Introduction Laxmikant Kale. 2 Course objectives and outline You will learn about: –Parallel programming models Emphasis on 3: message.
1 Lecture 1: Parallel Architecture Intro Course organization:  ~18 parallel architecture lectures (based on text)  ~10 (recent) paper presentations 
3/12/2013Computer Engg, IIT(BHU)1 PARALLEL COMPUTERS- 2.
3/12/2013Computer Engg, IIT(BHU)1 INTRODUCTION-1.
Jan. 5, 2000Systems Architecture II1 Machine Organization (CS 570) Lecture 1: Overview of High Performance Processors * Jeremy R. Johnson Wed. Sept. 27,
Compsci Today’s topics l Operating Systems  Brookshear, Chapter 3  Great Ideas, Chapter 10  Slides from Kevin Wayne’s COS 126 course l Performance.
CDA-5155 Computer Architecture Principles Fall 2000 Multiprocessor Architectures.
Background Computer System Architectures Computer System Software.
High Performance Computing1 High Performance Computing (CS 680) Lecture 2a: Overview of High Performance Processors * Jeremy R. Johnson *This lecture was.
Lecture 1: Introduction CprE 585 Advanced Computer Architecture, Fall 2004 Zhao Zhang.
VU-Advanced Computer Architecture Lecture 1-Introduction 1 Advanced Computer Architecture CS 704 Advanced Computer Architecture Lecture 1.
Lecture 13 Parallel Processing. 2 What is Parallel Computing? Traditionally software has been written for serial computation. Parallel computing is the.
Feeding Parallel Machines – Any Silver Bullets? Novica Nosović ETF Sarajevo 8th Workshop “Software Engineering Education and Reverse Engineering” Durres,
CMSC 611: Advanced Computer Architecture
Parallel Computers Definition: “A parallel computer is a collection of processing elements that cooperate and communicate to solve large problems fast.”
CS775: Computer Architecture
Lecture 1: Parallel Architecture Intro
Performance of computer systems
Course Description: Parallel Computer Architecture
Parallel Processing Architectures
CS 258 Parallel Computer Architecture
Chapter 4 Multiprocessors
Performance of computer systems
The University of Adelaide, School of Computer Science
Presentation transcript:

CS 258 Parallel Computer Architecture Lecture 1 Introduction to Parallel Architecture January 23, 2002 Prof John D. Kubiatowicz

1/23/02 John Kubiatowicz Slide 1.2 Computer Architecture Is … the attributes of a [computing] system as seen by the programmer, i.e., the conceptual structure and functional behavior, as distinct from the organization of the data flows and controls the logic design, and the physical implementation. Amdahl, Blaaw, and Brooks, 1964 SOFTWARE

1/23/02 John Kubiatowicz Slide 1.3 Instruction Set Architecture (ISA) Great picture for uniprocessors…. –Rapidly crumbling, however! Can this be true for multiprocessors??? –Much harder to say. instruction set software hardware

1/23/02 John Kubiatowicz Slide 1.4 What is Parallel Architecture? A parallel computer is a collection of processing elements that cooperate to solve large problems fast Some broad issues: –Models of computation: PRAM? BSP? Sequential Consistency? –Resource Allocation: »how large a collection? »how powerful are the elements? »how much memory? –Data access, Communication and Synchronization »how do the elements cooperate and communicate? »how are data transmitted between processors? »what are the abstractions and primitives for cooperation? –Performance and Scalability »how does it all translate into performance? »how does it scale?

1/23/02 John Kubiatowicz Slide 1.5 Computer Architecture Topics (252+) Instruction Set Architecture Pipelining, Hazard Resolution, Superscalar, Reordering, Prediction, Speculation, Vector, Dynamic Compilation Addressing, Protection, Exception Handling L1 Cache L2 Cache DRAM Disks, WORM, Tape Coherence, Bandwidth, Latency Emerging Technologies Interleaving Bus protocols RAID VLSI Input/Output and Storage Memory Hierarchy Pipelining and Instruction Level Parallelism Network Communication Other Processors

1/23/02 John Kubiatowicz Slide 1.6 Computer Architecture Topics (258) M Interconnection Network S PMPMPMP ° ° ° Topologies, Routing, Bandwidth, Latency, Reliability Network Interfaces Shared Memory, Message Passing, Data Parallelism Processor-Memory-Switch Multiprocessors Networks and Interconnections Everything in previous slide but more so!

1/23/02 John Kubiatowicz Slide 1.7 What will you get out of CS258? In-depth understanding of the design and engineering of modern parallel computers –technology forces –Programming models –fundamental architectural issues »naming, replication, communication, synchronization –basic design techniques »cache coherence, protocols, networks, pipelining, … –methods of evaluation from moderate to very large scale across the hardware/software boundary Study of REAL parallel processors –Research papers, white papers Natural consequences?? –Massive Parallelism  Reconfigurable computing? –Message Passing Machines  NOW  Peer-to-peer systems?

1/23/02 John Kubiatowicz Slide 1.8 Will it be worthwhile? Absolutely! –even through few of you will become PP designers The fundamental issues and solutions translate across a wide spectrum of systems. –Crisp solutions in the context of parallel machines. Pioneered at the thin-end of the platform pyramid on the most-demanding applications –migrate downward with time Understand implications for software Network attached storage, MEMs, etc? SuperServers Departmenatal Servers Workstations Personal Computers Workstations

1/23/02 John Kubiatowicz Slide 1.9 Role of a computer architect: To design and engineer the various levels of a computer system to maximize performance and programmability within limits of technology and cost. Parallelism: Provides alternative to faster clock for performance Applies at all levels of system design Is a fascinating perspective from which to view architecture Is increasingly central in information processing How is instruction-level parallelism related to course-grained parallelism?? Why Study Parallel Architecture?

1/23/02 John Kubiatowicz Slide 1.10 Is Parallel Computing Inevitable? This was certainly not clear just a few years ago Today, however: Application demands: Our insatiable need for computing cycles Technology Trends: Easier to build Architecture Trends: Better abstractions Economics: Cost of pushing uniprocessor Current trends: –Today’s microprocessors have multiprocessor support –Servers and workstations becoming MP: Sun, SGI, DEC, COMPAQ!... –Tomorrow’s microprocessors are multiprocessors

1/23/02 John Kubiatowicz Slide 1.11 Can programmers handle parallelism? Humans not as good at parallel programming as they would like to think! –Need good model to think of machine –Architects pushed on instruction-level parallelism really hard, because it is “transparent” Can compiler extract parallelism? –Sometimes How do programmers manage parallelism?? –Language to express parallelism? –How to schedule varying number of processors? Is communication Explicit (message-passing) or Implicit (shared memory)? –Are there any ordering constraints on communication?

1/23/02 John Kubiatowicz Slide 1.12 Granularity: Is communication fine or coarse grained? –Small messages vs big messages Is parallelism fine or coarse grained –Small tasks (frequent synchronization) vs big tasks If hardware handles fine-grained parallelism, then easier to get incremental scalability Fine-grained communication and parallelism harder than coarse-grained: –Harder to build with low overhead –Custom communication architectures often needed Ultimate course grained communication: –GIMPS (Great Internet Mercenne Prime Search) –Communication once a month

1/23/02 John Kubiatowicz Slide 1.13 CS258: Staff Instructor:Prof John D. Kubiatowicz Office: 673 Soda Hall, Office Hours: Thursday 1:30 - 3:00 or by appt. Class: Wed, Fri, 1:00 - 2:30pm 310 Soda Hall Administrative: Veronique Richard, Office: 676 Soda Hall, Web page: Lectures available online <11:30AM day of lecture Clip signup link on web page (as soon as it is up)

1/23/02 John Kubiatowicz Slide 1.14 Why Me? The Alewife Multiprocessor Cache-coherence Shared Memory –Partially in Software! User-level Message-Passing Rapid Context-Switching Asynchronous network One node/board

1/23/02 John Kubiatowicz Slide 1.15 TextBook: Two leaders in field Text:Parallel Computer Architecture: A Hardware/Software Approach, By: David Culler & Jaswinder Singh Covers a range of topics We will not necessarily cover them in order.

1/23/02 John Kubiatowicz Slide 1.16 Lecture style 1-Minute Review 20-Minute Lecture/Discussion 5- Minute Administrative Matters 25-Minute Lecture/Discussion 5-Minute Break (water, stretch) 25-Minute Lecture/Discussion Instructor will come to class early & stay after to answer questions Attention Time 20 min.Break“In Conclusion,...”

1/23/02 John Kubiatowicz Slide 1.17 Research Paper Reading As graduate students, you are now researchers. Most information of importance to you will be in research papers. Ability to rapidly scan and understand research papers is key to your success. So: you will read lots of papers in this course! –Quick 1 paragraph summaries will be due in class –Students will take turns discussing papers Papers will be scanned and on web page.

1/23/02 John Kubiatowicz Slide 1.18 How will grading work? No TA This term! Tentative breakdown: –20% homeworks / paper presentations –30% exam –40% project (teams of 2) –10% participation

1/23/02 John Kubiatowicz Slide 1.19 Application Trends Application demand for performance fuels advances in hardware, which enables new appl’ns, which... –Cycle drives exponential increase in microprocessor performance –Drives parallel architecture harder » most demanding applications Programmers willing to work really hard to improve high-end applications Need incremental scalability: –Need range of system performance with progressively increasing cost New Applications More Performance

1/23/02 John Kubiatowicz Slide 1.20 Speedup Speedup (p processors) = Common mistake: –Compare parallel program on 1 processor to parallel program on p processors Wrong!: –Should compare uniprocessor program on 1 processor to parallel program on p processors Why? Keeps you honest –It is easy to parallelize overhead. Time (1 processor) Time (p processors)

1/23/02 John Kubiatowicz Slide 1.21 Amdahl's Law Speedup due to enhancement E: ExTime w/o EPerformance w/ E Speedup(E) = = ExTime w/ E Performance w/o E Suppose that enhancement E accelerates a fraction F of the task by a factor S, and the remainder of the task is unaffected

1/23/02 John Kubiatowicz Slide 1.22 Amdahl’s Law for parallel programs? Best you could ever hope to do: Worse: Overhead may kill your performance!

1/23/02 John Kubiatowicz Slide 1.23 Metrics of Performance Compiler Programming Language Application Datapath Control TransistorsWiresPins ISA Function Units (millions) of Instructions per second: MIPS (millions) of (FP) operations per second: MFLOP/s Cycles per second (clock rate) Megabytes per second Answers per month Operations per second

1/23/02 John Kubiatowicz Slide 1.24 Commercial Computing Relies on parallelism for high end –Computational power determines scale of business that can be handled Databases, online-transaction processing, decision support, data mining, data warehousing... TPC benchmarks (TPC-C order entry, TPC-D decision support) –Explicit scaling criteria provided –Size of enterprise scales with size of system –Problem size not fixed as p increases. –Throughput is performance measure (transactions per minute or tpm)

1/23/02 John Kubiatowicz Slide 1.25 TPC-C Results for March 1996 Parallelism is pervasive Small to moderate scale parallelism very important Difficult to obtain snapshot to compare across vendor platforms

1/23/02 John Kubiatowicz Slide 1.26 Scientific Computing Demand

1/23/02 John Kubiatowicz Slide 1.27 Engineering Computing Demand Large parallel machines a mainstay in many industries –Petroleum (reservoir analysis) –Automotive (crash simulation, drag analysis, combustion efficiency), –Aeronautics (airflow analysis, engine efficiency, structural mechanics, electromagnetism), –Computer-aided design –Pharmaceuticals (molecular modeling) –Visualization »in all of the above »entertainment (films like Toy Story) »architecture (walk-throughs and rendering) –Financial modeling (yield and derivative analysis) –etc.

1/23/02 John Kubiatowicz Slide 1.28 Also CAD, Databases, processors gets you 10 years, 1000 gets you 20 ! Applications: Speech and Image Processing

1/23/02 John Kubiatowicz Slide 1.29 Is better parallel arch enough? AMBER molecular dynamics simulation program Starting point was vector code for Cray MFLOP on Cray90, 406 for final version on 128- processor Paragon, 891 on 128-processor Cray T3D

1/23/02 John Kubiatowicz Slide 1.30 Summary of Application Trends Transition to parallel computing has occurred for scientific and engineering computing In rapid progress in commercial computing –Database and transactions as well as financial –Usually smaller-scale, but large-scale systems also used Desktop also uses multithreaded programs, which are a lot like parallel programs Demand for improving throughput on sequential workloads –Greatest use of small-scale multiprocessors Solid application demand exists and will increase

1/23/02 John Kubiatowicz Slide 1.31 Technology Trends Today the natural building-block is also fastest!

1/23/02 John Kubiatowicz Slide 1.32 Microprocessor performance increases 50% - 100% per year Transistor count doubles every 3 years DRAM size quadruples every 3 years Huge investment per generation is carried by huge commodity market IntegerFP Sun MIPS M/120 IBM RS MIPS M2000 HP DEC alpha Can’t we just wait for it to get faster?

1/23/02 John Kubiatowicz Slide 1.33 Proc$ Interconnect Technology: A Closer Look Basic advance is decreasing feature size (  ) –Circuits become either faster or lower in power Die size is growing too –Clock rate improves roughly proportional to improvement in –Number of transistors improves like   (or faster) Performance > 100x per decade –clock rate < 10x, rest is transistor count How to use more transistors? –Parallelism in processing »multiple operations per cycle reduces CPI –Locality in data access »avoids latency and reduces CPI »also improves processor utilization –Both need resources, so tradeoff Fundamental issue is resource distribution, as in uniprocessors

1/23/02 John Kubiatowicz Slide % per year Growth Rates 40% per year

1/23/02 John Kubiatowicz Slide 1.35 Architectural Trends Architecture translates technology’s gifts into performance and capability Resolves the tradeoff between parallelism and locality –Current microprocessor: 1/3 compute, 1/3 cache, 1/3 off- chip connect –Tradeoffs may change with scale and technology advances Understanding microprocessor architectural trends => Helps build intuition about design issues or parallel machines => Shows fundamental role of parallelism even in “sequential” computers

1/23/02 John Kubiatowicz Slide 1.36 Phases in “VLSI” Generation

1/23/02 John Kubiatowicz Slide 1.37 Architectural Trends Greatest trend in VLSI generation is increase in parallelism –Up to 1985: bit level parallelism: 4-bit -> 8 bit -> 16-bit »slows after 32 bit »adoption of 64-bit now under way, 128-bit far (not performance issue) »great inflection point when 32-bit micro and cache fit on a chip –Mid 80s to mid 90s: instruction level parallelism »pipelining and simple instruction sets, + compiler advances (RISC) »on-chip caches and functional units => superscalar execution »greater sophistication: out of order execution, speculation, prediction to deal with control transfer and latency problems –Next step: thread level parallelism? Bit-level parallelism?

1/23/02 John Kubiatowicz Slide 1.38 How far will ILP go? Infinite resources and fetch bandwidth, perfect branch prediction and renaming –real caches and non-zero miss latencies

1/23/02 John Kubiatowicz Slide 1.39 No. of processors in fully configured commercial shared-memory systems Threads Level Parallelism “on board” Micro on a chip makes it natural to connect many to shared memory –dominates server and enterprise market, moving down to desktop Alternative: many PCs sharing one complicated pipe Faster processors began to saturate bus, then bus technology advanced –today, range of sizes for bus-based systems, desktop to large servers Proc MEM

1/23/02 John Kubiatowicz Slide 1.40 What about Multiprocessor Trends?

1/23/02 John Kubiatowicz Slide 1.41 Bus Bandwidth

1/23/02 John Kubiatowicz Slide 1.42 What about Storage Trends? Divergence between memory capacity and speed even more pronounced –Capacity increased by 1000x from , speed only 2x –Gigabit DRAM by c. 2000, but gap with processor speed much greater Larger memories are slower, while processors get faster –Need to transfer more data in parallel –Need deeper cache hierarchies –How to organize caches? Parallelism increases effective size of each level of hierarchy, without increasing access time Parallelism and locality within memory systems too –New designs fetch many bits within memory chip; follow with fast pipelined transfer across narrower interface –Buffer caches most recently accessed data –Processor in memory? Disks too: Parallel disks plus caching

1/23/02 John Kubiatowicz Slide 1.43 Economics Commodity microprocessors not only fast but CHEAP –Development costs tens of millions of dollars –BUT, many more are sold compared to supercomputers –Crucial to take advantage of the investment, and use the commodity building block Multiprocessors being pushed by software vendors (e.g. database) as well as hardware vendors Standardization makes small, bus-based SMPs commodity Desktop: few smaller processors versus one larger one? Multiprocessor on a chip?

1/23/02 John Kubiatowicz Slide 1.44 Can anyone afford high-end MPPs??? ASCI (Accellerated Strategic Computing Initiative) ASCI White: Built by IBM –12.3 TeraOps, 8192 processors (RS/6000) –6TB of RAM, 160TB Disk –2 basketball courts in size –Program it??? Message passing

1/23/02 John Kubiatowicz Slide 1.45 Consider Scientific Supercomputing Proving ground and driver for innovative architecture and techniques –Market smaller relative to commercial as MPs become mainstream –Dominated by vector machines starting in 70s –Microprocessors have made huge gains in floating-point performance »high clock rates »pipelined floating point units (e.g., multiply-add every cycle) »instruction-level parallelism »effective use of caches (e.g., automatic blocking) –Plus economics Large-scale multiprocessors replace vector supercomputers

1/23/02 John Kubiatowicz Slide 1.46 Raw Uniprocessor Performance: LINPACK

1/23/02 John Kubiatowicz Slide 1.47 Raw Parallel Performance: LINPACK Even vector Crays became parallel –X-MP (2-4) Y-MP (8), C-90 (16), T94 (32) Since 1993, Cray produces MPPs too (T3D, T3E)

1/23/02 John Kubiatowicz Slide 1.48 Number of systems u u u u n n n n s s s s 11/9311/9411/9511/ n PVP u MPP s SMP Fastest Computers

1/23/02 John Kubiatowicz Slide 1.49 Summary: Why Parallel Architecture? Increasingly attractive –Economics, technology, architecture, application demand Increasingly central and mainstream Parallelism exploited at many levels –Instruction-level parallelism –Multiprocessor servers –Large-scale multiprocessors (“MPPs”) Focus of this class: multiprocessor level of parallelism Same story from memory system perspective –Increase bandwidth, reduce average latency with many local memories Spectrum of parallel architectures make sense –Different cost, performance and scalability

1/23/02 John Kubiatowicz Slide 1.50 Where is Parallel Arch Going? Application Software System Software SIMD Message Passing Shared Memory Dataflow Systolic Arrays Architecture Uncertainty of direction paralyzed parallel software development! Old view: Divergent architectures, no predictable pattern of growth.

1/23/02 John Kubiatowicz Slide 1.51 Today Extension of “computer architecture” to support communication and cooperation –Instruction Set Architecture plus Communication Architecture Defines –Critical abstractions, boundaries, and primitives (interfaces) –Organizational structures that implement interfaces (hw or sw) Compilers, libraries and OS are important bridges today