CS 300 – Lecture 24 Intro to Computer Architecture / Assembly Language The LAST Lecture!

Slides:

Advertisements

Similar presentations

CMSC 611: Advanced Computer Architecture Pipelining Some material adapted from Mohamed Younis, UMBC CMSC 611 Spr 2003 course slides Some material adapted.

Advertisements

Lecture Objectives: 1)Define pipelining 2)Calculate the speedup achieved by pipelining for a given number of instructions. 3)Define how pipelining improves.

Pipelining Hwanmo Sung CS147 Presentation Professor Sin-Min Lee.

CS-447– Computer Architecture Lecture 12 Multiple Cycle Datapath

Mary Jane Irwin ( ) [Adapted from Computer Organization and Design,

Chapter XI Reduced Instruction Set Computing (RISC) CS 147 Li-Chuan Fang.

ENEE350 Ankur Srivastava University of Maryland, College Park Based on Slides from Mary Jane Irwin ( )

1 Recap (Pipelining). 2 What is Pipelining? A way of speeding up execution of tasks Key idea : overlap execution of multiple taks.

Datorsystem 1 och Datorarkitektur 1 – föreläsning 10 måndag 19 November 2007.

Prof. John Nestor ECE Department Lafayette College Easton, Pennsylvania Computer Organization Pipelined Processor Design 1.

1  2004 Morgan Kaufmann Publishers Chapter Six. 2  2004 Morgan Kaufmann Publishers Pipelining The laundry analogy.

L18 – Pipeline Issues 1 Comp 411 – Spring /03/08 CPU Pipelining Issues Finishing up Chapter 6 This pipe stuff makes my head hurt! What have you.

Computer ArchitectureFall 2007 © October 24nd, 2007 Majd F. Sakr CS-447– Computer Architecture.

L17 – Pipeline Issues 1 Comp 411 – Fall /1308 CPU Pipelining Issues Finishing up Chapter 6 This pipe stuff makes my head hurt! What have you been.

331 Lec18.1Fall :332:331 Computer Architecture and Assembly Language Fall 2003 Lecture 18 Introduction to Pipelined Datapath [Adapted from Dave.

CS 300 – Lecture 23 Intro to Computer Architecture / Assembly Language Virtual Memory Pipelining.

ENEE350 Ankur Srivastava University of Maryland, College Park Based on Slides from Mary Jane Irwin ( )

CS 152 L10 Pipeline Intro (1)Fall 2004 © UC Regents CS152 – Computer Architecture and Engineering Fall 2004 Lecture 10: Basic MIPS Pipelining Review John.

Computer ArchitectureFall 2007 © October 22nd, 2007 Majd F. Sakr CS-447– Computer Architecture.

1 CSE SUNY New Paltz Chapter Six Enhancing Performance with Pipelining.

Pipelining - II Adapted from CS 152C (UC Berkeley) lectures notes of Spring 2002.

Computer ArchitectureFall 2008 © October 6th, 2008 Majd F. Sakr CS-447– Computer Architecture.

Prof. John Nestor ECE Department Lafayette College Easton, Pennsylvania ECE Computer Organization Lecture 17 - Pipelined.

Spring W :332:331 Computer Architecture and Assembly Language Spring 2005 Week 11 Introduction to Pipelined Datapath [Adapted from Dave Patterson’s.

CSE431 L05 Basic MIPS Architecture.1Irwin, PSU, 2005 CSE 431 Computer Architecture Fall 2005 Lecture 05: Basic MIPS Architecture Review Mary Jane Irwin.

Lecture 15: Pipelining and Hazards CS 2011 Fall 2014, Dr. Rozier.

CS1104: Computer Organisation School of Computing National University of Singapore.

1 Pipelining Reconsider the data path we just did Each instruction takes from 3 to 5 clock cycles However, there are parts of hardware that are idle many.

Pipelining (I). Pipelining Example  Laundry Example  Four students have one load of clothes each to wash, dry, fold, and put away  Washer takes 30.

Analogy: Gotta Do Laundry

CSE431 Chapter 4A.1Irwin, PSU, 2008 CSE 431 Computer Architecture Fall 2008 Chapter 4A: The Processor, Part A Mary Jane Irwin ( )

CSE 340 Computer Architecture Summer 2014 Basic MIPS Pipelining Review.

Basic Pipelining & MIPS Pipelining Chapter 6 [Computer Organization and Design, © 2007 Patterson (UCB) & Hennessy (Stanford), & Slides Adapted from: Mary.

CS.305 Computer Architecture Enhancing Performance with Pipelining Adapted from Computer Organization and Design, Patterson & Hennessy, © 2005, and from.

Computer Organization CS224 Chapter 4 Part b The Processor Spring 2010 With thanks to M.J. Irwin, T. Fountain, D. Patterson, and J. Hennessy for some lecture.

CMPE 421 Parallel Computer Architecture

1 Designing a Pipelined Processor In this Chapter, we will study 1. Pipelined datapath 2. Pipelined control 3. Data Hazards 4. Forwarding 5. Branch Hazards.

Computer Architecture and Design – ELEN 350 Part 8 [Some slides adapted from M. Irwin, D. Paterson. D. Garcia and others]

Electrical and Computer Engineering University of Cyprus LAB3: IMPROVING MIPS PERFORMANCE WITH PIPELINING.

ECE 232 L18.Pipeline.1 Adapted from Patterson 97 ©UCBCopyright 1998 Morgan Kaufmann Publishers ECE 232 Hardware Organization and Design Lecture 18 Pipelining.

Chapter 6 Pipelined CPU Design. Spring 2005 ELEC 5200/6200 From Patterson/Hennessey Slides Pipelined operation – laundry analogy Text Fig. 6.1.

CECS 440 Pipelining.1(c) 2014 – R. W. Allison [slides adapted from D. Patterson slides with additional credits to M.J. Irwin]

Winter 2002CSE Topic Branch Hazards in the Pipelined Processor.

Cs 152 L1 3.1 DAP Fa97,  U.CB Pipelining Lessons °Pipelining doesn’t help latency of single task, it helps throughput of entire workload °Multiple tasks.

Chap 6.1 Computer Architecture Chapter 6 Enhancing Performance with Pipelining.

CSIE30300 Computer Architecture Unit 04: Basic MIPS Pipelining Hsin-Chou Chi [Adapted from material by and

1  1998 Morgan Kaufmann Publishers Chapter Six. 2  1998 Morgan Kaufmann Publishers Pipelining Improve perfomance by increasing instruction throughput.

Oct. 18, 2000Machine Organization1 Machine Organization (CS 570) Lecture 4: Pipelining * Jeremy R. Johnson Wed. Oct. 18, 2000 *This lecture was derived.

CMSC 611: Advanced Computer Architecture Pipelining Some material adapted from Mohamed Younis, UMBC CMSC 611 Spr 2003 course slides Some material adapted.

Pipelining CS365 Lecture 9. D. Barbara Pipeline CS465 2 Outline  Today’s topic  Pipelining is an implementation technique in which multiple instructions.

LECTURE 7 Pipelining. DATAPATH AND CONTROL We started with the single-cycle implementation, in which a single instruction is executed over a single cycle.

ECE-C355 Computer Structures Winter 2008 The MIPS Datapath Slides have been adapted from Prof. Mary Jane Irwin ( )

CSE431 L06 Basic MIPS Pipelining.1Irwin, PSU, 2005 MIPS Pipeline Datapath Modifications  What do we need to add/modify in our MIPS datapath? l State registers.

Lecture 9. MIPS Processor Design – Pipelined Processor Design #1 Prof. Taeweon Suh Computer Science Education Korea University 2010 R&E Computer System.

CPE432 Chapter 4B.1Dr. W. Abu-Sufah, UJ Chapter 4B: The Processor, Part B-1 Read Sections 4.7 Adapted from Slides by Prof. Mary Jane Irwin, Penn State.

CMSC 611: Advanced Computer Architecture Pipelining Some material adapted from Mohamed Younis, UMBC CMSC 611 Spr 2003 course slides Some material adapted.

CSCI-365 Computer Organization Lecture Note: Some slides and/or pictures in the following are adapted from: Computer Organization and Design, Patterson.

Lecture 18: Pipelining I.

Computer Organization

CMSC 611: Advanced Computer Architecture

Single Clock Datapath With Control

ECE232: Hardware Organization and Design

Pipelining in more detail

CS-447– Computer Architecture Lecture 14 Pipelining (2)

The Processor Lecture 3.6: Control Hazards

The Processor Lecture 3.4: Pipelining Datapath and Control

Instruction Execution Cycle

Presentation transcript:

CS 300 – Lecture 24 Intro to Computer Architecture / Assembly Language The LAST Lecture!

Final Exam Tuesday, 8am. The exam will be comprehensive I'll post a worksheet (practice exam) Thursday evening that covers topics since the last exam.

Homework The last homework is due Friday. "darcs pull" is your friend – minor bugs have been found (and fixed) in the assembler. Extra credit: if you want it, come see me after class. All EC stuff is due Friday of finals week.

Speeding Up An Instruction Typically, an instruction is executed in stages: * Fetch (brings the instruction in from cache / memory) * Decode (figure out what the instruction will do) * Execute (the actual operation, like add or multiply) * Store (place results in registers / memory) This varies a lot from processor to processor but the idea is always the same – break up execution into smaller chunks that overlap.

Other Speedup Strategies * Vector instructions: explicit parallelism in the instruction set to feed sequences of data to a functional unit (Cray-1 and successors) * Multiple instructions at once: pack instruction words with lots of independent operations that execute at the same time (VLIW) * Replicated CPU, one instruction stream (SIMD) * Multicore machines (MIMD) – various coupling strategies

Pipeline Hazards Structural: lack of computational resources to perform operations in parallel Data: dependencies among instructions (write-read) Control: conditional branching prevents you from knowing which instruction comes next

Memory Hazards * It's "obvious" which registers an instruction uses * It's hard to figure out how memory accesses interact. Much harder to find out which memory an instruction touches. What can we do? * Reorder reads without worry * Reads and writes can't be switched unless we know that they don't interfere * Separate "memory banks" (instruction / data) reduces hazards

Compiler / Programmer Help The compiler and programmer have access to information that the CPU doesn't. We can determine whether or not aliasing is possible – this is what makes it hard to reorganize memory access. A CPU isn't able to understand that two pointers must point to different things, for example.

Delayed Branching One interesting feature of the MIPS was the delayed branch. The instruction after a branch is always executed – that is, you want to have at least one instruction in the pipeline before you have to wait to see where the branch goes. You can't see this – the assembler is able to fill branch slots for you behind your back.

User-Level Pipelining One of the big ideas in software is to pipeline at the application level: * Anticipate future I/O * Use non-blocking reads (threading) * Understand whether writers may alter information that is prefetched Note that web browsers pipline web page display!

Pipeline Metaphors The idea of keeping lots of workers busy at once is pervasive. Managers have to deal with similar issues when scheduling work: * Not all workers can perform the same task * Workflow requires that things be passed from one worker to another * Speculative work is often done (like a restaurant preparing food that may or may not be ordered)

Examples The following slides are stolen from Mary Jane Irwin ( )

Single Cycle vs. Multiple Cycle Timing Clk Cycle 1 Multiple Cycle Implementation: IFetchDecExecMemWB Cycle 2Cycle 3Cycle 4Cycle 5Cycle 6Cycle 7Cycle 8Cycle 9Cycle 10 IFetchDecExecMem lwsw IFetch R-type Clk Single Cycle Implementation: lwsw Waste Cycle 1Cycle 2 multicycle clock slower than 1/5 th of single cycle clock due to stage register overhead

How Can We Make It Even Faster? Split the multiple instruction cycle into smaller and smaller steps –There is a point of diminishing returns where as much time is spent loading the state registers as doing the work Start fetching and executing the next instruction before the current one has completed –Pipelining – (all?) modern processors are pipelined for performance

A Pipelined MIPS Processor Start the next instruction before the current one has completed –improves throughput - total amount of work done in a given time –instruction latency (execution time, delay time, response time - time from the start of an instruction to its completion) is not reduced Cycle 1Cycle 2Cycle 3Cycle 4Cycle 5 IFetchDecExecMemWB lw Cycle 7Cycle 6Cycle 8 sw IFetchDecExecMemWB R-type IFetchDecExecMemWB clock cycle (pipeline stage time) is limited by the slowest stage for some instructions, some stages are wasted cycles

Single Cycle, Multiple Cycle, vs. Pipeline Multiple Cycle Implementation: Clk Cycle 1 IFetchDecExecMemWB Cycle 2Cycle 3Cycle 4Cycle 5Cycle 6Cycle 7Cycle 8Cycle 9Cycle 10 IFetchDecExecMem lwsw IFetch R-type lw IFetchDecExecMemWB Pipeline Implementation: IFetchDecExecMemWB sw IFetchDecExecMemWB R-type Clk Single Cycle Implementation: lwsw Waste Cycle 1Cycle 2

Pipelining the MIPS ISA What makes it easy –all instructions are the same length (32 bits) can fetch in the 1 st stage and decode in the 2 nd stage –few instruction formats (three) with symmetry across formats can begin reading register file in 2 nd stage –memory operations can occur only in loads and stores can use the execute stage to calculate memory addresses –each MIPS instruction writes at most one result (i.e., changes the machine state) and does so near the end of the pipeline (MEM and WB) What makes it hard –structural hazards: what if we had only one memory? –control hazards: what about branches? –data hazards: what if an instruction’s input operands depend on the output of a previous instruction?

Can Pipelining Get Us Into Trouble? Yes: Pipeline Hazards –structural hazards: attempt to use the same resource by two different instructions at the same time –data hazards: attempt to use data before it is ready An instruction’s source operand(s) are produced by a prior instruction still in the pipeline –control hazards: attempt to make a decision about program control flow before the condition has been evaluated and the new PC target address calculated branch instructions Can always resolve hazards by waiting –pipeline control must detect the hazard –and take action to resolve hazards

I n s t r. O r d e r Time (clock cycles) lw Inst 1 Inst 2 Inst 4 Inst 3 ALU Mem Reg MemReg ALU Mem Reg MemReg ALU Mem Reg MemReg ALU Mem Reg MemReg ALU Mem Reg MemReg A Single Memory Would Be a Structural Hazard Reading data from memory Reading instruction from memory Fix with separate instr and data memories (I$ and D$)

Register Usage Can Cause Data Hazards I n s t r. O r d e r add $1, sub $4,$1,$5 and $6,$1,$7 xor $4,$1,$5 or $8,$1,$9 ALU IM Reg DMReg ALU IM Reg DMReg ALU IM Reg DMReg ALU IM Reg DMReg ALU IM Reg DMReg Dependencies backward in time cause hazards Read before write data hazard

To Know for the Final * How caching, virtual memory, and pipelines work in broad terms * What problems are being solved * Basic approaches to these problems * What is done in software and hardware