RM0004 STMICROELECTRONICS | Alldatasheet

Document overview

  • Manufacturer or author: Provided By ALLDATASHEET.COM(FREE DATASHEET DOWNLOAD SITE)
  • PDF pages: 1176

Technical content

Datasheet sections

  • 1 Overview
  • 1.1 Overview Book E and the Book E implementation standards (EIS)
  • 1.1.1 Auxiliary processing units (APUs)
  • 1.2 Instruction set
  • 1.3 Register set
  • 1.4 Interrupts and exception handling
  • 1.4.1 Exception handling
  • 1.4.2 Interrupt classes
  • 1.4.3 Interrupt categories
  • 1.4.4 Interrupt registers
  • 1.5 Memory management
  • 1.5.1 Address translation
  • 1.5.2 MMU assist registers (MAS1–MAS7)
  • 1.5.3 Process ID registers (PID0–PID2)
  • 1.5.4 TLB coherency
  • 1.5.5 Atomic update memory references
  • 1.5.6 Memory access ordering
  • 1.5.7 Cache control instructions
  • 1.5.8 Programmable page characteristics

Datasheet sections

  • 1.6 Performance monitoring
  • 1.6.1 Global control register
  • 1.6.2 Performance monitor counter registers
  • 1.6.3 Local control registers
  • 1.7 Legacy support of PowerPC architecture
  • 1.7.1 Instruction set compatibility
  • 1.7.2 Memory subsystem
  • 1.7.3 Interrupt handling
  • 1.7.4 Memory management
  • 1.7.5 Requirements for system reset generation
  • 1.7.6 Little-endian mode
  • 2 Register model
  • 2.1 Overview
  • 2.2 Register model for 32-bit Book E implemen tations
  • 2.2.1 Special-purpose registers (SPRs)
  • 2.3 Registers for integer operations
  • 2.3.1 General purpose registers (GPRs)
  • 2.3.2 Integer exception register (XER)
  • 2.4 Registers for floating-point operations
  • 2.4.1 Floating-point registers (FPRs)
  • 2.4.2 Floating-point status and control register (FPSCR)
  • 2.5 Registers for branch operations
  • 2.5.1 Condition register (CR)
  • 2.5.2 Link register (LR)
  • 2.5.3 Count register (CTR)
  • 2.6 Processor control registers
  • 2.6.1 Machine state register (MSR)
  • 2.7 Hardware implementation-dependent registers
  • 2.7.1 Hardware implementation dependent register 0 (HID0)
  • 2.7.2 Hardware implementation dependent register 1 (HID1)
  • 2.7.3 Processor ID register (PIR)
  • 2.7.4 Processor version register (PVR)
  • 2.7.5 System version register (SVR)
  • 2.8 Timer registers
  • 2.8.1 Timer control register (TCR)

Datasheet sections

  • 2.16.6 User local control B registers (UPMLCb0–UPMLCb3)
  • 2.16.7 Performance monitor counter registers (PMC0–PMC3)
  • 2.16.8 User performance monitor counter registers (UPMC0–UPMC3)
  • 2.17 Device control registers (DCRs)
  • 2.18 Book E SPR model
  • 2.18.1 Invalid SPR references
  • 2.18.2 Synchronization requirements for SPRs
  • 2.18.3 Reserved SPRs
  • 2.18.4 Allocated SPRs
  • 3 Instruction model
  • 3.1 Operand conventions
  • 3.1.1 Data organization in memory and data transfers
  • 3.1.2 Alignment and misaligned accesses
  • 3.2 Instruction set summary
  • 3.2.1 Classes of instructions
  • 3.2.2 Instruction forms
  • 3.2.3 Addressing modes
  • 3.3 Instruction set overview
  • 3.3.1 Book E user-level instructions
  • 3.3.2 Supervisor level instructions
  • 3.3.3 Recommended simplified mnemonics
  • 3.3.4 Book E instructions with implementation-specific features
  • 3.3.5 EIS instructions
  • 3.3.6 Context synchronization
  • 3.4 Instruction fetching
  • 3.5 Memory synchronization
  • 3.6 EIS-specific instructions
  • 3.6.1 SPE and embedded fl oating-point APUs
  • 3.6.2 Integer select ( isel) APU
  • 3.6.3 Performance monitor APU
  • 3.6.4 Cache locking APU
  • 3.6.5 Machine check APU
  • 3.6.6 VLE extension
  • 3.7 Instruction listing

Datasheet sections

  • 5.2 Memory and cache coherency
  • 5.2.1 Memory/Cache access attributes
  • 5.2.2 Shared memory
  • 5.3 Cache model
  • 5.3.1 Cache programming model
  • 5.3.2 Primary (L1) cache model
  • 5.4 Storage model
  • 5.4.1 Storage programming model
  • 5.4.2 The storage architecture
  • 5.4.3 Virtual address (VA)
  • 5.4.4 Address spaces
  • 5.4.5 Process ID
  • 5.4.6 Address translation
  • 5.4.7 Address translation and the ST EIS
  • 5.4.8 Permission attributes
  • 5.4.9 Translation lookaside buffer (TLB) arrays
  • 5.4.10 TLB management
  • 5.4.11 MAS registers and exception handling
  • 6 Instruction set
  • 6.1 Notation
  • 6.2 Instruction fields
  • 6.3 Description of instruction operations
  • 6.3.1 SPE APU saturation and bi t-reverse models
  • 6.3.2 Embedded floating-point conversion models
  • 6.3.3 Integer saturation models
  • 6.3.4 Embedded floating-point results
  • 6.4 Instruction set
  • 7 Auxiliary processing units (APUs)

Datasheet sections

  • 9.1 Compatibility with PowerPC Book E
  • 9.2 Instruction mnemonics and operands
  • 10 VLE storage addressing
  • 10.1 Data memory addressing modes
  • 10.2 Instruction memory addressing modes
  • 11 VLE compatibility with the EIS
  • 11.1 Overview
  • 11.2 VLE extension processor and storage control extensions
  • 11.2.1 EIS instruction extensions
  • 11.2.2 Book E instruction extensions
  • 11.2.3 EIS MMU extensions
  • 11.2.4 EIS debug APU extensions
  • 12 VLE instruction classes
  • 12.1 Processor control instructions
  • 12.1.1 System linkage instructions
  • 12.1.2 Processor control register manipulation instructions
  • 12.1.3 Instruction synchronization instruction
  • 12.2 Branch operation instructions
  • 12.2.1 Registers for branch operations
  • 12.2.2 Branch instructions
  • 12.3 Condition register instructions
  • 12.4 Integer instructions
  • 12.4.1 Integer load instructions
  • 12.4.2 Integer store instructions
  • 12.4.3 Integer arithmetic instructions
  • 12.4.4 Integer logical and move instructions
  • 12.4.5 Integer compare and bit test instructions
  • 12.4.6 Integer select instruction
  • 12.4.7 Integer trap instructions
  • 12.4.8 Integer rotate and shift instructions
  • 12.5 Storage control instructions
  • 12.5.1 Storage synchronization instructions
  • 12.5.2 Cache management instructions

Datasheet sections

November 2007 Rev 1 1/1176 RM0004 Reference manual Programmer’s reference manual for Book E processors Introduction This reference manual gives an overview of Book E, a version of the PowerPC architecture intended for embedded processors. To ensure application level compatibility with the PowerPC architecture developed by Apple, IBM, and Freescale, Book E incorporates the user level resources defined in the user instruction set architecture (UISA), Book I, of the AIM architectural definition.

2.14.1 Signal processing, embedded floating-point status, control register

7.4 Embedded vector and scalar single-precision floating-point APUs

A.2 Instructions sorted by primary opc odes (decimal and hexadecimal) . . 1048 B.4.6 Simplified mnemonics that incorporate CR conditions (eliminates BO and

Table 293. Simplified mnemonics for bc and bca without comparison conditions or LR Update. . . 1126 Table 294. Simplified mnemonics for bclr and bcctr without comparison conditions or LR update 1126

The primary objective of this reference is to provide a view of the programming model defined by Book E and the Book E implementation standards (EIS). Book E is a PowerPC™ architecture definition for embedded processors that ensures binary compatibility with the user instruction set architecture (UISA) portion of the PowerPC architecture as it was jointly developed by Apple, IBM, and Motorola (now Freescale Semiconductor, Inc.). This book should be used with the user documentation for individual implementations; such documents provide a high-level summary of the information that appears here, as well as implementation-specific features and implementation differences that are not described here. This document distinguishes between the three levels of the architectural and implementation definition, as follows:

  • The Book E architecture —Book E defines a se t of user-level instructions and registers that are drawn from the UISA portion of the AIM definition of the PowerPC architecture. Book E also include numerous other supervisor-level registers and instructions as they were defined in the AIM version of the PowerPC architecture for the virtual environment architecture (VEA) and the operating environment architecture (OEA). Because Book E defines a much different model for operating system resources such as the MMU and interrupts, it defines many new registers and instructions.
  • Book E implementation standards (EIS). In many cases, the Book E architecture definition provides a very general framework, leaving many higher-level details up to the implementation. To ensure consistency among its Book E implementations, working standards were defined, providing an additional layer of architecture between Book E and actual devices. This layer includes more specific definitions of Book E features as well as extensions to the architecture, typically in the form of auxiliary processing units (APUs), which define additional registers, instructions, and interrupts that provide specially targeted capabilities. Note that some APUs are implementation-specific and are available only on individual devices. The APUs described here are those that are implemented on multiple processors or families of processors. The EIS guarantees that if an APU is implemented, it conforms to the EIS architecture described here. Information in this book is subject to change without notice, as described in the disclaimers on the title page of this book. As with any technical documentation, it is the readers’ responsibility to be sure they are using the most recent version of the documentation. Audience It is assumed that the reader has the appropriate general knowledge regarding operating systems, microprocessor system design, and the basic principles of RISC processing to use the information in this manual.

Following is a summary and a brief description of the major sections of this manual:

  • Part I: Book E and Book E implementation standards,” describes the programming model defined by the PowerPC Book E architecture and the EIS. It consists of the following chapters: – Chapter 1: Overview,” provides a general discussion of the programming, interrupt, cache, and memory management models as they are defined by Book E and the EIS. – Chapter 2: Register model,” is useful for software engineers who need to understand the programming model in general and the functionality of each register. – Chapter 3: Instruction model,” provides an overview of the addressing modes and a description of the instructions. Instructions are organized by function. – Chapter 4: Interrupts and exceptions,” provides an overview of the Book E– and EIS-defined interrupts and exception conditions that can cause them. – Chapter 5: Storage architecture,” describes the cache and MMU portions of the EIS. – Chapter 6: Instruction set,” functions as a handbook for the instruction set. Instructions are sorted by mnemonic. Each instruction description includes the instruction formats and an individualized legend that provides such information as the level or levels of the architecture in which the instruction may be found and the privilege level of the instruction.
  • Part II: EIS-defined extensions to the Book E architecture,” describes the auxiliary procession units (APUs) defined by the EIS. It consists of the following chapters: – Chapter 7: Auxiliary processing units (APUs),” describes extensions to the Book E architecture defined by the EIS. These include the following: - Chapter 7.1: Integer select APU” - Chapter 7.2: Performance monitor APU” - Chapter 7.3: Signal processing engine APU (SPE APU)” - Chapter 7.4: Embedded vector and scalar single-precision floating-point APUs (SPFP APUs)” - Chapter 7.5: Machine check APU” - Chapter 7.6: Debug APU” – Chapter 8: Storage-related APUs.” describes the following APUs defined by the storage architecture: - Chapter 8.1: Cache line locking APU” - Chapter 8.2: Direct cache flush APU” - Chapter 8.3: Cache way partitioning APU”

Subsequent chapters describe the VLE extension - Chapter 9: VLE introduction” - Chapter 10: VLE storage addressing” - Chapter 11: VLE compatibility with the EIS” - Chapter 12: VLE instruction classes” - Chapter 13: VLE instruction set” - Chapter 14: VLE instruction index”

  • The following appendixes are included: – Appendix A: Instruction set listings,” lists all instructions except those defined by the VLE extension instructions by both mnemonic and opcode, and includes a quick reference table with general information, such as the architecture level, privilege level, form, and whether the instruction is optional. VLE instruction opcodes are listed in Section 13: VLE instruction set.” – Appendix B: Simplified mnemonics for PowerPC instructions,” describes simplified mnemonics, which are provided for easier coding of assembly language programs. Simplified mnemonics are defined for the most frequently used forms of branch conditional, compare, trap, rotate and shift, and certain other instructions defined by the PowerPC™ architecture and by implementations of and extensions to the PowerPC architecture. – Appendix C: Programming examples,” gives examples of how memory synchronization instructions can be used to emulate various synchronization primitives and to provide more complex forms of synchronization. It also describes multiple precision shifts. – Appendix D: Guidelines for 32-bit book E,” provides guidelines used by 32-bit Book E implementations; a set of guidelines is also outlined for software developers. Application software written to these guidelines can be labeled 32-bit Book E applications and can be expected to execute properly on all implementations of Book E, both 32-bit and 64-bit implementations. – Appendix E: Embedded floating-point results,” provides guidelines used by 32-bit Book E implementations; a set of guidelines is also outlined for software developers. Application software written to these guidelines can be labeled 32-bit Book E applications and can be expected to execute properly on all implementations of Book E, both 32-bit and 64-bit implementations. This book includes a glossary and an index.

This section lists additional reading that provides background for the information in this manual as well as general information about the architecture. General information The following documentation, published by Morgan-Kaufmann Publishers, 340 Pine Street, Sixth Floor, San Francisco, CA, provides useful information about the PowerPC architecture and computer architecture in general:

  • The PowerPC Architecture: A Specification for a New Family of RISC Processors, Second Edition, by International Business Machines, Inc. Related documentation ST documentation is available from the sources listed on the back cover of this manual; the document order numbers are included in parentheses for ease in ordering:
  • Reference manuals—These books (formerly called user’s manuals) provide details about individual implementations and are intended for use with the EREF.
  • Addenda/errata to reference manuals—Because some processors have follow-on parts an addendum is provided that describes the additional features and functionality changes. These addenda are intended for use with the corresponding reference manuals.
  • Hardware specifications—Hardware specificat ions provide specific data regarding bus timing, signal behavior, and AC, DC, and thermal characteristics, as well as other design considerations.
  • Technical summaries—Each device has a technical summary that provides an overview of its features. This document is roughly the equivalent to the overview (Chapter 1) of an implementation’s reference manual.
  • Application notes—These short documents address specific design issues useful to programmers and engineers working with ST processors. Additional literature is published as new processors become available.

Table 1. Conventions Additional conventions used with instruction encodings are described in Table 195. takes a value of one, it is said to be set. mnemonics Instruction mnemonics are shown in lowercase bold. Italics indicate variable command parameters, for example, bcctrx. Book titles in text are set in italics. x An italicized x indicates an alphanumeric variable. n An italicized n indicates an numeric variable.

Table 2 contains acronyms and abbreviations that are used in this document. Table 2. Acronyms and abbreviated terms

Table 2. Acronyms and abbreviated terms (continued)

Table 4 describes instruction field notation conventions used in this manual. Table 3. Terminology conventions Table 4. Instruction field conventions

RM0004 Part I: Book E and Book E implementation standards Part I: Book E and Book E implementation standards Part I describes the registers and instructions defined by the Book E architecture and by the Book E implementation standards (EIS). It contains the following chapters:

  • Chapter 1: Overview,” provides a general discussion of the programming, interrupt, cache, and memory management models as they are defined by Book E and the EIS.
  • Chapter 2: Register model,” is useful for software engineers who need to understand the programming model in general and the functionality of each register.
  • Chapter 3: Instruction model,” provides an overview of the addressing modes and a description of the instructions. Instructions are organized by function.
  • Chapter 4: Interrupts and exceptions,” provides an overview of the Book E– and EIS– defined interrupts and exception conditions that can cause them.
  • Chapter 5: Storage architecture,” describes the cache and MMU portions of the EIS.

1 Overview

This document describes the Book E version of the PowerPC™ architecture as it is further defined by the Book E implementation standards (EIS) and implemented on Book E cores. This chapter includes overviews of the following:

  • Features of the Book E version of the PowerPC architecture and implementation- details defined by the EIS
  • The Book E and EIS programming model
  • The Book E and EIS interrupt model
  • The Book E and EIS memory management model
  • Architectural compatibility and migration from the original version of the PowerPC architecture as defined by Apple, IBM, and Motorola (referred to as the AIM version of the PowerPC architecture)

1.1 Overview Book E and the Bo ok E implementation standards

(EIS) Book E is a version of the PowerPC architecture intended for embedded processors. To ensure application-level compatibility with the PowerPC architecture developed by Apple, IBM and Freescale, Book E incorporates the user-level resources defined in the user instruction set architecture (UISA), Book I, of the AIM architectural definition. Because operating systems for embedded processors have different needs than those for desktop systems, Book E defines more flexible interrupt and memory management models. Instead of the segmented memory model defined by the AIM architecture, Book E provides a page-based memory system that supports multiple variable-sized pages managed through translation lookaside buffers (TLBs). Interrupt offsets can be programmed through interrupt-specific interrupt vector offset registers (IVORs). Book E defines the interrupt vector prefix register (IVPR), which is programmed with a prefix value that is concatenated with the IVOR values to place the interrupt vector table anywhere in memory. As a consequence, some resources defined by the AIM version of the architecture are no longer supported and new ones are provided. For example, segment and block address translation (BAT) registers are gone, and new instructions, registers, and interrupts have been defined for managing page translation and protection through TLBs. Moreover, the Book E architecture allows greater flexibility. For example, Book E defines the TLB Write Entry (tlbwe) and TLB Read Entry (tlbre) instructions only very generally, leaving details of their execution and behavior up to the implementation. However, to ensure compatibility among Book E implementations, the Book E implementation standard (EIS) defines more specifically how these instructions work.

1.1.1 Auxiliary processing units (APUs)

Book E supports the use of auxiliary processing units (APUs), which allocate opcode and register space for extending the instruction set without affecting the instruction set defined by Book E. This facilitates the development of special-purpose resources that are useful to some embedded environments but impractical for others. Note that instructions from multiple APUs may be assigned the same opcode numbers of the allocated opcode space. The EIS defines many APUs. These APUs are not required on all devices, but devices that implement them do so strictly following the EIS architectural definition. In addition, an implementation may also provide an APU that is not a part of the EIS. APUs may consist of any combination of instructions, optional behavior of Book E–defined instructions, registers, register files, fields within Book E–defined registers, interrupts, or exception conditions within Book E–defined interrupts. Chapter 7: Auxiliary processing units (APUs),” provides an overview of specific APUs.

1.2 Instruction set

The instruction set of a ST 32-bit Book E–compliant device includes the following:

  • The Book E instruction set for 32-bit implementations. This is composed primarily of the user-level instructions defined by the UISA. Some implementations do not include the Book E floating-point instructions or the Load String Word Indexed instruction (lswx).
  • Instructions defined by EIS APUs. These include the following: – Integer select APU. This APU consists of the Integer Select instruction ( isel), which incorporates an if-then-else statement that selects between two source registers by comparison to a CR bit. This instruction eliminates conditional branches, decreases band latency, and reduces the code footprint. – SPE (signal processing engine) APU instructions. SPE instructions treat 64-bit GPRs as a vector of two 32-bit elements (some instructions also read or write 16- bit elements). Chapter 3.6.1: SPE and embedded floating-point APUs on page 186,” lists SPE APU vector instructions. – The embedded vector floating-point APU provides instructions that use the upper and lower words of the 64-bit GPRs for single-precision, vector floating-point calculations. – The embedded scalar single-precision APU provides instructions that use the lower 32 bits of the GPRs for single-precision, scalar floating-point calculations. – The embedded scalar double-precision APU instructions use the 64-bit GPRs for floating-point calculations. – Performance monitor APU—This APU defines two instructions, mfpmr and mtpmr, used for reading and writing the performance monitor registers (PMRs). – Cache block lock and unlock APU, co nsisting of the following instructions: - Data Cache Block Lock Clear (dcblc) - Data Cache Block Touch and Lock Set (dcbtls) - Data Cache Block Touch for Store and Lock Set (dcbtstls) - Instruction Cache Block Lock Clear (icblc) - Instruction Cache Block Touch and Lock Set (icbtls)

1.3 Register set

Note: Devices that implement a particular core may not implement all registers defined by that core. Figure 1. EIS programming model register set registers; not part of the Book E architecture.

3 L1 Cache (Read-Only)

1.4 Interrupts and exception handling

Book E and the EIS support an extended exception handling model, with nested interrupt capability and extensive interrupt vector programmability. The following sections define the exception model, including an overview of exception handling as implemented in a ST Book E device, a brief description of the exception classes, and an overview of the registers involved.

1.4.1 Exception handling

In general, interrupt processing begins with an exception that occurs due to external conditions, errors, or program execution problems. When the exception occurs, the processor checks to verify that interrupt processing is enabled for that particular exception. If enabled, the interrupt causes the state of the processor to be saved in the appropriate registers, and prepares to begin execution of the handler located at the associated vector address for that particular exception. Once the handler is executing, the implementation may need to check one or more bits in the exception syndrome register (ESR) or the SPEFSCR, depending on the exception type, to verify the specific cause of the exception and take appropriate action. The interrupts are described in Chapter 1.4.4: Interrupt registers,” and in Table 6.

1.4.2 Interrupt classes

All interrupts may be categorized as asynchronous/synchronous and critical/noncritical.

  • Asynchronous interrupts are caused by events that are independent of instruction execution. The address reported in the save/restore register is that of the instruction that would have executed next had the asynchronous interrupt not occurred.
  • Synchronous interrupts are caused directly by the execution or attempted execution of instructions. Synchronous inputs can be precise or imprecise: – Synchronous precise interrupts are those that precisely indicate the address of the instruction causing the exception that generated the interrupt or, in some cases, the address of the next instruction in program order. The interrupt type and status bits allow determination of which of the two instructions has been addressed in the appropriate save/restore register. – Synchronous imprecise interrupts may indicate the address of the instruction causing the exception that generated the interrupt or some instruction after the instruction causing the interrupt. If the interrupt was caused by either the context synchronizing mechanism or the execution synchronizing mechanism, the address in the appropriate save/restore register is the address of the interrupt forcing instruction. If the interrupt was not caused by either of those mechanisms, the address in the save/restore register is the last instruction to start execution and may not have completed. No instruction following the instruction in the save/restore register has executed.

1.4.3 Interrupt categories

Book E defines critical and noncritical interrupt categories, and the EIS defines the machine check and debug interrupt categories. Each category has a separate set of save and restore registers to which machine state and a return address are automatically written when an interrupt is taken. Each category has a return from interrupt instruction that uses the save and restore registers to reestablish the machine state of the interrupted process and

  • Debug APU interrupt (if present)—Although Book E defines debug as a critical interrupt, the EIS defines a separate debug APU. Debug save and restore registers (DSRR0/DSRR1) save state when a debug interrupt is taken; rdci restores state at the end of the interrupt handler. These interrupts are masked by setting the machine check enable bit, MSR[DE].
  • Machine check APU interrupt (if present)—Although Book E defines machine check as a critical interrupt, the EIS defines a separate machine check APU. Machine check save and restore registers (MCSRR0/MCSRR1) save state when a machine check interrupt is taken; rfmci restores state at the end of the interrupt handler. These interrupts are masked by setting the machine check enable bit, MSR[ME].
  • Noncritical interrupts—First-level interrupts that allow the processor to change program flow to handle conditions generated by external signals, errors, or unusual conditions arising from program execution or from programmable timer-related events. These interrupts are largely identical to those defined by the OEA portion of the Power PC architecture. They use save and restore registers (SRR0/SRR1) to save processor state and the rfi instruction to restore state. Asynchronous noncritical interrupts can be masked by the external interrupt enable bit, MSR[EE].
  • Critical interrupts—Can be taken during a noncritical interrupt or during regular program flow. They use the critical save and restore registers (CSRR0/CSRR1) to save state when they are taken; they use the rfci instruction to restore state. These interrupts can be masked by the critical enable bit, MSR[CE]. Book E defines the critical input and watchdog timer interrupts as critical interrupts. One interrupt of each category can be reported at a time; when it is taken, no program state is lost. Save/restore register pairs are serially reusable, so program state may be lost when an unordered interrupt is taken. See Section 4.10: Interrupt ordering and masking.”

1.4.4 Interrupt registers

The registers associated with interrupt and exception handling are described in Table 5. Table 5. Interrupt registers the instruction that will execute after the rfi instruction. after an rfi instruction is executed. machine state after an rfci instruction is executed.

Table 6 lists IVOR registers and associated interrupts. and restores machine state (if recoverable) after an rfmci instruction is executed. DSRR0 Debug save/restore register 0—Stores the address of the instruction that executes after rfdi executes. machine state (if recoverable) after rfmci executes. interrupts and restores machine state after an rfmci instruction is executed. associated bit is set and all other bits are cleared. and the embedded floating-point APUs. cache management instruction that caused an alignment, data TLB miss, or data storage interrupt. exception processing routines defined in the IVOR registers. exception processing routines defined in the IVOR registers. See Table 6. Table 5. Interrupt registers (continued)

Table 6. Interrupt vector registers and exception conditions

1.5 Memory management

The EIS supports demand-paged virtual memory as well other memory management schemes that depend on precise control of effective-to-physical address translation and flexible memory protection as defined by Book E. The mapping mechanism consists of software-managed TLBs that support variable-sized pages with per-page properties and permissions. The following properties can be configured for each TLB:

  • User mode page execute access
  • User mode page read access
  • User mode page write access
  • Supervisor mode page execute access
  • Supervisor mode page read access
  • Supervisor mode page write access
  • Write-through required (W)
  • Caching inhibited (I)
  • Memory coherence required (M)
  • Guarded (G)
  • Endianness (E)
  • User-definable (U0–U3), a 4-bit implementation-specific field

1.5.1 Address translation

Figure 2 shows a typical translation flow, although each implementation may differ in the specific details. The MMU translates 32-bit effective addresses generated by loads, stores, and instruction fetches into 32-bit real addresses (used for memory bus accesses) using an interim 41-bit virtual address.

Figure 2. Effective-to-Real Address Translation Flow from the value of MSR[IS] or MSR[DS], for instruction or data accesses, respectively. The appropriate L1 MMU (instruction or data) is checked for a matching address translation. parallel, so that hits for instruction accesses and data accesses can occur in the same clock. entries are replaced from their L2 TLB counterparts using a true LRU algorithm.

1.5.2 MMU assist registers (MAS1–MAS7)

defined by the Book E standard; more specific details are left to individual implementations.

2 TLBs 2 TLBs

permission bits (UX, SX, UW, SW, UR, SR) that specify user and supervisor read, write, and execute permissions. Some cores may not does not implement all of the MAS registers. MAS registers are affected by the following instructions:

  • MAS registers are accessed with the mtspr and mfspr instructions.
  • The TLB Read Entry instruction (tlbre) causes the contents of a single TLB entry from the L2 MMU to be placed in defined locations in MAS0–MAS3. The TLB entry to be extracted is determined by information written to MAS0 and MAS2 before the tlbre instruction is executed.
  • The TLB Write Entry instruction (tlbwe) causes the information stored in certain locations of MAS0–MAS3 to be written to the TLB specified in MAS0.
  • The TLB Search Indexed instruction (tlbsx) updates MAS registers conditionally, based on success or failure of a lookup in the L2 MMU. The lookup is specified by the instruction encoding and specific search fields in MAS6. The values placed in the MAS registers may differ, depending on a successful or unsuccessful search. For TLB miss and certain MMU-related DSI/ISI exceptions, MAS4 provides default values for updating MAS0–MAS2.

1.5.3 Process ID regi sters (PID0–PID2)

The Book E architecture identifies a single process ID register (PID). The EIS defines additional PIDs to hold values used to construct the virtual addresses for each access. Among these PIDs, PID0 is the Book E–defined PID. These process IDs provide an extended page sharing capability. Which of these three virtual addresses is used for translation is controlled by the TID field of a matching TLB entry, and when TID = 0x00 (identifying a page as globally shared), the PID values are ignored. A hit to multiple TLB entries in the L1 MMU (even if they are in separate arrays) or a hit to multiple entries in the L2 MMU is considered to be a programming error.

1.5.4 TLB coherency

TLB entries can be invalidated as defined in the Book E architecture. The tlbivax instruction invalidates a matching local TLB entry.

1.5.5 Atomic update memory references

Book E supports atomic update memory references for both aligned word forms of data using the load and reserve and store conditional instruction pair, lwarx and stwcx.. Typically, a load and reserve instruction establishes a reservation and is paired with a store conditional instruction to achieve the atomic operation. However, the programmer is responsible for preserving reservations across context switches and for protecting reservations in multiprocessor implementations.

1.5.6 Memory access ordering

To optimize performance, Book E supports weakly ordered references to memory. Thus, a processor manages the order and synchronization of instructions to ensure proper execution when memory is shared between multiple processes or programs. The cache and data memory control attributes, along with msync and mbar, provide the required access

control; msync and mbar are also broadcast to provide the appropriate control in the case of multiprocessor or shared memory systems.

1.5.7 Cache control instructions

Book E cache control instructions perform a full range of cache control functions, including cache locking by line. The EIS defines the following cache locking instructions:

  • Data Cache Block Lock Clear (dcblc)
  • Data Cache Block Touch and Lock Set (dcbtls)
  • Data Cache Block Touch for Store and Lock Set (dcbtstls)
  • Instruction Cache Block Lock Clear (icblc)
  • Instruction Cache Block Touch and Lock Set (icbtls)

1.5.8 Programmable page characteristics

Cache and memory attributes are programmable on a per-page basis. In addition to the write-through, caching-inhibited, memory coherency enforce, and guarded characteristics defined by the WIMG bits, Book E defines an endianness bit, E, that selects big- or little- endian byte ordering on a per-page basis.

1.6 Performance monitoring

The EIS provides a performance monitoring capability that supports counting of events such as processor clocks, instruction cache misses, data cache misses, mispredicted branches, and others. The count of these events may be configured to trigger a performance monitor exception. This interrupt is assigned to vector offset register IVOR35. The register set associated with performance monitoring consists of counter registers, a global control register, and local control registers. These registers are read/write from supervisor mode, and each register is reflected to a corresponding read-only register for user mode. The mtpmr and mfpmr instructions move data to and from these registers. An overview of the performance monitoring registers is provided in the following sections. For more information, see Chapter 7.2: Performance monitor APU.”

1.6.1 Global control register

The performance monitor global control register 0 (PMGC0) provides global control of the performance monitor from supervisor mode. From this register all counters may be frozen, unfrozen, or configured to freeze on an enabled condition or event. Additionally, the performance monitoring facility may be disabled or enabled from this register. The PMGC0 contents are reflected to UPMGC0, which may be read from user mode using mfpmr.

1.6.2 Performance moni tor counter registers

There are four counter registers (PCM0–PCM3) provided in the performance monitor facility. These 32-bit registers hold the current count for software-selectable events and can be programmed to generate an exception on overflow. They can be accessed from supervisor mode using mtpmr and mfpmr. Their contents are reflected to UPCM0–UPCM3, which can be read from user mode with mfpmr. The exception generated on overflow can be masked by clearing MSR[EE].

1.6.3 Local control registers

For each counter register, there are two corresponding local control registers. These two registers specify which of the 128 available events is to be counted, the action to be taken on overflow, and options for freezing a counter value under given modes or conditions.

  • PMLCa0–PMLCa3 provide fields that allow freezing of the corresponding counter in user mode, supervisor mode, or under software control. The overflow condition may be enabled or disabled from these registers. Register contents are reflected to UPMCLa0– UPMLCa3, which can be read from user mode with mfpmr.
  • PMLCb0–PMLCb3 provide count scaling for each counter register using configurable threshold and multiplier values. The threshold is a 6-bit value and the multiplier is a 3- bit encoded value, allowing 8 multiplier values in the range of 1 to 128. Any counter may be configured to increment only when an event occurs more than [threshold × multiplier] times. The contents of these registers are reflected to UPMCLb0–UPMLCb3, which can be read from user mode with mfpmr.

1.7 Legacy support of PowerPC architecture

In general, ST Book E processors support the user-level portion of the AIM architecture. The following subsections highlight the main differences. For specific details, refer to the relevant chapter.

1.7.1 Instruction set compatibility

The following sections generally describe compatibility between Book E and AIM PowerPC instruction sets. User instruction set The user mode instruction set defined by the AIM version of the PowerPC architecture is compatible with ST Book E processors with the following exceptions:

  • Floating-point functionality provided by the embedded floating-point APUs differs from the AIM defined floating-point ISA. Also, the vector and double-precision floating-point APUs use 64-bit GPRs rather than the FPRs defined by the UISA. Most porting of floating-point operations can be handled by recompiling; however, there are new instructions specific to the APUs.
  • String instructions are typically not implemented; therefore, trap emulation must be provided to ensure backward compatibility. Supervisor instruction set The supervisor mode instruction set defined by the AIM version of the PowerPC architecture is compatible with the EIS with the following exceptions:
  • The MMU architecture is different, so some TLB manipulation instructions have different semantics.
  • Instructions that support the BATs and segment registers are not implemented.
  • Interrupt vectors are defined by the Book E IVORn and IVPR SPRs.
  • Additional instructions are defined for returning from Book E–defined critical interrupts (rfci) and APU-specific interrupts.

1.7.2 Memory subsystem

Both Book E and the AIM version of the PowerPC architecture provide separate instruction and data memory resources. The EIS provides additional cache control features, including cache locking.

1.7.3 Interrupt handling

Interrupt handling is generally the same as that defined in the AIM version of the PowerPC architecture, with the following differences: (see Chapter 1.4)

  • Book E defines a new critical interrupt, providing an extra level of interrupt nesting. The critical interrupt includes external critical and watchdog timer time-out inputs.
  • The machine check APU implements the machine check exception differently from the Book E and from the AIM definition. It defines the Return from Machine Check Interrupt instruction, rfmci, and two machine check save/restore registers, MCSRR0 and MCSRR1.
  • Book E processors can use IVPR and IVORs to set exception vectors individually. To provide compatibility, they can be set to the address offsets defined in the OEA.
  • Unlike the AIM version of the PowerPC architecture, Book E does not define a reset vector; execution begins at a fixed virtual address, 0xFFFF_FFFC.
  • Some SPRs are different from those defined in the AIM version of the PowerPC architecture, particularly those related to the MMU functions. Much of this information has been moved to a new exception syndrome register (ESR).
  • Timer services are generally compatible, although Book E defines a new decrementer auto reload feature and the fixed-interval timer critical interrupt.

1.7.4 Memory management

ST Book E processors implement a straightforward virtual address space that complies with the Book E MMU definition, which eliminates segment registers and block address translation resources. Book E defines resources for fixed 4-Kbyte pages and multiple, variable page sizes that can be configured in a single implementation. TLB management is provided with new instructions and SPRs.

1.7.5 Requirements for sy stem reset generation

Book E does not specify a system reset interrupt as was defined in the AIM version of the PowerPC architecture, but typically, system reset is initiated either by asserting a signal or by software (for example, writing a 1 to DBCR0[34], if MSR[DE] = 1 At reset, instead of invoking a reset interrupt, fetching at address 0xFFFF_FFFC, as defined by Book E. In addition to the Book E reset definition, the EIS and the implementation define specific aspects of MMU page translation and protection mechanisms. Unlike the AIM version of the PowerPC core, as soon as instruction fetching begins, the core is in virtual mode with a hardware-initialized TLB entry.

1.7.6 Little-endian mode

Unlike the AIM version of the PowerPC, where the little-endian mode is controlled on a system basis, Book E supports control of byte ordering on a memory page basis. Additionally, true little-endian mode is supported by byte swapping.

2 Register model

This chapter describes the register model and indicates the architecture level at which each register is defined.

2.1 Overview

Although this chapter organizes registers according to their functionality, they can be differentiated according to how they are accessed, as follows:

  • Register files. These user-level registers are accessed explicitly through source and destination operands of computational, load/store, logical, and other instructions. Book E defines two types of register files: – General-purpose registers (GPRs), used as source and destination operands for most operations (except Book E–defined floating-point instructions, which use FPRs). See Chapter 2.3.1: General purpose registers (GPRs).” – Floating-point registers (FPRs), used for Book E–defined floating-point instructions. See Chapter 2.4.1: Floating-point registers (FPRs).”
  • Special-purpose registers (SPRs)—SPRs are accessed by using the Book E–defined Move to Special-Purpose Register (mtspr) and Move from Special-Purpose Register (mfspr) instructions. Chapter 2.2.1: Special-purpose registers (SPRs),” lists SPRs.
  • System-level registers that are not SPRs. These are as follows: – Machine state register (MSR). MSR is accessed with the Move to Machine State Register (mtmsr) and Move from Machine State Register (mfmsr) instructions. See Chapter 2.6.1: Machine state register (MSR).” – Condition register (CR) bits are grouped into eight 4-bit fields, CR0–CR7, which are set as follows (see Chapter 2.5.1: Condition register (CR)”): - Specified CR fields can be set by a move to the CR from a GPR (mtcrf). - A specified CR field can be set by a move to the CR from another CR field (mcrf), from the FPSCR (mcrfs), or from the XER (mcrxr). - CR0 can be set as the implicit result of an integer instruction. - CR1 can be set as the implicit result of a floating-point instruction. - A specified CR field can be set as the result of an integer or floating-point compare instruction (including SPE and SPFP compare instructions). – The floating-point status and control register (FPSCR). See Chapter 2.4.2: Floating-point status and control register (FPSCR).” – The EIS-defined accumulator, which is accessed by signal processing engine (SPE) APU instructions that update the accumulator. See Chapter 2.14.2: Accumulator (ACC).”
  • Device control registers (DCRs). Book E defines the existence of a DCR address space and the instructions to access them, but does not define particular DCRs. The on-chip DCRs exist architecturally outside the processor core and thus are not part of Book E. The contents of DCR DCRN can be read into a GPR using mfdcr rD,DCRN. GPR contents can be written into DCR DCRN using mtdcr DCRN,rS. See Chapter 2.17: Device control registers (DCRs).”
  • Performance monitor registers (PMRs). (Performance monitor APU) Similar to SPRs, PMRs are accessed by using the EIS-defined Move to Performance Monitor Register

(mtpmr) and Move from Performance Monitor Register (mfspr) instructions. See Chapter 2.16: Performance monitor registers (PMRs).”

2.2 Register model for 32-bi t Book E implementations

Book E implementations include the following types of software-accessible registers:

  • Registers that are accessed as part of instruction execution. These include the following: – The following registers are used for integer operations and are described in Chapter 2.3: Registers for integer operations”: - General-purpose registers (GPRs)—Book E defines a set of 32 GPRs used to hold source and destination operands for load, store, arithmetic, and computational instructions, and to read and write to other registers. - Integer exception register (XER)—XER bits are set based on the operation of an instruction considered as a whole, not on intermediate results. (For example, the Subtract from Carrying instruction (subfc), the result of which is specified as the sum of three values, sets bits in the XER based on the entire operation, not on an intermediate sum.) – Registers for floating-point operations. These include the following: - Floating-point registers (FPRs)—32 registers used to hold source and destination operands for Book E defined floating-point operations. Note that the embedded floating-point APUs do not implement FPRs; they use GPRs for floating-point operands. - Floating-point status and control register (FPSCR)—Used with floating-point operations. These registers are described in Chapter 2.4: Registers for floating- point operations.” – Condition register (CR)—Used to record conditions such as overflows and carries that occur as a result of executing arithmetic instructions (including those implemented by the SPE and SPFP APUs). The CR is described in Chapter 2.5: Registers for branch operations.” – Machine state register (MSR)—Used by the operating system to configure parameters such as user/supervisor mode, address space, and enabling of

asynchronous interrupts. MSR is described in Chapter 2.6.1: Machine state register (MSR).”

  • Special-purpose registers (SPRs). – Book E–defined special-purpose registers (SPRs) that are accessed explicitly using mtspr and mfspr instructions. These registers are listed in Table 7 in Chapter 2.2.1: Special-purpose registers (SPRs).” – EIS–defined SPRs that are ac cessed explicitly using the mtspr and mfspr instructions. These registers are listed in Table 8 in Chapter 2.2.1: Special- purpose registers (SPRs).” – SPRs are described by function in the following sections: - Chapter 2.5: Registers for branch operations” - Chapter 2.6: Processor control registers” - Chapter 2.7: Hardware implementation-dependent registers” - Chapter 2.8: Timer registers” - Chapter 2.9: Interrupt registers” - Chapter 2.10: Software use sprs (SPRG0–SPRG7 and USPRG0)” - Chapter 2.11: L1 cache registers” - Chapter 2.12: MMU registers” - Chapter 2.13: Debug registers” - Chapter 2.14: SPE and SPFP APU registers” - Chapter 2.15: Alternate time base registers (ATBL and ATBU)”
  • EIS-defined performance monitor registers, described in Chapter 2.16: Performance monitor registers (PMRs).” PMRs are like SPRs, but are accessed with EIS-defined move to and move from PMR instructions (mtpmr and mfpmr).
  • EIS-defined device control registers (DCRs). Book E defines a format for implementing device-specific device-control registers. See Chapter 2.17: Device control registers (DCRs).” Book E defines 32- and 64-bit registers. However, except for the 64-bit FPRs, only bits 32– 63 of Book E’s 64-bit registers (such as LR, CTR, the GPRs, SRR0, and CSRR0) are required to be implemented in hardware in a 32-bit Book E implementation. Likewise, all Book E integer instructions defined to return a 64-bit result return only bits 32– 63 of the result on a 32-bit Book E implementation. SPE APU vector instructions return 64- bit values; SPFP APU instructions return single-precision 32-bit values. As with the instruction set and other aspects of the architecture, Book E defines some features very specifically, for example, resources that ensure compatibility with implementations of the PowerPC ISA. Other resources are either defined as optional or are defined in a very general way, leaving specific details up to the implementation.

Figure 3. Register model vector instructions can access the upper word. (2.) USPRG0 is a separate physical register from SPRG0. (3.) EIS-defined registers; not part of the Book E architecture.

2.2.1 Special-purpose registers (SPRs)

architected processor resources and are accessed with the mtspr and mfspr instructions. Unlisted encodings are reserved for future use.

  • If the invalid SPR falls within the range specified as user mode (SPR[5] = 0), an illegal exception is taken.
  • If supervisor software attempts to access an invalid supervisor-level SPR (SPR[5] = 1), results are undefined.
  • If user software attempts to access an invalid supervisor-level SPR, a privilege exception is taken.

Table 7. Book E special purpose registers (by SPR abbreviation)

Table 7. Book E special purpose register s (by SPR abbreviation) (continued)

this table when parsing instructions.

  1. The DBSR is read using mfspr. It cannot be directly written to. Instead, DBSR bits corresponding to 1 bits

in the GPR can be cleared using mtspr.

  1. The TSR is read using mfspr. It cannot be directly written to. Instead, TSR bits corresponding to 1 bits in

the GPR can be cleared using mtspr.

  1. User-mode read access to SPRG3 is implementation-dependent

Table 8. EIS–defined SPRs (by SPR abbreviation)

Table 8. EIS–defined SPRs (by SPR abbreviation) (continued)

2.3 Registers for integer operations

The following sections describe registers defined for integer computational instructions.

2.3.1 General purpose registers (GPRs)

  • The signal processing engine (SPE) APU and the embedded vector single-precision floating-point APU treat the 64-bit operands as consisting of two, 32-bit elements, as shown in Figure 4.
  • The embedded scalar double-precision floating-point APU treats the GPRs as single 64-bit operands that accommodate IEEE double-precision values. PID0 Process ID register 0. Book E defines only this PID register and refers to as PID, not PID0. 48 Read/Write Y es Section 2.12.1 on page 97 PID1 Process ID register 1 633 Read/Write Y es Section 2.12.1 on page 97 PID2 Process ID register 2 634 Read/Write Y es Section 2.12.1 on page 97 SPEFSCR Signal processing and embedded floating- point status and control register 512 Read/Write No Section 2.14.1 on page 119 SVR System version register 1023 Read-only Y es Section 2.7.5 on page 75 TLB0CFG TLB configuration register 0 688 Read-only Y es Section 2.12.4 on page 100 TLB1CFG TLB configuration register 1 689 Read-only Y es Section 2.12.4 on page 100

Figure 4. SPE and floating point APU GPR usage Gray text indicates that the APU does not use this register or register field. Formatting of floating-point operands is as defined by IEEE 754, as described in the APU chapter of the EREF. instructions use and return 32-bit values in GPR bits 32–63.

2.3.2 Integer exception register (XER)

Interrupt RegistersSingle-prec. Single-prec. Interrupt RegistersSingle-prec.

Table 9 describes XER bit definitions. Table 9. XER field descriptions instructions or by other instructions (except mtspr[XER] and mcrxr) that cannot overflow. instructions set CA if any 1 bits are shifted out of a negative operand and clear CA otherwise. mtspr[XER], and mcrxr) do not affect CA. 35–56 — Reserved, should be cleared. transferred by a load string indexed or store string indexed instruction.

2.4 Registers for floating-point operations

This section details floating-point registers and their field descriptions.

2.4.1 Floating-point registers (FPRs)

Book E defines 32 floating-point registers (FPR0–FPR31). Floating-point instruction formats provide 5-bit fields for specifying FPRs used in instruction execution. Each FPR contains 64 bits that support the floating-point format. Instructions that interpret FPR contents as floating-point values use double-precision format for this interpretation. The computational instructions and the move and select instructions operate on data in FPRs and, except for compare instructions, place the result into an FPR, and optionally place status information into the CR. Load and store double instructions are provided that transfer 64 bits of data between memory and the FPRs with no conversion. Load single instructions are provided to transfer and convert floating-point values in floating-point single format from memory to the same value in floating-point double format in the FPRs. Store single instructions are provided to transfer and convert floating-point values in floating-point double format from the FPRs to the same value in floating-point single format in memory. Instructions are provided that manipulate the FPSCR and the CR explicitly. Some of these instructions copy data between an FPR and the FPSCR. The computational instructions and the select instruction accept values from the FPRs in double format. For single-precision arithmetic instructions, all input values must be representable in single format; if they are not, the result placed into the target FPR, and the setting of status bits in the FPSCR and in the CR (if Rc = 1), are undefined.

2.4.2 Floating-point status and control register (FPSCR)

The FPSCR, shown below, controls how floating-point exceptions are handled and records status resulting from floating-point operations. FPSCR[32–55] are status bits; FPSCR[56– 63] are control bits. Floating-point status and control register (FPSCR) Access: User read/write 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 R FX FEX VX OX UX ZX XX VXSNAN VXISI VXIDI VXZDZ VXIMZ VXVC FR FI C W Reset All zeros 48 51 52 53 54 55 56 57 58 59 60 61 62 63 R FPCC — VXSOFT VXSQRT VXCVI VE OE UE ZE XE NI RN W Reset All zeros

FPSCR[FX,FEX,VX] are not considered to be exception bits, and only FX is sticky. FPSCR bits affected by the various instructions. FPSCR fields are described in Table 10. Table 10. FPSCR field descriptions change from 0 to 1. mcrfs, mtfsfi, mtfsf, mtfsb0, and mtfsb1 can alter FPSCR[FX] explicitly.

33 FEX

exception bits. mcrfs, mtfsfi, mtfsf, mtfsb0, and mtfsb1 cannot alter FPSCR[VX] explicitly.

35 OX Floating-point overflow exception

36 UX Floating-point underflow exception

37 ZX Floating-point zero divide exception

Floating-point inexact exception. value of FPSCR[XX] with the new value of FPSCR[FI]. If the instruction does not affect FPSCR[FI], the value of FPSCR[XX] is unchanged.

39 VXSNAN Floating-point invalid operation exception (SNaN)

40 VXISI Floating-point invalid operation exception ( ∞ − ∞)

41 VXIDI Floating-point invalid operation exception ( ∞ ÷ ∞)

42 VXZDZ Floating-point invalid operation exception (0 ÷ 0)

43 VXIMZ Floating-point invalid operation exception ( ∞ × 0)

44 VXVC Floating-point invalid operation exception (invalid compare). incremented the fraction during rounding. This bit is not sticky.

51 FPRF

that if any portion of the result is undefined, the value placed into FPRF is undefined. this bit with the FPCC bits, to indicate the class of the result.

51 FPCC

52 — Reserved, should be cleared. mtfsfi, mtfsf, mtfsb0, or mtfsb1.

54 VXSQRT

Floating-point invalid operation exception (invalid square root).

55 VXCVI Floating-point invalid operation exception (invalid integer convert)

56 VE Floating-point invalid operation exception enable

57 OE Floating-point overflow exception enable

58 UE Floating-point underflow exception enable

59 ZE Floating-point zero divide exception enable

60 XE Floating-point inexact exception enable

implementation are met. The other effects of setting NI may differ among implementations. Floating-point rounding control (RN). Table 10. FPSCR field descriptions (continued)

Table 11 describes floating-point result flags.

2.5 Registers for branch operations

This section describes registers used by Book E branch and CR operations.

2.5.1 Condition register (CR)

  • Specified CR fields can be set by a move to the CR from a GPR (mtcrf).
  • A specified CR field can be set by a move to the CR from another CR field (mcrf), from the FPSCR (mcrfs), or from the XER (mcrxr).
  • CR0 can be set as the implicit result of an integer instruction.
  • CR1 can be set as the implicit result of a floating-point instruction.
  • A specified CR field can be set as the result of either an integer or a floating-point compare instruction (including SPE and SPFP compare instructions). Instructions are provided to perform logical operations on individual CR bits and to test individual CR bits (see Condition register instructions on page 204”). Note that instructions that access CR bits (for example, Branch Conditional (bc), CR logicals, and Move to Condition Register Field (mtcrf)) determine the bit position by adding

Table 11. Floating-point result flags

10001 Q u i e t N a N

accesses bit BI + 32, as shown in Table 12. Table 12. BI operand settings for CR fields Negative (LT)—Set when the result is negative. Positive (GT)—Set when the result is positive (and not zero). Zero (EQ)—Set when the result is zero. Set to the OR of the result of the compare of the high and low elements. Set to the AND of the result of the compare of the high and low elements. Set to the OR of the result of the compare of the high and low elements. Set to the AND of the result of the compare of the high and low elements. Less than or floating-point less than (LT, FL). rA < SIMM or rB (signed comparison) or rA < UIMM or rB (unsigned comparison). For floating-point compare instructions: frA < frB.

first three bits of CR0 is undefined. CR0 bits are interpreted as described in Table 13. Note that CR0 may not reflect the true (infinitely precise) result if overflow occurs. Greater than or floating-point greater than (GT, FG). rA > SIMM or rB (signed comparison) or rA > UIMM or rB (unsigned comparison). For floating-point compare instructions: frA > frB. Equal or floating-point equal (EQ, FE). For integer compare instructions: rA = SIMM, UIMM, or rB. For floating-point compare instructions: frA = frB. Set to the OR of the result of the compare of the high and low elements. Summary overflow or floating-point unordered (SO, FU). For floating-point compare instructions, one or both of frA and frB is a NaN. Set to the AND of the result of the compare of the high and low elements. Table 12. BI operand settings for CR fields (continued) Table 13. CR0 bit descriptions 32 Negative (LT) Bit 32 of the result is equal to one. 34 Zero (EQ) Bits 32–63 of the result are equal to zero.

35 Summary overflow

This is a copy of the final state of XER[SO] at the completion of the instruction.

descriptions in Chapter 3,” for detailed descriptions of how CR0 is set. copied from FPSCR[32–35]. These bits are interpreted as shown in Table 14. the result of the comparison, as shown in Table 15. Table 14. CR setting for floating-point instructions Table 15. CR setting for compare instructions Less than or floating-point less than (LT, FL). UIMM or rB (unsigned comparison). UIMM or rB (unsigned comparison).

  • Specified CR fields can be set by a move to the CR from a GPR (mtcrf).
  • A specified CR field can be set by a move to the CR from another CR field (e_mcrf).
  • CR field 0 can be set as the implicit result of an integer instruction.
  • A specified CR field can be set as the result of an integer compare instruction.
  • CR field 0 can be set as the result of an integer bit test instruction. Instructions are provided to perform logical operations on individual CR bits and to test individual CR bits. CR settings for integer instructions For all integer word instructions in which the Rc bit is defined and set, and for addic., the first three bits of CR field 0 (CR[32–34]) are se t by signed comparison of bits 32–63 of the result to zero, and the fourth bit of CR field 0 (CR[35]) is copied from the final state of XER[SO]. if (target_register) 32:63 < 0 then c ← 0b100 else if (target_register)32:63 > 0 then c ← 0b010 else c ← 0b001 CR0 ← c || XERSO CRn[2] 4 * cr0 + eq (or eq) 4 * cr1 + eq 4 * cr2 + eq 4 * cr3+ eq 4 * cr4 + eq 4 * cr5 + eq 4 * cr6 + eq 4 * cr7 + eq 000 001 010 011 100 101 110 111 Equal or floating-point equal (EQ, FE). For integer compare instructions: rA = SIMM, UIMM, or rB. For floating-point compare instructions: frA = frB. CRn[3] 4 * cr0 + so/un (or so/un) 4 * cr1 + so/un 4 * cr2 + so/un 4 * cr3 + so/un 4 * cr4 + so/un 4 * cr5 + so/un 4 * cr6 + so/un 4 * cr7 + so/un 000 001 010 011 100 101 110 111 Summary overflow or floating-point unordered (SO, FU). For integer compare instructions, this is a copy of XER[SO] at instruction completion. For floating-point compare instructions, one or both of frA and frB is a NaN.

Table 15. CR setting for compare instructions (continued)

is undefined. The bits of CR field 0 are interpreted as shown in Table 16.

2.5.2 Link register (LR)

to LR (bclrx) instruction, and it holds the return address after branch and link instructions. Table 16. CR0 encodings 32 Negative (LT). Bit 32 of the result is equal to 1. 34 Zero (EQ). Bits 32–63 of the result are equal to 0. Table 17. Condition register setting for compare instructions For unsigned-integer compare, GPR(rA or rX) <u SCI8 or UI or UI5 or GPR(rB or rY). For unsigned-integer compare, GPR(rA or rX) >u SCI8 or UI or UI5 or GPR(rB or rY). XER[SO] at the completion of the instruction.

the LR using mtspr. LR[62–63] are ignored by bclr instructions. subset of all variants of Book E conditional branches involving the LR, as shown in Table 18. LR[30] is examined when the LR holds an instruction address.

2.5.3 Count register (CTR)

branch target address for a Branch Conditional to CTR (bcctrx) instruction. CTR[30] is examined when the CTR holds an instruction address. Table 18. Branch to link register instruction comparison

2.6 Processor control registers

This section addresses machine state, processor ID, and processor version registers.

2.6.1 Machine stat e register (MSR)

processor is in supervisor or user mode). SRR1. If a critical interrupt is taken, MSR contents are automatically copied into CSRR1. When an rfi or rfci is executed, MSR contents are restored from SRR1 or CSRR1. used to set or clear MSR[EE] without affecting other MSR bits. Table 19. Branch to count register instruction comparison

Table 20. MSR field descriptions

37 UCLE

cache-line locking by the operating system. to manage and track the locking/unlocking of cache lines by user-mode tasks.

38 SPE

1Software can execute any of the SPE APU instructions. SPE APU unavailable interrupt. 0The processor cannot execute APU instructions. 1The processor can execute APU instructions. management request is signaled to external logic. processor behaves in the wait state are implementation-dependent. 0Critical input and watchdog timer interrupts are disabled. 1Critical input and watchdog timer interrupts are enabled. any resource (for example, GPRs, SPRs, and the MSR). cannot access any privileged resource. PR also affects memory access control.

The floating-point exception mode bits FE0 and FE1 are described in Table 21. 1The processor can execute floating-point instructions. 0Machine check interrupts are disabled. 1Machine check interrupts are enabled. 53 — Allocated for implementation-dependent use. 0Debug interrupts are disabled. 1Debug interrupts are enabled if DBCR0[IDM] = 1. See the description of the DBSR[UDE] in Chapter 2.13.2. 56 — Reserved, should be cleared.

61 PMM

the state for which monitoring is enabled, counting is enabled. 63 — Preserved for Book III RI and LE, respectively.

  1. An MSR bit that is reserved may be alter ed by return from interrupt instructions.

Table 20. MSR field descriptions (continued)

2.7 Hardware implementation-dependent registers

that are defined by the EIS.

2 An integrated device may not use all HID fields implemented on an embedded core or may

descriptions in the reference manual for the integrated device.

2.7.1 Hardware implementatio n dependent register 0 (HID0)

compliant device implement all HID0 fields; see the user documentation. HID0 fields are described in Table 22. Table 21. Floating-point exception bits—MSR[FE0,FE1]

11 P r e c i s e

Table 22. HID0 field descriptions

32 EMCP

delivered to the core from the machine check input. 0Machine check exceptions from the machine check signal are disabled. 33 — Implementation dependent.

34 SFR

they would have when the processor is in 64-bit mode. 39 — Implementation dependent.

43 DPM

0Dynamic power management is disabled. 1Dynamic power management is enabled.

44 EDPM

have adverse effects on performance. 0Enhanced dynamic power management is disabled. 1Enhanced dynamic power management is enabled. 45 — Implementation dependent.

46 ICR

critical input interrupts cause an established reservation to be cleared. 0External and critical input interrupts do not affect reservation status.

47 EN_MAS7_UP

physical addressing but have since extended past 32 bits. 0Hardware updates of MAS7 are disabled. 1Hardware updates of MAS7 are enabled.

48 EIEC

errors cause a machine check exception. do not generate a machine check interrupt. generate a machine check interrupt.

49 TBEN

Time base enable. Used to control whether the time base increments. 0 The time base is not enabled and will not increment. base increments is determined by the value of HID0[SEL_TBCLK].

50 SEL_TBCLK

Select time base clock. Used to select the source of the time base clock. 0 The time base is updated based on a core implementation specific rate.

1 The time base is updated based on an external signal to the core

54 — Implementation dependent.

55 DAPUEN

save state and the rfci instruction to return from the debug interrupt.

1 The debug APU is enabled; debug interrupts use DSRR0 and DSRR1 to

save state and the rfdi instruction to return from the debug interrupt.

56 SGE

are gathered is implementation dependent. 0 Store gathering is disabled. 1 Store gathering is enabled. 57 — Implementation dependent.

58 EIEIO_EN

synchronization semantics as the eieio instruction. 0 Synchronization provided by mbar is performed in the Book E manner. 1 Synchronization provided by mbar is equivalent to eieio synchronization.

59 LWSYNC_EN

0 The synchronization provided by the msync instruction is performed in

1 The synchronization provided by the msync instruction is based on the L

field defined in PowerPC 2.xx architecture sync instruction. 60 — Implementation dependent.

2.7.2 Hardware implementatio n dependent register 1 (HID1)

distinguish the processor from other processors in the system.

61 NOPTST

for store instructions perform no operation. instructions defined in the cache line locking APU operate as defined.

62 NOPDST

stream prefetching through the dst instructions produce no-operation. prefetch streams are terminated.

63 NOPTI

instructions perform no operations. defined by the EIS and Book E unless disabled by NOPDST or NOPTST. locking APU operate as defined.

2.7.4 Processor versi on register (PVR)

attributes that may affect software. Table 23 describes PVR fields.

2.7.5 System version register (SVR)

documentation for the implementation.

2.8 Timer registers

Table 23. PVR field descriptions facilities and instructions are supported. A 16-bit number that distinguishes between implementations of the version. the same version number, such as clock rate and engineering change level.

Figure 5. Relationship of timer facilities to the time base

  • The TB is a long-period counter driven at an implementation-dependent frequency.
  • The decrementer, updated at the same rate as the TB, provides a way to signal an exception after a specified period unless one of the following occurs: – DEC is altered by software in the interim. – The TB update frequency changes.
  • The DEC is typically used as a general-purpose software timer.
  • The time base for the TB and DEC is selected by the time base enable (TBEN) and select time base clock (SEL_TBCLK) bits in HID0, as follows: – If HID0[TBEN] = 1 and HID0[SEL_TBCLK] = 0, the time base is updated every 8 bus clocks. – If HID0[TBEN] = 1 and HID0[SEL_TBCLK] = 1, the time base is updated by an implementation-specific clock input).
  • Software can select one from of four TB bits to signal a fixed-interval interrupt whenever the bit transitions from 0 to 1. It is typically used to trigger periodic system maintenance functions. Bits that may be selected are implementation-dependent.
  • The watchdog timer, also a selected TB bit, provides a way to signal a critical exception when the selected bit transitions from 0 to 1. It is typically used for system error recovery. If software does not respond in time to the initial interrupt by clearing the associated status bits in the TSR before the next expiration of the watchdog timer interval, a watchdog timer-generated processor reset may result, if so enabled. All timer facilities must be initialized during start-up.

2.8.1 Timer control register (TCR)

complex implements two fields not specified in Book E: TCR[WPEXT] and TCR[FPEXT].

Table 24 describes the TCR fields. Table 24. TCR field descriptions by software, it cannot be cleared by software (except by a software-induced reset). Once written to a non-zero value, WRC may no longer be altered by software. software, but cannot be cleared by software (except by a software-induced reset). xx Other values: Force processor to be reset on second time-out of watchdog timer. The exact function of any of these settings is implementation-dependent.

36 WIE Watchdog timer interrupt enable

37 DIE Decrementer interrupt enable

0 Decrementer interrupts disabled

1 Decrementer interrupts enabled

40 FIE Fixed interval interrupt enable

0 Fixed interval interrupts disabled

1 Fixed interval interrupts enabled

when the DEC value reaches 0000_0001.

0 Auto-reload disabled

1 Auto-reload enabled

42 — Reserved, should be cleared.

2.8.2 Timer status register (TSR)

watchdog timer-initiated processor reset. All TSR bits function as write-1-to-clear. Note: Register fields designated as write-1-to-clear are cleared only by writing ones to them. Writing zeros to them has no effect. Table 25 describes TSR fields. — Reserved, should be cleared. Table 24. TCR field descriptions (continued)

336 Access: supervisor w1c

Table 25. TSR field descriptions

32 ENW

0 Action on next watchdog timer time-out is to set TSR[ENW]. 1 Action on next watchdog timer time-out is governed by TSR[WIS].

33 WIS

0 A watchdog timer event has not occurred. caused by the watchdog timer. 00 No watchdog timer reset has occurred. xx All other values are implementation-dependent.

2.8.3 Time base (TBU and TBL)

functions for the system. TB is a volatile resource and must be initialized during start-up. any other indication) when this occurs.

  • Loading a GPR from the TB has no effect on the accuracy of the TB.
  • Storing a GPR to the TB replaces the value in the TB with the value in the GPR. Book E does not specify a relationship between the frequency at which the TB is updated and other frequencies, such as the CPU clock or bus clock in a Book E system. The TB

36 DIS

Decrementer interrupt status. 0 A decrementer event has not occurred.

37 FIS

Fixed-interval timer interrupt status. 0 A fixed-interval timer event has not occurred. 38–63 — Reserved, should be cleared. Table 25. TSR field descriptions (continued)

update frequency is not required to be constant. One of the following is required to ensure that system software can keep time of day and operate interval timers:

  • The system provides an (implementation-dependent) interrupt to software whenever the update frequency of the TB changes and a way to determine the current update frequency.
  • The update frequency of the TB is under the control of system software. Note: 1 Disabling the TB or making reading the time base privileged prevents the TB from being used to implement a covert channel in a secure system.

2 If the operating system initializes the TB on power-on to some reasonable value and the

update frequency of the TB is constant, the TB can be used as a source of values that increase at a constant rate, such as for time stamps in trace entries. Even if the update frequency is not constant, values read from the TB are monotonically increasing (except when the TB wraps from 2 64 – 1 to 0). If a trace entry is recorded each time the update frequency changes, the sequence of TB values can be post-processed to become actual time values. Successive readings of the TB may return identical values. It is intended that the TB be useful for timing reasonably short sequences of code (a few hundred instructions) and for low-overhead time stamps for tracing.

2.8.4 Decrementer register

The 32-bit decrementer (DEC), shown below, is a decrementing counter that is updated at the same rate as the TB. It provides a way to signal a decrementer interrupt after a specified period unless one of the following occurs:

  • DEC is altered by software in the interim.
  • The TB update frequency changes. DEC is typically used as a general-purpose software timer. The decrementer auto-reload register is used to automatically reload a programmed value into DEC, as described in Section 2.8.5: Decrementer auto-reload register (DECAR).” Decrementer register (DEC)2.8.5 Decrementer auto-r eload register (DECAR) The decrementer auto-reload register is shown in figure below. If the auto-reload function is enabled (TCR[ARE] = 1), the auto-reload value in DECAR is written to DEC when DEC decrements from 0x0000_0001 to 0x0000_0000. Note that writing DEC with zeros by using an mtspr[DEC] does not automatically generate a decrementer exception. SPR 222 Access: Supervisor read/write 32 63 R Decrementer value W Reset All zeros

Decrementer auto-reload register (DECAR)

2.9 Interrupt registers

Chapter 2.9.1: Interrupt registers defined by book E on page 81,” describes registers used for interrupt handling.

2.9.1 Interrupt regist ers defined by book E

This section describes the following register bits and their fields:

  • Save/restore register 0 (SRR0) on page 81”
  • Save/restore register 1 (SRR1) on page 81”
  • Critical save/restore register 0 (CSRR0) on page 82”
  • Critical save/restore register 1 (CSRR1) on page 82”
  • Data exception address register (DEAR) on page 82”
  • Interrupt vector prefix register (IVPR) on page 83”
  • Interrupt vector offset registers (IVORs) on page 83”
  • Exception syndrome register (ESR) on page 84” Save/restore register 0 (SRR0) On a noncritical interrupt, SRR0, shown in figure below, holds the address of the instruction where the interrupted process should resume. The instruction is interrupt-specific, although for instruction-caused exceptions, it is typically the address of the instruction that caused the interrupt. When rfi executes, instruction execution continues at the address in SRR0. Save/restore register 0 (SRR0) Save/restore register 1 (SRR1) SRR1 is provided to save and restore machine state on noncritical interrupts. When a noncritical interrupt is taken, MSR contents are placed in SRR1. When rfi executes, SRR1 contents are placed into MSR. SRR1 bits that correspond to reserved MSR bits are also reserved. These registers are not affected by rfci or rfmci. Reserved MSR bits may be altered by rfi, rfci, or rfmci. SPR 544 Access: supervisor write-only 32 63 R W Decrementer auto-reload value Reset All zeros SPR 2626 Access: sup[ervisor read/write 32 63 R Next instruction address W Reset All zeros

Save/restore register 1 (SRR1) Critical save/restore register 0 (CSRR0) CSRR0, is provided to save and restore machine state on critical interrupts. It is used by critical interrupts like SRR0 is used for standard interrupts: to hold the address of the instruction to which control is passed at the end of the interrupt handler. When rfci executes, instruction execution continues at the address in CSRR0. Critical save/restore register 0 (CSRR0)Critical save/restore register 1 (CSRR1) CSRR1, is used to save and restore machine state on critical interrupts. When a critical interrupt is taken, MSR contents are placed into CSRR1. When rfci executes, CSRR1 contents are restored into the MSR. CSRR1 bits that correspond to reserved MSR bits are also reserved; reserved MSR bits may be altered. Critical save/restore register 1 (CSRR1)Data exception address register (DEAR) DEAR, is loaded with the effective address of a data access (caused by a load, store, or cache management instruction) that results in an alignment, data TLB miss, or DSI exception. Data exception address register (DEAR) SPR 277 Access: supervisor read/write 32 63 R MSR state information W Reset All zeros SPR 587 Access: supervisor read/write 32 63 R Next instruction address W Reset All zeros SPR 597 Access: supervisor read/write 32 63 R MSR state information W Reset All zeros SPR 617 Access: supervisor read/write 32 63 R Exception address W Reset All zeros

processing routine. IVPR[48–63] are reserved. implementation-dependent use. IVOR assignments are shown in Table 26. Table 26. IVOR assignments

need to be cleared by software. Table 27 shows ESR bit definitions. The ESR is defined in Book E. Bits architected by EIS storage are defined here. Table 27 describes ESR bit definitions. Table 26. IVOR assignments (continued) Table 27. Exception syndrome register (ESR) definition

36 PIL Illegal instruction exception Program

37 PPR Privileged instruction exception Program

38 PTR Trap exception Program

42 DLK

Defined by cache line locking APU. Instruction cache locking attempt. executed in user mode (MSR[PR] = 1) while MSR[UCLE] = 0.

1 DSI occurred on an attempt to lock line in data cache when

43 ILK

Defined by cache line locking APU. Instruction cache locking attempt. user mode (MSR[PR] = 1) while MSR[UCLE] = 0.

1 DSI occurred on an attempt to lock line in instruction cache when

44 APU

56 SPE

1 Any exception caused by an SPE/embedded floating-point

Table 27. Exception syndrome register (ESR) definition (continued)

protection bits to determine if a protection violation also occurred. This section describes machine check save/store and syndrome registers.

58 VLEMI

with execution or attempted execution of a VLE instruction.

0 The instruction page associated with the instruction causing the

1 The instruction page associated with the instruction causing the

62 MIF

instruction caused an instruction TLB error. instruction caused an instruction TLB error.

63 XTE

contain the address of the instruction that initiated the transaction. 0 Default. No external transaction error was precisely detected.

1 An external transaction reported an error that was precisely

Debug save/restore register 0 (DSRR0) Debug Save/restore register 1 (DSRR1) DSRR1, is provided to save and restore machine state on debug interrupts. When a debug interrupt is taken, MSR contents are placed into DSRR1. When rfdi executes, the contents of DSRR1 are restored into MSR. DSRR1 bits that correspond to reserved MSR bits are also reserved. (See Section 2.6.1: Machine state register (MSR),” for more information.) DSRR0 and DSRR1 are not affected by rfi or rfci. Reserved MSR bits may be altered by rfi, rfci, or rfdi. Debug save/restore register 1 (DSRR1) Machine check save/restore register 0 (MCSRR0) When a machine check interrupt is taken, MCSRR0, is set to the address of the instruction where the interrupted process should resume. The instruction is interrupt-specific, although typically MCSRR0 holds address of the instruction that caused the interrupt. When rfmci is executed, instruction execution continues at this address. Machine check save/restore register 0 (MCSRR0) Machine check save/restore register 1 (MCSRR1) MCSRR1 is used to save and restore machine state on machine check interrupts. When a machine check interrupt is taken, MSR contents are placed into MCSRR1. When rfmci executes, MCSRR1 contents are restored to MSR. MCSRR1 bits that correspond to reserved MSR bits are also reserved; reserved MSR bits may be altered. SPR 574574 Access: Supervisor read/write 32 63 R Next instruction address W Reset Undefined SPR 575574 Access: Supervisor read/write 32 63 R MSR state information W Reset Implementation-specific SPR 570574 Access: Supervisor read/write 32 63 R Next instruction address W Reset All zeros

Machine check save/restore register 1 (MCSRR1) Machine check address register (MCAR/MCARU) When the core complex takes a machine check interrupt, it updates MCAR, to indicate the address of the data associated with the machine check. Note that if a machine check interrupt is caused by a signal, MCAR contents are not meaningful. Errors that cause MCAR contents to be updated are implementation-dependent. If MCSR[MAV] = 1, the address is an effective address; if MAV = 0, the address is a real address. Machine check address register (MCAR/MCARU) For 32-bit implementations that support physical addresses greater than 32 bits, MCARU provides an alias to the upper address bits that reside in MCAR[0–31]. Machine check syndrome register (MCSR) The MCSR, is used to record the cause of the machine check interrupt. In general, machine check syndrome bits correlating to specific hardware error conditions are implementation dependent. Consult the users manual for a complete definition of machine check error syndromes for a specific processor. Machine check syndrome register 1 (MCSR) Table 28 describes the MCSR fields. SPR 571574 Access: Supervisor read/write 32 63 R MSR state information W Reset All zeros SPR MCAR: 573 MCARU: 569 Access: Supervisor read-only MCARU 32 6332 63 R Machine check address 0–31 Machine check address 32–63 W Reset All zeros SPR 572574 Access: Read/w1c 32 43 44 45 46 47 63 R MCP — NMI MAV MEA — W Reset All zeros

MCSR bits causes an asynchronous machine check interrupt when MSR[ME] is set.

2.10 Software use sprs (S PRG0–SPRG7 and USPRG0)

  • SPRG0–SPRG2—can be accessed only in supervisor mode.
  • SPRG3—can be written only in supervisor mode. It is readable in supervisor mode, but whether it can be read in user mode is implementation-dependent.
  • SPRG4–SPRG7—can be written only in supervisor mode; readable in supervisor or user mode.
  • USPRG0—can be accessed in supervisor or user mode.

Table 28. MCSR field descriptions

32 MCP

dependent and may be tied to a an external pin on the IC package. 42 — Implementation-dependent.

44 MAV

not placed in MCAR unless MAV is 0 when the error is logged. 0 The address in MCAR is not valid. 1 The address in MCAR is valid.

45 MEA

0 The address in MCAR is a physical address. 1 The address in MCAR is an effective address (untranslated). 63 — Implementation-dependent.

Software-use sprs (SPRG0–SPRG7 and USPRG0) Software-use SPRs are read into a GPR by using mfspr and are written by using mtspr.

2.11 L1 cache registers

The EIS defines registers that provide control and configuration and status information for the L1 cache implementation.

2.11.1 L1 cache control and stat us register 0 (L1CSR0)

The L1CSR0, is defined by the EIS. It is used for general control and status of the L1 data cache. L1 cache control and status register 0 (L1CSR0) Table 29 describes the L1CSR0 fields. SPR SPRG0 SPRG1 SPRG2 SPRG3 SPRG4 SPRG5 SPRG6 SPRG7 USPRG0 272 273 274 259 275 260 276 261 277 262 278 263 279 256 Read/write Read/write Read/write Read-only Read/write Read-only Read/write Read-only Read/write Read-only Read/write Read-only Read/write Read/write Supervisor Supervisor Supervisor User (Implementation-dependent)/supervisor Supervisor User/supervisor Supervisor User/supervisor Supervisor User/supervisor Supervisor User/supervisor Supervisor User/supervisor R MSR state information W Reset All zeros SPR 1010 Supervisor read/write Cache way partitioning APU Bits 32 35 36 39 40 41 42 43 46 47 R WID WDD AWID AWDD WAM — CPE W Reset All zeros Cache Line Locking APU Bits 48 49 51 52 53 54 55 56 57 60 61 62 63 R CPI — CSLC CUL CLO CLFR CLOA — CABT CFI CE W Reset All zeros

Table 29. L1CSR0 field descriptions 0 The corresponding way is available for replacement by instruction miss line refills. 1 The corresponding way is not available for replacement by instruction miss line refills. 0 The corresponding way is available for replacement by data miss line refills.

1 The corresponding way is not available for replacement by data miss line refills

40 AWID

Cache way partitioning APU. Additional ways instruction disable. 0 Additional ways beyond 0–3 are available for replacement by instruction miss line fills. 1 Additional ways beyond 0–3 are not available for replacement by instruction miss line fills.

41 AWDD

Cache way partitioning APU. Additional ways data disable. 0 Additional ways beyond 0–3 are available for replacement by data miss line fills. 1 Additional ways beyond 0–3 are not available for replacement by data miss line fills.

42 WAM

Cache way partitioning APU. Way access mode. 0 All ways are available for access. 1 Only ways partitioned for the specific type of access are used for a fetch or read operation. 43-46 — Reserved for implementation dependent use.

47 CPE

0 Parity checking of the cache disabled

1 Parity checking of the cache enabled

48 CPI

[Data] Cache parity error injection enable.

0 Parity error injection disabled

0 cause the bit not to be set (that is, L1CSR0[CPI] = L1CSR0[CPE] & L1CSR0[CPI]). 49–51 — Reserved, should be cleared.

52 CSLC

whenever the line is invalidated. This bit can be cleared only by software. 0 The cache has not encountered a snoop that invalidated a locked line. 1 The cache has encountered a snoop that invalidated a locked line.

53 CUL

[Data]Cache unable to lock. Sticky bit set by hardware. This bit can be cleared only by software.

0 Indicates a lock set instructi on was effective in the cache

1 Indicates a lock set instruction was not effective in the cache

54 CLO

[Data]Cache lock overflow. Sticky bit set by hardware. This bit can be cleared only by software.

0 Indicates a lock overflow condition was not encountered in the cache

1 Indicates a lock overflow condition was encountered in the cache

2.11.2 L1 cache control and stat us register 1 (L1CSR1)

Table 30 describes the L1CSR1 fields.

55 CLFC

[Data]Cache lock bits flash clear. Clearing occurs regardless of the enable (L1CSR0[CE]) value.

56 CLOA

line when a lock overflow situation exists. Implementation of this bit is optional.

0 Indicates a lock overflow condition does not replace an existing locked line with the

1 Indicates a lock overflow condition replaces an existing locked line with the requested line

57–60 — Reserved, should be cleared.

61 CABT

[Data]Cache operation aborted.

0 No cache operation completed improperly

1 Cache operation did not complete properly

62 CFI

[Data]Cache flash invalidate. Invalidation occurs regardless of the enable (L1CSR0[CE]) value. 1 Cache flash invalidate operation. A cache invalidation operation is initiated by hardware. Once complete, this bit is cleared. During an invalidation operation, writing a 1 causes undefined results; writing a 0 has no effect. Table 29. L1CSR0 field descriptions (continued)

Table 30. L1CSR1 field descriptions 32–42 — Reserved, should be cleared. 43-46 — Reserved for implementation dependent use.

47 ICPE

48 ICPI

Instruction cache parity error injection enable. not to be set (that is, L1CSR0[ICPI] = L1CSR0[ICPE] & L1CSR0[ICPI]). 49–51 — Reserved, should be cleared.

52 ICSLC

line is cleared whenever the line is invalidated. This bit can be cleared only by software. 0 The cache has not encountered a snoop that invalidated a locked line. 1 The cache has encountered a snoop that invalidated a locked line.

53 ICUL

be cleared only by software.

0 Indicates a lock set instruction was effective in the cache

1 Indicates a lock set instructio n was not effective in the cache

54 ICLO

55 ICLFC

During a flash clear operation, writing a 1 causes undefined results; writing a 0 has no effect.

56 ICLOA

0 Indicates a lock overflow condition replaces an existing locked line with the requested line

1 Indicates a lock overflow condition does not replace an existing locked line with the requested

57–60 — Reserved, should be cleared.

61 ICABT

Instruction cache operation aborted.

2.11.3 L1 cache configurati on register 0 (L1CFG0)

unified cache, L1CFG0 applies to the unified cache and L1CFG1 is not implemented.

62 ICFI

complete, this bit is cleared. During an invalidation operation, writing a 1 causes undefined results; writing a 0 has no effect.

63 ICE

Table 30. L1CSR1 field descriptions (continued) Table 31. L1CFG0 field descriptions

00 Harvard

01 Unified

34 CWPA

Cache way partitioning APU available.

0 Unavailable

1 Available

35 CFAHA

36 CFISWA

37–38 — Reserved, should be cleared.

2.11.4 L1 cache configurati on register 1 (L1CFG1)

00 True LRU

01 Pseudo LRU

43 CLA

44 CPA

45–52 CNWAY Cache number of ways minus 1. 53–63 CSIZE Cache size in Kbytes. Table 31. L1CFG0 field descriptions (continued) Table 32. L1CFG1 field descriptions 32–38 — Reserved, should be cleared.

2.11.5 L1 flush and invalidate c ontrol register 0 (L1FINV0)

match in the cache is required.

43 ICLA

44 ICPA

45–52 ICNWAY Cache number of ways minus 1. 53–63 ICSIZE Cache size in Kbytes. Table 32. L1CFG1 field descriptions (continued) Table 33. L1FINV0 fields—L1 direct cache flush 0–31 — Reserved, should be cleared. 32–39 CWAY Cache way. Specifies the cache way to be selected. 40–41 — Reserved, should be cleared. 42–58 CSET Cache set. Specifies the cache set to be selected.

2.12 MMU registers

  • Process ID registers (PID0–PID2)
  • MMU control and status register 0 (MMUCSR0)
  • MMU configuration register (MMUCFG)
  • TLB configuration registers (TLBnCFG)
  • MMU assist registers (MAS0–MAS7)

2.12.1 Process ID re gisters (PID0–PIDn)

effective address (instruction or data) generated by the processor. processors may not implement all 14 bits of the process ID field. determine if other PID registers are implemented. mapped at the same virtual address in each process. 59–61 — Reserved, should be cleared. should be synonymous with a dcbi instruction that references the same line. a dcbst instruction that references the same line. be synonymous with a dcbf instruction that references that line.

2.12.2 MMU control and stat us register 0 (MMUCSR0)

The MMUCSR0 register is used for general control of the L1 and L2 MMUs. Table 34. MMUCSR0 field descriptions 32–60 — Reserved, should be cleared.

61 L2TLB0_FI

0 No flash invalidate. Writing a 0 to this bit during an invalidation operation is ignored.

62 L2TLB1_FI

0 No flash invalidate. Writing a 0 to this bit during an invalidation operation is ignored. an undefined operation. This invalidation typically takes 1 cycle. 63 — Reserved, should be cleared.

2.12.3 MMU configurati on register (MMUCFG)

MMUCFG, shown below, gives configuration information about the implementation’s MMU. Table 35. MMUCFG field descriptions 32–48 — Reserved, should be cleared. PIDSIZE+1 bits in the PID registers. 58–59 — Reserved, should be cleared. of the MMU implemented by the processor.

01 Reserved

10 Reserved

11 Reserved

2.12.4 TLB configuration registers (TLB nCFG)

Table 36. TLB nCFG field descriptions Associativity of TLBn. Number of ways of associativity of TLB array.

0001 Indicates smallest page size is 4 Kbytes

0002 Indicates smallest page size is 8 Kbytes

0001 Indicates maximum page size is 4 Kbytes

0002 Indicates maximum page size is 8 Kbytes

48 IPROT

Invalidate protect capability of TLBn array. 0 Indicates invalidate protection capability not supported. 1 Indicates invalidate protection capability supported.

49 AVAIL

Page size availability of TLBn array.

0 Fixed selectable page size from MINSIZE to MAXSIZE (all TLB entries are

1 Variable page size from MINSIZE to MAXSIZE (each TLB entry can be sized

50–51 — Reserved, should be cleared.

2.12.5 MMU assist registers (MAS0–MAS7)

TLBs. Note that some fields in these registers are redefined by implementations. MAS0, is used for MMU read/write and replacement control. Table 37. MAS0 field descriptions 32–33 — Reserved, should be cleared.

00 TLB0

01 TLB1

10 TLB2

11 TLB3

also updated on TLB error exceptions (misses) and tlbsx hit and miss cases. 48–51 — Reserved, should be cleared. the TLB selected by MAS0[TLBSEL] does not support NV, this field is undefined.

Below is the format of MAS1. Table 38. MAS1 field descriptions—descriptor context and configuration control 0 This TLB entry is invalid.

33 IPROT

such in the TLB configuration registers. 0 Entry is not protected from invalidation. 1 Entry is protected from invalidation. 34–35 — Reserved, should be cleared. matches with all process IDs. 48–50 — Reserved, should be cleared. MSR[DS], depending on the type of access) to select a TLB entry. supports all 16 page sizes defined in Book E. 56–63 — Reserved, should be cleared.

Table 39. MAS2 field descriptions—EPN and page attributes accessible only in 64-bit implementations as the upper 32 bits of the logical address of the page. 52–55 — Reserved, should be cleared. coherency domain (or protocol) used. ACM values are implementation dependent. Note: Some previous implementations may have a storage bit in the bit 57 position labeled as X0.

58 VLE

0 Instructions fetched from the page are decoded and executed as PowerPC (and associated EIS

1 Instructions fetched from the page are decoded and executed as VLE (and associated EIS APUs)

instructions.Implementation-dependent page attribute. 0 This page is cons idered write-back with respect to the caches in the system. 1 All stores performed to this page are written through the caches to main memory. 0 Accesses to this page are considered cacheable. the memory element specified by the operation. 0 Memory coherence is not required. participating in the coherence protocol.

0 Accesses to this page are not guarded and can be performed before it is known if they are

required by the sequential execution model.

1 Loads and stores to this page are performed without speculation (that is, they are known to be

previous devices that implement the PowerPC architecture. 0 The page is accessed in big-endian byte order. 1 The page is accessed in true little-endian byte order. Table 39. MAS2 field descriptions—EPN and page attributes (continued) Table 40. MAS3 field descriptions–RPN and access control 32 bits, RPN[0–31] are accessed through MAS7. 52–53 — Reserved, should be cleared. scanning algorithm or be used to mark more abstract page attributes.

The MAS4 fields are described in Table 41. Table 41. MAS4 field descriptions—hardware replacement assist configuration 32–33 — Reserved, should be cleared. 36–43 — Reserved, should be cleared. should be used to load MAS1[TID] on a TLB miss exception. 48–51 — Reserved, should be cleared.

searching TLB entries with the tlbsx instruction. entries with the tlbsx instruction. Table 42. MAS5 field descriptions—extended search pIDs 32–33 — Reserved, should be cleared. number of bits implemented for PID registers. 48–49 — Reserved, should be cleared. number of bits implemented for PID registers. Table 43. MAS 6 field descriptions—search pids and search AS 32–33 — Reserved, should be cleared. implemented for PID registers. 48 — Reserved, should be cleared.

support more than 32 bits of physical address.

2.13 Debug registers

software, and not by general application or operating system code. only the number of bits implemented for PID registers. executing tlbsx to search the TLB. Table 44. MAS 7 field descriptions—high order RPN 32–63 RPN[0–31] Real page number (bits 0–31). RPN[32–63] are accessed through MAS3. Table 43. MAS 6 field descriptions—search pids and search AS (continued)

2.13.1 Debug control r egisters (DBCR0–DBCR3)

timer operation during debug events, and set the debug mode of the processor. Table 45. DBCR0 field descriptions

32 EDM

0 The processor is not in external debug mode.

33 IDM

1 If MSR[DE] = 1, the occurrence of a debug event or the recording of an earlier

1x A hard reset is performed on the processor.

36 ICMP

0 ICMP debug events are disabled. 1 ICMP debug events are enabled. Note: Instruction completion does not cause an ICMP debug event if MSR[DE]=0.

37 BRT

0 BRT debug events are disabled. 1 BRT debug events are enabled. Note: Taken branches do not cause a BRT debug event if MSR[DE]=0.

38 IRPT

Interrupt taken debug event enable. 0 IRPT debug events are disabled.

1 IRPT debug events are enabled

39 TRAP

0 TRAP debug events cannot occur. 1 TRAP debug events can occur.

40 IAC1

0 IAC1 debug events cannot occur. 1 IAC1 debug events can occur.

41 IAC2

Instruction address compare 2 debug event enable. 0 IAC2 debug events cannot occur. 1 IAC2 debug events can occur.

42 IAC3

0 IAC3 debug events cannot occur. 1 IAC3 debug events can occur.

43 IAC4

0 IAC4 debug events cannot occur. 1 IAC4 debug events can occur. 00 DAC1 debug events cannot occur. 01 DAC1 debug events can occur only if a store-type data storage access. 10 DAC1 debug events can occur only if a load-type data storage access. 11 DAC1 debug events can occur on any data storage access. 00 DAC2 debug events cannot occur. 01 DAC2 debug events can occur only if a store-type data storage access. 10 DAC2 debug events can occur only if a load-type data storage access. 11 DAC2 debug events can occur on any data storage access.

48 RET

0 RET debug events cannot occur. 1 RET debug events can occur. 49–56 — Reserved, should be cleared. Table 45. DBCR0 field d escriptions (continued)

Table 46 provides bit definitions for the DBCR1.

57 CIRPT

uses the critical class, that is, uses CSRR0 and CSRR1) occurs. 0 Critical interrupt taken debug events are disabled. 1 Critical interrupt taken debug events are enabled.

58 CRET

instruction is executed) occurs. 0 Critical interrupt return debug events are disabled. 1 Critical interrupt return debug events are enabled.

59 VLES

0 CRET debug events are disabled.

1 An ICMP , BRT, TRAP , RET, CRET, IAC, or DAC debug event occurred on a

0 Enable clocking of timers. 1 Disable clocking of timers if any DBSR bit is set (except MRR). Table 46. DBCR1 field descriptions 00 IAC1 debug events can occur. 10 IAC1 debug events can occur only if MSR[PR]=0. 11 IAC1 debug events can occur only if MSR[PR]=1.

00 IAC1 debug events are based on effective addresses. 01 IAC1 debug events are based on real addresses.

10 IAC1 debug events are based on effective addresses and can occur only if

11 IAC1 debug events are based on effective addresses and can occur only if

00 IAC2 debug events can occur. 10 IAC2 debug events can occur only if MSR[PR]=0. 11 IAC2 debug events can occur only if MSR[PR]=1. 00 IAC2 debug events are based on effective addresses. 01 IAC2 debug events are based on real addresses.

10 IAC2 debug events are based on effective addresses and can occur only if

11 IAC2 debug events are based on effective addresses and can occur only if

the instruction fetch address equals the value in IAC2. in IAC1, also ANDed with the contents of IAC2. If IAC1US≠IAC2US or IAC1ER≠IAC2ER, results are boundedly undefined. If IAC1US≠IAC2US or IAC1ER≠IAC2ER, results are boundedly undefined. If IAC1US≠IAC2US or IAC1ER≠IAC2ER, results are boundedly undefined. 42–47 — Reserved, should be cleared. 00 IAC3 debug events can occur. 10 IAC3 debug events can occur only if MSR[PR]=0. 11 IAC3 debug events can occur only if MSR[PR]=1. Table 46. DBCR1 field d escriptions (continued)

00 IAC3 debug events are based on effective addresses. 01 IAC3 debug events are based on real addresses.

10 IAC3 debug events are based on effective addresses and can occur only if

11 IAC3 debug events are based on effective addresses and can occur only if

00 IAC4 debug events can occur. 10 IAC4 debug events can occur only if MSR[PR]=0. 11 IAC4 debug events can occur only if MSR[PR]=1. 00 IAC4 debug events are based on effective addresses. 01 IAC4 debug events are based on real addresses.

10 IAC4 debug events are based on effective addresses and can occur only if

11 IAC4 debug events are based on effective addresses and can occur only if

the instruction fetch address equals the value in IAC4. in IAC3, also ANDed with the contents of IAC4. If IAC3US≠IAC4US or IAC3ER≠IAC4ER, results are boundedly undefined. If IAC3US≠IAC4US or IAC3ER≠IAC4ER, results are boundedly undefined. If IAC3US≠IAC4US or IAC3ER≠IAC4ER, results are boundedly undefined. 58–63 — Reserved, should be cleared.

Table 47. DBCR2 field descriptions 00 DAC1 debug events can occur. 10 DAC1 debug events can occur only if MSR[PR]=0. 11 DAC1 debug events can occur only if MSR[PR]=1. 00 DAC1 debug events are based on effective addresses. 01 DAC1 debug events are based on real addresses.

10 DAC1 debug events are based on effective addresses and can occur only if

11 DAC1 debug events are based on effective addresses and can occur only if

00 DAC2 debug events can occur. 10 DAC2 debug events can occur only if MSR[PR]=0. 11 DAC2 debug events can occur only if MSR[PR]=1. 00 DAC2 debug events are based on effective addresses. 01 DAC2 debug events are based on real addresses.

10 DAC2 debug events are based on effective addresses and can occur only if

11 DAC2 debug events are based on effective addresses and can occur only if

only if the data access address equals the value in DAC2. DAC1, also ANDed with the DAC2 contents.

42 DAC1LNK

0 No effect

whether the instruction also generated an IAC1 debug event.

43 DAC2LNK

debug events are not generated in the other compare modes. 00 DAC1 debug events can occur.

01 DAC1 debug events can occur only when all bytes in DBCR2[DVC1BE] in

10 DAC1 debug events can occur only when at least one of the bytes in

11 DAC1 debug events can occur only when all bytes in DBCR2[DVC1BE]

access match their corresponding bytes in DVC1. Table 47. DBCR2 field desc riptions (continued)

The debug APU defines the DBCR3, however its contents are implementation specific. 00 DAC2 debug events can occur.

01 DAC2 debug events can occur only when all bytes in DBCR2[DVC2BE] in

10 DAC2 debug events can occur only when at least one of the bytes in

11 DAC2 debug events can occur only when all bytes in DBCR2[DVC2BE]

access match their corresponding bytes in DVC2. corresponding bytes in DVC1. corresponding bytes in DVC2.

2.13.2 De bug status register (DBSR)

The DBSR, provides status debug events information for the most recent processor reset. by writing ones to them; writing zeros has no effect. Table 48. DBSR field descriptions respective DBSR bit to be set.

33 UDE

1 1 DBSR[UDE] is set and a debug interrupt is taken. implementation documentation. occurred and DBCR0[ICMP] = 1.

2.13.3 Instruction address co mpare registers (IAC1–IAC4)

occurred (DBCR0[DAC1]=10 or 11). event occurred (DBCR0[DAC1]=01 or 11). occurred (DBCR0[DAC2]=10 or 11). event occurred (DBCR0[DAC2] =01 or 11). 48 RET Return debug event. Set if a return debug event occurred (DBCR0[RET]=1). 49–56 — Reserved, should be cleared. uses the critical class, that is, uses CSRR0 and CSRR1) occurs. 0 No critical interrupt taken debug event has occurred. 1 A critical interrupt taken debug event occurred. instruction is executed) occurs. 0 No critical interrupt return debug event has occurred. 1 A critical interrupt return debug event occurred. 59–63 — Reserved, should be cleared. Table 48. DBSR field descriptions (continued)

combination of the IAC1 and IAC2, or to blocks of addresses specified by the combination of the IAC3 and IAC4. Because all instruction addresses are required to be word-aligned, the two low-order bits of the IACs are reserved and do not participate in the comparison to the instruction address.

2.13.4 Data address compare registers (DAC1–DAC2)

The data address compare registers (DAC1 and DAC2), are each 32 bits. A debug event may be enabled to occur upon loads, stores, or cache operations to an address specified in either DAC1 or DAC2, inside or outside a range specified by the DAC1 and DAC2, or to blocks of addresses specified by the combination of the DAC1 and DAC2. Data address compare registers (DAC1–DAC2) The contents of DAC1 or DAC2 are compared to the address generated by a data storage access instruction.

2.13.5 Data value compare registers (DVC1 and DVC2)

The data value compare registers (DVC1 and DVC2) are shown below. A DAC1R, DAC1W, DAC2R, or DAC2W debug event may be enabled to occur upon loads or stores of a specific data value specified in either or both of DVC1 and DVC2. DBCR2[DVC1M] and DBCR2[DVC1BE] control how the contents of DVC1 is compared with the value and DBCR2[DVC2M] and DBCR2[DVC2BE] control how the contents of DVC2 is compared with the value. Table 47 describes the modes provided. Data value compare registers (DVC1–DVC2)

2.14 SPE and SPFP APU registers

The SPE and SPFP include the signal processing and embedded floating-point status and control register (SPEFSCR), which is described in Chapter 2.14.1 on page 119.”, and the SPE implements a 64-bit accumulator, described in Chapter 2.14.2 on page 122.” SPR 316 (DAC1) 317 (DAC2) Access: Supervisor read/write 32 63 R Data address W Reset All zeros SPR 318 (DVC1) 319 (DVC2) Access: Supervisor read/write 32 63 R Data value W Reset All zeros

2.14.1 Signal processing, em bedded floating-point status, control register

Table 49. SPEFSCR field descriptions OVH. This is a sticky bit that remains set until it is cleared by an mtspr instruction. the upper word of the result of an SPE instruction.

34 FGH

vector floating-point instruction. Execution of a scalar floating-point instruction leaves FGH undefined.

35 FXH

is detected on the high element of a vector floating-point instruction. Execution of a scalar floating-point instruction leaves FXH undefined.

36 FINVH

A conversion to integer or fractional value overflows. Execution of a scalar floating-point instruction leaves FINVH undefined.

37 FDBZH

operand and the dividend is a finite non-zero number. Execution of a scalar floating-point instruction leaves FDBZH undefined.

38 FUNFH

Execution of a scalar floating-point instruction leaves FUNFH undefined.

39 FOVFH

vector floating-point instruction results in an overflow on the high word operation. Execution of a scalar floating-point instruction leaves FOVFH undefined. 40–41 — Reserved, should be cleared.

42 FINXS

floating-point overflow exceptions are disabled (FOVFE=0). point data interrupt occurs. FINXS remains set until it is cleared by software.

43 FINVS

44 FDBZS

FDBZH | FDBZ. FDBZS remains set until it is cleared by software.

45 FUNFS

Table 49. SPEFSCR field d escriptions (continued)

46 FOVFS

47 MODE

embedded floating-point APUs.

0 Default hardware results operating mode

48 SOV (SPE APU) Summary integer overflow low. Set when an SPE instruction sets OV. This sticky bit remains set until an mtspr writes a 0 to this bit. in the lower word of the result of an SPE instruction. instruction or any scalar floating-point instruction.

52 FINV

53 FDBZ

the low word operand and the dividend is a finite non-zero number.

54 FUNF

55 FOVF

56 — Reserved, should be cleared.

2.14.2 Accumulator (ACC)

57 FINXE

0 Exception disabled

result of a floating-point operation.

58 FINVE

instruction sets FINV or FINVH.

59 FDBZE

instruction sets FDBZ or FDBZH.

60 FUNFE

instruction sets FUNF or FUNFH.

61 FOVFE

instruction sets FOVF or FOVFH.

00 Round to Nearest

01 Round toward Zero

which rounding is indicated. which rounding is indicated.

  1. Software note: Software can detect hardware that manages this sticky bit by performing an operation on a

2.15 Alternate time base registers (ATBL and ATBU)

frequency. ATB registers are accessible in both user and supervisor mode. Like the TB implementation, ATBL is an aliased name for ATB. counter. It is accessible in both user and supervisor mode. Table 50. ATBL field descriptions 32–63 ATBCL Alternate time base counter lower. Table 51. ATBU field descriptions 32–63 ATBCU Alternate time base counter upper.

2.16 Performance monitor registers (PMRs)

The EIS defines a set of register resources used exclusively by the performance monitor. mtpmr and mfpmr, which are also defined by the EIS. Table 52 lists supervisor-level PMRs. registers in supervisor or user mode causes an illegal instruction exception. Table 52. Performance monitor registers—supervisor level Table 53. Performance monitor registers—user level (read-only)

2.16.1 Global control register 0 (PMGC0)

PMGC0 is cleared by a hard reset. Reading this register does not change its contents. Table 53. Performance monitor registers—user level (read-only) (continued) Table 54. PMGC0 field descriptions

32 FAC

maintains its current value until it is changed by software. 0 The PMCs are incremented (if permitt ed by other PM control bits). 1 The PMCs are not incremented.

33 PMIE

0 Performance monitor interrupts are disabled.

1 Performance monitor interrupts are enabled and occur when an enabled

34 FCECE

0 The PMCs can be incremented (if permitted by other PM control bits).

1 The PMCs can be incremented (if permitted by other PM control bits) only until

occurs, PMGC0[FAC] is set. It is up to software to clear FAC. 35–50 — Reserved, should be cleared.

2.16.2 User global control register 0 (UPMGC0)

The contents of PMGC0 are reflected to UPMGC0, which is read by user-level software. UPMGC0 is read with the mfpmr instruction using PMR384. transition event (the event occurs when the selected bit changes from 0 to 1).

00 TB[63] (TBL[31])

01 TB[55] (TBL[23])

10 TB[51] (TBL[19])

11 TB[47] (TBL[15])

53–54 — Reserved, should be cleared.

55 TBEE

0 Exceptions from time base transition events are disabled. 0, the interrupt cannot be taken until MSR[EE] = 1. 55–63 — Reserved, should be cleared. Table 54. PMGC0 field descriptions (continued)

2.16.3 Local control A registers (PMLCa0–PMLCa3)

corresponding PMLCb register. Table 55. PMLCa0–PMLCa3 field descriptions 0 The PMC is incremented (if perm itted by other PM control bits). 1 The PMC is not incremented.

33 FCS

0 The PMC is incremented (if perm itted by other PM control bits). 1 The PMC is not incremented if MSR[PR] = 0.

34 FCU

0 The PMC is incremented (if perm itted by other PM control bits). 1 The PMC is not incremented if MSR[PR] = 1.

35 FCM1

0 The PMC is incremented (if perm itted by other PM control bits). 1 The PMC is not incremented if MSR[PMM] = 1.

36 FCM0

0 The PMC is incremented (if perm itted by other PM control bits). 1 The PMC is not incremented if MSR[PMM] = 0.

1 Overflow conditions occur when the most-significant-bit of PMCx is equal to

38–40 — Reserved, should be cleared. 41–47 EVENT Event selector. Up to 128 events selectable. 48–63 — Reserved, should be cleared.

2.16.4 User local control A registers (UPMLCa0–UPMLCa3)

by user-level software with mfpmr using PMR numbers in Table 53.

2.16.5 Local control B registers (PMLCb0–PMLCb3)

apply to a threshold event selected for the corresponding performance monitor counter. PMLCb works with the corresponding PMLCa. Table 56. PMLCb0 –PMLCb3 field descriptions 32–52 — Reserved, should be cleared.

000 Threshold field is multiplied by 1 (PMLCbn[THRESHOLD] * 1)

001 Threshold field is multiplied by 2 (PMLCbn[THRESHOLD] * 2)

010 Threshold field is multiplied by 4 (PMLCbn[THRESHOLD] * 4)

011 Threshold field is multiplied by 8 (PMLCbn[THRESHOLD] * 8)

100 Threshold field is multiplied by 16 (PMLCbn[THRESHOLD] * 16)

101 Threshold field is multiplied by 32 (PMLCbn[THRESHOLD] * 32)

110 Threshold field is multiplied by 64 (PMLCbn[THRESHOLD] * 64)

111 Threshold field is multiplied by 128 (PMLCbn[THRESHOLD] * 128)

56–57 — Reserved, should be cleared. the threshold value is interpreted. By varying the threshold value, software can profile event characteristics. a different threshold value each time.

2.16.6 User local control B registers (UPMLCb0–UPMLCb3)

by user-level software with mfpmr using the PMR numbers in Table 53.

2.16.7 Performance monitor counter registers (PMC0–PMC3)

PMGC0[PMIE] and PMLCax[CE] are also set as appropriate. to stop counting when an enabled condition or event occurs. may be generated without an event counting having taken place. PMC registers are accessed with mtpmr and mfpmr using the PMR numbers in Table 52.

2.16.8 User performance monitor count er registers (UPMC0–UPMC3)

level software with the mfpmr instruction using the PMR numbers in Table 53. Table 57. PMC0–PMC3 field descriptions 33–63 Counter Value Indicates the number of occurrences of the specified event.

2.17 Device control registers (DCRs)

processor core and thus are not part of Book E. definitions would be provided in the implementation’s user’s manual). can be written into DCR DCRN using mtdcr DCRN,rS. If DCRs are implemented, they are described as part of the implementation documentation.

2.18 Book E SPR model

preserved, reserved, and allocated registers.

2.18.1 Invalid SPR references

System behavior when an invalid SPR is referenced depends on the privilege level.

  • If the invalid SPR is accessible in user mode (SPR[5] = 0), an illegal instruction exception is taken.
  • If the invalid SPR is accessible only in supervisor mode (SPR[5] = 1) and the core complex is in supervisor mode (MSR[PR] = 0), the results of the attempted access are boundedly undefined.
  • If the invalid SPR address is accessible only in supervisor mode (bit 5 of an SPR number = 1) and the core complex is not in supervisor mode (MSR[PR] = 1), a privilege exception is taken. These results are summarized in Table 58.

2.18.2 Synchronization requirements for SPRs

Table 58. System response to an invalid spr reference Table 59. Synchronization requirements for sprs

2.18.3 Reserved SPRs

2.18.4 Allocated SPRs

SPR numbers allocated for implementation-dependent use are 0x200–0x3FF (512–1023). Table 59. Synchronization requirements for sprs (continued) Table 60. Allocated SPRs defined by the EIS defines only this PID register and refers to it as PID rather than PID0.

512 SPEFSCR Signal processing and embedded float ing-point status and control register

515 L1CFG0 L1 cache configuration register 0

516 L1CFG1 L1 cache configuration register 1

528 IVOR32 SPE APU unavailable exception

529 IVOR33 Embedded floating-point data exception

530 IVOR34 Embedded floating-point round exception

531 IVOR35 Performance monitor Inte rrupt vector offset register

570 MCSRR0 Machine-check save/restore register 0

571 MCSRR1 Machine-check save/restore register 1

572 MCSR Machine check syndrome register

573 MCAR Machine check address register

624 MAS0 MMU assist register 0

625 MAS1 MMU assist register 1

626 MAS2 MMU assist register 2

627 MAS3 MMU assist register 3

628 MAS4 MMU assist register 4

629 MAS5 MMU assist register 5

630 MAS6 MMU assist register 6

633 PID1 Process ID register 1

634 PID2 Process ID register 2

688 TLB0CFG TLB configuration register 0

689 TLB1CFG TLB configuration register 1

944 MAS7 MMU assist register 7

1008 HID0 Hardware implementation dependent register 0

1009 HID1 Hardware implementation dependent register 1

1010 L1CSR0 L1 cache control and status register 0

1011 L1CSR1 L1 cache control and status register

1012 MMUCSR0 MMU control and status register 0

1015 MMUCFG MMU configuration register

1023 SVR System version register

  1. An update to a PID register must always be followed by an isync.

Table 60. Allocated SPRs defined by the EIS (continued)

3 Instruction model

The architecture specifications allow for different processor implementations, which may provide extensions to or deviations from the architectural descriptions. This chapter provides information about the Book E architecture and the Book E implementation standards (EIS), which defines auxiliary processing units (APUs) and other architectural extensions that define additional instructions, registers, and interrupts. For more information, see Chapter 7: Auxiliary processing units (APUs) on page 823.”

3.1 Operand conventions

This section describes operand conventions as they are represented in the Book E architecture. These conventions follow the basic descriptions in the classic PowerPC architecture with some changes in terminology. For example, distinctions between user and supervisor-level instructions are maintained, but the designations—UISA, VEA, and OEA— do not apply. Detailed descriptions are provided of conventions used for storing values in registers and memory, accessing processor registers, and representing data in these registers.

3.1.1 Data organization in memory and data transfers

Bytes in memory are numbered consecutively starting with 0. Each number is the address of the corresponding byte. Memory operands can be bytes, half words, words, or double words or, for the load/store multiple instruction type and load/store string instructions, a sequence of bytes or words. The address of a memory operand is the address of its first byte (that is, of its lowest- numbered byte). Operand length is implicit for each instruction.

3.1.2 Alignment and misaligned accesses

The operand of a single-register memory access instruction has an alignment boundary equal to its length. An operand’s address is misaligned if it is not a multiple of its width. The concept of alignment is also applied more generally to data in memory. For example, a 12-byte data item is said to be word-aligned if its address is a multiple of four. Some instructions require their memory operands to have certain alignment. In addition, alignment can affect performance. For single-register memory access instructions, the best performance is obtained when memory operands are aligned. Instructions are 32 bits (one word) long and must be word-aligned. Note, however, that the VLE extension provides both 16- and 32-bit instructions. See VLE instruction alignment and byte ordering on page 217.” Table 61 lists characteristics for memory operands for single-register memory access instructions.

3.2 Instruction set summary

  • Integer instructions—These include arit hmetic and logical instructions. See Integer instructions on page 146.”
  • Floating-point instructions—These include floating-point vector and scalar arithmetic instructions. See Embedded vector and scalar floating-point APU instructions.” Note that some implementations do not support Book E–defined floating-point instructions or registers.
  • Load and store instructions—See Load and store instructions on page 156.”
  • Flow control instructions—These incl ude branching instructions, CR logical instructions, trap instructions, and other instructions that affect the instruction flow. See Branch and flow control instructions on page 163.”
  • Processor control instructions—These instructions are used for synchronizing memory accesses. See Processor control instructions on page 201.”
  • Memory synchronization instructions—These instructions are used for memory synchronizing. See Memory synchronization instructions on page 175.”
  • Memory control instructions—These instructions provide control of caches and TLBs. See Memory control instructions,” and Supervisor-level memory control instructions.”
  • Signal processing instructions—These inclu de a set of vector arithmetic and logic instructions optimized for signal processing. See Chapter 3.6.1 on page 186. Note: Instruction groupings used here do not indicate the execution unit that processes a particular instruction or group of instructions. This information, which is useful for scheduling instructions most effectively, is provided in the execution chapter for the implementation.” Integer instructions operate on word operands. Book E floating-point instructions operate on single-precision and double-precision floating-point operands. The PowerPC architecture uses instructions that are 4 bytes long and word-aligned. It provides for byte, half-word, and word operand loads and stores between memory and a set of 32 general-purpose registers (GPRs). It provides for word and double-word operand loads and stores between memory and a set of 32 floating-point registers (FPRs). Arithmetic and logical instructions do not read or modify memory. To use the contents of a memory location in a computation and then modify the same or another location, the memory contents must be loaded into a register, modified, and then written to the target location using load and store instructions.

Table 61. Address characteristics of aligned operands

  1. An x in an address bit position indicates that the bi t can be 0 or 1 independent of the state of other bits in

The description of each instruction includes the mnemonic and a formatted list of operands. To simplify assembly language programming, a set of simplified mnemonics and symbols is provided for some of the frequently used instructions; see Appendix B: Simplified mnemonics for PowerPC instructions on page 1110,” for a complete list of simplified mnemonics. Programs written to be portable across the various assemblers for the PowerPC architecture should not assume the existence of mnemonics not described in that document.

3.2.1 Classes of instructions

Instructions belong to one of the following four classes:

  • Defined instructions (See Defined instruction class on page 135.)
  • Allocated instructions (See Allocated instruction class on page 136.)
  • Preserved instructions (See Preserved instruction class on page 137.)
  • Reserved (illegal or no-op) instructions (See Reserved instruction class on page 138.) The class is determined by examining the primary opcode and any extended opcode. If the opcode, or combination of opcode and extended opcode, is not that of a defined, allocated, preserved, or reserved instruction, the instruction is illegal. Definition of boundedly undefined If instructions are encoded with incorrectly set bits in reserved fields, the results on execution can be said to be boundedly undefined. If a user-level program executes the incorrectly coded instruction, the resulting undefined results are bounded in that a spurious change from user to supervisor state is not allowed, and the level of privilege exercised by the program in relation to memory access and other system resources cannot be exceeded. Boundedly undefined results for a given instruction can vary between implementations and between execution attempts in the same implementation. Defined instruction class This class of instructions consists of all the instructions defined in Book E. In general, defined instructions are guaranteed to be supported within a Book E system as specified by the architecture, either within the processor implementation itself or within emulation software supported by the system operating software. For implementations that only provide the 32-bit subset of Book E, emulation of the 64-bit behavior of the defined instructions is not supported. See Appendix D: Guidelines for 32-bit book E on page 1154.
  • An illegal instruction exception-type program interrupt, if an implementation does not recognize the instruction
  • An unimplemented instruction exception-type program interrupt, if the instruction is recognized but not supported by the implementation and is not a floating-point instruction
  • An unimplemented instruction exception-type program interrupt, if the instruction is recognized but not supported by the implementation, and is a floating-point instruction and floating-point processing is enabled
  • The floating-point unavailable interrupt if the instruction is recognized but is not supported by the implementation or is a floating-point instruction and floating-point processing is disabled
  • The floating-point unavailable interrupt when floating-point processing is disabled and a floating-point instruction is recognized and is not supported by the implementation
  • If an instruction is recognized and supported by the implementation, the processor performs the actions described in the rest of this document. The architected behavior may cause other exceptions. A defined instruction may be retained by future versions of Book E as a defined instruction, or may be reclassified as a preserved instruction (process of removal from the architecture) and eventually classified as reserved-illegal. Allocated instruction class This class of instructions contains the set of instructions (a set of primary opcodes, as well as a set of extended opcodes for certain primary opcodes) used for implementation-specific instructions. Table 62 lists blocks of opcodes allocated for implementation-dependent use. Allocated instructions are allocated to purposes that are outside the scope of Book E for implementation-dependent and application-specific use. Any attempt to execute an allocated instruction results in one of the following:
  • An illegal instruction exception-type program interrupt, if the instruction is not recognized by the implementation
  • An unimplemented instruction exception-type program interrupt, if the instruction is recognized and enabled for execution but the implementation does not support direct

Table 62. Allocated instructions 0 All instruction encodings (bits 6–31) except 0x0000_0000 (1).

  1. Instruction encoding 0x0000_0000 is and always will be reserved-illegal.

4 All instruction encodings (bits 6–31)

19 Extended opcodes (bits 21–30) 0buuuuu_0u11u

31 Extended opcodes (bits 21–30) uuuuu_0u11u

59 Extended opcodes (bits 21–30) uuuuu_0u10u

63 Extended opcodes (bits 21–30) uuuuu_0u10u (except 00000_01100 frsp)

unsupported allocated instructions.

  • A floating-point unavailable interrupt, if an allocated instruction that extends the floating-point capabilities is recognized and floating-point processing is disabled
  • If an allocated instruction is implemented, the processor performs the actions described in the user’s manual. Implementation-dependent behavior may cause other exceptions. An allocated instruction is guaranteed by Book E to remain allocated. Note: Some allocated instructions may have associated new process state, and, therefore, may provide an associated enable bit, similar in function to MSR[FP] for floating-point instructions. For such instructions, being enabled for execution implies that any associated enable bit is set to allow, or enable, instruction execution. For such instructions, the architecture provides an auxiliary processor unavailable interrupt vector in case execution of such an instruction is attempted when execution is disabled. For example, MSR[SPE] enables the SPE unavailable interrupt. Other allocated instructions may not have any associated new state and therefore may not require an associated enable bit. If supported by an implementation, such instructions are assumed to be always enabled for execution. Preserved instruction class The preserved instruction class supports backward compatibility with the PowerPC architecture. An attempt to execute a preserved instruction results in one of the following:
  • If the implementation does not recognize the instruction, an illegal instruction exception-type program interrupt occurs.
  • If the instruction is recognized and supported by the implementation, the processor performs the actions defined in the previous version of the architecture. Future versions of Book E may retain a preserved instruction as a preserved instruction, may reclassify it as an allocated instruction, or may adopt it as part of Book E. Preserved opcodes are listed in Table 63.

Table 63. Preserved instructions

0 No preserved extended opcodes

4 No preserved extended opcodes

19 No preserved extended opcodes

59 No preserved extended opcodes

63 No preserved extended opcodes

Reserved instruction class This class of instructions consists of all instruction primary opcodes (and associated extended opcodes, if applicable) that do not belong to either the defined, allocated, or preserved instruction classes. Reserved instructions are available for future extensions of Book E. That is, some future version of Book E may define any of these instructions to perform new functions or make them available for implementation-dependent use as allocated instructions. There are two types of reserved instructions, reserved-illegal and reserved-nop. Attempts to execute a reserved-illegal instruction cause an illegal instruction exception-type program interrupt (see Chapter 4.7.6: Alignment interrupt on page 263) on implementations conforming to the current version of Book E. Reserved-illegal instructions are, therefore, available for future extensions to Book E that would affect architected state. Such extensions might include new forms of integer or floating-point arithmetic or new forms of load or store instructions that write their result in an architected register. Attempts to execute a reserved-nop instruction either do not affect implementations conforming to the current version of Book E (that is, treated as a no-operation instruction), or cause an illegal instruction exception-type program interrupt (see Chapter 4.7.7: Program interrupt on page 265”). Reserved-nop instructions are available for future architecture extensions that do not affect architected state. Such extensions might include performance- enhancing hints such as new forms of cache touch instructions and could be added while remaining functionally compatible with implementations of previous versions of Book E A reserved-illegal instruction may be retained by future versions of Book E as a reserved- illegal instruction, may be subsequently reclassified as an allocated instruction, or may even be employed in the role of a subsequently defined instruction. A reserved-nop instruction may be retained by future versions of Book E as a reserved-nop instruction, may be subsequently reclassified as an allocated instruction, or may even be employed in the role of a subsequently defined instruction that has no effect on architected state.

3.2.2 Instruction forms

This section describes preferred instruction forms, addressing modes, and synchronization. Preferred instruction forms (no-op) The Or Immediate (ori) instruction has the following preferred form for expressing a no-op: ori 0,0,0 Invalid instruction forms Some of the defined instructions have invalid forms. An instruction form is invalid if one or more fields of the instruction, excluding the opcode field(s), are coded incorrectly in a manner that can be deduced by examining only the instruction encoding. Attempts to execute an invalid form of an instruction either causes an illegal instruction type program interrupt or yields boundedly undefined results. Any exceptions to this rule are stated in the instruction descriptions.

Some kinds of invalid form instructions can be deduced just from examining the instruction layout. These are listed below.

  • Field shown as reserved but coded as nonzero
  • Field shown as containing a particular value but coded as some other value These invalid forms are not discussed further. Other invalid instruction forms can be deduced by detecting an invalid encoding of one or more of the instruction operand fields. These kinds of invalid form are identified in the instruction descriptions.
  • Branch conditional and branch conditional extended instructions (undefined encoding of BO field)
  • Load with update instructions (rD= rA or rA=0 )
  • Store with update instructions (rA=0 )
  • Load multiple instruction (rA or rB in range of registers to be loaded)
  • Load string immediate instructions (rA in range of registers to be loaded)
  • Load string indexed instructions (rD= rA or rD= rB)
  • Load/store floating-point with update instructions (rA=0 )

3.2.3 Addressing modes

This section describes conventions for addressing memory and for calculating effective addresses (EAs) as defined by the Book E architecture for 32-bit implementations. Memory addressing A program references memory using the effective address computed by the processor when it executes a memory access or branch instruction (or other instructions as described in Chapter : User-level cache instructions on page 180,” and Chapter : Supervisor-level cache instruction on page 183,” or when it fetches the next sequential instruction. Memory operands Bytes in memory are numbered consecutively starting with 0. Each number is the address of the corresponding byte. Memory operands may be bytes, half words, words or, for the load/store multiple and load/store string instructions, a sequence of words or bytes. The address of a memory operand is the address of its first byte (that is, of its lowest-numbered byte). Byte ordering can be either big endian or little endian (see Chapter : Byte ordering on page 141”). The default byte and bit ordering is big endian. Operand length is implicit for each instruction with respect to memory alignment. The operand of a scalar memory access instruction has a natural alignment boundary equal to the operand length. In other words, the natural address of an operand is an integral multiple of the operand length. A memory operand is said to be aligned if it is aligned at its natural boundary; otherwise it is said to be misaligned. For more information about alignment, see Chapter 3.1.2: Alignment and misaligned accesses on page 133.” Effective address calculation The 32-bit address computed by the processor when executing a memory access or branch instruction (or certain other instructions described in User-level cache instructions on page 180,” Supervisor-level cache instruction,” and Supervisor-level tlb management

instructions on page 183”), or when fetching the next sequential instruction, is called the effective address (EA) and specifies a byte in memory. For a memory access instruction, if the sum of the EA and the operand length exceeds the maximum EA, the memory access is considered to be undefined. Effective address arithmetic, except for next sequential instruction address computations, wraps around from the maximum address, 2 32– 1, to address 0. Data memory addressing modes Book E supports the following data memory addressing modes:

  • Base+displacement addressing mode—The 16-bit D field is sign-extended and added to the contents of the GPR designated by rA or to zero if rA = 0. Instructions that use this addressing mode are of the D instruction format.
  • Base+index addressing mode—The contents of the GPR designated by rB (or the value 0 for lswi and stswi) are added to the contents of the GPR designated by rA or to zero if rA = 0. Instructions that use this addressing mode are of the X instruction format.
  • Base+displacement extended addressing mode—The 12-bit DE field is sign-extended and added to the contents of the GPR designated by rA or to zero if rA = 0 to produce the 32-bit EA. Instructions that use this addressing mode are of the DE instruction format.
  • Base+displacement extended scaled addressing mode—The 12-bit DES field is concatenated on the right with zeros, sign-extended, and added to the contents of the GPR designated by rA or to zero if rA = 0 to produce the 32-bit EA. Instructions that use this addressing mode are of the DES instruction format. In addition, APUs may provide additional addressing modes. Instruction memory addressing modes Instruction memory addressing modes correspond with instructions forms, as follows:
  • I-form branch instructions—The 24-bit LI field is concatenated on the right with 0b00, sign-extended, and then added to either the address of the branch instruction if AA = 0, or to 0 if AA = 1.
  • Taken B-form branch instructions—The 14-bit BD field is concatenated on the right with 0b00, sign-extended, and then added to either the address of the branch instruction if AA = 0, or to 0 if AA = 1.
  • Taken XL-form branch instructions—The contents of bits LR[32–61] or CR[32–61] are concatenated on the right with 0b00.
  • Sequential instruction fetching (or non-taken branch instructions)—The value 4 is added to the address of the current instruction to form the 32-bit EA of the next instruction. If the address of the current instruction is 0xFFFF_FFFC, the address of the next sequential instruction is undefined.
  • Any branch instruction with LK = 1—The value 4 is added to the address of the current instruction and the 32-bit result is placed into the LR. If the address of the current instruction is 0xFFFF_FFFC, the result placed into the LR is undefined. Although some implementations may support next sequential instruction address computations wrapping from the highest address 0xFFFF_FFFC to 0x0000_0000 as part of the instruction flow, users are strongly encouraged not to depend on this behavior. Doing so can reduce the portability of their software. If code must span this boundary, software should place a non-linking branch at address 0xFFFF_FFFC, which always branches to address

0x0000_0000 (either absolute or relative branches work). See also Appendix D: Guidelines for 32-bit book E on page 1154.” Byte ordering If scalars (individual data items and instructions) were indivisible, there would be no such concept as byte ordering. It is meaningless to consider the order of bits or groups of bits within the smallest addressable unit of memory, because nothing can be observed about such order. Only when scalars, which the programmer and processor regard as indivisible quantities, can comprise more than one addressable unit of memory does the question of order arise. For a machine in which the smallest addressable unit of memory is the 64-bit double word, there is no question of the ordering of bytes within double words. All transfers of individual scalars between registers and memory are of double words, and the address of the byte containing the high-order 8 bits of a scalar is no different from the address of a byte containing any other part of the scalar. For Book E, as for most computer architectures currently implemented, the smallest addressable unit of memory is the 8-bit byte. Many scalars are half words and words (double words in 64-bit implementations) which consist of groups of bytes. When a word-length scalar is moved from a register to memory, the scalar occupies four consecutive byte addresses. It thus becomes meaningful to discuss the order of the byte addresses with respect to the value of the scalar: which byte contains the highest-order eight bits of the scalar, which byte contains the next-highest-order 8 bits, and so on. Given a scalar that contains multiple bytes, the choice of byte ordering is essentially arbitrary. There are 4! = 24 ways to specify the ordering of 4 bytes within a word but only two of these orderings are sensible:

  • The ordering that assigns the lowest address to the highest-order (left-most) 8 bits of the scalar, the next sequential address to the next-highest-order eight bits, and so on. This ordering is called big endian because the big (most-significant) end of the scalar, considered as a binary number, comes first in memory. The 68000 is an example of a processor using this byte ordering.
  • The ordering that assigns the lowest address to the lowest-order (right-most) 8 bits of the scalar, the next sequential address to the next-lowest-order eight bits, and so on. This ordering is called little endian because the little (least-significant) end of the scalar, considered as a binary number, comes first in memory. The Intel 8086 is an example of a processor using this byte ordering. Book E provides support for both big- and little-endian byte ordering in the form of a memory Memory/Cache access attributes on page 283.” Synchronization requirements This section describes synchronization requirements for special registers and TLBs. Changing the value in certain system registers and invalidating TLB entries can have the side effect of altering the context in which data addresses and instruction addresses are interpreted, and in which instructions are executed. For example, changing MSR[IS] = 0 to and MSR[IS] = 1 has the side effect of changing address space. Such effects need not occur in program order (program order refers to the execution of instructions in the strict order in which they occur in the program), and therefore may require explicit synchronization by software.

altering instruction” should be interpreted as meaning the context-altering instruction itself. fetched or executed in either context. is required within the sequence. Table 64. Synchronization requirements

Table 64. Synchronization requirements (continued)

An instruction or event is context synchronizing if it satisfies the requirements listed below. Context-synchronizing operations include instructions isync, sc, rfi, rfci, rfdi, and rfmci, and most interrupts. 1. The operation is not initiated or, in the case of isync, does not complete until all instructions already in execution have completed to a point at which they have reported all exceptions they cause. 2. The instructions that precede the operation complete execution in the context (including such parameters as privilege level, address space, and memory protection) in which they were initiated. 3. If the operation directly causes an interrupt (for example, sc directly causes a system call interrupt) or is an interrupt, the operation is not initiated until no interrupt-causing 1. A context synchronizing instruction is required after altering MSR[ME] to ensure that the alteration takes effect for subsequent machine check interrupts, which may not be recoverable and therefore may not be context synchronizing. 2. Synchronization requirements fo r changing any of the debug registers are implementation-dependent and are specified in the user’s manual for the implementation. 3. For data accesses, the context sy nchronizing instruction before the tlbwe or tlbivax instruction ensures that all storage accesses due to preceding instructions have completed to a point at which they have reported all exceptions they cause. The context synchronizing instruction after the tlbwe or tlbivax ensures that subsequent storage accesses (data and instruction) use the updated value in any affected TLB entries. It does not ensure that all storage accesses previously translated by the TLB entries being updated have completed with respect to storage; if these completions must be ensured, the tlbwe or tlbivax must be followed by an msync instruction as well as by a context synchronizing instruction. The following sequence shows why it is necessary for data accesses to ensure that all storage accesses due to instructions before a tlbwe or tlbivax have completed to a point at which they have reported all exceptions they will cause. Assume that valid TLB entries exist for the target storage location when the sequence starts. 1. A program issues a load or store to a page. 2. The same program executes a tlbwe or tlbivax that invalidates the corresponding TLB entry. 3. The load or store instruction finally executes, and gets a TLB miss exception. The TLB miss exception is semantically incorrect. In order to prevent it, a context synchronizing instruction must be executed between steps 1 and 2. 4. Multiprocessor systems have other requirements to synch ronize what is called TLB shoot down’ (that is, to invalidate one or more TLB entries on all processors in the multiprocessor system and to be able to determine that the invalidations have completed and that all side effects of the invalidations have taken effect). 5. The effect of changing MSR[EE] or MSR[CE] is immediate. If an mtmsr, wrtee, or wrteei clears MSR[EE], an external input, decrementer, or fixed-interval timer interrupt does not occur after the instruction executes. If an mtmsr, wrtee, or wrteei changes MSR[EE] from 0 to 1 when an external input, decrementer, fixed- interval timer, or higher priority enabled exception exists, the corresponding interrupt occurs immediately after the mtmsr, wrtee, or wrteei executes and before the next instruction is executed in the program that sets MSR[EE]. If an mtmsr clears MSR[CE], a critical input, or watchdog timer interrupt does not occur after the instruction is executed. If an mtmsr changes MSR[CE] from 0 to 1 when a critical input, watchdog timer, or higher priority enabled exception exists, the corresponding interrupt occurs immediately after mtmsr executes, and before the next instruction is executed in the program that set MSR[CE]. 6. The alteration must not cause an implicit branch in real address space. Thus the real address of the context-altering instruction and of each subsequent instruction, up to and including the next context synchronizing instruction, must be independent of whether the alteration has taken effect. 7. Synchronization requirements for changing the wait state enable are implementation-dependent, and are specified in the user’s manual for the implementation. 8. The elapsed time between the decrementer reaching zero, or the transition of the selected time base bit for the fixed-interval timer or the watchdog timer, and the signalling of the decrementer, fixed-interval timer or the watchdog timer exception is not defined.

exception exists having higher priority than the exception associated with the interrupt. See Chapter 4.11: Exception priorities on page 278.” 4. The instructions that follow the operation are fetched and executed in the context established by the operation as required by the sequential execution model. (This requirement dictates that any prefetched instructions be discarded and that any effects and side effects of executing them speculatively may also be discarded, except as described in Memory access ordering on page 290.” A context-synchronizing operation is necessarily execution synchronizing. Unlike msync and mbar, such operations do not affect the order of memory accesses with respect to other mechanisms. Execution synchronization An instruction is execution synchronizing if it satisfies items 1 and 2 of the definition of context synchronization .msync is treated like isync with respect to item 1 (that is, the conditions described in item 1 apply to completion of msync). Execution synchronizing instructions include msync, mtmsr, wrtee, and wrteei. All context-synchronizing instructions are execution synchronizing. Unlike a context-synchronizing operation, an execution synchronizing instruction need not ensure that the instructions following it execute in the context established by that execution synchronizing instruction. This new context becomes effective sometime after the execution synchronizing instruction completes and before or at a subsequent context-synchronizing operation. Instruction-related interrupts Interrupts are caused either directly by the execution of an instruction or by an asynchronous event. In either case, an exception may cause one of several types of interrupts to be invoked. Examples of interrupts that can be caused directly by the execution of an instruction include but are not limited to the following:

  • An attempt to execute a reserved-illegal instruction (illegal instruction exception-type program interrupt)
  • An attempt by an application program to execute a privileged instruction (privileged instruction exception-type program interrupt)
  • An attempt by an application program to access a privileged SPR (privileged instruction exception-type program interrupt)
  • An attempt by an application program to access an SPR that does not exist (unimplemented operation instruction exception-type program interrupt)
  • An attempt by a system program to access an SPR that does not exist (boundedly undefined)
  • Execution of a defined instruction using an invalid form (illegal instruction exception- type program interrupt, unimplemented operation exception-type program interrupt, or privileged instruction exception-type program interrupt)
  • An attempt to access a memory location that is either unavailable (instruction TLB error interrupt or data TLB error interrupt) or not permitted (instruction storage interrupt or data storage interrupt)
  • An attempt to access memory with an EA alignment not supported by the implementation (alignment interrupt)
  • Execution of a system call instruction (system call interrupt)
  • Execution of a trap instruction whose trap condition is met (trap type program interrupt)
  • Execution of a floating-point instruction when floating-point instructions are unavailable (floating-point unavailable interrupt)
  • Execution of a floating-point instruction that causes a floating-point enabled exception to exist (floating-point enabled exception-type program interrupt)
  • Execution of a defined instruction that is not implemented by the implementation (illegal instruction exception or unimplemented operation exception-type program interrupt)
  • Execution of an allocated instruction that is not implemented by the implementation (illegal instruction exception or unimplemented operation exception-type program interrupt)
  • Execution of an allocated instruction when the auxiliary instruction is unavailable (auxiliary processor unavailable interrupt).
  • Execution of an allocated instruction that causes an auxiliary enabled exception (enabled exception-type program interrupt). APUs, such as the SPE, may define additional instruction-caused exceptions and interrupts. The invocation of an interrupt is precise, except that if one of the imprecise modes for invoking the floating-point enabled exception-type program interrupt is in effect the invocation of the floating-point enabled exception-type program interrupt may be imprecise. When the interrupt is invoked imprecisely, the excepting instruction does not appear to complete before the next instruction starts (because one of the effects of the excepting instruction, namely the invocation of the interrupt, has not yet occurred). Chapter 4: Interrupts and exceptions on page 244 describes interrupt conditions in detail.

3.3 Instruction set overview

This section provides a brief overview of the Book E and Book E instructions. Note: some instructions have the following optional features:

  • CR update—The dot ( .) suffix on the mnemonic enables the update of the CR.
  • Overflow option—The o suffix indicates that the overflow bit in the XER is enabled.

3.3.1 Book E user-l evel instructions

This section discusses the user-level instructions defined in the Book E architecture. Integer instructions This section describes the integer instructions. These consist of the following:

  • Integer arithmetic instructions
  • Integer compare instructions
  • Integer logical instructions
  • Integer rotate and shift instructions Integer instructions use the content of the GPRs as source operands and place results into GPRs and the XER and CR fields. Integer arithmetic instructions Table 65 lists the integer arithmetic instructions for the PowerPC processors.

instructions on page 1110,” for examples. an overflow condition of a 32-bit result only if the instruction’s OE bit is set. Table 65. Integer arithmetic instructions

must be 0 for 32-bit implementations. The crD operand can be omitted if the result of the comparison is to be placed in CR0. Otherwise the target CR field must be specified in crD by using an explicit field number. operation. Logical instructions do not affect XER[SO], XER[OV], or XER[CA]. See Appendix B,” for simplified mnemonic examples for integer logical operations. Table 66. Integer 32-Bit compare instructions (L = 0) Table 67. Integer logical instructions

register, left or right justifying an arbitrary field, and simple rotates and shifts. PowerPC instructions”) are provided to simplify coding of such shifts. page 1148.” The integer shift instructions are summarized in Table 69. This section describes the floating-point instructions as they are defined by Book E. Table 68. Integer rotate instructions Table 69. Integer shift instructions Table 67. Integer logical instructions (continued)

The rules followed in assigning new primary and extended opcodes.

  • Primary opcode 63 is used for the double-precision arithmetic instructions as well as miscellaneous instructions (for example, FPSCR manipulation instructions). Primary opcode 59 is used for the single-precision arithmetic instructions.
  • The single-precision instructions for which there is a corresponding double-precision instruction have the same format and extended opcode as that double-precision instruction.
  • In assigning new extended opcodes for primary opcode 63, the following regularities are maintained. In addition, all new X-form instructions in primary opcode 63 have bits 21–22 = 11. – Bit 26 = 1 if and only if the instruction is A-form. – Bits 26–29 = 0b0000 if and only if the instruction is a comparison or mcrfs (if and only if the instruction sets an explicitly designated CR field). – Bits 26–28 = 0b001 if and only if the instruction explicitly refers to or sets the FPSCR (that is, is an FPSCR instruction) and is not mcrfs. – Bits 26–30 = 0b01000 if and only if the instruction is a move register instruction, or any other instruction that does not refer to or set the FPSCR.
  • In assigning extended opcodes for primary opcode 59, the following regularities have been maintained. They are based on those rules for primary opcode 63 that apply to the instructions having primary opcode 59. In particular, primary opcode 59 has no FPSCR instructions, so the corresponding rule does not apply. – If there is a corresponding instruction with primary opcode 63, its extended opcode is used. – Bit 26 = 1 if and only if the instruction is A form. – Bits 26–30 = 0b01000 if and only if the instruction is a move register instruction, or any other instruction that does not refer to or set the FPSCR. Floating-point load instructions There are two basic forms of load instruction: single-precision and double-precision. Because the FPRs support only floating-point double format, single-precision load floating- point instructions convert single-precision data to double format prior to loading the operand into the target FPR. The conversion and loading steps are as follows. Let WORD 0:31 be the floating-point single-precision operand accessed from memory. Normalized Operand if WORD1:8 > 0 and WORD1:8 < 255 then FPR(frD)0:1 ← WORD0:1 FPR(frD)2 ← ¬WORD1 FPR(frD)3 ← ¬WORD1 FPR(frD)4 ← ¬WORD1 FPR(frD)5:63 ← WORD2:31 || 290 Denormalized Operand if WORD1:8 = 0 and WORD9:31 ≠ 0 then sign ← WORD0 exp ← -126 frac0:52 ← 0b0 || WORD9:31 || 290 normalize the operand do while frac0 = 0 frac ← frac1:52 || 0b0

data from memory is copied directly into the FPR. rA=0 or rA=rD, the instruction form is invalid. instruction set is shown in Table 70. Table 70. Floating-point load instruction set

storing the operand. The conversion steps are as follows. 0:31 be the word in memory written to. the original source register). with the EA. For these forms, if rA≠0, the EA is placed into GPR(rA). attempts to access memory that is unavailable. Store instructions are shown in Table 71. Book E supports both big-endian and little-endian byte ordering. Table 71. Floating-point store instructions

fabs, and fnabs). These instructions do not alter the FPSCR. Table 72. Floating-point move instructions Table 71. Floating-point store instructions (continued)

compare, and status/control instructions. Table 73 lists mnemonics and syntax of floating-point elementary arithmetic instructions.

  • Overflow, underflow, and inexact exception bits, the FR, FI, and FPRF fields are set based on the final result of the operation, not on the result of the multiplication.
  • Invalid operation exception bits are set as if the multiplication and the addition were performed using two separate instructions (fmul[s], followed by fadd[s] or fsub[s]). That is, any of the following actions will cause appropriate exception bits to be set: – Multiplication of infinity by 0 – Multiplication of anything by an SNaN – Addition of anything with an SNaN

Table 73. Floating-point elementary arithmetic instructions Table 74. Floating-point multiply-add instructions

the other three. The floating-point condition code, FPSCR[FPCC], is set in the same way. The CR field and the FPCC are set as described in Table 76. The floating-point compare and select instruction set is shown in Table 77. Table 75. Floating-point rounding and conversion instructions Table 76. CR field settings Table 77. Floating-point compare and select instructions Table 74. Floating-point multiply-add instructions (continued)

  • All exceptions caused by the previously initiated instructions are recorded in the FPSCR before the FPSCR instruction is initiated.
  • All invocations of floating-point enabled exception-type program interrupt that will be caused by the previously initiated instructions have occurred before the FPSCR instruction is initiated.
  • No subsequent floating-point instruction that depends on or alters the settings of any FPSCR bits is initiated until the FPSCR instruction has completed. Floating-point load and floating-point store instructions (Table 78) are not affected. Load and store instructions Load and store instructions are issued and translated in program order; however, the accesses can occur out of order. Synchronizing instructions are provided to enforce strict ordering. The following load and store instructions are defined:
  • Integer load instructions
  • Integer store instructions
  • Integer load and store with byte-reverse instructions
  • Integer load and store multiple instructions
  • Memory synchronization instructions
  • SPE APU load and store instructions for reading and writing 64-bit GPRs. Some of these instructions are also implemented by processors that support the embedded vector single-precision and embedded scalar double-precision floating-point APUs, which use the extended 64-bit GPRs. See Chapter 3.6.1 on page 186.” Self-modifying code When a processor modifies any memory location that can contain an instruction, software must ensure that the instruction cache is made consistent with data memory and that the

Table 78. Floating-point status and control register instructions

modifications are made visible to the instruction fetching mechanism. This must be done even if the cache is disabled or if the page is marked caching-inhibited. The following instruction sequence can be used to accomplish this when the instructions being modified are in memory that is memory-coherence required and one processor both modifies the instructions and executes them. (Additional synchronization is needed when one processor modifies instructions that another processor will execute.) The following sequence synchronizes the instruction stream (using either dcbst or dcbf): dcbst (or dcbf)|update memory msync |wait for update icbi |remove (invalidate) copy in instruction cache msync |ensure the ICBI invalidate is complete isync |remove copy in own instruction buffer These operations are required because the data cache is a write-back cache. Because instruction fetching bypasses the data cache, changes to items in the data cache cannot be reflected in memory until the fetch operations complete. The msync after the icbi is required to ensure that the icbi invalidation has completed in the instruction cache. Special care must be taken to avoid coherency paradoxes in systems that implement unified secondary caches, and designers should carefully follow the guidelines for maintaining cache coherency discussed in the user’s manual. Integer load and store address generation Integer load and store operations generate EAs using register indirect with immediate index mode, register indirect with index mode, or register indirect mode, which are described as follows:

  • Register indirect with immediate index addressing for integer loads and stores. Instructions using this addressing mode contain a signed 16-bit immediate index (d operand), which is sign extended and added to the contents of a general-purpose register specified in the instruction (rA operand), to generate the EA. If r0 is specified, a value of zero is added to the immediate index (d operand) in place of the contents of r0. The option to specify rA or 0 is shown in the instruction descriptions as (rA|0). Figure 6 shows how an EA is generated using this mode.

Figure 8. Register indirect addressing for integer loads/stores information about load and store address alignment interrupts. rA = 0 or rA= rD as invalid forms. Table 79. Integer load instructions

  • If rA ≠ 0, the EA is placed into rA.
  • If rS= rA, the contents of register rS are copied to the target memory element and the generated EA is placed into rA (rS). The Book E architecture defines store with update instructions with rA = 0 as an invalid form. In addition, it defines integer store instructions with the CR update option enabled (Rc field, bit 31, in the instruction encoding = 1) to be an invalid form. Table 80 summarizes integer store instructions. Integer load and store with byte-reverse instructions Load Half Word and Zero with Update lhzu rD,d(rA) Load Half Word and Zero with Update Indexed lhzux rD,rA,rB Load Half Word Algebraic lha rD,d(rA) Load Half Word Algebraic Indexed lhax rD,rA,rB Load Half Word Algebraic with Update lhau rD,d(rA) Load Half Word Algebraic with Update Indexed lhaux rD,rA,rB Load Word and Zero lwz rD,d(rA) Load Word and Zero Indexed lwzx rD,rA,rB Load Word and Zero with Update lwzu rD,d(rA) Load Word and Zero with Update Indexed lwzux rD,rA,rB

Table 80. Integer store instructions Table 79. Integer load instructions (continued)

The load/store multiple instructions are used to move blocks of data to and from the GPRs. must be word aligned; otherwise, they cause an alignment exception. sequence of individual load or store instructions that produce the same results. Table 83 summarizes the integer load and store string instructions. Table 81. Integer load and store with byte-reverse instructions Table 82. Integer load and store multiple instructions Table 83. Integer load and store string instructions

Load string and store string instructions can involve operands that are not word-aligned. of floating-point loads and stores for direct-store accesses results in an alignment interrupt. data to double-precision format before loading an operand into an FPR. stfdx, and stfdux) are invalid when the Rc bit is one. The PowerPC architecture defines load with update with rA = 0 as an invalid form. convert double-precision data to single-precision format before storing the operands. Table 85 summarizes the floating-point store instructions. Table 84. Floating-point load instructions Table 85. Floating-point store instructions

conversions the LSU makes when executing a Store Floating-Point Single instruction. instruction. Most entries in the table indicate that the floating-point value is simply stored. Only in a few cases are any other actions taken.

  1. The stfiwx instruction is optional to the Book E architecture.

Table 86. Store floating-point single behavior Table 87. Store floating-point double behavior Table 85. Floating-point store instructions (continued)

  • Branch relative
  • Branch conditional to relative address
  • Branch to absolute address
  • Branch conditional to absolute address
  • Branch conditional to link register (LR)
  • Branch conditional to count register (CTR) Branch relative addressing mode Instructions that use branch relative addressing generate the next instruction address by sign extending and appending 0b00 to the immediate displacement operand LI, and adding the resultant value to the current instruction address. Branches using this mode have the absolute addressing option disabled (AA field, bit 30, in the instruction encoding = 0). The LR update option can be enabled (LK field, bit 31, in the instruction encoding = 1). This causes the EA of the instruction following the branch instruction to be placed in the LR. Figure 9 shows how the branch target address is generated using this mode.

Figure 9. Branch relative addressing branch target address is generated using this mode.

18 LI AA LK

Figure 10. Branch conditional relative addressing Figure 11. Branch to absolute addressing

16 BO BI BD AA LK

shows how the branch target address is generated using this mode. Figure 12. Branch conditional to absolute addressing zero. The LR update option can be enabled (LK field, bit 31, in the instruction encoding = 1). the LR. Figure 13 shows how the branch target address is generated using this mode. Figure 13. Branch conditional to link register addressing

19 BO BI 0 0 0 0 0 16 LK

address is generated when using this mode. Figure 14. Branch conditional to count register addressing described here. For those processors, the BO operand is ignored for branch prediction. value y, is used by some implementations for branch prediction as described below. The encodings for the BO operands are shown in Table 89.

19 BO BI 00000 528 LK

Table 88. BO bit descriptions 0 Setting this bit causes the CR bit to be ignored.

1 Bit value to test against

2 Setting this causes the decrement to not be decremented. 3 Setting this bit reverses the sense of the CTR test.

4 Used for the y bit, which provides a hint about whether a conditional branch is likely to be

taken and may be used by some implementations to improve performance.

The branch always encoding of the BO operand does not have a y bit.

  • For bcx with a negative value in the displacement operand, the branch is taken.
  • In all other cases (bcx with a non-negative value in the displacement operand, bclrx, or bcctrx), the branch is not taken. Setting the y bit reverses the preceding indications. The sign of the displacement operand is used as described above even if the target is an absolute address. The default value for the y bit should be 0 and should be set to 1 only if software has determined that the prediction corresponding to y = 1 is more likely to be correct than the prediction corresponding to y = 0. Software that does not compute branch predictions should clear the y bit. In most cases, the branch should be predicted to be taken if the value of the following expression is 1, and predicted to fall through if the value is 0. In the expression above, S (bit 16 of the branch conditional instruction coding) is the sign bit of the displacement operand if the instruction has a displacement operand and is 0 if the operand is reserved. BO[4] is the y bit, or 0 for the branch always encoding of the BO operand. (Advantage is taken of the fact that, for bclrx and bcctrx, bit 16 of the instruction is part of a reserved operand and therefore must be 0.) The 5-bit BI operand in branch conditional instructions specifies which CR bit represents the condition to test. The CR bit selected is BI +32, as shown in Table 17. If the branch instructions contain immediate addressing operands, the target addresses can be computed sufficiently ahead of the branch instruction that instructions can be fetched along the target path. If the branch instructions use the link and count registers, instructions along the target path can be fetched if the link or count register is loaded sufficiently ahead of the branch instruction.

Table 89. BO operand encodings 0000y Decrement the CTR, then branch if the decremented CTR ≠ 0 and the condition is FALSE. 0001y Decrement the CTR, then branch if the decremented CTR = 0 and the condition is FALSE. 001zy Branch if the condition is FALSE. 0100y Decrement the CTR, then branch if the decremented CTR ≠ 0 and the condition is TRUE. 0101y Decrement the CTR, then branch if the decremented CTR = 0 and the condition is TRUE. 011zy Branch if the condition is TRUE. 1z00y Decrement the CTR, then branch if the decremented CTR ≠ 0. 1z01y Decrement the CTR, then branch if the decremented CTR = 0. a meaning in some future version of the architecture. some implementations to improve performance.

instruction are also defined as flow control instructions. architecture defines these forms of the instructions as invalid. Table 90. Branch instructions Table 91. Condition register logical instructions

register (MSR), and special-purpose registers (SPRs). Table 94 summarizes the instructions for reading from or writing to the CR. Table 95 lists the mtspr and mfspr instructions. Table 92. Trap instructions Table 93. System linkage instruction Table 94. Move to/from condition register instructions Table 95. Move to/from special-purpose register instructions

Table 96 summarizes all SPRs defined in Book E, indicating which are user-level access. The SPR number column lists register numbers used in the instruction mnemonics. Table 96. Book E special-purpose registers (by SPR abbreviation)

Table 96. Book E special-purpose registers (by SPR abbreviation) (continued)

Table 97 lists EIS-specific SPRs, indicating which can be accessed by user-level software. Compilers should recognize SPR names when parsing instructions.

  1. The DBSR is read using mfspr. It cannot be directly written to. Instead, DBSR bits corresponding to 1 bits in the GPR can
  2. Implementations may support more than one PID. If mult iple PIDs are implemented, the Book E–defined PID is
  3. User-mode read access to SPRG3 is implementation-dependent.
  4. The TSR is read using mfspr. It cannot be directly written to. Instead, TSR bits corresponding to 1 bits in the GPR can be
  5. USPRG0 is a separate physi cal register from SPRG0.

Table 97. Implementation-specific SPRs (by SPR abbreviation)

Table 97. Implementation-specific SPRs (by SPR abbreviation) (continued)

are seen by other mechanisms that access memory. See Table 98 for a summary. Table 98. Memory synchronization instructions branch target is the next sequential instruction to be executed. fetch and add. Both instructions must use the same EA. Reservation granularity is implementation-dependent. MO = 0— mbar behaves identically to msync. user’s manual for implementation-specific behavior.

Atomic update primitives using lwarx and stwcx. The lwarx and stwcx. instructions together permit atomic update of a memory location. is the operation of lwarx and stwcx. cause the system data storage error handler to be invoked.

  • The memory coherence required attribute on other processors and mechanisms ensures that their stores to the specified location will cause the reservation created by the lwarx to be cancelled.
  • Warning: Support for load and reserve and store conditional instructions for which the specified location is in caching-inhibited memory is being phased out of Book E. It is likely not to be provided on future implementations. New programs should not use these instructions to access caching inhibited memory. A lwarx instruction is a load from a word-aligned location with the following side effects.
  • A reservation for a subsequent stwcx. instruction is created.
  • The memory coherence mechanism is notified that a reservation exists for the location accessed by the lwarx. Memory synchronize msync — Provides an ordering function for the effects of all instructions executed by the processor executing the msync. Executing an msync ensures that all previous instructions complete before it completes and that no subsequent instructions are initiated until after it completes. It also creates a memory barrier, which orders the storage accesses associated with these instructions. msync cannot complete before storage accesses associated with previous instructions are performed. msync is execution synchronizing. Note the following: msync is used to ensure that all stores into a data structure caused by store instructions executed in a critical section of a program are performed with respect to another processor before the store that releases the lock is performed with respect to that processor. mbar is preferable in many cases. On ST Book E devices: Unlike a context-synchronizing operations, msync does not discard prefetched instructions. Store word conditional indexed stwcx. rS,rA,rB lwarx with stwcx. can emulate semaphore operations such as test and set, compare and swap, exchange memory, and fetch and add. Both instructions must use the same EA. Reservation granularity is implementation-dependent. Executing lwarx and stwcx. to a page marked write-through (WIMG = 10xx) or cache-inhibited (WIMG = 01xx) when the data cache is locked may cause a data storage interrupt. If the location is not word-aligned, an alignment interrupt occurs.

Table 98. Memory synchronization instructions (continued)

The stwcx. is a store to a word-aligned location that is conditioned on the existence of the reservation created by the lwarx and on whether both instructions specify the same location. To emulate an atomic operation, both lwarx and stwcx. must access the same location. lwarx and stwcx. are ordered by a dependence on the reservation, and the program is not required to insert other instructions to maintain the order of memory accesses caused by these two instructions. A stwcx. performs a store to the target location only if the location accessed by the lwarx that established the reservation has not been stored into by another processor or mechanism between supplying a value for the lwarx and storing the value supplied by the stwcx.. If the instructions specify different locations, the store is not necessarily performed. CR0 is modified to indicate whether the store was performed, as follows: CR0[LT,GT,EQ,SO] = 0b00 || store_performed || XER[SO] If a stwcx. completes but does not perform the store because a reservation no longer exists, CR0 is modified to indicate that the stwcx. completed without altering memory. A stwcx. that performs its store is said to succeed. Examples using lwarx and stwcx. are given in Appendix C: Programming examples on page 1143.” A successful stwcx. to a given location may complete before its store has been performed with respect to other processors and mechanisms. As a result, a subsequent load or lwarx from the given location on another processor may return a stale value. However, a subsequent lwarx from the given location on the other processor followed by a successful stwcx. on that processor is guaranteed to have returned the value stored by the first processor’s stwcx. (in the absence of other stores to the given location). Reservations The ability to emulate an atomic operation using lwarx and stwcx. is based on the conditional behavior of stwcx., the reservation set by lwarx, and the clearing of that reservation if the target location is modified by another processor or mechanism before the stwcx. performs its store. A reservation is held on an aligned unit of real memory called a reservation granule. The size of the reservation granule is implementation-dependent, but is a multiple of 4 bytes for lwarx. The reservation granule associated with EA contains the real address to which the EA maps. (‘real_addr(EA)’ in the RTL for the load and reserve and store conditional instructions stands for ‘real address to which EA maps.’) When one processor holds a reservation and another processor performs a store, the first processor’s reservation is cleared if the store affects any bytes in the reservation granule. Note: One use of lwarx and stwcx. is to emulate a compare and swap primitive like that provided by the IBM System/370 compare and swap instruction, which checks only that the old and current values of the word being tested are equal, with the result that programs that use such a compare and swap to control a shared resource can err if the word has been modified and the old value is subsequently restored. The use of lwarx and stwcx. improves on such a compare and swap, because the reservation reliably binds lwarx and stwcx. together. The reservation is always lost if the word is modified by another processor or mechanism between the lwarx and stwcx., so the stwcx. never succeeds unless the word has not been stored into (by another processor or mechanism) since the lwarx.

A processor has at most one reservation at any time. Book E states that a reservation is established by executing a lwarx and is lost (or may be lost, in the case of the fourth and fifth bullets) if any of the following occurs.

  • The processor holding the reservation executes another lwarx; this clears the first reservation and establishes a new one.
  • The processor holding the reservation executes any stwcx., regardless of whether the specified address matches that of the lwarx.
  • Another processor executes a store or dcbz to the same reservation granule.
  • Another processor executes a dcbtst, dcbst, or dcbf to the same reservation granule; whether the reservation is lost is undefined.
  • Another processor executes a dcba to the reservation granule. The reservation is lost if the instruction causes the target block to be newly established in the data cache or to be modified; otherwise, whether the reservation is lost is undefined.
  • Some other mechanism modifies a location in the same reservation granule.
  • Other implementation-specific conditions may also cause the reservation to be cleared, See the core reference manual. Interrupts are not guaranteed to clear reservations. (However, system software invoked by interrupts may clear reservations.) In general, programming conventions must ensure that lwarx and stwcx. specify addresses that match; a stwcx. should be paired with a specific lwarx to the same location. Situations in which a stwcx. may erroneously be issued after some lwarx other than that with which it is intended to be paired must be scrupulously avoided. For example, there must not be a context switch in which the processor holds a reservation in behalf of the old context, and the new context resumes after a lwarx and before the paired stwcx.. The stwcx. in the new context might succeed, which is not what was intended by the programmer. Such a situation must be prevented by issuing a stwcx. to a dummy writable word-aligned location as part of the context switch, thereby clearing any reservation established by the old context. Executing stwcx. to a word-aligned location is enough to clear the reservation, regardless of whether it was set by lwarx. Forward progress Forward progress in loops that use lwarx and stwcx. is achieved by a cooperative effort among hardware, operating system software, and application software. Book E guarantees one of the following when a processor executes a lwarx to obtain a reservation for location X and then a stwcx. to store a value to location X: 1. The stwcx. succeeds and the value is written to location X. 2. The stwcx. fails because some other processor or mechanism modified location X. 3. The stwcx. fails because the processor’s reservation was lost for some other reason. In cases 1 and 2, the system as a whole makes progress in the sense that some processor successfully modifies location X. Case 3 covers reservation loss required for correct operation of the rest of the system. This includes cancellation caused by some other processor writing elsewhere in the reservation granule for X, as well as cancellation caused by the operating system in managing certain limited resources such as real memory or context switches. It may also include implementation-dependent causes of reservation loss. An implementation may make a forward progress guarantee, defining the conditions under which the system as a whole makes progress. Such a guarantee must specify the possible causes of reservation loss in case 3. Although Book E alone cannot provide such a

guarantee, the conditions in cases 1 and 2 are necessary for a guarantee. An implementation and operating system can build on them to provide such a guarantee. Note that Book E does not guarantee fairness. In competing for a reservation, two processors can indefinitely lock out a third. Reservation loss due to granularity Lock words should be allocated such that contention for the locks and updates to nearby data structures do not cause excessive reservation losses due to false indications of sharing that can occur due to the reservation granularity. A processor holding a reservation on any word in a reservation granule loses its reservation if some other processor stores anywhere in that granule. Such problems can be avoided only by ensuring that few such stores occur. This can most easily be accomplished by allocating an entire granule for a lock and wasting all but one word. Reservation granularity may vary for each implementation. There are no architectural restrictions bounding the granularity implementations must support, so reasonably portable code must dynamically allocate aligned and padded memory for locks to guarantee absence of granularity-induced reservation loss. Memory control instructions Memory control instructions can be classified as follows:

  • User- and supervisor-level cache management instructions.
  • Supervisor-level–only translation lookaside buffer management instructions This section describes the user-level cache management instructions. See Supervisor-level memory control instructions on page 183,” for information about supervisor-level cache and translation lookaside buffer management instructions. This section does not describe the cache-locking APU instructions, which are described in Chapter 3.6.4: Cache locking APU on page 200.” Cache management instructions Cache management instructions obey the sequential execution model except as described in the example in this section of managing coherence between the instruction and data caches. In the instruction descriptions the statements. “this instruction is treated as a load” and “this instruction is treated as a store,” mean that the instruction is treated as a load from or a store to the addressed byte with respect to address translation, memory protection, and the memory access ordering done by msync, mbar, and the other means described in Memory access ordering on page 290.” If caches are combined, the same value should be given for an instruction cache attribute and the corresponding data cache attribute. Each implementation provides an efficient way for software to ensure that all blocks that are considered to be modified in the data cache have been copied to main memory before the processor enters any power-saving mode in which data cache contents are not maintained. The means are described in the reference manual for the implementation. It is permissible for an implementation to treat any or all of the cache touch instructions (icbt, dcbt, or dcbtst) as no-operations, even if a cache is implemented.

coherence required and one program both modifies the instructions and executes them. systems provide a system service to perform the function described above. translation resources used for instruction fetches.

  • CT = 0 indicates the L1 cache.
  • CT = 1 indicates the I/O cache. (Note that some versions of the e500 documentation refer to the I/O cache as a frontside L2 cache.)
  • CT = 2 indicates a backside L2 cache. As with other memory-related instructions, the effects of cache management instructions on memory are weakly-ordered. If the programmer must ensure that cache or other instructions have been performed with respect to all other processors and system mechanisms, an msync must be placed after those instructions. Chapter 3.6.4,” describes cache-locking APU instructions.

Table 99. User-level cache instructions An implementation may chose to no-op the instruction.

An alignment interrupt is taken. An implementation may chose to no-op the instruction. translation and protection, and debug address comparisons. An implementation may chose to no-op the instruction. Table 99. User-level cache instructions (continued)

3.3.2 Supervisor level instructions

level instructions defined by the EIS. defines the rfmci for machine check interrupts and rfdi for debug APU interrupts. Table 101 lists instructions for accessing the MSR. An implementation may chose to no-op the instruction.

  1. A program that uses dcbt and dcbtst improperly is less efficient. To improve performance, HID0[NOPTI] can be set, which

latency. The default state of this bit is zero, which enables the use of these instructions. Table 100. System linkage instructions—supervisor-level are restored. rfdi is context-synchronizing. System call sc — The sc instruction is context-synchronizing.

page 1143,” describes context synchronization requirements when altering certain SPRs.

  • Cache management instructions (supervisor-level and user-level)
  • Translation lookaside buffer management instructions This section describes supervisor-level memory control instructions. Memory control instructions on page 179,” describes user-level memory control instructions. Supervisor-level cache instruction Table 102 lists the only supervisor-level cache management instruction. See User-level cache instructions on page 180,” for cache instructions that provide user- level programs the ability to manage the on-chip caches. Supervisor-level tlb management instructions The address translation mechanism is defined in terms of TLBs and page table entries (PTEs) Book E processors use to locate the logical-to-physical address mapping for a particular access. See Chapter 5.4: Storage model on page 301,” for more information about TLB operations. Table 103 summarizes the operation of the TLB instructions.

Table 101. Move to/from machine state register instructions wrteei E The value of E is placed into MSR[EE]. Other MSR bits are unaffected. Table 102. Supervisor-Level cache management instruction coherency required (WIMG = xx1x).

Table 103. TLB management instructions unsuccessful search occurred. msync instruction’s memory barrier is created.

3.3.3 Recommended si mplified mnemonics

The description of each instruction includes the mnemonic and a formatted list of operands. the existence of mnemonics not described in this document.

3.3.4 Book E instr uctions with implementation-specific features

implementation-specific behavior.

3.3.5 EIS instructions

embedded floating-point APU instructions are listed in Table 108 and Table 117. Table 104. Implementation-specific instructions summary Table 105. EIS-defined in structions (except SPE and SPFP instructions)

3.3.6 Context synchronization

instructions until the post-synchronized instruction is completely finished.

3.4 Instruction fetching

redirected after instructions are fetched and before they are scheduled for execution.

3.5 Memory synchronization

See Memory access ordering on page 290,” for detailed information.

3.6 EIS-specific instructions

3.6.1 SPE and embedded floating-point APUs

IEEE-compliant double-precision operands. Table 105. EIS-defined instru ctions (except SPE and SPFP instructions) (continued)

However, like 32-bit Book E instructions, scalar SPFP APU floating-point instructions use bits 32–63 of the GPRs to hold 32-bit single-precision operands, as described in Embedded vector and scalar floating-point APU instructions on page 196.” There is no record form of SPE or embedded floating-point instructions. Vector compare instructions store the result of the comparison into the CR. The meaning of the CR bits is now overloaded for vector operations. Vector compare instructions specify a CR field and two source registers as well as the type of compare: greater than, less than, or equal. Two bits in the CR field are written with the result of the vector compare, one for each element. The two defined bits could be used either by a vector select instruction or by a UISA branch instruction. A partially visible accumulator register is architected for the integer and fractional multiply accumulate SPE instructions. It is described in Chapter 2.14.2 on page 122.” Full descriptions of these instructions can be found in Chapter 13 on page 891.” SPE APU instruction architecture This section describes the instruction formats and instructions defined by the SPE APU. Signed fractions In signed fractional format, the N-bit operand is represented in a 1.[N–1] format (1 sign bit, N–1 fraction bits). Signed fractional numbers are in the following range: The real value of the binary operand SF[0:N-1] is as follows: The most negative and positive numbers representable in fractional format are as follows:

  • The most negative number is represented by SF(0) = 1 and SF[1:N–1] = 0 (that is,
  • The most positive number is represented by SF(0) = 0 and SF[1:N–1] = all 1s (that is, N=32; 0x7FFF_FFFF = 1.0 - 2 –(N–1) ). SPE APU—integer and fractional operations Figure 15 shows data formats for signed integer and fractional multiplication. Note that low word versions of signed saturate and signed modulo fractional instructions are not supported. Attempting to execute an opcode corresponding to these instructions causes boundedly undefined results. SF 1.0 SF 0 ()•–= SF i() 2 i–• i1= N1–

Figure 15. Integer and fractional operations because integer and fractional forms are identical for unsigned operands. Table 106 shows how SPE APU vector multiply instruction mnemonics are structured. Table 107 defines mnemonic extensions for these instructions. Table 106. SPE APU vector multiply instruction mnemonic structure

  1. Low word versions of signed saturate and signed modulo fracti onal instructions are not supported. Attempting to execute

an opcode corresponding to these instructions causes boundedly undefined results.

Table 108 lists SPE APU instructions. Table 107. Mnemonic extensions for multiply-accumulate instructions Table 108. SPE APU vector instructions

Table 108. SPE APU vector in structions (continued)

GPRs allows vector floating-point instructions to use SPE load and store instructions.

  • Vector SPFP instructions operate on a vector of two 32-bit, single-precision floating- point numbers that reside in the upper and lower halves of the 64-bit GPRs. These instructions are listed in Table 117 alongside their scalar equivalents.
  • Scalar SPFP instructions operate on single 32-bit operands that reside in the lower 32- bits of the GPRs. These instructions are listed in Table 117.
  • Scalar DPFP instructions operate on single 64-bit double-precision operands that reside in the 64-bit GPRs. These instructions are listed in Table 109. Note: Note that the vector and scalar versions of the instructions have the same syntax. Vector Subtract Signed, Modulo, Integer to Accumulator Word evsubfsmiaaw r D,rA Vector Subtract Signed, Saturate, Integer to Accumulator Word evsubfssiaaw r D,rA Vector Subtract Unsigned, Modulo, Integer to Accumulator Word evsubfumiaaw r D,rA Vector Subtract Unsigned, Saturate, Integer to Accumulator Word evsubfusiaaw r D,rA Vector XOR evxor r D,rA,rB

Table 109. Vector and scalar floating-point APU instructions

3.6.2 Integer select (isel) APU

helps eliminate branches. Section 7.1: Integer select APU,” describes the use of isel.

3.6.3 Performanc e monitor APU

On some cores, floating-point operations that produce a result of zero may generate an incorrect sign.

  1. Exception detection for these instruct ions is implementation dependent. On some devices, Infinities, NaNs, and Denorms

are always be treated as Norms. No exceptions are taken if SPEFSCR[FINVE] = 1. Table 109. Vector and scalar floating-point APU instructions (continued) Table 110. Integer select APU instruction

monitor APU instructions are described in Table 111. architecture and are accessed by mtpmr and mfpmr, which are also defined by the EIS. supervisor-level PMRs causes a privilege exception. Table 111. Performance monitor APU instructions Table 112. Performance monitor registers—supervisor level

write user-level registers in supervisor or user mode causes an illegal instruction exception. Table 113. Performance monitor registers—user level (read-only) Table 112. Performance monitor registers—supervisor level (continued)

3.6.4 Cache locking APU

  • dcbtls—Data Cache Block Touch and Lock Set
  • dcbtstls—Data Cache Block Touch for Store and Lock Set
  • icbtls—Instruction Cache Block Touch and Lock Set The rA and rB operands to these instructions form a EA identifying the line to be locked. The CT field indicates which cache in the cache hierarchy should be targeted. These instructions are similar to the dcbt, dcbtst, and icbt instructions, but locking instructions can not execute speculatively and may cause additional exceptions. For unified caches, both the instruction lock set and the data lock set target the same cache. Similarly, lines are unlocked from the cache by software using a series of lock-clear instructions. The following instructions are provided to lock instructions into the instruction cache:
  • dcblc—Data Cache Block Lock Clear
  • icblc—Instruction Cache Block Lock Clear The rA and rB operands to these instructions form an EA identifying the line to be unlocked. The CT field indicates which cache in the cache hierarchy should be targeted. Additionally, software may clear all the locks in the cache. For the primary cache, this is accomplished by setting the CLFC (DCLFC, ICLFC) bit in L1CSR0 (L1CSR1). Cache lines can also be implicitly unlocked in the following ways:
  • A locked line is invalidated if it is targeted by a dcbi, dcbf, or icbi instruction.
  • A snoop hit on a locked line that requires the line to be invalidated. This can occur because the data the line contains has been modified external to the processor, or another processor has explicitly invalidated the line.
  • The entire cache containing the locked line is flash invalidated. An implementation is not required to unlock lines if data is invalidated in the cache. Although the data may be invalidated (and thus not in the cache), the cache can keep the lock associated with that cache line present and fill the line from the memory subsystem when the next access occurs. If the implementation does not clear locks when the associated line is invalidated, the method of locking is said to be persistent. An implementation may choose to implement locks as persistent or not persistent; the preferred method is persistent.

Table 114. Cache locking APU instructions protection, and debug address comparisons. protection, and debug address comparisons.

3.6.5 Machine check APU

Machine Check Interrupt instruction (rfmci), is described in Table 115.

3.6.6 VLE extension

  • System linkage instructions on page 201”
  • Processor control register manipulation instructions on page 202”
  • Instruction synchronization instruction on page 202” System linkage instructions Data Cache Block Touch for Store and Lock Set dcbtstls CT,rA,rB It is implementation dependent whether this instruction is treated as a load or store with respect to any memory barriers, synchronization, translation and protection, and debug address comparisons. Instruction Cache Block Lock Clear icblc CT,rA,rB Treated as a load with respect to any memory barriers, synchronization, translation and protection, and debug address comparisons. Instruction Cache Block Touch and Lock Set icbtls CT,rA,rB Treated as a load with respect to any memory barriers, synchronization, translation and protection, and debug address comparisons.

Table 114. Cache locking APU instructions (continued) Table 115. Machine check APU instruction

Table 118 lists the VLE-defined se_isync instruction. Table 116. System linkage instruction set index Table 117. System register manipulation instruction set index Table 118. Instruction Synchronization Instruction Set Index

mode. It also describes the registers that support them.

  • Chapter 2.5.1: Condition register (CR) on page 61”
  • Chapter 2.5.2: Link register (LR) on page 66”
  • Chapter 2.5.3: Count register (CTR) on page 67” Branch instructions The sequence of instruction execution can be changed by the branch instructions. Because VLE instructions must be aligned on half-word boundaries, the low-order bit of the generated branch target address is forced to 0 by the processor in performing the branch. The branch instructions compute the EA of the target in one of the following ways, as described in Chapter 10.2: Instruction memory addressing modes on page 854 .” 1. Adding a displacement to the address of the branch instruction. 2. Using the address contained in the LR (Branch to Link Register [and Link]). 3. Using the address contained in the CTR (Branch to Count Register [and Link]). Branching can be conditional or unconditional, and the return address can optionally be provided. If the return address is to be provided (LK = 1), the EA of the instruction following the branch instruction is placed into the LR after the branch target address has been computed: this is done whether or not the branch is taken. In branch conditional instructions, the BI32 or BI16 instruction field specifies the CR bit to be tested. For 32-bit instructions using BI32, CR[32–47] (corresponding to bits in CR0–CR3) may be specified. For 16-bit instructions using BI16, only CR[32–35] (bits within CR0) may be specified. In branch conditional instructions, the BO32 or BO16 field specifies the conditions under which the branch is taken and how the branch is affected by or affects the CR and CTR. Note that VLE instructions also have different encodings for the BO32 and BO16 fields than in Book E’s BO field. If the BO32 field specifies that the CTR is to be decremented, CTR[32–63] are decremented. If BO[16,32] specifies a condition that must be TRUE or FALSE, that condition is obtained from the contents of CR[BI+32]. (Note that CR bits are numbered 32– 63. BI refers to the BI field in the branch instruction encoding. For example, specifying BI = 2 refers to CR[34].) Encodings for the BO32 field for the VLE extension are shown in Table 120.

Table 119. VLE extension BO32 encodings 00 Branch if the condition is FALSE. 01 Branch if the condition is TRUE.

The encoding for the BO16 field for the VLE extension is shown in Table 120. The various branch instructions supported by the VLE extension are shown in Table 121. 10 Decrement CTR[32–63] , then branch if the decremented CTR[32–63]≠0. 11 Decrement CTR[32–63], then branch if the decremented CTR[32–63] = 0. Table 120. VLE extension BO16 encodings 0 Branch if the condition is FALSE. 1 Branch if the condition is TRUE. Table 121. Branch instruction set index Table 122. Condition register instruction set index

This section lists the integer instructions supported by the VLE extension. The VLE extension supports both big- and little-endian byte ordering for data accesses. load with update instructions in Book E. Table 123. Basic integer load instruction set index

Integer load byte-reversed instructions are listed in Table 124. The VLE-defined integer load multiple instruction is listed in Table 125. The VLE-defined integer load and reserve instruction is listed in Table 126. The VLE extension supports both big- and little-endian byte ordering for data accesses. Table 124. Integer load byte-reverse instruction set index Table 125. Integer load multiple instruction set index Table 126. Integer load and reserve instruction set index

EA. For these forms, the following rules (from Book E) apply.

  • If rA ≠ 0, the EA is placed into GPR(rA).
  • If rS = rA, the contents of GPR(rS) are copied to the target memory element and then EA is placed into GPR(rA). The basic integer store instructions are listed in Table 127. The integer store byte-reverse instructions are listed in Table 128. The integer store multiple instruction is listed in Table 129. The integer store conditional instruction is listed in Table 130.

Table 127. Basic integer store instruction set index Table 128. Integer store byte-reverse instruction set index Table 129. Integer store multiple instruction set index Table 130. Integer store conditional instruction set index

place results into GPRs, into status bits in the XER and into CR0. integers unless the instruction is explicitly identified as performing an unsigned operation. 32–63 of the result to zero. e_addic[.] and e_subfic[.] always set CA to reflect the carry out of bit 32. The integer arithmetic instructions are listed in Table 131. Table 131. Integer arithmetic instruction set index

another GPR, or an immediate value. page 208.” The logical instructions do not change XER[SO,OV,CA]. The integer logical instructions are listed in Table 132. Table 131. Integer arithmetic instruction set index (continued)

Table 132. Integer logical instruction set index

  • The value of the SCI8 field
  • The zero-extended value of the UI field
  • The zero-extended value of the UI5 field
  • The sign-extended value of the SI field
  • The contents of GPR(rB) or GPR(rY). The following comparisons are signed: e_cmph, e_cmpi, e_cmp16i, e_cmph16i, se_cmp, se_cmph, and se_cmpi. The following comparisons are unsigned: e_cmphl, e_cmpli, e_cmphl16i, e_cmpl16i, se_cmpli, se_cmpl, and se_cmphl. When operands are treated as 32-bit signed quantities, GPRn[32] is the sign bit. When operands are treated as 16-bit signed quantities, GPRn[48] is the sign bit. For 32-bit implementations, the L field must be zero. Compare instructions set one of the left-most three bits of the designated CR field and clears the other two. XER[SO] is copied to bit 3 of the designated CR field. The CR field is set as shown in Table 133. The integer bit test instruction tests the bit specified by the UI5 instruction field and sets the CR0 field as shown in Table 134. orc rA,rS,rB orc. rA,rS,rB OR with Complement Book E e_ori[.] rA,rS,SCI8 e_or2i rD,UI OR Immediate Page -966 e_or2is rD,UI OR Immediate Shifted Page -966 xor rA,rS,rB xor. rA,rS,rB XOR Book E e_xori[.] rA,rS,SCI8 XOR Immediate Page -966

Table 133. CR settings for compare instructions

3 SO Summary overflow from the XER

Table 132. Integer logical instruction set index (continued)

Table 135 is an index for integer compare and bit test operations. destination register under the control of a predicate value supplied by a CR bit. The integer select instruction is listed in Table 136. Table 134. CR settings for integer bit test instructions

0 LT Always cleared

Table 135. Integer compare and bit test instruction set index Table 136. Integer select instruction set index

instruction execution continues normally. the contents of bits 32–63 of rA (and rB) participate in the comparison. The integer trap instruction is listed in Table 138. result, or a portion of the result, to a GPR. The rotation operations rotate a 32-bit quantity left by a specified number of bit positions. Bits that exit from position 32 enter at position 63. The rotate32 operation is used to rotate a given 32-bit quantity. There is no way to specify an all-zero mask. Table 137. Integer trap conditions

0 Less Than, using signed comparison

1 Greater Than, using signed comparison

2 Equal

3 Less Than, using unsigned comparison

4 Greater Than, using unsigned comparison

Table 138. Integer trap instruction set index

The use of the mask is described in following sections. and shift instructions, except algebraic right shifts, do not change the CA bit. type, the amount of the rotation is either specified as an immediate, or contained in a GPR. ANDed with a mask before being placed into the target register. (in concept) by a left-rotation of 32-n, where n is the number of bits by which to rotate right. concept) by a left-rotation of 32-n, where n is the number of bits by which to rotate right. The integer shift instructions are listed in Table 141. Table 139. Integer rotate instruction set index Table 140. Integer rotate with mask instruction set index Table 141. Integer shift instruction set index

The storage synchronization instructions are listed in Table 142. The cache management instructions are listed in Table 143. defined in Book E and in the EIS. The TLB management instructions are listed in Table 144. Table 142. Storage synchronization instruction set index Table 143. Cache management instruction set index

This section lists instructions either defined or supported by the VLE extension. Table 145 lists instructions by instruction name. Table 144. TLB management instruction set index Table 145. Instructions listed by name

Table 145. Instructions listed by name (continued)

Table 145 lists instructions that can be executed in VLE mode by mnemonic.

Table 146. Instructions listed by mnemonic

Table 146. Instructions listed by mnemonic (continued)

3.7 Instruction listing

Table 147 lists instructions defined in Book E, in the PowerPC architecture, and by the EIS. A check mark (√) or text in a column indicates that the instruction is defined or implemented. extension that defines the instruction. Table 147. List of instructions

Table 147. List of instructions (continued)

4 Interrupts and exceptions

implementation standards (EIS). interrupt handler address with a modified MSR.

  • An exception is the event that, if enabled, causes the processor to take an interrupt.

peripherals, instructions, the internal timer facility, debug events, or error conditions.

4.1 Overview

the interrupt handler returns control to the interrupted program. types which may be implemented on ST Book E devices. These are described in Table 148. Table 148. Interrupt types identical to those defined by the OEA. SRR0/SRR1 SPRs and rfi instruction. during regular program flow.

the Book E–defined critical interrupt type. machine check enable bit, MSR[ME]. defined critical interrupt type. machine check enable bit, MSR[DE].

RM0004 Interrupts and exceptions

4.2 EIs interrupt definitions

This section gives an overview of additions and modifications to the Book E interrupt model defined by the EIS. Specific details are also provided throughout this chapter. Except for the following, the core complex reports exceptions as specified in Book E:

  • The machine check exception differs as follows: – It is not processed as a critical interrupt, but uses MCSRR0 and MCSRR1 for saving the return address and the MSR in case the machine check is recoverable. – Return From Machine Check Interrupt instruction ( rfmci) is implemented to support the return to the address saved in MCSRR0. – A machine check syndrome register, MCSR, logs the cause of the machine check (instead of ESR). The core complex reports the machine check exception as described in Chapter 4.7.2.”
  • The following interrupts are defined for use with the embedded floating-point and signal-processing (SPE) APUs: – SPE/embedded floating-point unavailable interrupt. IVOR32 (SPR 528) contains the vector offset. See SPE/embedded floating-point APU unavailable interrupt on page 272.” – Embedded floating-point data interrupt. IVOR33 (SPR 529) contains the vector offset. See Embedded floating-point data interrupt on page 272.” – Embedded floating-point round interrupt. IVOR34 (SPR 530) contains the vector offset. See Embedded floating-point round interrupt on page 273.” The following additional bits are defined to support SPE and SPFP exceptions: – MSR[38] is defined as the ve ctor available bit (SPE). If this bit is clear and software attempts to execute any of the SPE instructions, the SPE unavailable interrupt is taken. If this bit is set, software can execute any SPE instructions. Note: SPFP instructions require MS R[SPE] to be set. An attempt to execute an SPFP instruction when MSR[SPE] is 0 causes an SPE APU unavailable interrupt. Embedded vector and scalar floating-point APU instructions on page 196,” lists affected instructions. – ESR[SPE], the SPE exception bit, is set wh en the processor reports an exception related to the execution of SPFP or SPE instructions.
  • The debug exception implementation does not support IAC3, IAC4, DAC3, and DAC4 comparisons.
  • The core complex supports instruction address compare (IAC1 and IAC2) and data address compare (DAC1 and DAC2) for effective addresses only. Real-address support is not provided.
  • Some implementations do not support the Book E-defined floating-point unavailable and auxiliary processor unavailable interrupts.
  • Data value compare (DVC) debug exceptions are not supported.
  • The interrupt priorities differ from those specified in Book E as described in Chapter 4.11.”
  • Alignment exceptions. Vector operations can cause alignment exceptions as described in Chapter 4.7.6.”
  • Book E and the machine check APU define sources of externally generated interrupts.

Interrupts and exceptions RM0004

4.2.1 Recoverability from interrupts

All interrupts except some machine check interrupts are recoverable. The state of the core complex (return address and MSR contents) is saved when a machine check interrupt is taken. The conditions that cause a machine check may or may not prohibit recovery.

4.3 Interrupt registers

Table 149 summarizes registers used for interrupt handling. These registers are described in detail in Chapter 2.”

Table 149. Interrupt registers defined by the PowerPC architecture noncritical interrupt is serviced. When a noncritical interrupt is taken, MSR contents are placed into SRR1. is reserved may be altered by rfi. return to after a critical interrupt is serviced. When a critical interrupt is taken, MSR contents are placed into CSRR1. that is reserved may be altered by rfci. routine. IVPR[48–63] are reserved.

register (ESR) on page 84,” shows ESR bit definitions. the machine check syndrome register (MCSR) rather than the ESR. Table 149. Interrupt registers defined by the PowerPC architecture (continued)

each interrupt type. IVOR0–IVOR15 are provided for defined interrupt types. by interrupts defined by the EIS.) IVOR assignments are shown below. word of a 64-bit GPR, an SPE APU unavailable interrupt is taken. 1: Software can execute any embedded floating-point or SPE instructions. instruction to return to after a machine check interrupt is serviced. MCSRR1. When rfmci is executed, MCSRR1 contents are restored to MSR. that an MSR bit that is reserved may be altered by rfmci.

check condition is recoverable. ABIST status is logged in MCSR[48–54]. defined ESR for this purpose.

RM0004 Interrupts and exceptions

4.4 Exceptions

Exceptions are caused directly by instruction execution or by an asynchronous event. In either case, the exception may cause one of several types of interrupts to be invoked. The following examples are of exceptions caused directly by instruction execution:

  • An attempt to execute a reserved-illegal instruction (illegal instruction exception-type program interrupt)
  • An attempt by an application program to execute a privileged instruction or to access a privileged SPR (privileged instruction exception-type program interrupt)
  • In general, an attempt by an application program to access a nonexistent SPR (unimplemented operation instruction exception-type program interrupt). Note the following behavior defined by the EIS: – If MSR[PR] = 1 (user mode), SPR bit 5 = 0 (user-accessible SPR), and the SPR number is invalid, an illegal instruction exception is taken. – If MSR[PR] = 0 (supervisor mode) and the SPR number is invalid, an illegal instruction exception is taken. – If MSR[PR] = 1, SPR bit 5 = 1, and invalid SPR address (supervisor-only SPR), a privileged instruction exception-type program interrupt is taken.
  • Execution of a defined instruction using an invalid form (illegal instruction exception- type program interrupt, unimplemented operation exception-type program interrupt, or privileged instruction exception-type program interrupt).
  • An attempt to access a location that is either unavailable (instruction or data TLB error interrupt) or not permitted (instruction or data storage interrupt)
  • An attempt to access a location with an effective address alignment not supported by the implementation (alignment interrupt)
  • Execution of a System Call (sc) instruction (system call interrupt)
  • Execution of a trap instruction whose trap condition is met (trap interrupt type)
  • Execution of a floating-point instruction when floating-point instructions are unavailable (floating-point unavailable interrupt)
  • Execution of a floating-point instruction that causes a floating-point enabled exception to exist (enabled exception-type program interrupt)
  • Execution of a defined instruction that is not implemented (illegal instruction exception or unimplemented operation exception-type program interrupt)
  • Execution of an allocated instruction that is not implemented (illegal instruction exception or unimplemented operation exception-type program interrupt)
  • Execution of an allocated instruction when the auxiliary instruction is unavailable (auxiliary unavailable interrupt)
  • Execution of an allocated instruction that causes an auxiliary enabled exception (enabled exception-type program interrupt) Invocation of an interrupt is precise, except that if one of the imprecise modes for invoking a floating-point enabled exception-type program interrupt is in effect, the invocation may be imprecise. When the interrupt is invoked imprecisely, the excepting instruction does not appear to complete before the next instruction starts (because the invocation of the interrupt required to complete execution has not occurred).

4.5 Interrupt classes

  • Critical/noncritical. Some interrupt types demand immediate attention even if other interrupt types being processed have not had the opportunity to save the machine state (that is, return address and captured state of the MSR). To enable taking a critical interrupt immediately after a noncritical interrupt is taken (that is, before the machine state is saved), two sets of save/restore register pairs are provided. Critical interrupts use CSRR0/CSRR1, and noncritical interrupts use SRR0/SRR1.
  • Asynchronous/synchronous. Asynchronous interrupts are caused by events external to instruction execution; synchronous interrupts are caused by instruction execution and are either precise or imprecise. Table 150 describes asynchronous and synchronous interrupts.

Table 150. Asynchronous and synchronous interrupts

4.5.1 Requirements for sy stem reset generation

  • Assertion of a signal that resets the internal state of the core complex
  • By writing a 1 to DBCR0[34], if MSR[DE] = 1 Synchronous, Precise Caused directly by instruction execution. Synchronous interrupts are precise or imprecise. These interrupts precisely indicate the address of the instruction causing the exception or, for certain synchronous, precise interrupt types, the address of the immediately following instruction. When the execution or attempted execution of an instruction causes a synchronous, precise interrupt, the following conditions exist at the interrupt point: Whether SRR0 or CSRR0 addresses the instruction causing the exception or the next instruction is determined by the interrupt type and status bits. An interrupt is generated such that all instructions before the instruction causing the exception appear to have completed with respect to the executing processor. However, some accesses associated with these preceding instructions may not have been performed with respect to other processors and mechanisms. The exception-causing instruction may appear not to have begun execution (except for causing the exception), may be partially executed, or may have completed, depending on the interrupt type. See Chapter 4.9.” Architecturally, no instruction beyond the exception-causing instruction executed. Synchronous, Imprecise Imprecise interrupts may indicate the address of the instruction causing the exception that generated the interrupt or some instruction after that instruction. When execution or attempted execution of an instruction causes an imprecise interrupt, the following conditions exist at the interrupt point. SRR0 or CSRR0 addresses either the exception-causing instruction or some instruction following the exception-causing instruction that generated the interrupt. An interrupt is generated such that all instructions preceding the instruction addressed by SRR0 or CSRR0 appear to have completed with respect to the executing processor. If context synchronization forces the imprecise interrupt due to an instruction that causes another exception that generates an interrupt (for example, alignment or data storage interrupt), SRR0 addresses the interrupt-forcing instruction, which may have partially executed (see Chapter 4.9”). If execution synchronization forces an imprecise interrupt due to an execution- synchronizing instruction other than msync or isync, SRR0 or CSRR0 addresses the interrupt-forcing instruction, which appears not to have begun execution (except for its forcing the imprecise interrupt). If the interrupt is forced by msync or isync, SRR0 or CSRR0 may address msync or isync, or the following instruction. If context or execution synchronization forces an imprecise interrupt, the instruction addressed by SRR0 or CSRR0 may have partially executed (see Chapter 4.9”). No instruction following the instruction addressed by SRR0 or CSRR0 has executed.

Table 150. Asynchronous and synchronous interrupts (continued)

Interrupts and exceptions RM0004

4.6 Interrupt processing

Associated with each kind of interrupt is an interrupt vector, the address of the initial instruction that is executed when an interrupt occurs. Interrupt processing consists of saving a small part of the processor’s state in certain registers, identifying the cause of the interrupt in another register, and continuing execution at the corresponding interrupt vector location. When an exception exists that causes an interrupt to be generated and it has been determined that the interrupt can be taken, the following steps are performed: 1. SRR0 (for noncritical class interrupts) or CSRR0 (for critical class interrupts) or MCSRR0 for machine check interrupts is loaded with an instruction address that depends on the type of interrupt; see the specific interrupt description for details. 2. The ESR or MCSR is loaded with informatio n specific to the exception type. Note that many interrupt types can only be caused by a single type of exception event, and thus do not need nor use an ESR setting to indicate the cause of the interrupt. 3. SRR1 (for noncritical class interrupts) or CSRR1 (for critical class interrupts) or MCSRR1 for machine check interrupts is loaded with a copy of the MSR contents. 4. New MSR values take effect beginning with the first instruction following the interrupt. The MSR is updated as follows: – MSR[SPE,WE,EE,PR,FP ,FE0,FE1,IS,DS] are cleared by all interrupts. – MSR[CE,DE] are cleared by critical class interrupts and unchanged by noncritical class interrupts. – MSR[ME] is cleared by machine check interrupts and unchanged by other interrupts. – Other defined MSR bits are unchanged by all interrupts. MSR fields are described in Chapter 2.6.1: Machine state register (MSR) on page 68.” 5. Instruction fetching and execution resumes, using the new MSR value, at a location specific to the interrupt type (IVPR[32–47] || IVORn[48–59] || 0b0000) The IVORn for the interrupt type is described in Table 151. IVPR and IVOR contents are indeterminate upon reset and must be initialized by system software. Interrupts do not clear reservations obtained with load and reserve instructions. The operating system should do so at appropriate points, such as at process switch. At the end of a noncritical interrupt handling routine, executing rfi causes the MSR to be restored from the SRR1 contents and instruction execution to resume at the address contained in SRR0. Likewise, rfci and rfmci perform the same function at the end of critical and machine check interrupt handling routines respectively, using the critical and machine check save/restore registers. Note: In general, at process switch, due to possible process interlocks and possible data availability requirements, the operating system needs to consider executing the following: stwcx.—Clears outstanding reservations to prevent pairing a lwarx in the old process with a stwcx. in the new one msync—Ensures that memory operations of an interrupted process complete with respect to other processors before that process begins executing on another processor rfi, rfci, rfmci, or isync—Ensures that instructions in the new process execute in the new context

4.7 Interrupt definitions

the interrupt type, and which IVOR is used to specify the vector address. Table 151. Interrupt and exception types

  1. A = asynchronous, C = critical, SI = synchronous , imprecise, SP = synchronous, precise
  2. In general, when an interrupt causes an ESR bit or bits to be set (or cleared) as indicated in the table, it also causes all

Legend: xxx (no brackets) means ESR[xxx] is set. [xxx] means ESR[xxx] could be set. [xxx,yyy] means either ESR[xxx] or ESR[yyy] may be set, but never both. {xxx,yyy} means either ESR[xxx] or ESR[yyy] may be set, or possibly both.

  1. Although not part of Book E, system interrupt controllers commonly provide independent mask and status bits for critical

input and external input interrupt sources. Table 151. Interrupt and exception types (continued)

4.7.1 Critical input interrupt

implementations may provide other ways to mask the critical input interrupt. CSRR0, CSRR1, and MSR are updated as shown in Table 152. Instruction execution resumes at address IVPR[32–47] || IVOR0[48–59] || 0b0000. on whether MSR[CE] is set when the critical interrupt signal is asserted. implementation to clear any critical input exception status before reenabling MSR[CE].

  1. Machine check status information is commonly provided as par t of the system implementation but is not part of Book E.
  2. Software must examine the instruction and the subject TL B entry to determine the exact cause of the interrupt.
  3. Cache locking and cache locking ex ceptions are implementation-dependent.
  4. The precision of the floating-point enabled exception ty pe is controlled by MSR[FE0,FE1], as described in <Cross

The precision of the auxiliary processor enabled exception type program interrupt is implementation-dependent.

  1. Auxiliary processor ex ception status is commonly provided as part of the implementation and is not part of Book E.
  2. Instruction complete and branch taken debug events ar e defined only for MSR[DE] = 1 for internal debug mode

events cannot occur, no DBSR status bits are set, and no subsequent imprecise debug interrupt can occur. Table 152. Critical input interrupt register settings MSR ME is unchanged. All other MSR bits are cleared.

4.7.2 Machine check interrupt

  • Book E defines machine check interrupts as critical interrupts, but the machine check APU treats them as a distinct interrupt type.
  • Machine check is no longer a critical interrupt but uses MCSRR0 and MCSRR1 to save the return address and the MSR in case the machine check is recoverable.
  • Return from machine check interrupt instruction (rfmci) is implemented to support the return to the address saved in MCSRR0.
  • An address related to the machine check may be stored in MCAR, according to Table 153.
  • A machine check syndrome register, MCSR, is used to log the cause of the machine check (instead of ESR). The MCSR is described in Table 153. The following general information applies to both the Book E and EIS definitions. A machine check interrupt occurs when no higher priority exception exists, a machine check exception is presented to the interrupt mechanism, and MSR[ME] = 1. Specific causes of machine check exceptions are implementation-dependent, as are the details of the actions taken on a machine check interrupt. Machine check interrupts are typically caused by a hardware or memory subsystem failure or by an attempt to access an invalid address. They may be caused indirectly by execution of an instruction, but may not be recognized or reported until long after the processor has executed past the instruction that caused the machine check. As such, machine check interrupts are not thought of as synchronous or asynchronous nor as precise or imprecise. The following general rules apply:
  • No instruction after the one whose address is reported to the machine check interrupt handler in MCSRR0 has begun execution.
  • The instruction whose address is reported to the machine check interrupt handler in MCSRR0 and all prior instructions may or may not have completed successfully. All instructions certain to complete appear to have done so within the context existing before the machine check interrupt. No further interrupts (other than possible additional machine check interrupts) occur as a result of those instructions. If MSR[ME] is cleared, the processor enters checkstop state immediately on detecting the machine check condition. When a machine check interrupt is taken, registers are updated as shown in Table 153.

Table 153. Machine check interrupt settings MSR UCLE, SPE, WE, CE, EE, PR, FP , ME, FE0, FE1, DE, IS, DS, PMM, and RI are cleared. ESR Implementation-dependent. The EIS us es the MCSR rather than the ESR.

Instruction execution resumes at address IVPR[32–47] || IVOR1[48–59] || 0b0000. incorrect data, which may be placed into registers or on-chip caches.

2 For implementations on which a machine check interrupt is caused by referring to an invalid

which the block is the target for replacement or as the result of executing dcbst or dcbf.

4.7.3 Data storage interrupt

conditions for a data storage interrupt as defined by Book E. Machine check address register (MCAR/MCARU) on page 88. MCARU is an alias to the upper 32 bits of MCAR. MCSR Set according to the machine check condition. See Table 20.

  1. These registers are us ed if the machine check APU is not implemented.

Table 153. Machine check interrupt settings (continued) Table 154. Data storage interrupt exception conditions

  • In user mode (MSR[PR] = 1), a load or load-class cache management instruction attempts to access a memory location that is not user-mode read enabled (page access control bit UR = 0).
  • In supervisor mode (MSR[PR] = 0), a load or load-class cache management instruction attempts to access a location that is not supervisor-mode read enabled (page access control bit SR = 0). Write access control exception Occurs when either of the following conditions exists:
  • In user mode (MSR[PR] = 1), a store or store-class cache management instruction attempts to access a location that is not user-mode write enabled (page access control bit UW = 0).
  • In supervisor mode (MSR[PR] = 0), a store or store-class cache management instruction attempts to access a location that is not supervisor-mode write enabled (page access control bit SW = 0).

cause a data storage interrupt, regardless of the effective address. causing the data storage exception. cause a byte-ordering exception. (EIS) The locked state of one or more cache lines has the potential to be altered.

  • For icbtls and icblc, ESR[ILK] is set.
  • For dcbtls, dcbtstls, or dcblc, ESR[DLK] is set. Book E refers to this as a cache-locking exception. Storage synchronization exception Occurs when either of the following conditions exists:
  • An attempt is made to execute a load and reserve or store conditional instruction from or to a location that is write-through required or caching inhibited. (If the interrupt does not occur, the instruction executes correctly.)
  • A store conditional instruction produces an effective address for which a normal store would cause a data storage interrupt but the processor does not have the reservation from a load and reserve instruction. Book E states that it is implementation-dependent whether a data storage interrupt occurs. The EIS defines that the data storage interrupt is taken.

Table 154. Data storage interrupt exception conditions (continued)

Instruction execution resumes at address IVPR[32–47] || IVOR2[48–59] || 0b0000.

4.7.4 Instruction storage interrupt

exception conditions are described in Table 156. class of accesses, or do not support misaligned accesses using a specific byte order. instruction causing the exception. SRR0, SRR1, MSR, and ESR are updated as shown in Table 157. Table 155. Data Storage Interrupt Register Settings All other defined ESR bits are cleared. MSR CE, ME, and DE are unchanged. All other MSR bits are cleared. Table 156. Instruction storage interrupt exception conditions not user mode execute enabled (page access control bit UX = 0). that is not supervisor mode execute enabled (page access control bit SX = 0). boundary such that endianness changes cause a byte-ordering exception.

determine whether a permissions violation also may have occurred. Instruction execution resumes at address IVPR[32–47] || IVOR3[48–59] || 0b0000.

4.7.5 External input interrupt

assertion of an asynchronous signal that is part of the processing system. taken depends on whether MSR[EE] is set when the external interrupt signal is asserted. In addition to MSR[EE], implementations may provide other ways to mask this interrupt. SRR0, SRR1, and MSR are updated as shown in Table 158. Instruction execution resumes at address IVPR[32–47] || IVOR4[48–59] || 0b0000. clear any external input exception status before reenabling MSR[EE].

4.7.6 Alignment interrupt

  • The operand of a load or store is not aligned.
  • The instruction is a move assist, load multiple, or store multiple.
  • A dcbz operand is in write-through-required or caching-inhibited memory, or dcbz is executed in an implementation with no data cache or a write-through data cache.
  • The operand of a store, except store conditional, is in write-through required memory.

Table 157. Instruction storage interrupt register settings MSR CE, ME, and DE are unchanged. All other MSR bits are cleared. other defined ESR bits are cleared. Table 158. External input interrupt register settings MSR CE, ME, and DE are unchanged. All other MSR bits are cleared.

  • Execution of a dcbz references a page marked as write-through or cache inhibited.
  • A load multiple word instruction (lmw) reads an address that is not a multiple of four.
  • A lwarx or stwcx. instruction references an address that is not a multiple of four.
  • SPFP and SPE APU instructions are not aligned on a natural boundary. A natural boundary is defined by the size of the data element being accessed.
  • A vector operation reports an exception if the physical address of the following instructions is not aligned to the 64-bit boundary: evldd, evlddx, evldw, evldwx, evldh, evldhx, evstdd, evstddx, evstdw, evstdwx, evstdh, and evstdhx. Table 159 describes additional ESR settings. For lmw and stmw with a non–word-aligned operand and for load and reserve and store conditional instructions with an misaligned operand, an implementation may yield boundedly undefined results instead of causing an alignment interrupt. A store conditional to a write- through required location may either cause an alignment or data storage interrupt or may correctly execute the instruction. For all other cases listed above, an implementation may execute the instruction correctly instead of causing an alignment interrupt. For dcbz, correct execution means clearing each byte of the block in main memory. Note: Book E does not support use of an misaligned effective address by load and reserve and store conditional instructions. If an misaligned effective address is specified, the alignment interrupt handler should treat the instruction as a programming error and must not attempt to emulate the instruction. When an alignment interrupt occurs, the processor suppresses the execution of the instruction causing the alignment exception. SRR0, SRR1, MSR, DEAR, and ESR are updated as shown in Table 159. Instruction execution resumes at address IVPR[32–47] || IVOR5[48–59] || 0b0000.

Table 159. Alignment interrupt register settings MSR CE, ME, and DE are unchanged. All other MSR bits are cleared. All other defined ESR bits are cleared.

4.7.7 Program interrupt

of the following exceptions occurs during execution of an instruction. Table 160. Program interrupt exception conditions processor is in supervisor mode (MSR[PR] = 0), results are undefined. parentheses. See the user’s manual for the implementation. – An instruction that is in invalid form (boundedly undefined results). – A reserved no-op instruction (no-operation performed is preferred). an illegal instruction exception. Trap exception Occurs when any of the conditi ons specified in a trap instruction are met. reference manual for the implementation.

MSR[FE0,FE1], as described in Table 161. SRR0, SRR1, MSR, and ESR are updated as shown in Table 162. Instruction execution resumes at address IVPR[32–47] || IVOR6[48–59] || 0b0000. Table 161. MSR[FE0,FE1] settings these bit settings and treat all affected interrupts as precise.

00 The interrupt is masked and the interrupt su bsequently occurs if and when floating-point

and also causes ESR[PIE] to be set. Table 162. Program interrupt register settings Table 164), set to the EA of the instruction that caused the interrupt. executed next, not with the EA of the instruction that modified MSR causing the interrupt. SRR1 Set to the MSR contents at the time of the interrupt. MSR CE, ME, and DE are unchanged. All other MSR bits are cleared. ESR PIL Set if an illegal instruction exception- type program interrupt; otherwise cleared. PPR Set if a privileged instruction exception- type program interrupt; otherwise cleared. PTR Set if a trap exception-type program interrupt; otherwise cleared. instruction that caused FPSCR[FEX] to be set); otherwise cleared. All other defined ESR bits are cleared.

4.7.8 Floating-point unavailable interrupt

and moves), and MSR[FP] = 0. the instruction causing the floating-point unavailable interrupt. SRR0, SRR1, and MSR are updated as shown in Table 163. Instruction execution resumes at address IVPR[32–47]||IVOR7[48–59]||0b0000.

4.7.9 System call interrupt

(sc) instruction is executed. SRR0, SRR1, and MSR are updated as shown in Table 164. Instruction execution resumes at address IVPR[32–47] || IVOR8[48–59] || 0b0000.

4.7.10 Auxiliary processor unavailable interrupt

manual for the implementation. execution of the instruction causing the auxiliary processor unavailable interrupt. Registers SRR0, SRR1, and MSR are updated as shown in Table 165. Table 163. Floating-point unavailable interrupt register settings SRR0 Set to the effective address of the instruction that caused the interrupt. SRR1 Set to the MSR contents at the time of the interrupt. MSR CE, ME, and DE are unchanged. All other MSR bits are cleared. Table 164. System call interrupt register settings SRR0 Set to the effective address of the instruction after the sc instruction. SRR1 Set to the MSR contents at the time of the interrupt. MSR CE, ME, and DE are unchanged. All other MSR bits are cleared.

Instruction execution resumes at address IVPR[32–47]||IVOR9[48–59]||0b0000.

4.7.11 Decrementer Interrupt

exception exists (TSR[DIS] = 1) & the interrupt is enabled (TCR[DIE] = 1 and MSR[EE] = 1). MSR[EE] also enables external input and fixed-interval timer interrupts. SRR0, SRR1, MSR, and TSR are updated as shown in Table 166. Instruction execution resumes at address IVPR[32–47] || IVOR10[48–59] || 0b0000. but a mask. Writing a 1 to this bit causes it to be cleared; writing a 0 has no effect.

4.7.12 Fixed-interv al timer interrupt

interval timer exception on a transition from 0 to 1. Note: MSR[EE] also enables external input and decrementer interrupts. SRR0, SRR1, MSR, and TSR are updated as shown in Table 167. Table 165. Auxiliary processor unavailable interrupt register settings SRR0 Set to the effective address of the instruction that caused the interrupt. SRR1 Set to the MSR contents at the time of the interrupt. MSR CE, ME, and DE are unchanged. All other MSR bits are cleared. Table 166. Decrementer interrupt register settings SRR0 Set to the effective address of the next instruction to be executed. SRR1 Set to the MSR contents at the time of the interrupt. MSR CE, ME, and DE are unchanged. All other MSR bits are cleared.

Instruction execution resumes at address IVPR[32–47] || IVOR11[48–59] || 0b0000. but a mask. Writing a 1 causes the bit to be cleared; writing a 0 has no effect.

4.7.13 Watchdog timer interrupt

MSR[CE] also enables the critical input interrupt. CSRR0, CSRR1, MSR, and TSR are updated as shown in Table 168. Instruction execution resumes at address IVPR[32–47] || IVOR12[48–59] || 0b0000. but a mask. Writing a 1 to this bit causes it to be cleared; writing a 0 has no effect.

4.7.14 Data tlb error interrupt

described in Table 169 is presented to the interrupt mechanism. Table 167. Fixed-interval timer interrupt register settings SRR0 Set to the effective address of the next instruction to be executed. SRR1 Set to the MSR contents at the time of the interrupt. MSR CE, ME, and DE are unchanged. All other MSR bits are cleared. Table 168. Watchdog timer interrupt register settings CSRR0 Set to the effective address of the next instruction to be executed. CSRR1 Set to the MSR contents at the time of the interrupt. MSR ME is unchanged; all other MSR bits are cleared. Table 169. Data tlb error interrupt exception conditions

whether a data TLB error interrupt occurs. The EIS defines that the interrupt is taken. instruction causing the data TLB error exception. SRR0, SRR1, MSR, DEAR, and ESR are updated as shown in Table 170. are set as defined by the EIS. Instruction execution resumes at address IVPR[32–47] || IVOR13[48–59] || 0b0000.

4.7.15 Instruction tlb error interrupt

exception described in Table 171 is presented to the interrupt mechanism. instruction causing the instruction TLB miss exception. SRR0, SRR1, and MSR are updated as shown in Table 172. Table 170. Data tlb error interrupt register settings SRR0 Set to the effective address of the inst ruction causing the data TLB error interrupt. SRR1 Set to the MSR contents at the time of the interrupt. MSR CE, ME, and DE are unchanged. All other MSR bits are cleared. caused the data TLB error exception. All other defined ESR bits are cleared. Table 171. Instruction TLB error interrupt exception conditions

Instruction execution resumes at address IVPR[32–47] || IVOR14[48–59] || 0b0000.

4.7.16 Debug interrupt

exception occurs when a debug event causes a corresponding DBSR bit to be set. CSRR0, CSRR1, MSR, and DBSR are updated as shown in Table 173. Instruction execution resumes at address IVPR[32–47] || IVOR15[48–59] || 0b0000.

4.7.17 EIS-defined interrupts

The interrupts in this section are defined by the EIS. Table 172. Instruction TLB error interrupt register settings SRR0 Set to the effective address of the instruct ion causing the instruction TLB error interrupt. SRR1 Set to the MSR contents at the time of the interrupt. MSR CE, ME, and DE are unchanged. All other MSR bits are cleared. Table 173. Debug interrupt register settings that would have executed after the one that caused the debug interrupt. instruction that would have executed next if the debug interrupt had not occurred. interrupt that caused the interrupt taken debug event. that would have executed after the rfi, rfci, or rfmci that caused the debug interrupt. modified DBCR0 or MSR and thus caused the interrupt. CSRR1 Set to the MSR contents at the time of the interrupt. MSR ME is unchanged. All other MSR bits are cleared. DBSR Set to indicate type of debug event.

executed. It is not used by the embedded scalar single-precision floating-point APU. registers are modified as shown in Table 174. Instruction execution resumes at address IVPR–47] || IVOR32[48–59] || 0b0000.

  • SPEFSCR[FINVE] = 1 and either SPEFSCR[FINVH,FINV] = 1
  • SPEFSCR[FDBZE] = 1and either SPEFSCR[FDBZH,FDBZ] = 1
  • SPEFSCR[FUNFE] = 1 and either SPEFSCR[FUNFH,FUNF] = 1
  • SPEFSCR[FOVFE] = 1 and either SPEFSCR[FOVFH,FOVF] = 1 Note that although SPEFSCR status bits can be updated by using mtspr, interrupts occur only if they are set as the result of an arithmetic operation. When an embedded floating-point data interrupt occurs, the processor suppresses execution of the instruction causing the interrupt. Table 175 shows register settings. Instruction execution resumes at address IVPR[32–47] || IVOR33[48–59] || 0b0000.

Table 174. SPE/embedded floati ng-point APU unavailable interrupt register settings SRR0 Set to the effective address of the instruction causing the interrupt. SRR1 Set to the MSR contents at the time of the interrupt. MSR CE, ME, and DE are unchanged . All other bits are cleared. ESR SPE (bit 24) is set. All other ESR bits are cleared. Table 175. Embedded floating-point data interrupt register settings SRR0 Set to the effective address of the instruction causing the interrupt. SRR1 Set to the MSR contents at the time of the interrupt. MSR CE, ME, and DE are unchanged. All other bits are cleared. ESR SPE (bit 24) is set. All other ESR bits are cleared. are set to indicate the interrupt type.

  • SPEFSCR[FINXE] = 1 and any of the SPEFSCR[FGH,FXH,FG,FX] bits = 1
  • SPEFSCR[FRMC] = 0b10 (+∞)
  • SPEFSCR[FRMC] = 0b11 (–∞) Note that although these SPEFSCR status bits can be updated by using an mtspr[SPEFSCR], interrupts occur only if they are set as the result of an arithmetic operation. When an embedded floating-point round interrupt occurs, the unrounded (truncated) result is placed in the target register. Table 176 describes register settings. Instruction execution resumes at address IVPR[32–47] || IVOR34[48–59] || 0b0000.

4.8 Performance monitor interrupt

  • PMLCax[CE] = 1; that is, for the given counter the overflow condition is enabled.
  • PMCx[32] = 1; that is, the given counter indicates an overflow. For a performance monitor interrupt to be signaled on an enabled condition or event, PMGC0[PMIE] must be set. The performance monitor can also freeze the performance monitor counters triggered by an enabled condition or event. For the performance monitor counters to freeze on an enabled condition or event, PMGC0[FCECE] must be set. Although the interrupt condition could occur with MSR[EE] = 0, the interrupt cannot be taken until MSR[EE] = 1. If a counter overflows while PMGC0[FCECE] = 0, PMLCa[CE] = 1, and MSR[EE] = 0, it is possible for the counter to wrap around to all zeros again without the performance monitor interrupt being taken. The priority of the performance monitor interrupt is below that of the fixed-interval interrupt and above that of the decrementer interrupt.

Table 176. Embedded floating-point round interrupt register settings SRR1 Set to the MSR contents at the time of the interrupt. MSR CE, ME, and DE are unchanged. All other MSR bits are cleared. ESR SPE (bit 24) is set. All other ESR bits are cleared. SPEFSCR FGH, FXH, FG, FX, and FRMC are set appropriately to indicate the interrupt type.

RM0004 Interrupts and exceptions

4.9 Partially executed instructions

In general, the PowerPC architecture permits load and store instructions to be partially executed, interrupted, and then restarted from the beginning upon return from the interrupt. To guarantee that a particular load or store instruction completes without being interrupted and restarted, software must mark the memory as guarded and use an elementary (non- string or non-multiple) load or store aligned on an operand-sized boundary. To guarantee that load and store instructions can, in general, be restarted and completed correctly without software intervention, the following rules apply when an execution is partially executed and then interrupted:

  • For an elementary load, no part of a target register rD or frD has been altered.
  • For update forms of load or store, the update register, rA, will not have been altered. The following effects are permissible when certain instructions are partially executed and then restarted:
  • For any store, bytes at the target location may have been altered (if write access to that page in which bytes were altered is permitted by the access control mechanism). In addition, for store conditional instructions, CR0 has been set to an undefined value, and it is undefined whether the reservation has been cleared or not.
  • For any load, bytes at the addressed location may have been accessed (if read access to that page in which bytes were accessed is permitted by the access control mechanism).
  • For load multiple or load string, some registers in the range to be loaded may have been altered. Including the addressing registers rA and possibly rB in the range to be loaded is a programming error, and thus the rules for partial execution do not protect these registers against overwriting. In no case is access control violated. As previously stated, elementary, aligned, guarded loads and stores are the only access instructions guaranteed not to be interrupted after being partially executed. The following list identifies the specific instruction types for which interruption after partial execution may occur, as well as the specific interrupt types that could cause the interruption: 1. Any load or store (except elementary, aligned, or guarded): – Any asynchronous interrupt – Machine check – Program (imprecise mode floating-point enabled) – Program (imprecise mode auxiliary processor enabled) – Decrementer – Fixed-interval timer – Watchdog timer – Debug (unconditional debug event) 2. Misaligned elementary load or store, or any multiple or string: All of the above listed under item 1, plus the following: – Alignment – Data storage (if the access crosses a protection boundary) – Debug (data address compare, data value compare)

Interrupts and exceptions RM0004 The mtcrf and mfcr instructions can also be partially executed due to the occurrence of any of the interrupts listed under item 1 at the time mtcrf or mfcr was executing.

  • All instructions before mtcrf or mfcr have completed execution. Some memory accesses generated by these preceding instructions may not have completed.
  • No subsequent instruction has begun execution.
  • The mtcrf or mfcr instruction, whose address was saved in SRR0/CSRR0 at the time of the interrupt, may appear not to have begun or may have partially executed.

4.10 Interrupt ordering and masking

Multiple exceptions that can each generate an interrupt can exist simultaneously. However, the PowerPC architecture does not provide for reporting multiple simultaneous interrupts of the same class (critical or noncritical). Therefore, the PowerPC architecture defines that interrupts must be ordered with one another and provides a way to mask certain persistent interrupt types. When an interrupt type is masked (disabled) and an event causes an exception that would normally generate an interrupt of that type, the exception persists as a status bit in a register (which register depends upon the exception type) but no interrupt is generated. Later, if the interrupt type is enabled (unmasked) and the exception status has not been cleared by software, the interrupt due to the original exception event is finally generated. (A typical implementation has such a mechanism for certain debug events. A signal that triggers an asynchronous interrupt, such as external input, must be asserted until they are taken. There is no mechanism for saving the external interrupt if the signal is negated before the interrupt is taken. All interrupts are level-sensitive except for machine check, which is edge- triggered.) All asynchronous interrupt types and some synchronous interrupt types can be masked. An example of a maskable synchronous interrupt type is the floating-point enabled exception- type program interrupt. The execution of a floating-point instruction that causes FPSCR[FEX] to be set is considered an exception event, regardless of the setting of MSR[FE0,FE1]. If MSR[FE0,FE1] are both 0, the floating-point enabled exception-type program interrupt is masked, but the exception persists in FPSCR[FEX]. Later, if MSR[FE0,FE1] are enabled, the interrupt is generated. The PowerPC architecture allows implementations to avoid situations in which an interrupt would cause state information (saved in save/restore registers) from a previous interrupt to be overwritten and lost. As a first step, upon any noncritical class interrupt, hardware automatically disables further asynchronous, noncritical class interrupts (external input) by clearing MSR[EE]. Likewise, upon any critical class interrupt, hardware automatically disables further asynchronous interrupts, both critical and noncritical, by clearing MSR[CE] and MSR[EE]. Critical input, watchdog timer, and debug interrupts are disabled by clearing MSR[CE,DE]. Note that machine check interrupts, while considered neither asynchronous nor synchronous, are not maskable by MSR[CE,DE,EE] and could be presented in a situation that could cause loss of state information. This first step of clearing MSR[EE] (and MSR[CE,DE] for critical class interrupts) prevents subsequent asynchronous interrupts from overwriting save/restore registers before software can save their contents. On any interrupt, hardware also automatically clears MSR[WE,PR,FP ,FE0,FE1,IS,DS], which helps avoid subsequent interrupts of certain other types. However, guaranteeing that these interrupt types do not occur also requires system software to avoid executing instructions that could cause (or enable) a subsequent interrupt, if SRR1 contents have not been saved.

4.10.1 Guidelines fo r system software

If the machine check APU is not implemented, machine check interrupts are a special case. Table 177. Operations to avoid instruction address overflow exceptions. required access permissions. Execution of any floating-point instruction Prevents floating-point unavailable interrupts. to the automatic clearing of MSR[FP]. any privileged instructions. elementary load or store instructions. instructions that cause alignment interrupts.

Interrupts and exceptions RM0004 within a normal critical interrupt handler before it saves the save/restore registers’ contents. In such a case, the interrupt may not be recoverable. It is unnecessary for hardware or software to avoid critical-class interrupts from within noncritical-class interrupt handlers (hence hardware does not automatically clear MSR[CE,ME,DE] on a noncritical interrupt), since the two interrupt classes use different save/restore registers. However, because a critical class interrupt can occur within a noncritical handler before the noncritical handler saves SRR0/SRR1, hardware and software must cooperate to avoid both critical and noncritical class interrupts from within critical class interrupt handlers. Therefore, within the critical class interrupt handler, both pairs of save/restore registers may contain data necessary to system software.

4.10.2 Interrupt order

Enabled interrupt types for which simultaneous exceptions can exist are prioritized as follows: 1. Synchronous (non-debug) interrupts: – Data storage – Instruction storage – Alignment –P r o g r a m – Floating-point unit unavailable – Auxiliary processor unavailable – System call – Data TLB error – Instruction TLB error Only one of the above synchronous interrupt types may have an existing exception generating it at a given time. This is guaranteed by the exception priority mechanism (see Chapter 4.11: Exception priorities”) and the sequential execution model. 2. Machine check 3. Debug 4. Critical input 5. Watchdog timer 6. External input 7. Fixed-interval timer 8. Decrementer Although, as indicated above, noncritical, synchronous exception types listed under item 1 are generated with higher priority than critical interrupt types in items 2–5, noncritical interrupts are immediately followed by the highest priority existing critical interrupt type, without executing any instructions at the noncritical interrupt handler. This is because noncritical interrupt types do not automatically disable MSR mask bits for critical interrupt types (CE and ME). In all other cases, a particular interrupt type listed above automatically disables subsequent interrupts of the same type, as well as all lower priority interrupt types.

RM0004 Interrupts and exceptions

4.11 Exception priorities

Book E requires all synchronous (precise and imprecise) exceptions to be reported in program order, as required by the sequential execution model. The one exception to this rule is the case of multiple synchronous imprecise exceptions. Upon a synchronizing event, all previously executed instructions are required to report any synchronous imprecise interrupt- generating exceptions, and the interrupt is then generated with all of those exception types reported cumulatively in the ESR and in any status registers associated with the particular exception type (such as the FPSCR). For any single instruction attempting to cause multiple exceptions for which the corresponding synchronous interrupt types are enabled, this section defines the priority order by which the instruction is permitted to cause a single enabled exception, thus generating a particular synchronous interrupt. Note that it is this exception priority mechanism, along with the requirement that synchronous interrupts be generated in program order, that guarantees that at any given time there exists for consideration only one of the synchronous interrupt types listed in item 1 of Chapter 4.10.2: Interrupt order.” The exception priority mechanism also prevents certain debug exceptions from existing in combination with certain other synchronous interrupt-generating exceptions. The EIS defines priorities for all exceptions including those defined in optional APUs. Interrupt types are defined as either synchronous (the interrupt is as a direct result of an instruction in execution) or asynchronous, (the interrupt results from an event external to the execution of a particular instruction or an instruction removes a gating condition to a pending exception). Except for machine check interrupts, which can be either synchronous or asynchronous, interrupts are either synchronous or asynchronous exclusively. Because asynchronous interrupts may temporally be sampled either before or after an instruction is completed, an implementation can order asynchronous interrupts among only asynchronous interrupts and order synchronous interrupts among only synchronous interrupt. The distinction is important because synchronous interrupts that require post- completion actions (such as system call or debug instruction complete exceptions) cannot be separated from the completion of the instruction. Therefore, asynchronous interrupts cannot be sampled during the completion and post-completion synchronous exceptions for a given instruction. The relative priorities for asynchronous exceptions is given in Table 178 and for synchronous exceptions in and Table 179. In many cases, certain exceptions cannot occur at the same time (for example, program-trap and program-Illegal cannot occur on the same instruction). In general those exceptions are grouped at the same relative priority.

Table 178. EIS asynchronous exception priorities

  1. The interrupt level defines the set of save/restore registers used when the interrupt is taken—base
  2. Pre- or post-completion refers to whether the exception occurs before an instruction completes (pre) and

(post) and the corresponding interrupt points to the next instruction to be executed.

0 Machine check Machine check Asynch/synch pre for

1 Debug—UDE Critical/debug Asynch N/A Generally used for an

2 Critical input Critical Asynch N/A

3 Watchdog Critical Asynch N/A

4 External input Base Asynch N/A

18 Fixed interval timer Base Asynch N/A

19 Decrementer Base Asynch N/A

20 Performance

Table 179. EIS synchronous exception priorities

5 Debug–instruction

8 Debug—trap Critical/debug Synch pre

15 Debug—DAC Critical/debug Synch pre or post Preferred method is pre-

16 Debug—DVC Critical/debug Synch pre or post Preferred method is pre-

  1. The interrupt level defines the set of save/restore registers used when the interrupt is taken—base
  2. Pre- or post-completion refers to whether the exception occurs before an instruction completes (pre) and

(post) and the corresponding interrupt points to the next instruction to be executed. Table 179. EIS synchronous exception priorities (continued)

RM0004 Storage architecture

5 Storage architecture

This chapter describes the cache and MMU portions of the Book E implementation standards (EIS). Note that not all features that are defined by the EIS storage architecture are supported on all ST EIS processors; consult the user documentation. This chapter is organized into three section:

  • Chapter 5.2: Memory and cache coherency”
  • Chapter 5.3: Cache model”
  • Chapter 5.4: Storage model”

5.1 Overview

The Book E architecture memory and cache definitions support a wide variety of embedded implementations. To provide such flexibility, Book E defines many features in a very general way, leaving specific details up to the implementation. To ensure consistency among its Book E cores and devices, ST has defined more specific implementation standards. However, these standards still leave many details up to individual implementations. To provide context for those features, this chapter describes aspects of the memory hierarchy and the memory management model defined by Book E; it also describes the ST EIS. Note: This chapter describes some features (in particular, registers) in a very general way that does not include some details that are important to the programmer. There are also small differences in how some features are defined here and how they are implemented. For implementation-specific details, see the user documentation. Throughout this chapter, references to load instructions include cache management and other instructions that are stated in the instruction descriptions to be treated as a load, and references to store instructions include the cache management and other instructions that are treated as a store. The following APUs, which are part of the EIS storage architecture, are defined in Chapter 8: Storage-related APUs on page 848”:

  • Cache line locking APU
  • Cache way partitioning APU
  • Direct cache flush APU These APUs may be implemented independently of each other. They are defined together in a single specification because it is likely that an implementation will include more than one of these APUs.

5.2 Memory and cache coherency

The primary objective of a coherent memory system is to provide the same image of memory to all devices using the system. Coherency allows synchronization and cooperative use of shared resources. Otherwise, multiple copies of data corresponding to a memory location, some containing outdated values, could exist in a system, resulting in errors when the outdated values are used. Each memory-sharing device must follow rules for managing the state of its cache. This section describes the coherency mechanisms of the Book E architecture and the cache coherency protocols that the ST Book E devices support.

Storage architecture RM0004 Unless specifically noted, the discussion of coherency in this section applies to the core complex data cache only. The instruction cache is not snooped for general coherency with other caches; however, it is snooped when the Instruction Cache Block Invalidate (icbi) instruction is executed by this processor or any processor in the system.

5.2.1 Memory/Cache access attributes

Some memory characteristics can be set on a page basis by using the WIMGE bits in the translation lookaside buffer (TLB) entries. These bits allow both uniprocessor and multiprocessor system designs to exploit numerous system-level performance optimizations. The WIMGE attributes control the following:

  • Write-through (W bit)
  • Caching-inhibited (I bit)
  • Memory-coherency-required (M bit)
  • Guarded (G bit)
  • Endianness (E bit) In addition to the WIMGE bits, the Book E MMU model defines the following attributes on a page basis:
  • User-definable (U0, U1, U2, U3) The EIS defines the following optional attributes, which are manipulated by software through MMU assist register 2 (MAS2):
  • Alternate coherency mode (ACM). The ACM attribute, programmed through MAS2[ACM], allows an implementation to employ multiple coherency methods and to participate in multiple coherency protocols. If the M attribute (memory coherence required) is not set for a page (M = 0), the page has no coherency associated with it and the ACM attribute is ignored. If the M attribute is set for a page (M = 1), the ACM attribute determines the coherency domain (or protocol) used. ACM values are implementation dependent.
  • Variable length encoding (VLE). The VLE attribute, MAS2[VLE], identifies pages that contain instructions from the VLE instruction set. If VLE = 0, instructions fetched from the page are decoded and executed as PowerPC (and associated EIS APUs) instructions. If VLE = 1, instructions fetched from the page are decoded and executed as Power Embedded instructions. Consult the user documentation to determine whether the EIS-defined attributes are implemented. The WIMGE attributes are programmed by the operating system for each page. The W and I attributes control how the processor performing an access uses its own cache. The M attribute ensures that coherency is maintained for all copies of the addressed memory location. The G attribute prevents speculative loading from the addressed memory location. (An operation is said to be performed speculatively if, at the time that it is performed, it is not known to be required by the sequential execution model.) The E attribute defines the order in which the bytes that comprise a multiple-byte data object are stored in memory (big- or little-endian). The WIMGE attributes occupy 5 bits in the TLB entries for page address translation. The operating system writes the WIMGE bits for each page into the TLB entries in system memory as it maps translations. For more information, see TLB entries on page 319.”

RM0004 Storage architecture All combinations of these attributes are supported except those that simultaneously specify a region as write-through and caching-inhibited. Write-through and caching-inhibited attributes are mutually exclusive because the write-through attribute permits the data to be in the data cache while the caching-inhibited attribute does not. Memory that is write-through or caching-inhibited is not intended for general-purpose programming. For example, lwarx and stwcx. instructions may cause the system DSI exception handler to be invoked if they specify a location in memory having either of these attributes. Some implementations take a data storage interrupt if the location is write- through but does not take the interrupt if the location is cache-inhibited. Note that, except that the guarded bit does not prevent instruction prefetches, the definitions of the WIMG bits are unchanged Write-through attribute A page marked W = 0 is considered to be write-back. If some store instructions executed by a given processor access locations in a block as write-through and other store instructions executed by the same processor access locations in that block as write-back, software must ensure that the block cannot be accessed by another processor or mechanism in the system. A store to a write-through (W = 1) memory location is performed in main memory and may cause additional memory locations to be accessed. If a copy of the block containing the specified location is retained in the data cache, the store is also performed in the data cache. A store to write-through memory cannot cause a block to be put in a modified state in the data cache. Also, if a store instruction that accesses a block in a location marked as write-through is executed when the block is already considered to be modified in the data cache, the block may continue to be considered to be modified in the data cache even if the store causes all modified locations in the block to be written to main memory. In some processors, accesses caused by separate store instructions that specify locations in write-through memory may be combined into one access. This is called store-gathering. Such combining does not occur if the store instructions are separated by an msync or an mbar. Caching-inhibited attribute A load instruction that specifies a location in caching-inhibited (I = 1) memory is performed to main memory and may cause additional locations in main memory to be accessed unless the specified location is also guarded. An instruction fetch from caching-inhibited memory may cause additional words in main memory to be accessed. No copy of the accessed locations is placed into the caches. In some processors, nonoverlapping accesses caused by separate load instructions that specify locations in caching-inhibited memory may be combined into one access, as may nonoverlapping accesses caused by separate store instructions to caching-inhibited memory (that is, store-gathering). Such combining does not occur if the load or store instructions are separated by an msync instruction, or by an mbar instruction if the memory is also guarded. Memory-coherence-required attribute Memory coherence refers to the ordering of stores to a single location. Atomic stores to a given location are coherent if they are serialized in some order, and no processor or mechanism is able to observe any subset of those stores as occurring in a conflicting order. This serialization order is an abstract sequence of values; the physical location need not

Storage architecture RM0004 assume each of the values written to it. For example, a processor may update a location several times before the value is written to physical memory. The result of a store operation is not available to every processor or mechanism at the same instant, and it may be that a processor or mechanism observes only some of the values that are written to a location. However, when a location is accessed atomically and coherently by all processors and mechanisms, the sequence of values loaded from the location by any processor or mechanism during any interval of time forms a sub-sequence of the sequence of values that the location logically held during that interval. That is, a processor or mechanism can never load a newer value first and then, later, load an older value. Memory coherence is managed in blocks called coherence blocks. Although a block’s size is implementation-dependent, it is usually larger than a word and is often the size of a cache block. When memory coherence is not required (M = 0), the hardware need not enforce data coherence for memory accesses initiated by the processor. When memory coherence is required (M = 1), the hardware must enforce data coherence for memory accesses initiated by the processor. Hardware support for the memory-coherence-required attribute is optional for implementations that do not support multiprocessing. Guarded attribute When the guarded bit is set, the page is designated as guarded. This setting can be used to protect certain memory areas from read accesses made by the processor that are not dictated directly by the program. If areas of physical memory are not fully populated (in other words, there are holes in the physical memory map within this area), this setting can protect the system from undesired accesses caused by speculative (referred to as ‘out of order’ in the architecture specification, and described in Definition of apeculative and out-of-order memory accesses on page 285”) load operations that could lead to the generation of the machine check exception. Also, the guarded bit can be used to prevent speculative load operations from occurring to certain peripheral devices that produce undesired results when accessed in this way. Definition of apeculative and out-of-order memory accesses In the architecture definition, the term ‘out of order’ replaced the term ‘speculative’ with respect to memory accesses to avoid a conflict between the word’s meaning in the context of execution of instructions past unresolved branches. The architecture’s use of out of order in this context could in turn be confused with the notion of loads and stores being reordered in a weakly ordered memory system. In the context of memory accesses, this document uses the terms ‘speculative’ and ‘out of order’ as follows:

  • Speculative memory access—An access to memory that occurs before it is known to be required by the sequential execution model.
  • Out-of-order memory access—A memory access performed ahead of one that may have preceded it in the sequential model, such as is allowed by a weakly ordered memory model. Performing operations speculatively An operation is said to be nonspeculative if it is guaranteed to be required by the sequential execution model. Any other operation is said to be performed speculatively, which the architecture specification refers to as out of order.

RM0004 Storage architecture Operations are performed speculatively by hardware on the expectation that the results will be needed by an instruction that will be required by the sequential execution model. Whether the results are needed depends on whether control flow is diverted away from the instruction by an event such as an exception, branch, trap, system call, return from interrupt instruction, or anything else that changes the context in which the instruction is executed. Typically, the hardware performs operations speculatively when it has resources that would otherwise be idle, so the operation incurs little or no cost. If subsequent events such as branches or exceptions indicate that the operation would not have been performed, the processor abandons any results of the operation except as described below. Most operations can be performed speculatively, as long as the machine appears to follow the sequential execution model. Certain speculative operations are restricted, as follows:

  • Stores—A store instruction cannot execute speculatively in a manner such that the alteration of the target location can be observed by other processors or mechanisms.
  • Accessing guarded memory—The restrictions for this case are given in Speculative accesses to guarded memory on page 286.” No error of any kind other than a machine check exception may be reported due to an operation that is performed speculatively, until such time as it is known that the operation is required by the sequential execution model. The only other permitted side effect (other than machine check) of performing an operation speculatively is that nonguarded memory locations that could be fetched into a cache by nonspeculative execution may be fetched speculatively into that cache. Guarded memory Memory is said to be well behaved if the corresponding physical memory exists and is not defective, and if the effects of a single access to it are indistinguishable from the effects of multiple identical accesses to it. Data and instructions can be fetched speculatively from well-behaved memory without causing undesired side effects. Memory is said to be guarded if the G bit is set for the page. In general, memory that is not well-behaved should be guarded. Because such memory may represent an I/O device or include nonexistent locations, a speculative access to such memory may cause an I/O device to perform incorrect operations or may cause a machine check. Note that if separate store instructions access memory that is both caching-inhibited and guarded, the accesses are performed in the order specified by the program. If an aligned load or store that is not a string or multiple access to caching-inhibited, guarded memory has accessed main memory and an external, decrementer, or imprecise-mode floating-point enabled exception is pending, the load or store is completed before the exception is taken. Speculative accesses to guarded memory Accesses for load instructions from guarded memory may be performed speculatively if a copy of the target location is in a cache; in this case, the location may be accessed from the cache or from main memory. Note that software should ensure that only well-behaved memory is loaded into a cache, either by marking as caching-inhibited (and guarded) all memory that may not be well- behaved or by marking such memory caching-allowed (and guarded) and referring only to cache blocks that are well-behaved.

Storage architecture RM0004 Instrubction accesses: guarded memory and no-execute memory The G bit is ignored for instruction fetches, and instructions are speculatively fetched from guarded pages. To prevent speculative fetches from pages that do not contain instructions and are not well-behaved, the page should be designated as no-execute (with the UX/SX page permission bits cleared). If the effective address of the current instruction is mapped to no-execute memory, an ISI exception is generated. Endianness Objects may be loaded from or stored to memory in byte, half-word, word, or double-word units. For a particular data length, the load and store operations are symmetrical; a store followed by a load of the same data object yields an unchanged value. Book E makes no guarantees about the order in which the bytes that comprise multiple-byte data objects are stored into memory. The endianness (E) page attribute distinguishes between memory that is big or little endian, as described in the following subsections. Except for instruction fetches, it is always permitted to access the same location using two effective addresses with different E bit settings. Instruction pages must be flushed from any caches before the E bit can be changed for those addresses. See Byte ordering on page 141,” for more information about endianness. Big-endian pages If a stored multiple-byte object is probed by reading its component bytes one at a time using load-byte instructions, the store order may be perceived. If such probing shows that the lowest memory address contains the highest-order byte of the multiple-byte scalar, the next- higher sequential address the next-least-significant byte, and so on, the multiple-byte object is stored in big-endian form. Big-endian memory is defined on a page basis by the memory/cache attribute, E = 0. Note that strings are not multiple-byte scalars but are interpreted as a series of single-byte scalars. Bytes in a string are loaded from memory using a load string word instruction, starting at the lowest-numbered address, and placed into the target register or registers starting at the left-most byte of the least-significant word. Bytes in a string are stored using a store string word instruction from the source register, starting at the left-most byte of the least-significant word, and placed into memory, starting at the lowest-numbered address. Little-endian pages Alternatively, if the probing shows that the lowest memory address contains the lowest-order byte of the multiple-byte scalar, the next-higher sequential address the next-most-significant byte, and so on, the multiple-byte object is stored in little-endian form. Little-endian memory is defined on a page basis by the memory/cache attribute, E = 1, and for Book E devices is defined as true little-endian memory. Structure mapping examples The following C programming example defines the data structure S used in this section to demonstrate how the bytes that comprise each element (a, b, c, d, e, and f) are mapped into memory. The structure contains scalars (shown in hexadecimal in the comments) and a sequence of characters, shown in single quotation marks. struct { int a; /* 0x1112_1314 word*/ double b; /* 0x2122_2324_2526_2728double word*/ char * c; /* 0x3132_3334 word*/ char d[7]; /* 'L','M','N','O','P ','Q','R' array of bytes*/

RM0004 Storage architecture short e; /* 0x5152 half word*/ int f; /* 0x6162_6364 word*/ } S; big-endian mapping of the structure ia shown below. Big-endian mapping of structure Note that the MSB of each scalar is at the lowest address. The mapping uses padding (indicated by (x)) to align the scalars—4 bytes between elements a and b, 1 byte between d and e, and 2 bytes between e and f. Note that the padding is determined by the compiler, not the architecture. The structure using little-endian mapping, showing double words laid out with addresses increasing from right to left. Little-endian mapping of structure S—alternate view Contents 11 12 13 14 (x) (x) (x) (x) A d d r e s s 0 00 10 20 30 40 50 60 7 Contents 21 22 23 24 25 26 27 28 Address 08 09 0A 0B 0C 0D 0E 0F Contents 31 32 33 34 ‘L ’ ‘M’ ‘N’ ‘O’ A d d r e s s 1 01 11 21 31 41 51 61 7 Contents ‘P’ ‘Q’ ‘R’ (x) 51 52 (x) (x) Address 18 19 1A 1B 1C 1D 1E 1F Contents 61 62 63 64 (x) (x) (x) (x) A d d r e s s 2 02 12 22 32 42 52 62 7 Contents (x) (x) (x) (x) 11 12 13 14 A d d r e s s 0 70 60 50 40 30 20 10 0 Contents 21 22 23 24 25 26 27 28 Address 0F 0E 0D 0C 0B 0A 09 08 Contents ‘O’ ‘N’ ‘M’ ‘L ’ 31 32 33 34 A d d r e s s 1 71 61 51 41 31 21 11 0 Contents (x) (x) 51 52 (x) ‘R’ ‘Q’ ‘P’ Address 1F 1E 1D 1C 1B 1A 19 18 Contents (x) (x) (x) (x) 61 62 63 64 A d d r e s s 2 72 62 52 42 32 22 12 0

Storage architecture RM0004 Mismatched memory cache attributes Accesses to the same memory location using two effective addresses for which the write- through required attribute (W bit) differs meet the memory coherence requirements described in Write-through attribute on page 284,” if the accesses are performed by a single processor. If the accesses are performed by two or more processors, coherence is enforced by the hardware only if the write-through attribute is the same for all the accesses. Loads, stores, dcbz instructions, and instruction fetches to the same memory location using two effective addresses for which the caching-inhibited attribute (I bit) differs must meet the requirement that a copy of the target location of an access to caching-inhibited memory not be in the cache. Violation of this requirement is considered a programming error; software must ensure that the location has not previously been brought into the cache or, if it has, that it has been flushed from the cache. If the programming error occurs, the result of the access is boundedly undefined. It is not considered a programming error if the target location of any other cache management instruction to caching-inhibited memory is in the cache. Accesses to the same memory location using two effective addresses for which the memory coherence attribute (M bit) differs may require explicit software synchronization before accessing the location with M = 1 if the location has previously been accessed with M = 0. Any such requirement is system-dependent. For example, in some systems that use bus snooping, no software synchronization may be required. In some directory-based systems, software may be required to execute dcbf instructions on each processor to flush all cache entries accessed with M = 0 before accessing those locations with M = 1. Accesses to the same memory location using two effective addresses for which the guarded attribute (G bit) differs are always permitted. Except for instruction fetches, accesses to the same memory location using two effective addresses for which the endian storage attribute (E bit) differs are always permitted as described in Endianness on page 287.” Instruction memory locations must be flushed before the endian attribute can be changed for those addresses. The requirements on mismatched user-defined memory attributes (U0–U3) is implementation-dependent. Coherency paradoxes and WIMGE Care must be taken with respect to the use of the WIMGE bits if coherent memory support is desired. Careless programming of these bits may create situations that present coherency paradoxes to the processor. These paradoxes can occur within a single processor or across several processors. It is important to note that, in the presence of a paradox, the operating system software is responsible for correctness. In particular, a coherency paradox can occur when the state of these bits is changed without appropriate precautions (such as flushing the pages that correspond to the changed bits from the caches of all processors in the system) or when the address translations of aliased real addresses specify different values for certain WIMGE bit values. For more information, see Mismatched memory cache attributes on page 289.” Support for M = 1 memory is optional. Cache attribute settings where both W = 1 and I = 1 are not supported. For all supported combinations of the W, I, and M bits, both G and E may be 0 or 1. The default setting of the WIMGE bits is 0b01000.

RM0004 Storage architecture Self-modifying code When a processor modifies any memory location that can contain an instruction, software must ensure that the instruction cache is made consistent with data memory and that the modifications are made visible to the instruction fetching mechanism. This must be done even if the cache is disabled or if the page is marked caching-inhibited. The following instruction sequence can be used to accomplish this when the instructions being modified are in memory that is memory-coherence required and one processor both modifies the instructions and executes them. (Additional synchronization is needed when one processor modifies instructions that another will execute.) The following sequence synchronizes the instruction stream (using either dcbst or dcbf): dcbst (or dcbf)|update memory msync |wait for update icbi |remove (invalidate) copy from instruction cache msync |ensure that ICBI invalidation at icache has completed isync |remove copy in own instruction buffer

5.2.2 Shared memory

The architecture supports sharing memory between programs, between different instances of the same program, and between processors and other mechanisms. It also supports access to a memory location by one or more programs using different effective addresses. In these cases, memory is shared in blocks that are an integral number of pages. When one physical memory location has different effective addresses, the addresses are said to be aliases. Each application can be granted separate access privileges to aliased pages. Lock acquisition and import barriers on page 294,” gives examples of how msync and mbar are used to control memory access ordering when memory is shared among programs. Memory access ordering The memory model in Book E for memory access ordering is weakly consistent. This provides an opportunity for improved performance over a model with stronger consistency rules but places the responsibility on the program to ensure that ordering or synchronization instructions are properly placed for correct execution of the program. The order in which a processor accesses memory, the order in which those accesses are performed with respect to other processors or mechanisms, and the order in which they are performed in main memory may all be different. Table 180 describes how the architecture defines requirements for ordering of loads and stores.

following the barrier-creating instruction. Table shows an example using a two-processor system. Table 180. Load and store ordering and guarded (G = 1) are not reordered with respect to one another. address specified by the second load). are weakly ordered with respect to one another. sequential ordering of loads with respect to stores. Table 181. Memory barrier when coherency is required (M = 1)

  • G1 includes all applicable memory accesses by any such processor or mechanism that have been performed with respect to P1 before the memory barrier is created.
  • G2 includes all applicable memory accesses by any such processor or mechanism that are performed after a load instruction executed by that processor or mechanism has returned the value stored by a store that is in G2. Table 182 shows an example of a cumulative memory barrier in a two-processor system.

Table 182. Cumulative memory barrier properties of the memory barrier created by that instruction. instruction O occurs before P1’s msync is executed). except those associated with fetching instructions following msync.

Storage architecture RM0004 between a path that includes a store instruction if the condition is met, that dependent store is not performed unless and until the condition determined by the load is met. Because instructions following an isync cannot execute until all instructions preceding isync have completed, if an isync follows a conditional branch instruction that depends on the value returned by a preceding load instruction, that load is performed before any loads caused by instructions following the isync. This is true even if the effects of the dependency are independent of the value loaded (for example, the value is compared to itself and the branch tests CRn[EQ]), and even if the branch target is the next sequential instruction. Except for the cases described above and earlier in this section, data and control dependencies do not order memory accesses. Examples include the following:

  • If a load specifies the same memory location as a preceding store and the location is not caching inhibited, the load may be satisfied from a store queue (a buffer into which the processor places stored values before presenting them to the memory subsystem) and not be visible to other processors and mechanisms. As a result, if a subsequent store depends on the value returned by the load, the two stores need not be performed in program order with respect to other processors and mechanisms.
  • Because a store conditional instruction may complete before its store is performed, a conditional branch instruction that depends on the CR0 value set by a store conditional instruction does not order that store with respect to memory accesses caused by instructions that follow the branch. For example, in the following sequence, the stw is the bc instruction’s target: stwcx. bc stw To complete, the stwcx. must update the architected CR0 value, even though its store may not have been performed. The architecture does not require that the store generated by the stwcx. must be performed before the store generated by the stw.
  • Because processors may predict branch target addresses and branch condition resolution, control dependencies (branches, for example) do not order memory accesses except as described above. For example, when a subroutine returns to its caller, the return address may be predicted, with the result that loads caused by instructions at or after the return address may be performed before the load that obtains the return address is performed. Some processors implement nonarchitected duplicates of architected resources such as GPRs, CR fields, and the LR, so resource dependencies (for example, specification of the same target register for two load instructions) do not force ordering of memory accesses. Examples of correct uses of dependencies, msync, and mbar to order memory accesses can be found in hi.” Because the memory model is weakly consistent, the sequential execution model as applied to instructions that cause memory accesses guarantees only that those accesses appear to be performed in program order with respect to the processor executing the instructions. For example, an instruction may complete, and subsequent instructions may be executed, before memory accesses caused by the first instruction have been performed. However, for a sequence of atomic accesses to the same memory location for which memory coherence is required, the definition of coherence guarantees that the accesses are performed in program order with respect to any processor or mechanism that accesses the location coherently, and similarly if the location is one for which caching is inhibited.

RM0004 Storage architecture Because caching-inhibited memory accesses are performed in main memory, memory barriers and dependencies on load instructions order such accesses with respect to any processor or mechanism even if the memory is not marked as requiring memory coherence. Programming examples Example 1 shows cumulative ordering of memory accesses preceding a memory barrier, Example 2 shows cumulative ordering of memory accesses following a memory barrier. In both examples, assume that locations X, Y , and Z initially contain the value 0. In both, cumulative ordering dictates that the value loaded from location X by processor C is 1. Example 1:

  • Processor A stores the value 1 to location X.
  • Processor B loads from location X obtaining the value 1, executes an msync, then stores the value 2 to location Y .
  • Processor C loads from location Y obtaining the value 2, executes an msync, then loads from location X. Example 2:
  • Processor A stores the value 1 to location X, executes an msync, then stores the value 2 to location Y .
  • Processor B loops, loading from location Y until the value 2 is obtained, then stores the value 3 to location Z.
  • Processor C loads from location Z obtaining the value 3, executes an msync, then loads from location X. Lock acquisition and import barriers An import barrier is an instruction or instruction sequence that prevents memory accesses caused by instructions following the barrier from being performed before memory accesses that acquire a lock have been performed. An import barrier can be used to ensure that a shared data structure protected by a lock is not accessed until the lock has been acquired. An msync can always be used as an import barrier, but the approaches shown below generally yield better performance because they order only the relevant memory accesses. Acquire lock and import shared memory If lwarx and stwcx. are used to obtain the lock, an import barrier can be constructed by placing an isync immediately following the loop containing the lwarx and stwcx.. The following example uses the compare and swap primitive (see Chapter C.1.1: Synchronization primitives on page 1144”) to acquire the lock. This example assumes that the address of the lock is in GPR 3, the value indicating that the lock is free is in GPR 4, the value to which the lock should be set is in GPR 5, the old value of the lock is returned in GPR 6, and the address of the shared data structure is in GPR 9. loop:lwarxr6,0,r3 # load lock and reserve cmpw r4,r6 # skip ahead if bne- wait # lock not free stwcx. r5,0,r3 # try to set lock bne- loop # loop if lost reservation isync # import barrier lwz r7,data1(r9) # load shared data wait: ... #wait for lock to free

Storage architecture RM0004 The second bne- does not complete until CR0 has been set by the stwcx.. The stwcx. does not set CR0 until it has completed (successfully or unsuccessfully). The lock is acquired when the stwcx. completes successfully. Together, the second bne- and the subsequent isync create an import barrier that prevents the load from data1 from being performed until the branch is resolved to be not taken. Obtain pointer and import shared memory If lwarx and stwcx. are used to obtain a pointer into a shared data structure, an import barrier is not needed if all the accesses to the shared data structure depend on the value obtained for the pointer. The following example uses the fetch and add primitive (see Section C.1.1: Synchronization primitives”) to obtain and increment the pointer. In this example, it is assumed that the address of the pointer is in GPR 3, the value to be added to the pointer is in GPR 4, and the old value of the pointer is returned in GPR 5. loop: lwarx r5,0,r3 # load pointer and reserve add r0,r4,r5 # increment the pointer stwcx. r0,0,r3 # try to store new value bne- loop # loop if lost reservation lwz r7,data1(r5) # load shared data The load from data1 cannot be performed until the lwarx loads the pointer value into GPR 5. The load from data1 may be performed out of order before the stwcx.. But if the stwcx. fails, the branch is taken and the value returned by the load from data1 is discarded. If the stwcx. succeeds, the value returned by the load from data1 is valid even if the load is performed out of order, because the load uses the pointer value returned by the instance of the lwarx that created the reservation used by the successful stwcx.. An isync could be placed between the bne- and the subsequent lwz, but no isync is needed if all accesses to the shared data structure depend on the value returned by the lwarx. Atomic memory references The Book E architecture defines the Load Word and Reserve Indexed (lwarx) and the store word conditional indexed (stwcx.) instructions to provide an atomic update function for a single, aligned word of memory. These instructions can be used to develop a rich set of multiprocessor synchronization primitives. Note that atomic memory references constructed using lwarx/stwcx. instructions depend on the presence of a coherent memory system for correct operation. These instructions should not be expected to provide atomic access to noncoherent memory. The lwarx instruction performs a load word from memory operation and creates a reservation for the same reservation granule that contains the accessed word. Reservation granularity is implementation-dependent. The lwarx instruction makes a nonspecific reservation with respect to the executing processor and a specific reservation with respect to other masters. This means that any subsequent stwcx. executed by the same processor, regardless of address, cancels the reservation. Also, any bus write or invalidate operation from another processor to an address that matches the reservation address cancels the reservation.

RM0004 Storage architecture

5.3 Cache model

A cache model in which there is one cache for instructions and another cache for data is called a ‘Harvard-style’ cache. This is the model assumed by Book E, for example in the descriptions of the cache management instructions in Chapter 3: Instruction model on page 133.” Book E allows the following additional cache models are defined by the EIS:

  • Unified cache, in which a cache is shared by both instructions and data
  • Multi-level caches, which must support the programming model implied by a Harvard- style cache. A processor is not required to maintain copies of storage locations in the instruction cache that are consistent with modifications to those storage locations (that is, modifications by store instructions). In general, a location in the data cache is considered to be modified in that cache if the location has been modified (for example, by a store instruction) and the modified data has not been written to main storage. The only exception to this rule is described in Write- through attribute on page 284.” Cache management instructions are provided so that programs can manage the caches when needed. For example, program management of the caches is needed when a program generates or modifies code that will be executed (i.e., when the program modifies data in storage and then attempts to execute the modified data as instructions). Cache management instructions are also useful in optimizing the use of memory bandwidth in such applications as graphics and numerically intensive computing. The functions performed by these instructions depend on the storage attributes associated with the specified storage location. Cache management instructions allow the program to do the following.
  • Give a hint that a block of storage should be copied to the instruction cache, so that the copy of the block is more likely to be in the cache when subsequent accesses to the block occur, thereby reducing delays (icbt)
  • Invalidate the copy of storage in an instruction cache block (icbi)
  • Discard prefetched instructions (isync)
  • Invalidate the copy of storage in a data cache block (dcbi)
  • Give a hint that a block of storage should be copied to the data cache, so that the copy of the block is more likely to be in the cache when subsequent accesses to the block occur, thereby reducing delays (dcbt, dcbtst)
  • Allocate a data cache block and set the contents of that block to zeros, but no-operation if no access is allowed to the data cache block and do not cause any exceptions (dcba)
  • Set the contents of a data cache block to zeros (dcbz)
  • Copy the contents of a modified data cache block to main storage (dcbst)
  • Copy the contents of a modified data cache block to main storage and make the copy of the block in the data cache invalid (dcbf).

5.3.1 Cache programming model

This section summarizes the register and instructions defined to support the cache model. Full descriptions of these resources are provided in Chapter 2: Register model on page 46, and Chapter 3: Instruction model on page 133.

  • Machine state register (MSR). Defines the processor state (that is, enabling and disabling of interrupts and debugging exceptions, enabling and disabling of address translation for instruction and data memory accesses, enabling and disabling some APUs, and specifying whether the processor is in supervisor or user mode). EIS storage defines the user cache locking enable bit (MSR[UCLE]) as part of the cache line locking APU. Book E and the EIS define the MSR fields described in Table 183. The MSR is described in detail in Chapter 2.6.1: Machine state register (MSR) on page 68.”
  • Exception syndrome register (ESR). The ESR provides a syndrome to differentiate between different kinds of exceptions that can generate the same interrupt type. When such an interrupt is generated, bits corresponding to the exception that generated the interrupt are set and all other ESR bits are cleared. Other interrupt types do not affect ESR contents. The ESR does not need to be cleared by software.
  • Book E and the EIS defines the storage-related ESR fields described in Table 184. The ESR is described in detail in Exception syndrome register (ESR) on page 84.”

Table 183. Storage related MSR fields mode cache-line locking by the operating system. for access violations though). (TS = 0 in the relevant TLB entry). (TS = 1 in the relevant TLB entry). 0 (TS = 0 in the relevant TLB entry). 1 (TS = 1 in the relevant TLB entry).

Table 184. Exception syndrome register (ESR) definition

40 ST (Book E–defined) Store operation Alignment, data

(MSR[PR] = 1) while MSR[UCLE] = 0.

  • L1 cache control and status registers (L1CSR0–L1CSR1). – L1CSR0 provides general control and status for the processor’s primary data cache. If a processor implements a unified L1 cache, L1CSR0 applies to the unified cache and L1CSR1 is not implemented. See Chapter 2.11.1: L1 cache control and status register 0 (L1CSR0) on page 90.” – L1CSR1 provides general control and status for the processor’s primary instruction cache. If a processor implements a unified L1 cache, L1CSR0 applies

Defined by SPE, embedded floating-point APU.

0 Default

1 Any exception caused by an SPE/embedded floating-

execution of a VLE instruction.

0 The instruction page associated with the instruction

or the VLE extension is not implemented.

1 The instruction page associated with the instruction

VLE extension is implemented. 32-bit VLE instruction caused an instruction TLB error. Table 184. Exception syndrome register (ESR) definition (continued)

RM0004 Storage architecture to the unified cache and L1CSR1 is not implemented. See Chapter 2.11.2: L1 cache control and status register 1 (L1CSR1) on page 92.”

  • L1 cache configuration registers (L1CFG0) – L1CFG0 provides configuration information for the processor’s primary data cache. If a processor implements a unified cache, L1CFG0 applies to the unified cache and L1CFG1 is not implemented. See Chapter 2.11.3: L1 cache configuration register 0 (L1CFG0) on page 94.” – L1CFG1 provides configuration information for the processor’s primary instruction cache. If a processor implements a unified cache, L1CFG0 applies to the unified cache and L1CFG1 is not implemented. L1CFG1 allows software to identify the organization and capabilities of the primary instruction cache. Chapter 2.11.4: L1 cache configuration register 1 (L1CFG1) on page 95.” Cache model instructions The Book E PowerPC architecture defines instructions for controlling both the instruction and data caches (when they exist).
  • Data Cache Block Touch (dcbt)
  • Data Cache Block Touch for Store (dcbtst)
  • Data Cache Block Zero (dcbz)
  • Data Cache Block Store (dcbst)
  • Data Cache Block Flush (dcbf)
  • Data Cache Block Allocate (dcba)
  • Data Cache Block Invalidate (dcbi)
  • Instruction Cache Block Invalidate (icbi)
  • Instruction Synchronize (isync)
  • Instruction Cache Block Touch (icbt) These instructions are described in User-level cache instructions on page 180,” and Supervisor-level cache instruction on page 183.” Note that the behavior of many of these instructions is determined by the value of the cache target operand (CT). See CT instruction field on page 301.” Permission control and cache management instructions on page 316,” describes conditions in which cache control instructions can generate protection violations. The cache block locking APU, defined by the EIS, adds the following instructions:
  • Data Cache Block Lock Clear (dcblc)
  • Data Cache Block Touch and Lock Set (dcbtls)
  • Data Cache Block Touch for Store and Lock Set (dcbtstls)
  • Instruction Cache Block Lock Clear (icblc)
  • Instruction Cache Block Touch and Lock Set (icbtls) These instructions are described in Chapter 8.1: Cache line locking APU on page 848.”

Storage architecture RM0004 CT instruction field Instructions having a CT (cache target) field for specifying a cache hierarchy use the value 0 to specify the primary cache. ST devices interpret this operand as follows:

  • CT = 0 indicates the L1 cache.
  • CT = 1 indicates the I/O cache. (Note that some versions of the e500 documentation refer to the I/O cache as a frontside L2 cache.)
  • CT = 2 indicates a backside L2 cache.

5.3.2 Primary (L1) cache model

This section describes the L1 cache model defined by the EIS. Types Primary caches may separate instruction and data caches into two separate structures (commonly known as Harvard architecture), or they may provide a unified cache combining instructions and data. Caches are physically tagged. Storage attributes and coherency Primary data caches must support the storage attributes defined by Book E with the following advisory: Note: The primary data cache may be implemented not to snoop (that is, not coherent with transactions outside the processor). System software is then responsible to maintain coherency. Thus the setting of the M attribute is meaningless. The preferred implementation provides snooping for primary data caches. Primary instruction caches must support the storage attributes defined by Book E with the following advisory:

  • The guarded attribute should be ignored for instruction fetch accesses. To prevent speculative fetch accesses to guarded memory, software should mark those pages as no-execute.
  • The cache may be implemented not to snoop (that is, not coherent with transactions outside the processor). System software is then responsible to maintain coherency. The preferred implementation does not provide snooping for primary instruction caches. As with other memory-related instructions, the effects of cache management instructions on memory are weakly-ordered. If the programmer must ensure that cache or other instructions have been performed with respect to all other processors and system mechanisms, an msync must be placed after those instructions.

5.4 Storage model

This section describes the storage model as it is defined by Book E and by the EIS.

5.4.1 Storage programming model

This section summarizes the register and instructions defined to support the cache model. Full descriptions of these resources are provided in Chapter 2: Register model on page 46,” and Chapter 3: Instruction model on page 133.”

RM0004 Storage architecture Storage model registers This section provides an overview of the registers used for programming the MMU. Full descriptions are provided in Chapter 2.12: MMU registers on page 97.” These registers consist of the following:

  • Process ID registers (PID0–PID2) are used by system software to identify TLB entries that are used by the processor to accomplish address translation for loads, stores, and instruction fetches. Book E defines one PID register (PID synonymous with PID0). The EIS defines 14 additional PID registers, PID1 through PID14. A implementation may choose to provide any number of PIDs up to a maximum of 15. The number of PIDs implemented is indicated by the value of MMUCFG[NPIDS] and the number of bits implemented in each PID register is indicated by the value of MMUCFG[PIDSIZE]. PID values are used to construct virtual addresses for accessing memory (see Chapter 5.4.6”).
  • MMU assist registers (MAS0–MAS7) are used to transfer data to and from the TLB arrays. Software uses mfspr and mtspr to read and write MAS registers. Executing tlbre causes the TLB entry specified by MAS0[TLBSEL,ESEL] and MAS2[EPN] to be copied to the MAS registers. Conversely, execution of a tlbwe instruction causes the TLB entry specified by MAS0[TLBSEL,ESEL] and MAS2[EPN] to be written with the MAS register contents. Hardware can also updated MAS registers on the occurrence of an instruction or data TLB error interrupt or as the result of a tlbsx. All MAS registers are supervisor level, and all except MAS5 and MAS7 must be implemented. MAS7 is not required if the processor supports 32 bits or less of physical address. Implementing MAS5 is implementation dependent. Processors are required to implement only the necessary bits of any multiple-bit MAS register field such that only the resources supplied by the processor are represented. Any non-implemented bits in a field should have no effect when writing and should

Storage architecture RM0004 always read as zero. For example, a processor that implements only two TLB arrays would likely implement only the lower-order MAS0[TLBSEL] bits. – MAS0, contains fields for identifying and selecting a TLB entry. – MAS1, contains fields for selecting a TLB entry during translation. – MAS2, contains fields for specifying the effective page address and the storage attributes for a TLB entry. – MAS3, contains fields for specifying the real page address and the permission attributes for a TLB entry. – MAS4, contains fields for specifying de fault information to be pre-loaded on certain MMU-related exceptions. – The optional MAS5 register, contains fields for specifying PID values to be used when searching TLB entries with the tlbsx instruction. – MAS6, contains fields for specifying PID and AS values used when the tlbsx instruction is used to search TLB entries. MAS7, contains the high-order address bits of the RPN for implementations that support more than 32 bits of physical address. Implementations that support 32 bits or fewer do not implement MAS7.

  • MMU configuration register (MMUCFG), provides configuration information about the MMU.
  • TLB configuration registers (TLBnCFG). One TLBnCFG register, is implemented to provide information about each TLB implemented. TLB0CFG corresponds to TLB0, TLB1CFG corresponds to TLB1, etc.
  • MMU control and status register (MMUCSR0), is used for general control of the MMU including flash invalidation of the TLB arrays and page sizes for programmable fixed size arrays. For TLB arrays with programmable fixed sizes, the TLBn_PS fields allow software to specify the page size. Storage model instructions The address translation mechanism is defined in terms of TLBs and page table entries (PTEs) Book E processors use to locate the logical-to-physical address mapping for a particular access. Table 103 on page 184 describes the operation of the TLB instructions, which are summarized as follows:
  • TLB Invalidate Virtual Address Indexed (tlbivax)
  • TLB Read Entry (tlbre)
  • TLB Search Indexed (tlbsx)
  • TLB Synchronize (tlbsync)
  • TLB Write Entry (tlbwe)

5.4.2 The storage architecture

This section describes the storage model as it is defined by Book E and by the ST EIS. Book E storage architecture The memory management approach defined by the Book E EIS is suited for desktop applications and has the simplicity and flexibility necessary for embedded applications. Book E supports demand-paged virtual memory as well as a variety of other management schemes that depend on precise control of effective-to-real address translation and flexible

RM0004 Storage architecture memory protection. Address translation misses and protection faults cause precise exceptions. Sufficient information is available to correct the fault and restart the faulting instruction. Each program on a 32-bit implementation can access 2 32 bytes of effective address (EA) space, subject to limitations imposed by the operating system. In a typical Book E system, each program’s EA space is a subset of a larger virtual address (VA) space managed by the operating system. Each effective (logical) address is translated to a real (physical) address before being used to access physical memory or an I/O device. Hardware does this by using the address translation mechanism described in Chapter 5.4.6.” The operating system manages the physically addressed resources of the system by setting up the tables used by the address translation mechanism. The Book E architecture divides the effective address space into pages. The page represents the granularity of effective address translation, permission control, and memory/cache attributes. Up to 12 page sizes (1, 4, 16, 64, or 256 Kbytes; 1, 4, 16, 64, or 256 Mbytes; or 1 Gbyte) may be simultaneously supported. For an effective-to-real address translation to exist, a valid entry for the page containing the effective address must be in a translation lookaside buffer (TLB). Addresses for which no TLB entry exists cause TLB miss exceptions (instruction or data TLB error interrupts). The instruction addresses generated by a program and the addresses used by load, store, and cache management instructions are effective addresses. However, in general, the physical memory space may not be large enough to map all the virtual pages used by the currently active applications. With support provided by hardware, the operating system can attempt to use the available real pages to map enough virtual pages for an application. If a sufficient set is maintained, paging activity is minimized, therefore maximizing performance. The operating system can restrict access to virtual pages by selectively granting permissions for user-state read, write, and execute, and supervisor-state read, write, and execute on a per-page basis. These permissions can be set up for a particular system (for example, program code might be execute-only, data structures may be mapped as read/write/no-execute) and can also be changed by the operating system based on application requests and operating system policies. EIS storage architecture The standard for Book E MMUs establishes a common way of implementing Book E processors to provide a programming model that is consistent across all products in the family. Having a standard reduces the software efforts required in porting to a new processor because the common programming model minimizes implementation differences. Thus, the standard defines configuration information for features such as TLBs, caches, and other entities that have standard forms, but differing attributes (like cache sizes and associativity) such that a single software implementation can be created that works efficiently for all implementations of a class.

  • The TLB, from a programming point of view, consists of zero or more TLB arrays, each of which may have differing characteristics.
  • The logical-to-physical address translation mechanism
  • Methods and effects of changing and manipulating TLB arrays
  • Configuration information available to the operating system that describes the structure and form of the TLB arrays and translation mechanism To assist or accelerate translation, implementations may contain other TLB structures not visible to the programming model. These structures and the methods for using them are not explicitly defined in the architecture or the ST standard, but they may be considered at the operating system level because they may affect an implementation’s performance.

5.4.3 Virtual address (VA)

an effective address, both used to construct the virtual address for an access. Figure 16. Virtual Address Space in Book E

5.4.4 Address spaces

cache management instructions. Figure 17. Current address space

RM0004 Storage architecture If the type of translation performed is an instruction fetch, the value of the AS bit is taken from the contents of MSR[IS]. If the type of translation performed is a load, store, or other data translation including target addresses of software-initiated instruction fetch hints and locks (icbt, icbtls, icbtlc) the value of the AS bit is taken from the contents of MSR[DS]. The address space indicator (MSR[IS] or MSR[DS], as appropriate) is used in addition to the effective address generated by the processor for translation into a physical address by the TLB mechanism. Because MSR[IS] and MSR[DS] are cleared when an interrupt occurs, an address space value of zero can be used to denote interrupt-related address spaces, or possibly all system software address spaces; an address space value of one can be used to denote non– interrupt-related address spaces, or possibly all user address spaces. Software Note: Although system software is free to use address space bits as it sees fit, on an interrupt, the MSR[IS] and MSR[DS] are cleared. This encourages software to use address space 0 for system software and address space 1 for user software. Instruction address spaces The two effective instruction address spaces are defined by the value of MSR[IS], and instruction fetch addresses are translated from the effective address space specified by the current value of MSR[IS]. Changing the value of MSR[IS] is considered a context-altering operation, requiring a context synchronization operation to follow it. When a context synchronizing event occurs, any prefetched instructions are discarded and instructions are refetched using the then-current state of MSR[IS] and the then-current program counter. See Context synchronization on page 144,” for more information on the definition of context synchronizing events. Instructions are not fetched from memory designated by the TLB mechanism as no-execute (UX = 0 or SX = 0). If the effective address of the current instruction is mapped to no- execute memory, an instruction storage interrupt (ISI) is generated. Note that mapping a page as no-execute does not affect instruction caches in the system (or any instructions resident in unified caches). Thus, if an instruction is loaded into a cache when its effective address is mapped to execute permitted memory, and the execute permissions for that page are later changed to no-execute, any instructions fetched before the no-execute mapping remain in the cache until explicitly evicted by an icbi instruction or through the cache’s replacement policy. However, attempted execution of such instructions still results in an ISI. Thus, for example, the operating system can change the designation of an application’s instruction pages to no-execute without having to first flush instruction cache blocks that map to these pages. Data address spaces The two effective data address spaces are defined by the value of MSR[DS] and data is accessed to/from the effective address space specified by the current value of MSR[DS]. As is the case with MSR[IS], changing the value of MSR[DS] is considered a context-altering operation, requiring a context synchronization operation to follow it. When a context synchronizing event occurs, subsequent accesses are made using the new state of MSR[DS] (see Context synchronization on page 144”). Data can be read from a page, provided the user read (UR) permission bit is set in the TLB for a user access, or the supervisor read (SR) bit is set for a supervisor access. Likewise, data write access permissions are determined by the user write (UW) and supervisor write (SW) permission bits. If permissions are violated, the appropriate interrupt is taken.

5.4.5 Process ID

Figure 18. Current PID Value processors may not implement all 14 bits of the process ID field. determine if other PID registers are implemented. process and for other PIDs to handle mappings that may be common to multiple processes. shared address space is mapped at the same virtual address in each process. (instruction or data) generated by the processor. processors may not implement all 14 bits of the process ID field. determine if other PID registers are implemented.

RM0004 Storage architecture assign PID0 to contain the unique process ID (for private mappings for the current processes) and may assign PID1 to contain the unique process ID for a common set of shared libraries. Note that Book E defines the value of all zeros for a TID field in a TLB entry as an entry that is globally shared. Thus, when PID values (up to 12 bits for ST devices) are compared to the TID fields in the TLB arrays for matches, if a TLB entry contains all zeros in the TID field, it globally matches all PID values. PID registers are more fully described in Chapter 2.12.1: Process ID registers (PID0–PIDn) on page 97.” Address space identifiers The AS bit is the address space identifier. Thus there are two possible address spaces, 0 and 1. The value of the AS bit is determined by the type of translation performed and from the contents of the MSR when an address is translated. If the type of translation performed is an instruction fetch, the value of the AS bit is taken from the contents of MSR[IS]. If the type of translation performed is a load, store, or other data translation including target addresses of software initiated instruction fetch hints and locks (icbt, icbtls, icbtlc) the value of the AS bit is taken from the contents of MSR[DS]. The AS bit is defined by Book E. Note: Although system software is free to use address space bits as it sees fit, it should be noted that on interrupt, the MSR[IS] and MSR[DS] bits are cleared. This encourages software to use address space 0 for system software and address space 1 for user software.

5.4.6 Address translation

The effective address (EA) is the untranslated address for an instruction fetch address or for a data address that is calculated as a result of a load, store, or cache management instruction. The EA, concatenated with the MSR[IS] or MSR[DS] address space (AS) value, is compared to the appropriate number of bits of the EPN field (depending on the page size) and the TS field of the TLB entry. If a match occurs, that TLB entry is a candidate for a translation match. In addition to a match in the EPN field and TS, a matching TLB entry must match with the current process ID of the access. Figure 19 shows the translation match logic for the effective address plus its attributes (collectively called the virtual address) and how it is compared with the corresponding fields in the TLB entries.

RM0004 Storage architecture Note that a PID register containing a 0 value (or the same value as another PID register) forms a non-unique VA. Duplicate VAs are ignored. Each of the unique VAs are compared to all the valid TLB entries by comparing specific fields of each TLB entry to each of the VAs. The fields of each valid (TLB[V] = 1) TLB entry are combined to form a set of matching TLB address (TAs): TA ← TLB Each TA is compared to all VAs under a mask based on the page size (TLB[SIZE]) of the TLB entry. The mask of the comparison of the EA and EPN portions of the virtual and translation addresses is computed as follows: mask ← ~(1024 << (2 * TLB SIZE)) - 1) where the number of bits in the mask is equal to the number of bits in a TA (or VA). If a TA matches any VA the TLB entry is said to match. If more than one TA/VA match occurs, it is considered a serious programming error and the results are undefined. The recommended behavior is that a machine check interrupt is taken. Once a match occurs the matching TLB is used for access control, storage attributes, and effective to real address translation. Access control, storage attributes, and address translation are defined by Book E (additional storage attributes are defined within this document).

5.4.7 Address transl ation and the ST EIS

Translating an effective address to a real address is defined by Book E to require four elements:

  • The address space value. Depending on the type of translation (instruction or data), MSR[IS] or MSR[DS] is used.
  • The TLB entries in the TLB arrays
  • The effective address being translated The following subsections describe these elements as they are further defined by the EIS. Match criteria for TLB entries TLB arrays contain TLB entries that are used to match any address presented for translation. All TLB entries for any given implementation are candidates for any given translation. The TLB itself is unordered with respect to the various elements used in address translations, and regardless of implementation, should be considered to perform the translation comparison with all entries in parallel. There should be only one valid matching translation for a given effective address, PID value, and address space value. If the TLB contains more than one matching entry, it is considered a programming error, and the behavior of any such translation is undefined. In this case, the processor is likely to enter checkstop state or take a machine check interrupt.

Storage architecture RM0004 The following fields are compared in the TLB entries:

  • V—The matching entry must have the V bit set.
  • TS—The address space identifier used for translation. The appropriate bit of MSR[IS] or MSR[DS] must match the TS bit for a matching entry.
  • TID—The contents of a PID register must match the TID field of a matching entry, or the TID field must be all zeros for a matching entry.
  • EPN—The appropriate number of bits (depending on the page size) of the effective address being translated is compared to the EPN field of the TLB entry. If a match occurs on all the fields listed above, the physical address is formed by replacing the effective page number in the effective address with the value in the RPN field of the matching TLB entry. The number of bits in the page number depends on the page size for that TLB entry. Translation algorithms The following algorithm describes how translation operates at the ST Book E level: ea = effective address if translation is an instruction address then as = MSR[IS] else // data address translation as = MSR[DS] for all TLB entries if ! TLBV then next // compare next TLB entry if as != TLBTS then next if TLBTID == 0 then goto pid_match for all PID registers if this PID register == TLBTID then goto pid_match endfor next // no PIDs matched pid_match: // translation match mask = ~((1024 << (2 * TLB TSIZE)) - 01) if (ea & mask) != TLBEPN then next // no address match real address = TLBRPN | (ea & ~mask) // real address computed end translation -- success endfor end translation -- tlbmiss The algorithm for the granting of permission is as follows: if MSR PR == 0 then x = TLBSX r = TLBSR w = TLBSW else x = TLBUX r = TLBUR

RM0004 Storage architecture w = TLBUW if instruction fetch address then if x == 0 then Instruction Storage Interrupt else // data access if data read (load) then if r == 0 then Data Storage Interrupt else // write access (store) if w == 0 then Data Storage Interrupt Access control If address translation results in a match (hit), the matching TLB entry is used to perform access control (permission checks). These checks are based on the privilege level of the access (MSR[PR]) and the type of access (fetch for execute, read for loads, and write for stores). The TLB entry’s permission bits (TLB[US,SX,UW,SW,UR,SR]) determine if the operation should succeed. If permission is denied, execution of the instruction is suppressed and an instruction storage interrupt or data storage interrupt occurs as defined in Book E. Software uses the ESR, SRR0, and the DEAR to determine the type of operation attempted and then must perform a TLB search if updating the TLB is desired. The algorithm for determining access control is as follows: if MSR PR = 0 then x ← TLBSX r ← TLBSR w ← TLBSW else x ← TLBUX r ← TLBUR w ← TLBUW if instruction_fetch & x = 0 then take instruction storage interrupt else if load & r = 0 then take data storage interrupt else if store & w = 0 then take data storage interrupt else access permitted Physical (real) address generation If permission checking is successful, the real address is formed by combining the TLB[RPN] with the lower order offset bits of the EA based on the page size of the TLB entry. mask ← ~(1024 << (2 * TLB SIZE)) - 1) real_address ← ((TLBRPN << 12) & mask) | (EA & ~mask)

TLB entry to determine how the location should be accessed. are compared with the corresponding EPN field in the TLB entry as shown in Table 185.

  • SR—Supervisor read permission
  • SW—Supervisor write permission
  • SX—Supervisor execute permission
  • UR—User read permission
  • UW—User write permission
  • UX—User execute permission If the virtual address translation comparison with TLB entries was successful, the permission bits for the matching entry are checked as shown in Figure 21. If the access is not allowed by the access permission mechanism, the processor generates an instruction or data storage interrupt (ISI or DSI).

Table 185. Page size and EPN field comparison

Figure 21. Granting of Access Permission If no virtual address match occurs, the translation fails and a TLB miss exception occurs. interrupt or the data TLB error interrupt is taken. Table 186. Real address generation

5.4.8 Permission attributes

The UX and SX bits of the TLB entry control execute access to the corresponding page. while the processor is in supervisor mode. an execute access control exception-type instruction storage interrupt (ISI) is taken. The UR and SR bits of the TLB entry control read access to the corresponding page. The UW and SW bits of the TLB entry control write access to the corresponding page. Table 187. Permission control for instruction, data read, and data write accesses

and a write access control exception-type data storage interrupt (DSI) is taken. access control exception-type DSIs. completes, but the allocate operation is merely cancelled (essentially, a no-op). can cause a read access control exception-type data storage interrupt. instruction execution completes, but the operation is cancelled (essentially, a no-op). The dcbf and dcbst instructions are treated as loads with respect to permissions checking. access control exception-type data storage interrupts. Table 188. Permission control and cache instructions

Storage architecture RM0004 Permissions control and string instructions When the string length is zero, neither lswx nor stswx can cause data storage interrupts due to permissions violations. Use of permissions to maintain page history The Book E architecture TLB entry definition does not define bits for maintaining page history information. The U0–U3 bits in the TLB entries can be used by software for storing history information, but implementations may ignore these bits internally. Page changed bit status can be implemented in the system software by disabling write permissions to all pages. The first attempt to write to the page results in a data storage interrupt. At this point, system software can record the page changed bit in memory, update the TLB entry permission to allow writes to that page, and return to the user program allowing further writes to the page to proceed without exception. Crossing page boundaries Care must be taken with single instruction accesses (load/stores) that cross page boundaries. Examples are lmw and stmw instructions and misaligned accesses on implementations that support misaligned load/stores. Architecturally, each of the parts of the access that cross the natural boundary of the access size (half word, word, double word) are treated separately with respect to exception conditions. Additionally, these types of instructions may optionally partially complete. For example, a store word instruction that crosses a page boundary because it is misaligned to the last half word of a page might actually store the first 16 bits because the access was permitted, but produce a DSI or data TLB error exception because the second 16 bits in the next page were not valid or they were protected. An implementation may choose to suppress the first 16-bit store or perform it.

5.4.9 Translation lookaside buffer (TLB) arrays

The MMU contains up to four TLB arrays, which are on-chip storage areas for holding TLB entries. A TLB entry contains effective-to-physical address mappings for loads, stores, and instruction fetches. A TLB array must contain TLB entries that share the same characteristics and contains zero or more TLB entries. Each TLB entry has specific fields that can be addressed by the corresponding fields in the MMU assist registers (see Chapter 2.12.5: MMU assist registers (MAS0–MAS7) on page 101”). Each implemented TLB array has an associated configuration register (TLBnCFG) describing the size and attributes of the TLB entries in that array. See Chapter 2.12.4: TLB configuration registers (TLBnCFG) on page 100.” The architected fields of a TLB entry are described in Table 189. 1. dcba, dcbt, dcbtst, and icbt may cause a read access control exception but does not result in a data storage interrupt (DSI).

5.4.10 TLB management

Table 189. TLB entry V Valid bit. A 1-bit entry that specifies whether this TLB entry is valid for translation. TID Translation ID. Identifies which process ID (PID) that this TLB entry is valid for. global and matches all PID values. when a transition from user mode to supervisor mode occurs. This is a 1 bit field. = 0) then this field is ignored. Kbyte page sizes defined in Book E. Page size encoding is defined by Book E. EPN Effective page number. Describes the logica l or effective starting address of the page. bit implementations, this field is 2 to 20 bits, depending on the page size (SIZE). page size and the number of bits of real address supported by the implementation. subsystem (caches and bus transactions). The WIMGE bits are defined by Book E. Coherence Mode is used only when the M bit (from WIMGE) is set. access to this page to decode and execute as VLE (and EIS APU) instructions. Permissions. User and supervisor read, write, and execute permission bits. Supervisor and user permission bits are defined by Book E. associated with a TLB entry to be used by system software. invalidation mechanisms except the explicit writing of a 0 to the V bit.

Storage architecture RM0004 default values when translation or protection faults occur. See Chapter 2.12.5: MMU assist registers (MAS0–MAS7) on page 101.” TLB configuration information Information about the configuration for a given TLB implementation is available to system software by reading the contents of the MMU configuration SPRs. These SPRs describe the architectural version of the MMU, the number of TLB arrays, and the characteristics of each TLB array. MMU architecture version number 1 is defined as these SPRs with the field definitions as described in this section.

  • MMU configuration register (MMUCFG), implemented by all ST Book E processors, contains basic information about the MMU architecture for each device. TLB configuration registers (TLBnCFG). Implemented by all ST Book E processors for each of the TLBs specified in MMUCFG[NTLBS]. They contain configuration information about each particular TLB. See Chapter 2.12.3: MMU configuration register (MMUCFG) on page 99.”
  • The TLBnCFG number assignment is the same as the value in MAS0[TLBSEL]. For example, TLB0CFG provides configuration information about TLB0, and TLB1CFG provides configuration information about TLB1. See Chapter 2.12.4: TLB configuration registers (TLBnCFG) on page 100.” TLB entries The software-visible TLB is subdivided into zero or more TLB arrays. Each array must contain TLB entries that share the same characteristics. Each TLB array contains one or more TLB entries. Each entry has specific fields that correspond to fields in the seven MMU assist (MAS) registers, described in Chapter 2.12.5: MMU assist registers (MAS0–MAS7) on page 101.” Some TLB fields are architected in Book E and others are architected in the EIS. Note that Book E architected fields may have restrictions or enhancements imposed by the EIS for the Book E implementations. The IPROT TLB entry, architected by the EIS, designates TLB entries as protected from certain kinds of invalidation. TLB invalidation and the IPROT field are described further in Invalidating TLB entries on page 321.” Reading and writing TLB entries All TLB entries are updated by executing tlbwe instructions. At the time of tlbwe execution, the MMU assist registers (MAS0–MAS6), a set of SPRs defined by the EIS, are used to index a specific TLB entry. The MAS registers also contain the information that is written to the indexed entry, such that they serve as the ports into the TLBs, as shown in Figure 22. The contents of the MAS registers are described in Chapter 2.12.5: MMU assist registers (MAS0–MAS7) on page 101.

Figure 22. TLBs accessed through MAS registers and TLB instructions and then the desired tlbre or tlbwe instructions must be executed. an illegal instruction exception program interrupt if RA ≠ 0. registers contain the contents of the indexed TLB entry. from 0 to associativity - 1. instruction is summarized in Table 191. dependent field should be treated as a reserved field. of the MAS registers are written to the indexed TLB entry.

Storage architecture RM0004 Selection of the TLB entry to write is performed by setting MAS0[TLBSEL], MAS0[ESEL] and MAS2[EPN] to indicate the entry to write. MAS0[TLBSEL] selects which TLB the entry should be written from (0 to 3) and MAS2[EPN] selects the set of entries from which MAS0[ESEL] selects an entry. For fully associative TLBs, MAS2[EPN] is not used to identify a TLB entry since the value in MAS0[ESEL] fully identifies the TLB entry. Valid values for MAS0[ESEL] are from 0 to associativity minus 1. The selected TLB entry is then written with following fields of the MAS registers: V, IPROT, TID, TS, TSIZE, EPN, ACM, VLE, WIMGE, RPN, U0 — U3, and permissions. If the TLB array supports NV, it is written with the NV value. The effects of updating the TLB entry are not guaranteed to be visible to the programming model until the completion of a context synchronizing operation. Writing a TLB entry that is used by the programming model prior to a context synchronizing operation produces undefined behavior. No operands are given for the tlbwe instruction and the Book E defined implementation dependent field should be treated as a reserved field. Specifying invalid values for MAS0[TLBSEL] and MAS0[ESEL] produce boundedly undefined results. Note: Writing TLB entries should be followed by an isync or an rfi before the new entries are to be used by the programming model. Invalidating TLB entries TLB entries may be invalidated by any of the following methods:

  • A TLB entry can be invalidated as the result of a tlbwe instruction that clears MAS0[V] in the entry.
  • As a result of a tlbivax instruction or from a received broadcast invalidation resulting from a tlbivax on another processor in an SMP system.
  • As a result of a flash invalidate. In both multiprocessor and uniprocessor systems, invalidations can occur on a wider set of TLB entries than intended. This are called generous invalidations That is, a virtual address presented for invalidation may invalidate not only the targeted TLB, but also may invalidate other TLB entries, depending on the implementation. This is because parts of the translation mechanism may not be fully specified to the hardware at invalidate time. This is especially true in SMP systems where the invalidation address must be broadcast globally to all processors in the system. Hardware may impose other limitations. The architecture ensures that the intended TLB is invalidated, but does not guarantee that it is the only one. A TLB entry invalidated by clearing the V bit of the TLB entry by use of a tlbwe is guaranteed to invalidate only the addressed TLB entry. However, invalidates occurring from tlbivax instructions or from the multiprocessor broadcasts as a result of tlbivax instructions may cause generous invalidates. The architecture provides a method to protect against generous invalidations. This is important, because certain virtual memory regions (most notably, the code memory region that serves as the exception handler for MMU faults) must be properly mapped to for forward progress to occur. If this region does not have a valid mapping, an MMU exception cannot be handled because the first address of the interrupt handler causes another MMU exception. To prevent this, the architecture specifies an IPROT bit for TLB entries. Setting the MAS0[PROT] protects the corresponding TLB entry from invalidations resulting from tlbivax instructions, as a result of broadcast invalidation from another processor in an SMP

RM0004 Storage architecture system, or from flash invalidations. TLB entries with the IPROT field set can be invalidated only by explicitly writing the TLB entry and specifying a 0 for MAS1[V]. Note: Software Note: Not all TLB arrays in a given implementation implement the IPROT attribute. It is likely that implementations that are suitable for demand page environments implement it for only a single array, while not implementing it for other arrays. Software Note: Operating systems must use great care when using protected (IPROT) TLB entries, particularly in SMP systems. An SMP system that contains TLB entries on other processors requires a cross-processor interrupt or some other synchronization mechanism to assure that each processor performs the required invalidation by writing its own TLB entries. Invalidations using tlbivax: The tlbivax instruction provides a virtual address as a target for invalidation. EA[0–51] are used to find a TLB entry with a matching EPN field. The page size of the TLB entry is used to mask the low order bits in the comparison. The comparison is performed only for TLB entries in the specified TLB array, do not have the IPROT attribute set (if supported by the TLB array), and are valid. The AS bit does not participate in the comparison. The EA specified by the rA and rB operands in the tlbivax instruction contains fields in the lower order bits to augment the invalidation to specific TLB arrays and to flash invalidate those arrays. Note that TLB entry invalidations resulting from tlbivax instructions do not invalidate any entry that has IPROT = 1 unless the specified TLB array does not support the IPROT attribute. The encoding of the EA used by tlbivax is shown in Table 190. Note: Software Note: To ensure a TLB entry that is not protected by IPROT is invalidated if software does not know which TLB array the entry is in, software should issue a tlbivax instruction targeting each TLB in the implementation with the EA to be invalidated. Software Note: The preferred form of the tlbivax instruction contains the entire EA in rB and zero in rA. Some implementations may take an Unimplemented Instruction exception if rA is non-zero. EA format for tlbivax ia shown below. EA Format for tlbivax Table 190 describes EA fields for tlbivax. 0 51 52 58 59 60 61 62 63 EA for tlbivax EA0:51 —T L B I A —

array does not support the IPROT attribute. Software may search the MMU by using the tlbsx instruction that is provided by Book E. summarizes the update of MAS registers as a result of a tlbsx instruction. unimplemented instruction exception or an illegal instruction exception if rA != 0. or instruction TLB error interrupt. Table 190. Fields for EA format of tlbivax 0–51 EA 0:51 The upper bits of the address to invalidate. 52–58 — Reserved, should be cleared. 59–60 TLB Selects TLB array for invalidation. 61 IA Invalidate all entries in selected TLB array. 62–63 — Reserved, should be cleared.

targeted for replacement is implementation dependent. requiring only the single MAS register manipulation by software before writing the TLB entry. explicitly search the TLB to find the appropriate entry. Table 191. MAS register update summary

5.4.11 MAS registers and exception handling

exception, some MAS register fields are loaded with default information specified in MAS4. System software should set up the default information in MAS4 before allowing exceptions. detail specific MAS register fields and the contents loaded for each exception type.

  • An instruction TLB error interrupt
  • A data TLB error interrupt Instruction TLB error interrupt settings An instruction TLB error interrupt occurs when the virtual address associated with an instruction address (fetch) does not match any valid entry in the TLB (that is, the address for the instruction cannot be translated). In addition to the values automatically written to the MAS2[W] MAS4[WD] TLB[ W] MAS4[WD] TLB[W] MAS2[I] MAS4[ID] TLB[I] MAS4[ID] TLB[I] MAS2[M] MAS4[MD] TLB[M] MAS4[MD] TLB[M] MAS2[G] MAS4[GD] TLB[G] MAS4[GD] TLB[G] MAS2[E] MAS4[ED] TLB[E] MAS4[ED] TLB[E] MAS3[RPN] 0 TLB[RPN] (bits 32:51) 0T L B [ R P N ] (bits 32:51) MAS3[U0,U1,U 2,U3] — TLB[U0,U1,U2,U3] — TLB[U0,U1,U2,U MAS3[UX,SX,U SW,UR,SR]

0 TLB[UX,SX,UW,

Table 191. MAS register update summary (continued)

RM0004 Storage architecture MAS registers (described in TLB miss exception MAS register settings on page 326”), SRR0 contains the address of the instruction that caused the instruction TLB error. This SRR0 value is used to identify the EA for handling the exception as well as the address to return to when system software has resolved the exception condition by writing a new TLB entry. Data TLB error interrupt settings A data TLB error interrupt occurs when the virtual address associated with a data reference from a load, store, or cache management instruction does not match any valid entry in the TLB (that is, the address of the data item of a load or store instruction cannot be translated). In addition to the values automatically written to the MAS registers (described in TLB miss exception MAS register settings on page 326”), the effective address of the data access that caused the exception is automatically loaded in the data exception address register (DEAR). Also, SRR0 contains the address of the instruction that caused the data TLB error and its value is used to identify the address to return to when system software has resolved the exception condition (by writing a new TLB entry). TLB miss exception MAS register settings When either an instruction or data TLB error interrupt occurs, the TLB information and selection fields of the MAS registers are loaded with default values from other MAS registers to assist in processing the exception. The intention is that the common case of a page fault generally requires only system software to load the RPN (corresponding to the physical address that will be used for this page), and the access permissions and the defaults can be used for the remaining MAS fields. The processor may use the next victim (NV) field from the TLB array to select which TLB entry should be used for the new translation. The method used to select the candidate TLB for replacement (the next victim) is implementation-dependent and may vary on different Book E implementations. In any case, software is free to choose any TLB entry for the replacement (software can overwrite the value in MAS0[ESEL]). The EIS defines the fields set in the MAS registers at exception time for an instruction or data TLB error interrupt as shown in Table 192.

All other MAS register values are unchanged. Table 192. MAS settings for an instruction or data TLB error interrupt the value loaded into ESEL is undefined. value 1, the contents of PID1 are written to the TID field. described the context that was running when the exception occurred). variable-sized pages, the value for EPN is undefined. Permissions SR, UR, UW, SW, UX, SX cleared to 0 (no permissions); note that U0–U3 are unchanged.

when system software has resolved the exception condition by writing a new TLB entry. fields are automatically loaded into the MAS registers to assist in processing the interrupt. the MAS registers at exception time for instruction or data storage interrupts. All other MAS register values are unchanged. Table 193. MAS settings for permissions violation ISI or DSI the context that was running when the exception occurred).

Table 194. MMU assist register field updates—EIS definition

6 Instruction set

  • Book E instructions defined for 32-bit implementations. This includes instructions not implemented in all Book E devices.
  • Instructions defined by the EIS, except for the instructions defined by the VLE extension. Full descriptions of these instructions are provided in Chapter 13: VLE instruction set on page 891.”

6.1 Notation

Table 195. Notation conventions n0 means a field of n bits with each bit equal to 0. Thus 50 is equivalent to 0b0_0000. n1 means a field of n bits with each bit equal to 1. Thus 51 is equivalent to 0b1_1111.

6.2 Instruction fields

Table 196 describes instruction fields. Table 196. Instruction field descriptions

0 The immediate field represents an address relative to the current instruction

target is the value CIA+EXTS(LI||0b00). target is the value CIA+EXTS(BD||0b00). 1 The immediate field represents an absolute address. target is the value EXTS(LI||0b00). target is the value EXTS(BD||0b00).

LINK bit. Indicates whether the link register (LR) is set. bits from bit MB+32 through bit ME+32 inclusive and 0 bits elsewhere. 0Do not alter the condition register. 1Set condition register field 0 or field 1. Table 196. Instruction field descriptions (continued)

6.3 Description of instruction operations

Figure 23. Some of this notation is used in the formal descriptions of instructions. are boundedly undefined, and may not cover all invalid forms. instruction. They do not imply any particular implementation. Table 197. RTL notation signed fractional result of x+y bits.

the block is removed from the data cache. cache, the block is removed from the instruction cache. copied into the instruction cache. Table 197. RTL notation (continued)

Contents of y bytes of memory starting at address x. address x+y–1 is the LSB of the value being accessed. address x+y–1 is the MSB of the value being accessed. between different executions on the same implementation. does not correspond to any architected register.

evaluated before serving as operands.

6.3.1 SPE APU saturati on and bit-reverse models

those functions that are referenced in the instruction pseudo RTL. leave Leave innermost do loop, or do loop described in leave statement. Table 198. Operator precedence

6.3.2 Embedded floating- point conversion models

called from the individual instruction pseudo RTL descriptions. Table 199. Conversion models

Common embedded floating-point functions This section includes common functions used by the functions in subsequent sections. 32-Bit NaN or Infinity Test // Determine if fp value is a NaN or Infinity Isa32NaNorInfinity(fp) return (fp exp = 255) Isa32NaN(fp) return ((fp exp = 255) & (fpfrac ≠ 0)) Isa32Infinity(fp) return ((fp exp = 255) & (fpfrac = 0)) // Determine if fp value is denormalized Isa32Denorm(fp) return ((fp exp = 0) & (fpfrac ≠ 0)) // Determine if fp value is a NaN or Infinity Isa64NaNorInfinity(fp) return (fp exp = 2047) Isa64NaN(fp) return ((fp exp = 2047) & (fpfrac ≠ 0)) Isa64Infinity(fp) return ((fp exp = 2047) & (fpfrac = 0)) // Determine if fp value is denormalized Isa64Denorm(fp) return ((fp exp = 0) & (fpfrac ≠ 0)) Signal Floating-Point Error // Signal a Floating-Point Error in the SPEFSCR SignalFPError(upper_lower, bits) if (upper_lower = UPPER) then bits ← bits << 15 SPEFSCR ← SPEFSCR | bits bits ← (FG | FX) if (upper_lower = UPPER) then bits ← bits << 15 SPEFSCR ← SPEFSCR & ¬bits Round a 32-Bit Value // Round a result Round32(fp, guard, sticky) FP32format fp; if (SPEFSCR FINXE = 0) then if (SPEFSCRFRMC = 0b00) then // nearest if (guard) then if (sticky | fpfrac[22]) then v[0:23] ← fpfrac + 1 if v[0] then if (fpexp >= 254) then // overflow fp ← fp sign || 0b11111110 || 231 else

fpexp ← fpexp + 1 fpfrac ← v1:23 else fpfrac ← v[1:23] else if ((SPEFSCRFRMC & 0b10) = 0b10) then // infinity modes // implementation dependent return fp Round a 64-Bit Value // Round a result Round64(fp, guard, sticky) FP32format fp; if (SPEFSCR FINXE = 0) then if (SPEFSCRFRMC = 0b00) then // nearest if (guard) then if (sticky | fpfrac[51]) then v[0:52] ← fpfrac + 1 if v[0] then if (fpexp >= 2046) then // overflow fp ← fp sign || 0b11111111110 || 521 else fpexp ← fpexp + 1 fpfrac ← v1:52 else fpfrac ← v1:52 else if ((SPEFSCRFRMC & 0b10) = 0b10) then // infinity modes // implementation dependent return fp Convert from single-precision floating-point to integer word with saturation // Convert 32-bit floating point to integer/factional // signed = SIGN or UNSIGN // upper_lower = UPPER or LOWER // round = ROUND or TRUNC // fractional = F (fractional) or I (integer) CnvtFP32ToI32Sat(fp, signed, upper_lower, round, fractional) FP32format fp; if (Isa32NaNorInfinity(fp)) then // SNaN, QNaN, +-INF SignalFPError(upper_lower, FINV) if (Isa32NaN(fp)) then return 0x00000000 // all NaNs if (signed = SIGN) then if (fp sign = 1) then return 0x80000000 else return 0x7fffffff else

if (fpsign = 1) then return 0x00000000 else return 0xffffffff if (Isa32Denorm(fp)) then SignalFPError(upper_lower, FINV) return 0x00000000 // regardless of sign if ((signed = UNSIGN) & (fp sign = 1)) then SignalFPError(upper_lower, FOVF) // overflow return 0x00000000 if ((fp exp = 0) & (fpfrac = 0)) then return 0x00000000 // all zero values if (fractional = I) then // convert to integer max_exp ← 158 shift ← 158 - fpexp if (signed = SIGN) then if ((fpexp ≠ 158) | (fpfrac ≠ 0) | (fpsign ≠ 1)) then max_exp ← max_exp - 1 else // fractional conversion max_exp ← 126 shift ← 126 - fpexp if (signed = SIGN) then shift ← shift + 1 if (fpexp > max_exp) then SignalFPError(upper_lower, FOVF) // overflow if (signed = SIGN) then if (fp sign = 1) then return 0x80000000 else return 0x7fffffff else return 0xffffffff result ← 0b1 || fpfrac || 0b00000000 // add U to frac guard ← 0 sticky ← 0 for (n ← 0; n < shift; n ← n + 1) do sticky ← sticky | guard guard ← result & 0x00000001 result ← result > 1 // Report sticky and guard bits if (upper_lower = UPPER) then SPEFSCR FGH ← guard SPEFSCRFXH ← sticky else SPEFSCRFG ← guard SPEFSCRFX ← sticky

if (guard | sticky) then SPEFSCRFINXS ← 1 // Round the integer result if ((round = ROUND) & (SPEFSCR FINXE = 0)) then if (SPEFSCRFRMC = 0b00) then // nearest if (guard) then if (sticky | (result & 0x00000001)) then result ← result + 1 else if ((SPEFSCRFRMC & 0b10) = 0b10) then // infinity modes // implementation dependent if (signed = SIGN) then if (fpsign = 1) then result ← ¬result + 1 return result Convert from double-precision floating-point to integer word with saturation // Convert 64-bit floating point to integer/fractional // signed = SIGN or UNSIGN // round = ROUND or TRUNC // fractional = F (fractional) or I (integer) CnvtFP64ToI32Sat(fp, signed, round, fractional) FP64format fp; if (Isa64NaNorInfinity(fp)) then // SNaN, QNaN, +-INF SignalFPError(LOWER, FINV) if (Isa64NaN(fp)) then return 0x00000000 // all NaNs if (signed = SIGN) then if (fp sign = 1) then return 0x80000000 else return 0x7fffffff else if (fpsign = 1) then return 0x00000000 else return 0xffffffff if (Isa64Denorm(fp)) then SignalFPError(LOWER, FINV) return 0x00000000 // regardless of sign if ((signed = UNSIGN) & (fp sign = 1)) then SignalFPError(LOWER, FOVF) // overflow return 0x00000000 if ((fp exp = 0) & (fpfrac = 0)) then return 0x00000000 // all zero values

if (fractional = I) then // convert to integer max_exp ← 1054 shift ← 1054 - fpexp if (signed ← SIGN) then if ((fpexp ≠ 1054) | (fpfrac ≠ 0) | (fpsign ≠ 1)) then max_exp ← max_exp - 1 else // fractional conversion max_exp ← 1022 shift ← 1022 - fpexp if (signed = SIGN) then shift ← shift + 1 if (fpexp > max_exp) then SignalFPError(LOWER, FOVF) // overflow if (signed = SIGN) then if (fp sign = 1) then return 0x80000000 else return 0x7fffffff else return 0xffffffff result ← 0b1 || fpfrac[0:30] // add U to frac guard ← fpfrac[31] sticky ← (fpfrac[32:63] ≠ 0) for (n ← 0; n < shift; n ← n + 1) do sticky ← sticky | guard guard ← result & 0x00000001 result ← result > 1 // Report sticky and guard bits SPEFSCRFG ← guard SPEFSCRFX ← sticky if (guard | sticky) then SPEFSCRFINXS ← 1 // Round the result if ((round = ROUND) & (SPEFSCR FINXE = 0)) then if (SPEFSCRFRMC = 0b00) then // nearest if (guard) then if (sticky | (result & 0x00000001)) then result ← result + 1 else if ((SPEFSCRFRMC & 0b10) = 0b10) then // infinity modes // implementation dependent if (signed = SIGN) then if (fpsign = 1) then result ← ¬result + 1 return result

Convert from double-precision floating-point to integer double word with saturation // Convert 64-bit floating point to integer/fractional // signed = SIGN or UNSIGN // round = ROUND or TRUNC CnvtFP64ToI64Sat(fp, signed, round) FP64format fp; if (Isa64NaNorInfinity(fp)) then // SNaN, QNaN, +-INF SignalFPError(LOWER, FINV) if (Isa64NaN(fp)) then return 0x00000000_00000000 // all NaNs if (signed = SIGN) then if (fp sign = 1) then return 0x80000000_00000000 else return 0x7fffffff_ffffffff else if (fpsign = 1) then return 0x00000000_00000000 else return 0xffffffff_ffffffff if (Isa64Denorm(fp)) then SignalFPError(LOWER, FINV) return 0x00000000_00000000 // regardless of sign if ((signed = UNSIGN) & (fp sign = 1)) then SignalFPError(LOWER, FOVF) // overflow return 0x00000000_00000000 if ((fp exp = 0) & (fpfrac = 0)) then return 0x00000000_00000000 // all zero values max_exp ← 1086 shift ← 1086 - fpexp if (signed = SIGN) then if ((fpexp ≠ 1086) | (fpfrac ≠ 0) | (fpsign ≠ 1)) then max_exp ← max_exp - 1 if (fpexp > max_exp) then SignalFPError(LOWER, FOVF) // overflow if (signed = SIGN) then if (fp sign = 1) then return 0x80000000_00000000 else return 0x7fffffff_ffffffff else return 0xffffffff_ffffffff result ← 0b1 || fpfrac || 0b00000000000 // add U to frac guard ← 0 sticky ← 0 for (n ← 0; n < shift; n ← n + 1) do

sticky ← sticky | guard guard ← result & 0x00000000_00000001 result ← result > 1 // Report sticky and guard bits SPEFSCRFG ← guard SPEFSCRFX ← sticky if (guard | sticky) then SPEFSCRFINXS ← 1 // Round the result if ((round = ROUND) & (SPEFSCR FINXE = 0)) then if (SPEFSCRFRMC = 0b00) then // nearest if (guard) then if (sticky | (result & 0x00000000_00000001)) then result ← result + 1 else if ((SPEFSCRFRMC & 0b10) = 0b10) then // infinity modes // implementation dependent if (signed = SIGN) then if (fpsign = 1) then result ← ¬result + 1 return result Convert to single-precision floating-point from integer word with saturation // Convert from integer/factional to 32-bit floating point // signed = SIGN or UNSIGN // upper_lower = UPPER or LOWER // fractional = F (fractional) or I (integer) CnvtI32ToFP32Sat(v, signed, upper_lower, fractional) FP32format result; result sign ← 0 if (v = 0) then result ← 0 if (upper_lower = UPPER) then SPEFSCRFGH ← 0 SPEFSCRFXH ← 0 else SPEFSCRFG ← 0 SPEFSCRFX ← 0 else if (signed = SIGN) then if (v0 = 1) then v ← ¬v + 1 resultsign ← 1 if (fractional = F) then // fractional bit pos alignment maxexp ← 127 if (signed = UNSIGN) then

maxexp ← maxexp - 1 else maxexp ← 158 // integer bit pos alignment sc ← 0 while (v0 = 0) v ← v << 1 sc ← sc + 1 v0 ← 0 // clear U bit resultexp ← maxexp - sc guard ← v24 sticky ← (v25:31 ≠ 0) // Report sticky and guard bits if (upper_lower = UPPER) then SPEFSCR FGH ← guard SPEFSCRFXH ← sticky else SPEFSCRFG ← guard SPEFSCRFX ← sticky if (guard | sticky) then SPEFSCRFINXS ← 1 // Round the result resultfrac ← v1:23 result ← Round32(result, guard, sticky) return result Convert to double-precision floating-point from integer word with saturation // Convert from integer/factional to 64-bit floating point // signed = SIGN or UNSIGN // fractional = F (fractional) or I (integer) CnvtI32ToFP64Sat(v, signed, fractional) FP64format result; result sign ← 0 if (v = 0) then result ← 0 SPEFSCRFG ← 0 SPEFSCRFX ← 0 else if (signed = SIGN) then if (v[0] = 1) then v ← ¬v + 1 resultsign ← 1 if (fractional = F) then // fractional bit pos alignment maxexp ← 1023 if (signed = UNSIGN) then maxexp ← maxexp - 1

maxexp ← 1054 // integer bit pos alignment sc ← 0 while (v0 = 0) v ← v << 1 sc ← sc + 1 v0 ← 0 // clear U bit resultexp ← maxexp - sc // Report sticky and guard bits SPEFSCRFG ← 0 SPEFSCRFX ← 0 resultfrac ← v1:31 || 210 return result Convert to double-precision floating-point from integer double word with saturation // Convert from 64 integer to 64-bit floating point // signed = SIGN or UNSIGN CnvtI64ToFP64Sat(v, signed) FP64format result; result sign ← 0 if (v = 0) then result ← 0 SPEFSCRFG ← 0 SPEFSCRFX ← 0 else if (signed = SIGN) then if (v0 = 1) then v ← ¬v + 1 resultsign ← 1 maxexp ← 1054 sc ← 0 while (v0 = 0) v ← v << 1 sc ← sc + 1 v0 ← 0 // clear U bit resultexp ← maxexp - sc guard ← v53 sticky ← (v54:63 ≠ 0) // Report sticky and guard bits SPEFSCRFG ← guard SPEFSCRFX ← sticky if (guard | sticky) then SPEFSCRFINXS ← 1

// Round the result resultfrac ← v1:52 result ← Round64(result, guard, sticky) return result

6.3.3 Integer saturation models

// Saturate after addition SATURATE(ovf, carry, neg_sat, pos_sat, value) if ovf then if carry then return neg_sat else return pos_sat else return value

6.3.4 Embedded floating-point results

Appendix E: Embedded floating-point results on page 1156,” summarizes results of various types of SPE and SPFP floating-point operations on various combinations of input operands.

6.4 Instruction set

The rest of this chapter describes individual instructions, which are listed in alphabetical order by mnemonic. Figure 23 shows the format for instruction description pages.

Figure 23. Instruction description Note: The execution unit that executes the instruction may not be the same for all processors. The sum of the contents of rA and rB is placed into rD.

add r D,rA,rB( O E = 0 , R c = 0 ) add. r D,rA,rB( O E = 0 , R c = 1 ) addo r D,rA,rB( O E = 1 , R c = 0 ) addo. r D,rA,rB( O E = 1 , R c = 1 ) carry0:63 ← Carry(rA + rB) sum0:63 ← rA + rB if OE=1 then do OV ← carry32 ⊕ carry33 SO ← SO | (carry32 ⊕ carry33) if Rc=1 then do LT ← sum32:63 < 0 GT ← sum32:63 > 0 EQ ← sum32:63 = 0 rD ← sum carry0:63 ← Carry(rA + rB) sum0:63 ← rA + rB if OE=1 then do OV ← carry32 ⊕ carry33 SO ← SO | (carry32 ⊕ carry33) if Rc=1 then do LT ← sum32:63 < 0 GT ← sum32:63 > 0 EQ ← sum32:63 = 0 rD ← sum The sum of the contents of rA and rB is placed into rD. Other registers altered:

  • CR0 (if Rc=1) SO OV (if OE=1) Book E User 0 5 6 1 01 1 1 51 6 2 02 12 2 3 03 1 011111 rD rA rB O E 100001010 R c

_addx _a d d x Add se_add ‘ r X,rY sum32:63 ← GPR(RX) + GPR(RY) GPR(RX) ← sum32:63 The sum of the contents of GPR(rX) and the contents of GPR(rY) is placed into GPR(rX). Special Registers Altered: None 05 6 1 0 1 1 1 5 000001 0 0 R Y R X VLE User

addc r D,rA,rB( O E = 0 , R c = 0 ) addc. r D,rA,rB( O E = 0 , R c = 1 ) addco r D,rA,rB( O E = 1 , R c = 0 ) addco. r D,rA,rB( O E = 1 , R c = 1 ) carry0:63 ← Carry(rA + rB) sum0:63 ← rA + rB if OE=1 then do OV ← carry32 ⊕ carry33 SO ← SO | (carry32 ⊕ carry33) if Rc=1 then do LT ← sum32:63 < 0 GT ← sum32:63 > 0 EQ ← sum32:63 = 0 rD ← sum CA ← carry32 The sum of the contents of rA and rB is placed into rD. Other registers altered:

  • CA CR0 (if Rc=1) SO OV (if OE=1) Book E User 0 5 6 1 01 1 1 51 6 2 02 12 2 3 03 1 011111 rD rA rB O E 000001010 R c

adde r D,rA,rB( O E = 0 , R c = 0 ) adde. r D,rA,rB( O E = 0 , R c = 1 ) addeo r D,rA,rB( O E = 1 , R c = 0 ) addeo. r D,rA,rB( O E = 1 , R c = 1 ) if E=0 then Cin ← CA carry0:63 ← Carry(rA + rB + Cin) sum0:63 ← rA + rB + Cin if OE=1 then do OV ← carry32 ⊕ carry33 SO ← SO | (carry32 ⊕ carry33) if Rc=1 then do LT ← sum32:63 < 0 GT ← sum32:63 > 0 EQ ← sum32:63 = 0 rD ← sum CA ← carry32 For adde[o][.], the sum of the contents of rA, the contents of rB, and CA is placed into rD. Other registers altered:

  • CA CR0 (if Rc=1) SO OV (if OE=1) Book E User 0 5 6 1 01 1 1 51 6 2 02 12 2 3 03 1 011111 rD rA rB O E 010001010 R c

Add immediate [shifted] addi r D,rA,SIMM (S=0) addis r D,rA,SIMM (S=1) if rA=0 then a ← 640 else a ← rA if s=0 then b ← EXTS(SIMM) if s=1 then b ← EXTS(SIMM || 160) rD ← a + b If addi and rA=0, the sign-extended value of the SIMM field is placed into rD. If addi and rA≠0, the sum of the contents of rA and the sign-extended value of field SIMM is placed into rD. If addis and rA=0, the sign-extended value of the SIMM field, concatenated with 16 zeros, is placed into rD. If addis and rA≠0, the sum of the contents of rA and the sign-extended value of the SIMM field concatenated with 16 zeros, is placed into rD. Other registers altered: None Book E User 0 4 5 6 10 11 15 16 31 00111S r D r A S I M M

_addix _addix Add [2 operand] Immediate [Shifted] [and Record] e_add16i r D,rA,SI a ← GPR(RA) b ← EXTS(SI) GPR(RD) ← a + b The sum of the contents of GPR(rA) and the sign-extended value of field SI is placed into GPR(rD). Special Registers Altered: None e_add2i. r A,SI SI ← SI 0:4 || SI5:15 sum32:63 ← GPR(RA) + EXTS(SI) LT ← sum32:63 < 0 GT ← sum32:63 > 0 EQ ← sum32:63 = 0 GPR(RA) ← sum32:63 The sum of the contents of GPR(rA) and the sign-extended value of SI is placed into GPR(rA). Special Registers Altered: CR0 e_add2is r A,SI SI ← SI 0:4 || SI5:15 sum32:63 ← GPR(RD) + (SI || 160) GPR(RA) ← sum32:63 The sum of the contents of GPR(rA) and the value of SI concatenated with 16 zeros is placed into GPR(rAarav2006). Special Registers Altered: None e_addi r D,rA,SCI8 (Rc = 0) e_addi. r D,rA,SCI8 (Rc = 1) VLE User 0 5 6 1 01 1 1 51 6 3 1

000111 R D R A S I

0 5 6 1 01 1 1 51 6 2 02 1 3 1

011100 SI0:4 R A 1 0001 SI5:15

0 5 6 1 01 1 1 51 6 2 02 1 3 1

011100 SI0:4 R A 1 0010 SI5:15

0 5 6 1 01 1 1 51 6 2 02 12 22 32 4 3 1

000110 R D R A 1000 R c F S C L U I 8

imm ← SCI8(F ,SCL,UI8) sum32:63 ← GPR(RA) + imm if Rc=1 then do LT ← sum 32:63 < 0 GT ← sum32:63 > 0 EQ ← sum32:63 = 0 GPR(RD) ← sum32:63 The sum of the contents of GPR(rA) and the value of SCI8 is placed into GPR(rD). Special Registers Altered: CR0 (if Rc = 1) se_addi r X,OIMM GPR(RX) ← GPR(RX) + (270 || OFFSET(OIM5)) The sum of the contents of GPR(rX) and the zero-extended offset value of OIM5 (a final value in the range 1–32), is placed into GPR(rX). Special Registers Altered: None 0 5 6 7 11 12 15

0010000 OIM5(1)

  1. OIMM = OIM5 +1 RX

Add immediate carrying [and record] addic r D,rA,SIMM (Rc=0) addic. r D,rA,SIMM (Rc=1) carry0:63 ← Carry(rA + EXTS(SIMM)) sum0:63 ← rA + EXTS(SIMM) if Rc=1 then do LT ← sum 32:63 < 0 GT ← sum32:63 > 0 EQ ← sum32:63 = 0 rD ← rA+EXTS(SIMM) CA ← carry32 The sum of the contents of rA and the sign-extended value of the SIMM field is placed into rD. Other registers altered:

  • CA CR0 (if Rc=1) Book E User 0 4 5 6 10 11 15 16 31

00110 R c rD rAS I M M

_addicx _addicx Add Immediate Carrying [and Record] e_addic r D,rA,SCI8 (Rc = 0) e_addic. r D,rA,SCI8 (Rc = 1) imm ← SCI8(F ,SCL,UI8) carry32:63 ← Carry(GPR(RA) + imm) sum32:63 ← GPR(RA) + imm if Rc=1 then do LT ← sum 32:63 < 0 GT ← sum32:63 > 0 EQ ← sum32:63 = 0 GPR(RD) ← sum32:63 CA ← carry32 The sum of the contents of GPR(rA) and the value of SCI8 is placed into GPR(rD). Special Registers Altered: CA, CR0 (if Rc=1) VLE User 0 5 6 1 01 1 1 51 6 1 92 02 12 22 32 4 3 1

000110 R D R A 1001 R c F S C L U I 8

addme r D,rA( O E = 0 , R c = 0 ) addme. r D,rA( O E = 0 , R c = 1 ) addmeo r D,rA( O E = 1 , R c = 0 ) addmeo. r D,rA( O E = 1 , R c = 1 ) if E=0 then Cin ← CA carry0:63 ← Carry(rA + Cin + 0xFFFF_FFFF_FFFF_FFFF) sum0:63 ← rA + Cin + 0xFFFF_FFFF_FFFF_FFFF if OE=1 then do OV ← carry32 ⊕ carry33 SO ← SO | (carry32 ⊕ carry33) if Rc=1 then do LT ← sum32:63 < 0 GT ← sum32:63 > 0 EQ ← sum32:63 = 0 rD ← sum CA ← carry32 For addme[o][.], the sum of the contents of rA, CA, and 641 is placed into rD. Other registers altered:

  • CA CR0 (if Rc=1) SO OV (if OE=1) Book E User 0 5 6 1 01 1 1 51 6 2 02 12 2 3 03 1 011111 rD rA / / / O E 011101010 R c

addze r D,rA( O E = 0 , R c = 0 ) addze. r D,rA( O E = 0 , R c = 1 ) addzeo r D,rA( O E = 1 , R c = 0 ) addzeo. r D,rA( O E = 1 , R c = 1 ) if E=0 then Cin ← CA carry0:63 ← Carry(rA + Cin) sum0:63 ← rA + Cin if OE=1 then do OV ← carry32 ⊕ carry33 SO ← SO | (carry32 ⊕ carry33) if Rc=1 then do LT ← sum32:63 < 0 GT ← sum32:63 > 0 EQ ← sum32:63 = 0 rD ← sum CA ← carry32 For addze[o][.], the sum of the contents of rA and CA is placed into rD. Other registers altered:

  • CA CR0 (if Rc=1) SO OV (if OE=1) Book E User 0 5 6 1 01 1 1 51 6 2 02 12 2 3 03 1 011111 rD rA / / / O E 011001010 R c

AND [Immediate [Shifted] | with Complement] and r A,rS,rB( R c = 0 ) and. r A,rS,rB( R c = 1 ) andi. r A,rS,UIMM (S=0, Rc=1) andis. r A,rS,UIMM (S=1, Rc=1) andc r A,rS,rB( R c = 0 ) andc. r A,rS,rB( R c = 1 ) if ‘andi.’ then b ← 480 || UIMM if ‘andis.’ then b ← 320 || UIMM || 160 if ‘and[.]’ then b ← rB if ‘andc[.]’ then b ← ¬rB result0:63 ← rS & b if Rc=1 then do LT ← result32:63 < 0 GT ← result32:63 > 0 EQ ← result32:63 = 0 rA ← result For andi., the contents of rS are ANDed with 480 || UIMM. For andis., the contents of rS are ANDed with 320 || UIMM || 160. For and[.], the contents of rS are ANDed with the contents of rB. For andc[.], the contents of rS are ANDed with the one’s complement of the contents of rB. The result is placed into rA. Other registers altered: CR0 (if Rc=1) Book E User 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1 011111 rS rA rB 0000011100 R c 0 4 5 6 10 11 15 16 31 01110S rS rAU I M M 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1 011111 rS rA rB 0000111100 R c

_andx _andx AND [2 operand] [Immediate | with Complement] [and Record] se_and r X,rY( R c = 0 ) se_and. r X,rY( R c = 1 ) e_and2i. r D,UI e_and2is. r D,UI e_andi r A,rS,SCI8 (Rc = 0) e_andi. r A,rS,SCI8 (Rc = 1) se_andi r X,UI5 se_andc r X,rY if ‘e_andi[.]’ then b ← SCI8(F ,SCL,UI8) if ‘se_andi’ then b ← UI5 if ‘se_and[.]’ then b ← GPR(RY) if ‘se_andc’ then b ← ¬GPR(RY) if ‘e_and2i.’ then b ← 160 || UI0:4 || UI5:15 if ‘e_and2is.’ then b ← UI0:4 || UI5:15 || 160 result32:63 ← GPR(RS or RD or RX) & b if Rc=1 then do LT ← result32:63 < 0 GT ← result32:63 > 0 EQ ← result32:63 = 0 if ‘se_and[ci]’ then GPR(RX) ← result32:63 else GPR(RA or RD) ← result32:63 For e_andi[.], the contents of GPR(rS) are ANDed with the value of SCI8. For e_and2i., the contents of GPR(rD) are ANDed with 160 || UI. For e_and2is., the contents of GPR(rD) are ANDed with UI || 160. For se_andi, the contents of GPR(rX) are ANDed with the value of UI5. For se_and[.], the contents of GPR(rX) are ANDed with the contents of GPR(rY). 05 6 7 8 1 1 1 2 1 5

0100011 R c R Y R X

0 5 6 1 01 1 1 51 6 2 02 1 3 1

011100 R D UI0:4 11001 UI5:15

0 5 6 1 01 1 1 51 6 2 02 1 3 1

011100 R D UI0:4 11101 UI5:15

0 5 6 1 01 1 1 51 6 2 02 12 22 32 4 3 1

000110 R S R A 1100 R c F S C L U I 8

0010111 U I 5 R X

01000101 R Y R X

For se_andc, the contents of GPR(rX) are ANDed with the one’s complement of the contents of GPR(rY). The result is placed into GPR(rA) or GPR(rX) (se_and[ic][.]) Special Registers Altered: CR0 (if Rc = 1)

Branch [and Link] [Absolute] b LI (AA=0, LK=0) ba LI (AA=1, LK=0) bl LI (AA=0, LK=1) bla LI (AA=1, LK=1) if AA=1 then a ← 640 else a ← CIA if E=0 then NIA ← 320 || (a + EXTS(LI||0b00))32:63 if LK=1 then LR ← CIA + 4 The branch target effective address (BTEA) is calculated as follows:

  • For 32-bit implementations, BTEA is bits 32–63 of the sum of the current instruction address (CIA), or 32 zeros if AA=1, and the sign-extended value of the LI instruction field concatenated with 0b00 BTEA is the address of the next instruction to be executed. If LK=1, the sum CIA+4 is placed into the LR. Other registers altered: LR (if LK=1) Book E User 05 6 29 30 31

010010 L I A A L K

_bx _bx Branch [and Link] e_b BD24 (LK = 0) e_bl BD24 (LK = 1) a ← CIA NIA ← (a + EXTS(BD24||0b0))32:63 if LK=1 then LR ← CIA + 4 Let the BTEA be calculated as follows:

  • For e_b[l], let BTEA be the sum of the CIA and the sign-extended value of the BD24 instruction field concatenated with 0b0. The BTEA is the address of the next instruction to be executed. If LK = 1, the sum CIA+4 is placed into the LR. Special Registers Altered: LR (if LK = 1) se_b BD8 (LK = 0) se_bl BD8 (LK = 1) a ← CIA NIA ← (a + EXTS(BD8||0b0)) 32:63 if LK=1 then LR ← CIA + 2 Let the BTEA be calculated as follows:
  • For se_b[l], let BTEA be the sum of the CIA and the sign-extended value of the BD8 instruction field concatenated with 0b0. The BTEA is the address of the next instruction to be executed. If LK = 1, the sum CIA+2 is placed into the LR. Special Registers Altered: LR (if LK = 1) VLE User 0 567 30 31

0111100 B D 2 4 L K

1110100 L K B D 8

Branch conditional [and link] [absolute] bc BO,BI,BD (AA=0, LK=0) bca BO,BI,BD (AA=1, LK=0) bcl BO,BI,BD (AA=0, LK=1) bcla BO,BI,BD (AA=1, LK=1) if ¬BO2 then CTR32:63 ← CTR32:63 : 1 ctr_ok ← BO2 | ((CTR32:63 ≠ 0) ⊕ BO3) cond_ok ← BO0 | (CRBI+32 ≡ BO1) if ctr_ok & cond_ok then if AA=1 then a ← 640 else a ← CIA if E=0 then NIA ← 320 || (a + EXTS(BD||0b00))32:63 else NIA ← CIA + 4 if LK=1 then LR ← CIA + 4 The branch target effective address (BTEA) is calculated as follows:

  • For 32-bit implementations, BTEA is bits 32–63 of the sum of the current instruction address (CIA), or 32 zeros if AA=1, and the sign-extended value of the LI instruction field concatenated with 0b00 The BO instruction field specifies any conditions that must be met for the branch to be taken, as defined in Conditional branch control on page 167.” The sum BI+32 specifies the CR bit to be used. The BI field specifies the CR bit used as the condition of the branch, as shown in Table 200. Book E User 0 5 6 1 01 1 1 51 6 2 93 03 1

010000 B O B I B D A A L K

If the branch conditions are met, the BTEA is the address of the next instruction to be executed. If LK=1, the sum CIA + 4 is placed into the LR.

  • CTR(if BO2=0) LR(if LK=1)

Table 200. BI operand settings for CR fields CR0[0] 32 00000 Negative (LT)—Set when the result is negative. CR0[1] 33 00001 Positive (GT)—Set when t he result is positive (and not zero). CR0[2] 34 00010 Zero (EQ)—Set when the result is zero. CR1[0] 36 00100 Copy of FPSCR[FX] at the instruction’s completion. CR1[1] 37 00101 Copy of FPSCR[FEX] at the instruction’s completion. CR1[2] 38 00110 Copy of FPSCR[VX] at the instruction’s completion. CR1[3] 39 00111 Copy of FPSCR[OX] at the instruction’s completion. Less than or floating-point less than (LT, FL). For floating-point compare instructions:frA < frB. Greater than or floating-point greater than (GT, FG). For floating-point compare instructions:frA > frB. Equal or floating-point equal (EQ, FE). For integer compare instructions: rA = SIMM, UIMM, or rB. For floating-point compare instructions: frA = frB. Summary overflow or floating-point unordered (SO, FU). completion of the instruction.

_bcx _bcx Branch Conditional [and Link] e_bc BO32,BI32,BD15 (LK = 0) e_bcl BO32,BI32,BD15 (LK = 1) if BO320 then CTR32:63 ← CTR32:63 – 1 ctr_ok ← ¬BO320 | ((CTR32:63 ≠ 0) ⊕ BO321) cond_ok ← BO320 | (CRBI32+32 ≡ BO321) if ctr_ok & cond_ok then NIA ← (CIA + EXTS(BD15 || 0b0)) 32:63 else NIA ← CIA + 4 if LK=1 then LR ← CIA + 4 Let the BTEA be calculated as follows:

  • For e_bc[l], let BTEA be the sum of the CIA and the sign-extended value of the BD15 instruction field concatenated with 0b0. BO32 specifies any conditions that must be met for the branch to be taken, as defined in Chapter 12.2.2: Branch instructions on page 864.” The sum BI32+32 specifies the CR bit. Only CR[32–47] may be specified. If the branch conditions are met, the BTEA is the address of the next instruction to be executed. If LK = 1, the sum CIA + 4 is placed into the LR. Special Registers Altered: CTR (if BO32 0 =1 ) LR (if LK = 1) se_bc BO16,BI16,BD8 cond_ok ← (CRBI16+32 ≡ BO16) if cond_ok then NIA ← (CIA + EXTS(BD8 || 0b0)) 32:63 else NIA ← CIA + 2 Let the BTEA be calculated as follows:
  • For se_bc, BTEA is the sum of the CIA and the sign-extended value of the BD8 instruction field concatenated with 0b0. BO16 specifies any conditions that must be met for the branch to be taken, as defined in Chapter 12.2.2: Branch instructions.” The sum BI16+32 specifies CR bit; only CR[32–35] may be specified. If the branch conditions are met, the BTEA is the address of the next instruction to be executed. Special Registers Altered: None VLE User 0 5 6 9 10 11 12 15 16 30 31 011110 1 0 0 0 B O 3 2 B I 3 2 B D 1 5 L K 0 4 5 678 1 5

11100 B O 1 6B I 1 6 B D 8

Branch conditional to count register [and link] bcctr BO,BI (LK=0) bcctrl BO,BI (LK=1) cond_ok ← BO0 | (CRBI+32 ≡ BO1) if cond_ok & E=0 then NIA ← 320 || CTR32:61 || 0b00 if ¬cond_ok then NIA ← CIA + 4 if LK=1 then LR ← CIA + 4 The branch target effective address (BTEA) is calculated as follows:

  • For bcctr[l], BTEA is the contents of CTR[32–61] concatenated with 0b00. BO specifies conditions that must be met for the branch to be taken. BI+32 specifies the CR bit to be used; see Table 201. Book E User 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1

010011 B O B I / / / 1000010000 L K

If the condition is met, the BTEA is the address of the next instruction to be executed. If LK=1, the sum CIA + 4 is placed into the LR. If the decrement and test CTR option is specified (BO[2]=0), the instruction form is invalid. Table 201. BI operand settings for CR fields CR0[0] 32 00000 Negative (LT)—Set when the result is negative. CR0[1] 33 00001 Positive (GT)—Set when t he result is positive (and not zero). CR0[2] 34 00010 Zero (EQ)—Set when the result is zero. CR1[0] 36 00100 Copy of FPSCR[FX] at the instruction’s completion. CR1[1] 37 00101 Copy of FPSCR[FEX] at the instruction’s completion. CR1[2] 38 00110 Copy of FPSCR[VX] at the instruction’s completion. CR1[3] 39 00111 Copy of FPSCR[OX] at the instruction’s completion. Less than or floating-point less than (LT, FL). For floating-point compare instructions:frA < frB. Greater than or floating-point greater than (GT, FG). For floating-point compare instructions: frA > frB. Equal or floating-point equal (EQ, FE). For integer compare instructions: rA = SIMM, UIMM, or rB. For floating-point compare instructions: frA = frB. Summary overflow or floating-point unordered (SO, FU). completion of the instruction.

Branch conditional to link register [and link] bclr BO,BI (LK=0) bclrl BO,BI (LK=1) if ¬BO2 then CTR32:63 ← CTR32:63 - 1 ctr_ok ← BO2 | ((CTR32:63 ≠ 0) ⊕ BO3) cond_ok ← BO0 | (CRBI+32 ≡ BO1) if ctr_ok & cond_ok & E=0 then NIA ← 320 || LR32:61 || 0b00 if ¬(ctr_ok & cond_ok) then NIA ← CIA + 4 if LK=1 then LR ← CIA + 4 The branch target effective address (BTEA) is calculated as follows:

  • For bclr[l], BTEA is the contents of LR[32–61] concatenated with 0b00. The BO field specifies any conditions that must be met for the branch to be taken, as defined in Conditional branch control on page 167.” The sum BI+32 specifies the CR bit to be used. The BI field specifies the CR bit used as the condition of the branch, as shown in Table 202. Book E User 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1

010011 B O B I / / / 0000010000 L K

If the condition is met, the BTEA is the address of the next instruction to be executed. If LK=1, the sum CIA + 4 is placed into the LR.

  • CTR (if BO2=0) LR (if LK=1)

Table 202. BI operand settings for CR fields CR0[0] 32 00000 Negative (LT)—Set when the result is negative. CR0[1] 33 00001 Positive (GT)—Set when the result is positive (and not zero). CR0[2] 34 00010 Zero (EQ)—Set when the result is zero. CR1[0] 36 00100 Copy of FPSCR[FX] at the instruction’s completion. CR1[1] 37 00101 Copy of FPSCR[FEX] at the instruction’s completion. CR1[2] 38 00110 Copy of FPSCR[VX] at the instruction’s completion. CR1[3] 39 00111 Copy of FPSCR[OX] at the instruction’s completion. Less than or floating-point less than (LT, FL). For floating-point compare instructions: frA < frB. Greater than or floating-point greater than (GT, FG). For floating-point compare instructions: frA > frB. Equal or floating-point equal (EQ, FE). For integer compare instructions: rA = SIMM, UIMM, or rB. For floating-point compare instructions: frA = frB. Summary overflow or floating-point unordered (SO, FU). completion of the instruction.

_bclri _bclri Bit Clear Immediate se_bclri r X,UI5 a ← UI5 result32:63 ← GPR(RX) & b GPR(RX) ← result32:63 For se_bclri, the bit of GPR(rX) specified by the value of UI5 is cleared and all other bits in GPR(rX) remain unaffected. Special Registers Altered: None 0 5 6 7 11 12 15

0110000 U I 5 R X

_bctrx _bctrx Branch to Count Register [and Link] se_bctr (LK = 0) se_bctrl (LK = 1) NIA ← CTR32:62 || 0b0 if LK=1 then LR ← CIA + 2 Let the BTEA be calculated as follows:

  • For se_bctr[l], let BTEA be bits 32–62 of the contents of the CTR concatenated with 0b0. The BTEA is the address of the next instruction to be executed. If LK = 1, the sum CIA + 2 is placed into the LR. Special Registers Altered: LR (if LK = 1) 01 4 1 5

000000000000011 L K

_bgeni _bgeni Bit Generate Immediate se_bgeni r X,UI5 a ← UI5 GPR(RX) ← b For se_bgeni, a constant value consisting of a single ‘1’ bit surrounded by ‘0’s is generated and the value is placed into GPR(rX). The position of the ‘1’ bit is specified by the UI5 field. Special Registers Altered: None 0 5 6 7 11 12 15

0110001 U I 5 R X

_blrx _blrx Branch to Link Register [and Link] se_blr (LK = 0) se_blrl (LK = 1) NIA ← LR32:62 || 0b0 if LK=1 then LR ← CIA + 2 Let the BTEA be calculated as follows:

  • For se_blr[l], let BTEA be bits 32–62 of the contents of the LR concatenated with 0b0. The BTEA is the address of the next instruction to be executed. If LK = 1, the sum CIA + 2 is placed into the LR. Special Registers Altered: LR (if LK = 1) 01 4 1 5

000000000000010 L K

_bmaski _bmaski Bit Mask Generate Immediate se_bmaski r X,UI5 a ← UI5 if a = 0 then b ← 321 else b ← 32-a0 || a1 GPR(RX) ← b For se_bmaski, a constant value consisting of a mask of low-order ’1’ bits that is zero- extended to 32 bits is generated, and the value is placed into GPR(rX). The number of low- order ’1’ bits is specified by the UI5 field. If UI5 is 0b00000, a value of all ’1’s is generated Special Registers Altered: None 0 5 6 7 11 12 15

0010110 U I 5 R X

Table 203. Data samples and sizes

_bseti _bseti Bit Set Immediate se_bseti r X,UI5 a ← UI5 result32:63 ← GPR(RX) | b GPR(RX) ← result32:63 For se_bseti, the bit of GPR(rX) specified by the value of UI5 is set, and all other bits in GPR(rX) remain unaffected. Special Registers Altered: None 0 5 6 7 11 12 15

0110010 U I 5 R X

_btsti _btsti Bit Test Immediate se_btsti r X,UI5 a ← UI5 c ← GPR(RX) & b if c = 320 then d ← 0b001 else d ← 0b010 CR0:3 ← d || XERSO For se_btsti, the bit of GPR(rX) specified by the value of UI5 is tested for equality to ’1’. The result of the test is recorded in the CR. EQ is set if the tested bit is clear, LT is cleared, and GT is set to the inverse value of EQ. Special Registers Altered: CR[0–3] 0 5 6 7 11 12 15

0110011 U I 5 R X

Compare [immediate] cmp cr D,L,rA,rB cmpi cr D,L,rA,SIMM if L=0 then a ← EXTS(rA32:63) else a ← rA if ‘cmpi’ then b ← EXTS(SIMM) if ‘cmp’ & L=0 then b ← EXTS(rB32:63) if ‘cmp’ & L=1 then b ← rB if a < b then c ← 0b100 if a > b then c ← 0b010 if a = b then c ← 0b001 CR4×crD+32:4×crD+35 ← c || XERSO If cmp and L=0, the contents of rA[32–63] are compared with the contents of rB[32–63], treating the operands as signed integers. If cmpi and L=0, the contents of rA[32–63] are compared with the sign-extended value of the SIMM field, treating the operands as signed integers. The result of the comparison is placed into CR field crD. Other registers altered: CR field crD Book E User 0 5 6 8 9 1 01 1 1 51 6 2 02 1 3 03 1 011111 crD/ L rA rB 0000000000 / 05 6 8 9 1 0 1 1 1 5 1 6 3 1 001011 crD/ L rAS I M M

_cmp cmp Compare [Immediate] e_cmp16i r A,SI e_cmpi cr D32,rA,SCI8 a ← GPR(RA)32:63 if ‘e_cmpi’ then b ← SCI8(F ,SCL,UI8) if ‘e_cmp16i’ then b ← EXTS(SI0:4 || SI5:15) if a < b then c ← 0b100 if a > b then c ← 0b010 if a = b then c ← 0b001 if ‘e_cmpi’ then CR4×CRD32+32:4×CRD32+35 ← c || XERSO // only CR0-CR3 if ‘e_cmp16i’ then CR32:35 ← c || XERSO // only CR0 If e_cmpi, GPR(rA) contents are compared with the value of SCI8, treating operands as signed integers. If e_cmp16i, GPR(rA) contents are compared with the sign-extended value of the SI field, treating operands as signed integers. The result of the comparison is placed into CR field crD (crD32). For e_cmpi, only CR0– CR3 may be specified. For e_cmp16i, only CR0 may be specified. Special Registers Altered: CR field crD (crD32) (CR0 for e_cmp16i) se_cmp r X,rY se_cmpi r X,UI5 a ← GPR(RX)32:63 if ‘se_cmpi’ then b ← 270 || UI5 if ‘se_cmp’ then b ← GPR(RY)32:63 if a < b then c ← 0b100 if a > b then c ← 0b010 if a = b then c ← 0b001 CR0:3 ← c || XERSO If se_cmp, the contents of GPR(rX) are compared with the contents of GPR(rY), treating the operands as signed integers. The result of the comparison is placed into CR field 0. If se_cmpi, the contents of GPR(rX) are compared with the value of the zero-extended UI5 field, treating the operands as signed integers. The result of the comparison is placed into CR field 0. Special Registers Altered: CR[0–3] VLE User 0 5 6 1 01 1 1 51 6 2 02 1 3 1

011100 SI0:4 R A 1 0011 SI5:15

0 5 6 8 9 1 01 1 1 51 6 2 02 12 22 32 4 3 1 0001100 00 C R D 3 2 R A 10101F S C L U I 8 05 6 7 8 1 1 1 2 1 5

00001100 R Y R X

0010101 U I 5 R X

_cmph _cmph Compare Halfword [Immediate] e_cmph cr D,rA,rB a ← EXTS(GPR(RA)48:63) b ← EXTS(GPR(RB)48:63) if a < b then c ← 0b100 if a > b then c ← 0b010 if a = b then c ← 0b001 CR4×CRD+32:4×CRD+35 ← c || XERSO For e_cmph, the contents of the low-order 16 bits of GPR(rA) and GPR(rB) are compared, treating the operands as signed integers. The result of the comparison is placed into CR field CRD. Special Registers Altered: CR field CRD se_cmph r X,rY a ← EXTS(GPR(RX) 48:63) b ← EXTS(GPR(RY)48:63) if a < b then c ← 0b100 if a > b then c ← 0b010 if a = b then c ← 0b001 CR0:3 ← c || XERSO For se_cmph, the contents of the low-order 16 bits of GPR(rX) and GPR(rY) are compared, treating the operands as signed integers. The result of the comparison is placed into CR field 0. Special Registers Altered: CR[0–3] e_cmph16i r A,SI a ← EXTS(GPR(RA) 48:63) b ← EXTS(SI0:4 || SI5:15) if a < b then c ← 0b100 if a > b then c ← 0b010 if a = b then c ← 0b001 CR32:35 ← c || XERSO // only CR0 The contents of the lower 16-bits of GPR(rA) are sign-extended and compared with the sign-extended value of the SI field, treating the operands as signed integers. The result of the comparison is placed into CR0. Special Registers Altered: CR0 VLE User 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1

011111 C R D / R A R B 0000001110 /

00001110 R Y R X

0 5 6 1 01 1 1 51 6 2 02 1 3 1

011100 SI0:4 R A 1 0110 SI5:15

_cmphl _cmphl Compare Halfword Logical [Immediate] e_cmphl cr D,rA,rB a ← EXTZ(GPR(RA)48:63) b ← EXTZ(GPR(RB)48:63) if a < b then c ← 0b100 if a > b then c ← 0b010 if a = b then c ← 0b001 CR4×CRD+32:4×CRD+35 ← c || XERSO For e_cmphl, the contents of the low-order 16 bits of GPR(rA) and GPR(rB) are compared, treating the operands as unsigned integers. The result of the comparison is placed into CR field CRD. Special Registers Altered: CR field CRD se_cmphl r X,rY a ← GPR(RX) 48:63 b ← GPR(RY)48:63 if a < b then c ← 0b100 if a > b then c ← 0b010 if a = b then c ← 0b001 CR0:3 ← c || XERSO For se_cmphl, the contents of the low-order 16 bits of GPR(rX) and GPR(rY) are compared, treating the operands as unsigned integers. The result of the comparison is placed into CR field 0. Special Registers Altered: CR[0–3] e_cmphl16i r A,UI a ← 160 || GPR(RA)48:63) if a < b then c ← 0b100 if a > b then c ← 0b010 if a = b then c ← 0b001 CR32:35 ← c || XERSO // only CR0 The contents of the lower 16-bits of GPR(rA) are zero-extended and compared with the zero-extended value of the UI field, treating the operands as unsigned integers. The result of the comparison is placed into CR0. Special Registers Altered: CR0 VLE User 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1

011111 C R D / R A R B 0000101110 /

00001111 R Y R X

0 5 6 1 01 1 1 51 6 2 02 1 3 1

011100 UI0:4 R A 1 0111 UI5:15

Compare logical [immediate] cmpl cr D,L,rA,rB cmpli cr D,L,rA,UIMM if L=0 then a ← 320 || rA32:63 else a ← rA if ‘cmpli’ then b ← 480 || UIMM if ‘cmpl’ & L=0 then b ← 320 || rB32:63 if ‘cmpl’ & L=1 then b ← rB if a <u b then c ← 0b100 if a >u b then c ← 0b010 if a = b then c ← 0b001 CR4×crD+32:4×crD+35 ← c || XERSO If cmpl and L=0, the contents of rA[32–63] are compared with the contents of rB[32–63], treating the operands as unsigned integers. If cmpli and L=0, the contents of rA[32–63] are compared with the zero-extended value of the UIMM field, treating the operands as unsigned integers. The result of the comparison is placed into CR field crD. Other registers altered: CR field crD Book E User 0 5 6 8 9 1 01 1 1 51 6 2 02 1 3 1 011111 crD/ L rA rB 0000100000 / 05 6 8 9 1 0 1 1 1 5 1 6 3 1 001010 crD/ L rAU I M M

_cmpl _cmpl Compare Logical [Immediate] e_cmpl16i r A,UI e_cmpli cr D32,rA,SCI8 a ← GPR(RA)32:63 if ‘e_cmpli’ then b ← SCI8(F ,SCL,UI8) if ‘e_cmpl16i’ then b ← 160 || UI0:4 || UI5:15 if a <u b then c ← 0b100 if a >u b then c ← 0b010 if a = b then c ← 0b001 if ‘e_cmpli’ then CR4×CRD32+32:4×CRD32+35 ← c || XERSO // only CR0-CR3 if ‘e_cmp16i’ then CR32:35 ← c || XERSO // only CR0 If e_cmpi, the contents of bits 32–63 of GPR( rA) are compared with the value of SCI8, treating the operands as unsigned integers. L must be 0 for 32-bit implementations If e_cmpl16i, the contents of GPR(rA) are compared with the zero-extended value of the UI field, treating the operands as unsigned integers. The result of the comparison is placed into CR field CRD (CRD32). For e_cmpli, only CR0– CR3 may be specified. For e_cmpl16i, only CR0 may be specified. Special Registers Altered: CR field CRD (CRD32) (CR0 for e_cmpl16i) se_cmpl r X,rY se_cmpli r X,OIMM a ← GPR(RX) 32:63 if ‘se_cmpli’ then b ← 270 || OFFSET(OIM5) if ‘se_cmpl’ then b ← GPR(RY)32:63 if a <u b then c ← 0b100 if a >u b then c ← 0b010 if a = b then c ← 0b001 CR0:3 ← c || XERSO If se_cmpl, the contents of GPR(rX) are compared with the contents of GPR(rY), treating the operands as unsigned integers. The result of the comparison is placed into CR field 0. VLE User 0 5 6 1 01 1 1 51 6 2 02 1 3 1

011100 UI0:4 R A 1 0101 UI5:15

0 5 6 8 9 1 01 1 1 51 6 2 02 12 22 32 4 3 1 0001100 01 C R D 3 2 R A 10101F S C L U I 8 05 6 7 8 1 1 1 2 1 5

00001101 R Y R X

0010001 OIM5

(1) 1. OIMM = OIM5 +1 RX

If se_cmpli, the contents of GPR(rX) are compared with the value of the zero-extended offset value of the OIM5 field (a final value in the range 1–32), treating the operands as unsigned integers. The result of the comparison is placed into CR field 0. Special Registers Altered: CR[0–3]

Count leading zeros (word) cntlzw r A,rS( Z = 0 , R c = 0 ) cntlzw. r S( Z = 0 , R c = 1 ) if ‘cntlzd’ then n ← 0 else n ← 32 i ← 0 do while n < 64 if rS n = 1 then leave n ← n + 1 i ← i + 1 rA ← i if Rc=1 then do GT ← i > 0 EQ ← i = 0 For cntlzw[.], a count of the number of consecutive zero bits starting at rS[32] is placed into rA. This number ranges from 0 to 32, inclusive. If Rc=1, CR field 0 is set to reflect the result. Other registers altered: CR0 (if Rc=1) Book E User 0 5 6 1 01 1 1 51 6 2 02 1 2 42 52 6 3 03 1 011111 rS rA / / / 0000Z11010 R c

crand crb D,crbA,crbB CRcrbD+32 ← CRcrbA+32 & CRcrbB+32 The content of bit crbA+32 of CR is ANDed with the content of bit crbB+32 of CR, and the result is placed into bit crbD+32 of CR. Other registers altered: CR Book E User 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1 010011 crbD crbA crbB 0100000001 /

_crand _crand Condition Register AND e_crand crb D,crbA,crbB CRBT+32 ← CRBA+32 & CRBB+32 The content of bit CRBA+32 of the CR is ANDed with the content of bit CRBB+32 of the CR, and the result is placed into bit CRBD+32 of the CR. Special Registers Altered: CR Condition Register AND with Complement e_crandc crb D,crbA,crbB CR BT+32 ← CRBA+32 & ¬CRBB+32 The content of bit CRBA+32 of the CR is ANDed with the one’s complement of the content of bit CRBB+32 of the CR, and the result is placed into bit CRBD+32 of the CR. Special Registers Altered: CR CR Equivalent e_creqv crb D,crbA,crbB CR BT+32 ← CRBA+32 ≡ CRBB+32 The content of bit CRBA+32 of the CR is XORed with the content of bit CRBB+32 of the CR, and the one’s complement of result is placed into bit CRBD+32 of the CR. Special Registers Altered: CR VLE User 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1

011111 C R B D C R B A C R B B 0100000001 /

0 5 6 1 01 1 1 51 6 2 02 1 3 03 1

011111 C R B D C R B A C R B B 0010000001 /

0 5 6 1 01 1 1 51 6 2 02 1 3 03 1

011111 C R B D C R B A C R B B 0100100001 /

Condition register AND with complement crandc crb D,crbA,crbB CRcrbD+32 ← CRcrbA+32 & ¬CRcrbB+32 The content of bit crbA+32 of CR is ANDed with the one’s complement of the content of bit crbB+32 of CR, and the result is placed into bit crbD+32 of CR. Other registers altered: CR Book E User 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1 010011 crbD crbA crbB 0010000001 /

Condition register equivalent creqv crb D,crbA,crbB CRcrbD+32 ← CRcrbA+32 ≡ CRcrbB+32 The content of bit crbA + 32 of CR is XORed with the content of bit crbB + 32 of CR, and the one’s complement of result is placed into bit crbD+32 of CR. Other registers altered: CR Book E User 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1 010011 crbD crbA crbB 0100100001 /

crnand crb D,crbA,crbB CRcrbD+32 ← ¬(CRcrbA+32 & CRcrbB+32) The content of bit crbA+32 of CR is ANDed with the content of bit crbB+32 of CR, and the one’s complement of the result is placed into bit crbD+32 of CR. Other registers altered: CR Book E User 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1 010011 crbD crbA crbB 0011100001 /

_crnand _crnand Condition Register NAND e_crnand crb D,crbA,crbB CRBT+32 ← ¬(CRBA+32 & CRBB+32) The content of bit CRBA+32 of the CR is ANDed with the content of bit CRBB+32 of the CR, and the one’s complement of the result is placed into bit CRBD+32 of the CR. Special Registers Altered: CR VLE User 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1

011111 C R B D C R B A C R B B 0011100001 /

crnor crb D,crbA,crbB CRcrbD+32 ← ¬(CRcrbA+32 | CRcrbB+32) The content of bit crbA+32 of CR is ORed with the content of bit crbB+32 of CR, and the one’s complement of the result is placed into bit crbD+32 of CR. Other registers altered: CR Book E User 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1 010011 crbD crbA crbB 0000100001 /

_crnor _crnor Condition Register NOR e_crnor crb D,crbA,crbB CRBT+32 ← ¬(CRBA+32 | CRBB+32) The content of bit CRBA+32 of the CR is ORed with the content of bit CRBB+32 of the CR, and the one’s complement of the result is placed into bit CRBD+32 of the CR. Special Registers Altered: CR VLE User 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1

011111 C R B D C R B A C R B B 0000100001 /

cror crb D,crbA,crbB CRcrbD+32 ← CRcrbA+32 | CRcrbB+32 The content of bit crbA+32 of CR is ORed with the content of bit crbB+32 of CR, and the result is placed into bit crbD+32 of CR. Other registers altered: CR Book E User 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1 010011 crbD crbA crbB 0111000001 /

_cror _cror Condition Register OR e_cror crb D,crbA,crbB CRBT+32 ← CRBA+32 | CRBB+32 The content of bit CRBA+32 of the CR is ORed with the content of bit CRBB+32 of the CR, and the result is placed into bit CRBD+32 of the CR. Special Registers Altered: CR VLE User 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1

011111 C R B D C R B A C R B B 0111000001 /

Condition register OR with complement crorc crb D,crbA,crbB CRcrbD+32 ← CRcrbA+32 | ¬CRcrbB+32 The content of bit crbA+32 of CR is ORed with the one’s complement of the content of bit crbB+32 of CR, and the result is placed into bit crbD+32 of CR. Other registers altered: CR Book E User 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1 010011 crbD crbA crbB 0110100001 /

_crorc _crorc Condition Register OR with Complement e_crorc crb D,crbA,crbB CRBT+32 ← CRBA+32 | ¬CRBB+32 The content of bit CRBA+32 of the CR is ORed with the one’s complement of the content of bit CRBB+32 of the CR, and the result is placed into bit CRBD+32 of the CR. Special Registers Altered: CR VLE User 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1

011111 C R B D C R B A C R B B 0110100001 /

crxor crb D,crbA,crbB CRcrbD+32 ← CRcrbA+32 ⊕ CRcrbB+32 The content of bit crbA+32 of CR is XORed with the content of bit crbB+32 of CR, and the result is placed into bit crbD+32 of CR. Other registers altered: CR Book E User 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1 010011 crbD crbA crbB 0011000001 /

_crxor _crxor Condition Register XOR e_crxor crb D,crbA,crbB CRcrbD+32 ← CRBA+32 ⊕ CRBB+32 The content of bit CRBA+32 of the CR is XORed with the content of bit CRBB+32 of the CR, and the result is placed into bit CRBD+32 of the CR. Special Registers Altered: CR VLE User 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1

011111 C R B D C R B A C R B B 0011000001 /

dcba r A,rB if rA=0 then a ← 640 else a ← rA AllocateDataCacheBlock(EA) EA calculation: Addressing ModeEA for rA=0EA for rA≠0 320 || rB32:63 dcba is a hint that performance would likely improve if the block containing the byte addressed by EA is established in the data cache without fetching the block from main memory, because the program is likely to soon store into a portion of the block and the contents of the rest of the block are not meaningful to the program. If the hint is honored, the contents of the block are undefined when the instruction completes. The hint is ignored if the block is caching-inhibited. If the block containing the byte addressed by EA is in memory that is memory-coherence required and the block exists in a data cache of any other processors, it is kept coherent in those caches. This instruction is treated as a storeexcept that an interrupt is not taken for a translation or protection violation. This instruction may establish a block in the data cache without verifying that the associated real address is valid. This can cause a delayed machine check interrupt. Other registers altered: None Book E User 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1 011111 / / / rA rB 1011110110 /

dcbf r A,rB if rA=0 then a ← 640 else a ← rA FlushDataCacheBlock( EA ) EA calculation: Addressing ModeEA for rA=0EA for rA≠0 320 || rB32:63 If the block containing the byte addressed by EA is in memory that is memory-coherence required, a block containing the byte addressed by EA is in the data cache of any processor, and any locations in the block are considered to be modified there, then those locations are written to main memory. Additional locations in the block may also be written to main memory. The block is invalidated in the data caches of all processors. If the block containing the byte addressed by EA is in memory that is not memory-coherence required, a block containing the byte addressed by EA is in the data cache of this processor and any locations in the block are considered to be modified there, then those locations are written to main memory. Additional locations in the block may also be written to main memory. The block is invalidated in the data cache of this processor. On some implementations, HID1[ABE] must be set to allow management of external L2 caches (for implementations with L2 caches) as well as other L1 caches in the system. The function of this instruction is independent of whether the block containing the byte addressed by EA is in memory that is write-through required or caching-inhibited. This instruction is treated as a load. See Cache management instructions on page 216.” Other registers altered: None Book E User 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1 011111 / / / rA rB 0001010110 /

Data cache block invalidate dcbi r A,rB if rA=0 then a ← 640 else a ← rA InvalidateDataCacheBlock( EA ) EA calculation: Addressing ModeEA for rA=0EA for rA≠0 320 || rB32:63 If the block containing the byte addressed by EA is in is coherence-required memory and any block containing the addressed byte is any processors’ data cache is invalidated in those caches. On some implementations, before the block is invalidated, if any locations in the block are considered to be modified in any such data cache, those locations are written to main memory and additional locations in the block may be written to main memory. If the block containing the byte addressed by EA is not coherence-required memory and a block containing the byte addressed by EA is in the data cache of this processor, then the block is invalidated in that data cache. On some implementations, before the block is invalidated, any locations in the block considered modified in that data cache are written to main memory; additional locations in the block may be written to main memory. dcbi is treated as a store on implementations that invalidate a block without first writing to main memory all locations in the block that are considered to be modified in the data cache, except that the invalidation is not ordered by mbar. On other implementations this instruction is treated as a load. Additional information about this instruction is as follows.

  • The data cache block size for dcbi is the same as for dcbf.
  • If a processor holds a reservation and some other processor executes a dcbi to the same reservation granule, whether the reservation is lost is undefined. Other registers altered: None Book E User 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1 011111 / / / rA rB 0111010110 /

Data cache block lock clear dcblc CT,rA,rB Form: X if rA = 0 then a ← 640 else a ← GPR(rA) if Mode32 then EA ← 320 || (a + GPR(rB))32:63 if Mode64 then EA ← a + GPR(rB) DataCacheBlockClearLock(CT, EA) EA calculation: EA for rA=0EA for rA≠0 320 || GPR(rB)32:63 320 || (GPR(rA)+GPR(rB))32:63 The data cache specified by CT has the cache line corresponding to EA unlocked allowing the line to participate in the normal replacement policy. Cache lock clear instructions remove locks previously set by cache lock set instructions. User-level cache instructions on page 180,” lists supported CT values. An implementation may use other CT values to enable software to target specific, implementation-dependent portions of its cache hierarchy or structure. The instruction is treated as a load with respect to translation and memory protection and can cause DSI and DTLB error interrupts accordingly. An unable-to-unlock condition is said to occur any of the following conditions exist:

  • The target address is marked cache-inhibited, or the storage attributes of the address uses a coherency protocol that does not support locking.
  • The target cache is disabled or not present.
  • The CT field of the instructions contains a value not supported by the implementation.
  • The target address is not in the cache or is present in the cache but is not locked. If an unable-to-unlock condition occurs, no cache operation is performed. EIS Specifics Clearing and then setting L1CSR0[CLFR] allows system software to clear all L1 data cache locking bits without knowing the addresses of the lines locked. Cache locking APU User 0 5 61 0 1 11 5 1 62 0 2 1 3 0 3 1

011111 C T rA rB 0110000110 /

dcbst r A,rB if rA=0 then a ← 640 else a ← rA StoreDataCacheBlock( EA ) EA calculation: Addressing ModeEA for rA=0EA for rA≠0 320 || rB32:63 If the block containing the byte addressed by EA is in memory that is memory-coherence required and a block containing the byte addressed by EA is in the data cache of any processor, and any locations in the block are considered to be modified there, those locations are written to main memory. Additional locations in the block may be written to main memory. The block ceases to be considered to be modified in that data cache. If the block containing the byte addressed by EA is in memory that is not memory-coherence required and a block containing the byte addressed by EA is in the data cache of this processor and any locations in the block are considered to be modified there, those locations are written to main memory. Additional locations in the block may be written to main memory. The block ceases to be considered to be modified in that cache. The function of this instruction is independent of whether the block containing the byte addressed by EA is in memory that is write-through required or caching-inhibited. This instruction is treated as a load. On some implementations, HID1[ABE] must be set to allow management of external L2 caches (for implementations with L2 caches) as well as other L1 caches in the system. Other registers altered: None Book E User 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1 011111 / / / rA rB 0000110110 /

dcbt CT,rA,rB if rA=0 then a ← 640 else a ← rA PrefetchDataCacheBlock( CT, EA ) EA calculation: Addressing ModeEA for rA=0EA for rA≠0 320 || rB32:63 User-level cache instructions on page 180,” lists supported CT values. An implementation may use other CT values to enable software to target specific, implementation-dependent portions of its cache hierarchy or structure. Implementations should perform no operation when CT specifies a value not supported by the implementation. The hint is ignored if the block is caching-inhibited. This instruction is treated as a load except that an interrupt is not taken for a translation or protection violation. Other registers altered: None Book E User 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1

011111 C T rA rB 0100010110 /

Data cache block touch and lock set dcbtls CT,rA,rB Form: X if rA = 0 then a ← 640 else a ← GPR(rA) if Mode32 then EA ← 320 || (a + GPR(rB))32:63 if Mode64 then EA ← a + GPR(rB) PrefetchDataCacheBlockLockSet(CT, EA) EA calculation: EA for rA=0EA for rA≠0 320 || GPR(rB)32:63 320 || (GPR(rA)+GPR(rB))32:63 The data cache specified by CT has the cache line corresponding to EA loaded and locked into the cache. If the line already exists in the cache, it is locked without being refetched. Cache touch and lock set instructions let software lock cache lines into the cache to provide lower latency for critical cache accesses and more deterministic behavior. Locked lines do not participate in the normal replacement policy when a line must be victimized for replacement. User-level cache instructions on page 180,” lists supported CT values. An implementation may use other CT values to enable software to target specific, implementation-dependent portions of its cache hierarchy or structure. The instruction is treated as a load with respect to translation and memory protection and can cause DSI and DTLB error interrupts accordingly. An unable to lock condition is said to occur any of the following conditions exist:

  • The target address is marked cache-inhibited, or the storage attributes of the address uses a coherency protocol that does not support locking.
  • The target cache is disabled or not present.
  • The CT field of the instructions contains a value not supported by the implementation. If an unable to lock condition occurs, no cache operation is performed and LICSR0[DCUL] is set appropriately. Overlocking is said to exist is all available ways for a given cache index are already locked. If overlocking occurs for dcbtls and if the lock was targeted for the primary cache (CT = 0), the requested line is not locked into the cache. When overlock occurs, L1CSR1[DCLO] is set. If L1CSR1[DCLOA] is set, the requested line is locked into the cache and implementation dependent line currently locked in the cache is evicted. The results of overlocking and unable to lock conditions for caches other than the primary cache and secondary cache are defined as part of the architecture for the specific cache hierarchy designated by CT. Other registers altered:
  • L1CSR0[DCUL] if unable to lock occurs
  • L1CSR0[DCLO] (L2CSR[L2CLO]) if lock overflow occurs Cache locking APU User 0 5 61 0 1 11 5 1 62 0 2 1 3 0 3 1

011111 C T rA rB 0010100110 /

Data cache block touch for store dcbtst CT,rA,rB if rA=0 then a ← 640 else a ← rA PrefetchForstoreDataCacheBlock( CT, EA ) EA calculation: Addressing ModeEA for rA=0EA for rA≠0 320 || rB32:63 If CT=0, this instruction is a hint that performance would likely be improved if the block containing the byte addressed by EA is fetched into the data cache, because the program will probably soon store into the addressed byte. User-level cache instructions on page 180,” lists supported CT values. An implementation may use other CT values to enable software to target specific, implementation-dependent portions of its cache hierarchy or structure. Implementations should perform no operation when CT specifies a value not supported by the implementation. The hint is ignored if the block is caching-inhibited. This instruction is treated as a load , except that an interrupt is not taken for a translation or protection violation. Other registers altered: None Book E User 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1

011111 C T rA rB 0011110110 /

Data cache block touch for store and lock set dcbtstls CT,rA,rB Form: X if rA = 0 then a ← 640 else a ← GPR(rA) if Mode32 then EA ← 320 || (a + GPR(rB))32:63 if Mode64 then EA ← a + GPR(rB) PrefetchDataCacheBlockLockSet(CT, EA) EA calculation: EA for rA=0EA for rA≠0 320 || GPR(rB)32:63 320 || (GPR(rA)+GPR(rB))32:63 The data cache specified by CT has the cache line corresponding to EA loaded and locked into the cache. If the line already exists in the cache, it is locked without refetching from memory. Cache touch and lock set instructions allow software to lock lines into the cache to shorten latency for critical cache accesses and more deterministic behavior. Lines locked in the cache do not participate in the normal replacement policy when a line must be victimized for replacement. User-level cache instructions on page 180,” lists supported CT values. An implementation may use other CT values to enable software to target specific, implementation-dependent portions of its cache hierarchy or structure. Table 114 describes how this instruction is treated with respect to translation and memory protection. For unable-to-lock conditions, described in Unable-to-lock conditions on page 849,” no cache operation is performed and LICSR0[DCUL] is set. Overlocking occurs when all available ways for a given cache index are already locked. If an overlocking condition occurs for a dcbtstls instruction and if the lock was targeted for the primary cache or secondary cache (CT = 0 or CT = 2), the requested line is not locked into the cache. When overlock occurs, L1CSR1[DCLO] (L2CSR[L2CLO] for CT = 2) is set. If L1CSR1[DCLOA] is set (or L2CSR[L2CLOA] for CT = 2), the requested line is locked into the cache and implementation dependent line currently locked in the cache is evicted. If system software wants to precisely determine if an overlock event has occurred in the L1 data cache, it must perform the following code sequence: dcbtstls msync mfspr (L1CSR0) (check L1CSR0[DCUL] bit for data cache index unable-to-lock condition) (check L1CSR0[DCLO] bit for data cache index overlock condition) Results of overlocking and unable-to-lock conditions for caches other than the primary and secondary cache are defined as part of the architecture for the cache hierarchy designated by CT. Other registers altered:

  • L1CSR0[DCUL] if unable to lock occurs
  • L1CSR0[DCLO] (L2CSR[L2CLO]) if lock overflow occurs Cache locking APU User 0 5 61 0 1 11 5 1 62 0 2 1 3 0 3 1

011111 C T rA rB 0010000110 /

EIS specifics: Clearing and then setting L1CSR0[CLFR] allows system software to clear all data cache locking bits without knowing the addresses of the lines locked.

Data cache block set to zero dcbz r A,rB if rA=0 then a ← 640 else a ← rA ZeroDataCacheBlock( EA ) EA calculation: Addressing ModeEA for rA=0EA for rA≠0 320 || rB32:63 If the block containing the addressed byte is in the data cache, all bytes of the block are cleared. If the block containing the byte addressed by EA is not in the data cache and is in memory that is not caching-inhibited, the block is established in the data cache without fetching the block from main memory, and all bytes of the block are cleared. If the block containing the byte addressed by EA is not in the data cache and is in storage that is not caching inhibited and cannot be established in the cache, then one of the following occurs:

  • All bytes of the area of main storage that corresponds to the addressed block are set to zero
  • An alignment interrupt is taken If the block containing the byte addressed by EA is in storage that is caching inhibited or write through required, one of the following occurs:
  • All bytes of the area of main storage that corresponds to the addressed block are set to zero
  • An alignment interrupt is taken. If the block containing the byte addressed by EA is in memory-coherence required memory and the block exists in any other processors’ data cache, it is kept coherent in those caches. dcbz may establish a block in the data cache without verifying that the associated real address is valid. This can cause a delayed machine check interrupt. dcbz is treated as a store.
  • On some implementations, HID1[ABE] must be set to allow management of external L2 caches (for implementations with L2 caches) as well as other L1 caches in the system.
  • dcbz may cause a cache-locking exception on some implementations. See the user documentation. Other registers altered: None Programming note: If the block containing the byte addressed by EA is in memory that is caching-inhibited or write-through required, the alignment interrupt handler should clear all bytes of the area of main memory that corresponds to the addressed block. Book E User 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1 011111 / / / rA rB 1111110110 /

divw r D,rA,rB( O E = 0 , R c = 0 ) divw. r D,rA,rB( O E = 0 , R c = 1 ) divwo r D,rA,rB( O E = 1 , R c = 0 ) divwo. r D,rA,rB( O E = 1 , R c = 1 ) dividend0:31 ← rA32:63 divisor0:31 ← rB32:63 quotient0:31 ← dividend ÷ divisor if OE=1 then do OV ← ( (rA SO ← SO | OV if Rc=1 then do LT ← quotient < 0 GT ← quotient > 0 EQ ← quotient = 0 rD32:63 ← quotient rD0:31 ← undefined The 32-bit quotient of the contents of rA[32–63] divided by the contents of rB[32–63] is placed into rD[32–63]. rD[0–31] are undefined. The remainder is not supplied as a result. Both operands and the quotient are interpreted as signed integers. The quotient is the unique signed integer that satisfies the following: dividend = (quotient × divisor) + r Here, 0 ≤ r < |divisor| if the dividend is nonnegative and –|divisor| < r ≤ 0 if it is negative. If any of the following divisions is attempted, the contents of rD are undefined as are (if Rc=1) the contents of the CR0[LT,GT,EQ]. In these cases, if OE=1, OV is set. 0x8000_0000 ÷ –1 <anything> ÷ 0 Other registers altered:

  • CR0 (if Rc=1) SO OV (if OE=1) Book E User 0 5 6 1 01 1 1 51 6 2 02 12 2 3 03 1 011111 rD rA rB O E 111101011 R c

divwu r D,rA,rB( O E = 0 , R c = 0 ) divwu. r D,rA,rB( O E = 0 , R c = 1 ) divwuo r D,rA,rB( O E = 1 , R c = 0 ) divwuo. r D,rA,rB( O E = 1 , R c = 1 ) dividend0:31 ← rA32:63 divisor0:31 ← rB32:63 quotient0:31 ← dividend ÷ divisor if OE=1 then do OV ← (rB 32:63=0) SO ← SO | OV if Rc=1 then do LT ← quotient < 0 GT ← quotient > 0 EQ ← quotient = 0 rD32:63 ← quotient rD0:31 ← undefined The 32-bit quotient of the contents of rA[32–63] divided by the contents of rB[32–63] is placed into rD[32–63]. rD[0–31] are undefined. The remainder is not supplied as a result. Both operands and the quotient are interpreted as unsigned integers, except that if Rc=1 the first three bits of CR field 0 are set by signed comparison of the result to zero. The quotient is the unique unsigned integer that satisfies the following: dividend = (quotient × divisor) + r Here, 0 ≤ r < divisor. If an attempt is made to perform the following division, the contents of rD are undefined as are (if Rc=1) the contents of the LT, GT, and EQ bits of CR0. In this case, if OE=1 OV is set. <anything> ÷ 0 Other registers altered:

  • CR0 (if Rc=1)
  • SO OV (if OE=1) Book E User 0 5 6 1 01 1 1 51 6 2 02 12 2 3 03 1 011111 rD rA rB O E 111001011 R c

Floating-point double-precision absolute value efdabs r D,rA rD0:63 ← 0b0 || rA1:63 The sign bit of rA is set to 0 and the result is placed into rD. Exceptions: Exception detection for embedded floating-point absolute value operations is implementation dependent. An implementation may choose to not detect exceptions and carry out the sign bit operation. If the implementation does not detect exceptions, or if exception detection is disabled, the computation can be carried out in one of two ways, as a sign bit operation ignoring the rest of the contents of the source register, or by examining the input and appropriately saturating the input prior to performing the operation. If an implementation chooses to handle exceptions, the exception is handled as follows: If rA is Infinity, Denorm, or NaN, SPEFSCR[FINV] is set, and FG and FX are cleared. If floating- point invalid input exceptions are enabled, an interrupt is taken and the destination register is not updated. Scalar DPFP APU User 0 5 6 1 01 1 1 51 6 2 02 1 3 1 000100 rD rA 0 00000 1011100100

Floating-point double-precision add efdadd r D,rA,rB rD0:63 ← rA0:63 +dp rB0:63 rA is added to rB and the result is stored in rD. If rA is NaN or infinity, the result is either pmax (asign==0), or nmax (asign==1). Otherwise, If rB is NaN or infinity, the result is either pmax (bsign==0), or nmax (bsign==1). Otherwise, if an overflow occurs, pmax or nmax (as appropriate) is stored in rD. If an underflow occurs, +0 (for rounding modes RN, RZ, RP) or - 0 (for rounding mode RM) is stored in rD. Exceptions: If the contents of rA or rB are Infinity, Denorm, or NaN, SPEFSCR[FINV] is set. If SPEFSCR[FINVE] is set, an interrupt is taken, and the destination register is not updated. Otherwise, if an overflow occurs, SPEFSCR[FOVF] is set, or if an underflow occurs, SPEFSCR[FUNF] is set. If either underflow or overflow exceptions are enabled and the corresponding bit is set, an interrupt is taken. If any of these interrupts are taken, the destination register is not updated. If the result of this instruction is inexact or if an overflow occurs but overflow exceptions are disabled, and no other interrupt is taken, SPEFSCR[FINXS] is set. If the floating-point inexact exception is enabled, an interrupt is taken using the floating-point round interrupt vector. In this case, the destination register is updated with the truncated result, the FG and FX bits are properly updated to allow rounding to be performed in the interrupt handler. FG and FX are cleared if an overflow, underflow, or invalid operation/input error is signaled, regardless of enabled exceptions. Scalar DPFP APU User 0 5 6 1 01 1 1 51 6 2 02 1 3 1 000100 rD rA rB 0 1011100000

Floating-point double-precision convert from single-precision efdcfs rD,rB FP32format f; FP64format result; f ← rB32:63 if (fexp = 0) & (ffrac = 0)) then result ← fsign || 630 // signed zero value else if Isa32NaNorInfinity(f) | Isa32Denorm(f) then SPEFSCRFINV ← 1 result ← fsign || 0b11111111110 || 521/ / m a x v a l u e else if Isa32Denorm(f) then SPEFSCRFINV ← 1 result ← fsign || 630 else resultsign ← fsign resultexp ← fexp - 127 + 1023 resultfrac ← ffrac || 290 rD0:63 = result The single-precision floating-point value in the low element of rB is converted to a double- precision floating-point value and the result is placed into rD. The rounding mode is not used since this conversion is always exact. Exceptions: If the low element of rB is Infinity, Denorm, or NaN, SPEFSCR[FINV] is set. If SPEFSCR[FINVE] is set, an interrupt is taken, and the destination register is not updated. FG and FX are always cleared. Note: Architecture Note: This instruction is optional if neither the embedded scalar single- precision floating-point APU or the embedded vector single-precision floating-point APU are implemented. Scalar DPFP APU User 0 5 6 1 01 1 1 51 6 2 02 1 3 1 000100 rD 00000 rB 0 1011101111

Convert floating-point double-precision from signed fraction efdcfsf rD,rB rD0:63 ← CnvtI32ToFP64(rB32:63, SIGN, F) The signed fractional low element in rB is converted to a double-precision floating-point value using the current rounding mode and the result is placed into rD. Exceptions: None. Scalar DPFP APU User 0 5 6 1 01 1 1 51 6 2 02 1 3 1 000100 rD 00000 rB 0 1011110011

Convert floating-point double-precision from signed integer efdcfsi rD,rB rD0:63 ← CnvtSI32ToFP64(rB32:63, SIGN, I) The signed integer low element in rB is converted to a double-precision floating-point value using the current rounding mode and the result is placed into rD. Exceptions: None. Scalar DPFP APU User 0 5 6 1 01 1 1 51 6 2 02 1 3 1 000100 rD 00000 rB 0 1011110001

Convert floating-point double-precision from signed integer doubleword efdcfsid rD,rB rD0:63 ← CnvtI64ToFP64(rB0:63, SIGN) The signed integer doubleword in rB is converted to a double-precision floating-point value using the current rounding mode and the result is placed into rD. Exceptions: This instruction can signal an inexact status and set SPEFSCR[FINXS] if the conversion is not exact. If the floating-point inexact exception is enabled, an interrupt is taken using the floating-point round interrupt vector. In this case, the destination register is updated with the truncated result, the FG and FX bits are properly updated to allow rounding to be performed in the interrupt handler. This instruction may only be implemented for 64-bit implementations. Scalar DPFP APU User 0 5 6 1 01 1 1 51 6 2 02 1 3 1 000100 rD 00000 rB 0 1011100011

Convert floating-point double-precision from unsigned fraction efdcfuf rD,rB rD0:63 ← CnvtI32ToFP64(rB32:63, UNSIGN, F) The unsigned fractional low element in rB is converted to a double-precision floating-point value using the current rounding mode and the result is placed into rD. Exceptions: None. Scalar DPFP APU User 0 5 6 1 01 1 1 51 6 2 02 1 3 1 000100 rD 00000 rB 0 1011110010

Convert floating-point double-precision from unsigned integer efdcfui rD,rB rD0:63 ← CnvtSI32ToFP64(rB32:63, UNSIGN, I) The unsigned integer low element in rB is converted to a double-precision floating-point value using the current rounding mode and the result is placed into rD. Exceptions: None. Scalar DPFP APU User 0 5 6 1 01 1 1 51 6 2 02 1 3 1 000100 rD 00000 rB 0 1011110000

Convert floating-point double-precision from unsigned integer doubleword efdcfuid rD,rB rD0:63 ← CnvtI64ToFP64(rB0:63, UNSIGN) The unsigned integer doubleword in rB is converted to a double-precision floating-point value using the current rounding mode and the result is placed into rD. Exceptions: This instruction can signal an inexact status and set SPEFSCR[FINXS] if the conversion is not exact. If the floating-point inexact exception is enabled, an interrupt is taken using the floating-point round interrupt vector. In this case, the destination register is updated with the truncated result, the FG and FX bits are properly updated to allow rounding to be performed in the interrupt handler. This instruction may only be implemented for 64-bit implementations. Scalar DPFP APU User 0 5 6 1 01 1 1 51 6 2 02 1 3 1 000100 rD 00000 rB 0 1011100010

Floating-point double-precision compare equal efdcmpeq crf D,rA,rB al ← rA0:63 bl ← rB0:63 if (al = bl) then cl ← 1 else cl ← 0 CR4*crD:4*crD+3 ← undefined || cl || undefined || undefined rA is compared against rB. If rA is equal to rB, the bit in the crfD is set, otherwise it is cleared. Comparison ignores the sign of 0 (+0 = -0). Exceptions: If the contents of rA or rB are Infinity, Denorm, or NaN, SPEFSCR[FINV] is set, and the FGH FXH, FG and FX bits are cleared. If floating-point invalid input exceptions are enabled, an interrupt is taken and the condition register is not updated. Otherwise, the comparison proceeds after treating NaNs, Infinities, and Denorms as normalized numbers, using their values of ‘e’ and ‘f’ directly. Scalar DPFP APU User 0 5 6 8 9 1 01 1 1 51 6 2 02 1 3 1 000100 crfD0 0 rA rB 0 1011101110

Floating-point double-precision compare greater than efdcmpgt crf D,rA,rB al ← rA0:63 bl ← rB0:63 if (al > bl) then cl ← 1 else cl ← 0 CR4*crD:4*crD+3 ← undefined || cl || undefined || undefined rA is compared against rB. If rA is greater than rB, the bit in the crfD is set, otherwise it is cleared. Comparison ignores the sign of 0 (+0 = -0). Exceptions: If the contents of rA or rB are Infinity, Denorm, or NaN, SPEFSCR[FINV] is set, and the FGH FXH, FG and FX bits are cleared. If floating-point invalid input exceptions are enabled, an interrupt is taken and the condition register is not updated. Otherwise, the comparison proceeds after treating NaNs, Infinities, and Denorms as normalized numbers, using their values of ‘e’ and ‘f’ directly. Scalar DPFP APU User 0 5 6 8 9 1 01 1 1 51 6 2 02 1 3 1 000100 crfD0 0 rA rB 0 1011101100

Floating-point double-precision compare less than efdcmplt crf D,rA,rB al ← rA0:63 bl ← rB0:63 if (al < bl) then cl ← 1 else cl ← 0 CR4*crD:4*crD+3 ← undefined || cl || undefined || undefined rA is compared against rB. If rA is less than rB, the bit in the crfD is set, otherwise it is cleared. Comparison ignores the sign of 0 (+0 = -0). Exceptions: If the contents of rA or rB are Infinity, Denorm, or NaN, SPEFSCR[FINV] is set, and the FGH FXH, FG and FX bits are cleared. If floating-point invalid input exceptions are enabled, an interrupt is taken and the condition register is not updated. Otherwise, the comparison proceeds after treating NaNs, Infinities, and Denorms as normalized numbers, using their values of ‘e’ and ‘f’ directly. 0 5 6 8 9 1 01 1 1 51 6 2 02 1 3 1 000100 crfD0 0 rA rB 0 1011101101

Convert floating-point double-precision to signed fraction efdctsf rD,rB rD32:63 ← CnvtFP64ToI32Sat(rB0:63, SIGN, ROUND, F) The double-precision floating-point value in rB is converted to a signed fraction using the current rounding mode and the result is saturated if it cannot be represented in a 32-bit fraction. NaNs are converted as though they were zero. Exceptions: If the contents of rB are Infinity, Denorm, or NaN, or if an overflow occurs, SPEFSCR[FINV] is set, and the FG, and FX bits are cleared. If SPEFSCR[FINVE] is set, an interrupt is taken, and the destination register is not updated. This instruction can signal an inexact status and set SPEFSCR[FINXS] if the conversion is not exact. If the floating-point inexact exception is enabled, an interrupt is taken using the floating-point round interrupt vector. In this case, the destination register is updated with the truncated result, the FG and FX bits are properly updated to allow rounding to be performed in the interrupt handler. Scalar DPFP APU User 0 5 6 1 01 1 1 51 6 2 02 1 3 1 000100 rD 00000 rB 0 1011110111

Convert floating-point double-precision to signed integer efdctsi rD,rB rD32:63 ← CnvtFP64ToI32Sat(rB0:63, SIGN, ROUND, I) The double-precision floating-point value in rB is converted to a signed integer using the current rounding mode and the result is saturated if it cannot be represented in a 32-bit integer. NaNs are converted as though they were zero. Exceptions: If the contents of rB are Infinity, Denorm, or NaN, or if an overflow occurs, SPEFSCR[FINV] is set, and the FG, and FX bits are cleared. If SPEFSCR[FINVE] is set, an interrupt is taken, the destination register is not updated, and no other status bits are set. This instruction can signal an inexact status and set SPEFSCR[FINXS] if the conversion is not exact. If the floating-point inexact exception is enabled, an interrupt is taken using the floating-point round interrupt vector. In this case, the destination register is updated with the truncated result, the FG and FX bits are properly updated to allow rounding to be performed in the interrupt handler. Scalar DPFP APU User 0 5 6 1 01 1 1 51 6 2 02 1 3 1 000100 rD 00000 rB 0 1011110101

Convert floating-point double-precision to signed integer doubleword with round toward zero efdctsidz rD,rB rD0:63 ← CnvtFP64ToI64Sat(rB0:63, SIGN, TRUNC) The double-precision floating-point value in rB is converted to a signed integer doubleword using the rounding mode Round toward Zero and the result is saturated if it cannot be represented in a 64-bit integer. NaNs are converted as though they were zero. Exceptions: If the contents of rB are Infinity, Denorm, or NaN, or if an overflow occurs, SPEFSCR[FINV] is set, and the FG, and FX bits are cleared. If SPEFSCR[FINVE] is set, an interrupt is taken, the destination register is not updated, and no other status bits are set. This instruction can signal an inexact status and set SPEFSCR[FINXS] if the conversion is not exact. If the floating-point inexact exception is enabled, an interrupt is taken using the floating-point round interrupt vector. In this case, the destination register is updated with the truncated result, the FG and FX bits are properly updated to allow rounding to be performed in the interrupt handler. This instruction may only be implemented for 64-bit implementations. Scalar DPFP APU User 0 5 6 1 01 1 1 51 6 2 02 1 3 1 000100 rD 00000 rB 0 1011101011

Convert floating-point double-precision to signed integer with round toward zero efdctsiz rD,rB rD32:63 ← CnvtFP64ToI32Sat(rB0:63, SIGN, TRUNC, I The double-precision floating-point value in rB is converted to a signed integer using the rounding mode Round toward Zero and the result is saturated if it cannot be represented in a 32-bit integer. NaNs are converted as though they were zero. Exceptions: If the contents of rB are Infinity, Denorm, or NaN, or if an overflow occurs, SPEFSCR[FINV] is set, and the FG, and FX bits are cleared. If SPEFSCR[FINVE] is set, an interrupt is taken, the destination register is not updated, and no other status bits are set. This instruction can signal an inexact status and set SPEFSCR[FINXS] if the conversion is not exact. If the floating-point inexact exception is enabled, an interrupt is taken using the floating-point round interrupt vector. In this case, the destination register is updated with the truncated result, the FG and FX bits are properly updated to allow rounding to be performed in the interrupt handler. Scalar DPFP APU User 0 5 6 1 01 1 1 51 6 2 02 1 3 1 000100 rD 00000 rB 0 1011111010

Convert floating-point double-precision to unsigned fraction efdctuf rD,rB rD32:63 ← CnvtFP64ToI32Sat(rB0:63, UNSIGN, ROUND, F) The double-precision floating-point value in rB is converted to an unsigned fraction using the current rounding mode and the result is saturated if it cannot be represented in a 32-bit unsigned fraction. NaNs are converted as though they were zero. Exceptions: If the contents of rB are Infinity, Denorm, or NaN, or if an overflow occurs, SPEFSCR[FINV] is set, and the FG, and FX bits are cleared. If SPEFSCR[FINVE] is set, an interrupt is taken, and the destination register is not updated. This instruction can signal an inexact status and set SPEFSCR[FINXS] if the conversion is not exact. If the floating-point inexact exception is enabled, an interrupt is taken using the Floating-Point Round Interrupt vector. In this case, the destination register is updated with the truncated result, the FG and FX bits are properly updated to allow rounding to be performed in the interrupt handler. Scalar DPFP APU User 0 5 6 1 01 1 1 51 6 2 02 1 3 1 000100 rD 00000 rB 0 1011110110

Convert floating-point double-precision to unsigned integer efdctui rD,rB rD32:63 ← CnvtFP64ToI32Sat(rB0:63, UNSIGN, ROUND, I The double-precision floating-point value in rB is converted to an unsigned integer using the current rounding mode and the result is saturated if it cannot be represented in a 32-bit integer. NaNs are converted as though they were zero. Exceptions: If the contents of rB are Infinity, Denorm, or NaN, or if an overflow occurs, SPEFSCR[FINV] is set, and the FG, and FX bits are cleared. If SPEFSCR[FINVE] is set, an interrupt is taken, and the destination register is not updated. This instruction can signal an inexact status and set SPEFSCR[FINXS] if the conversion is not exact. If the floating-point inexact exception is enabled, an interrupt is taken using the floating-point round interrupt vector. In this case, the destination register is updated with the truncated result, the FG and FX bits are properly updated to allow rounding to be performed in the interrupt handler. Scalar DPFP APU User 0 5 6 1 01 1 1 51 6 2 02 1 3 1 000100 rD 00000 rB 0 1011110100

Convert floating-point double-precision to unsigned integer doubleword with round toward zero efdctuidz rD,rB rD0:63 ← CnvtFP64ToI64Sat(rB0:63, UNSIGN, TRUNC) The double-precision floating-point value in rB is converted to an unsigned integer doubleword using the rounding mode Round toward Zero and the result is saturated if it cannot be represented in a 64-bit integer. NaNs are converted as though they were zero. Exceptions: If the contents of rB are Infinity, Denorm, or NaN, or if an overflow occurs, SPEFSCR[FINV] is set, and the FG, and FX bits are cleared. If SPEFSCR[FINVE] is set, an interrupt is taken, and the destination register is not updated. This instruction can signal an inexact status and set SPEFSCR[FINXS] if the conversion is not exact. If the floating-point inexact exception is enabled, an interrupt is taken using the floating-point round interrupt vector. In this case, the destination register is updated with the truncated result, the FG and FX bits are properly updated to allow rounding to be performed in the interrupt handler. This instruction may only be implemented for 64-bit implementations. Scalar DPFP APU User 0 5 6 1 01 1 1 51 6 2 02 1 3 1 000100 rD 00000 rB 0 1011101010

Convert floating-point double-precision to unsigned integer with round toward zero efdctuiz rD,rB rD32:63 ← CnvtFP64ToI32Sat(rB0:63, UNSIGN, TRUNC, I) The double-precision floating-point value in rB is converted to an unsigned integer using the rounding mode Round toward Zero and the result is saturated if it cannot be represented in a 32-bit integer. NaNs are converted as though they were zero. Exceptions: If the contents of rB are Infinity, Denorm, or NaN, or if an overflow occurs, SPEFSCR[FINV] is set, and the FG, and FX bits are cleared. If SPEFSCR[FINVE] is set, an interrupt is taken, and the destination register is not updated. This instruction can signal an inexact status and set SPEFSCR[FINXS] if the conversion is not exact. If the floating-point inexact exception is enabled, an interrupt is taken using the floating-point round interrupt vector. In this case, the destination register is updated with the truncated result, the FG and FX bits are properly updated to allow rounding to be performed in the interrupt handler. Scalar DPFP APU User 0 5 6 1 01 1 1 51 6 2 02 1 3 1 000100 rD 00000 rB 0 1011111000

Floating-point double-precision divide efddiv r D,rA,rB rD0:63 ← rA0:63 ÷dp rB0:63 rA is divided by rB and the result is stored in rD. If rB is a NaN or infinity, the result is a properly signed zero. Otherwise, if rB is a zero (or a denormalized number optionally transformed to zero by the implementation), or if rA is either NaN or infinity, the result is either pmax (asign==bsign), or nmax (asign!=bsign). Otherwise, if an overflow occurs, pmax or nmax (as appropriate) is stored in rD. If an underflow occurs, +0 or -0 (as appropriate) is stored in rD. Exceptions: If the contents of rA or rB are Infinity, Denorm, or NaN, or if both rA and rB are +/-0, SPEFSCR[FINV] is set. If SPEFSCR[FINVE] is set, an interrupt is taken, and the destination register is not updated. Otherwise, if the content of rB is +/-0 and the content of rA is a finite normalized non-zero number, SPEFSCR[FDBZ] is set. If floating-point divide by zero Exceptions are enabled, an interrupt is then taken. Otherwise, if an overflow occurs, SPEFSCR[FOVF] is set, or if an underflow occurs, SPEFSCR[FUNF] is set. If either underflow or overflow exceptions are enabled and the corresponding bit is set, an interrupt is taken. If any of these interrupts are taken, the destination register is not updated. If the result of this instruction is inexact or if an overflow occurs but overflow exceptions are disabled, and no other interrupt is taken, SPEFSCR[FINXS] is set. If the floating-point inexact exception is enabled, an interrupt is taken using the floating-point round interrupt vector. In this case, the destination register is updated with the truncated result, the FG and FX bits are properly updated to allow rounding to be performed in the interrupt handler. FG and FX are cleared if an overflow, underflow, divide by zero, or invalid operation/input error is signaled, regardless of enabled exceptions. Scalar DPFP APU User 0 5 6 1 01 1 1 51 6 2 02 1 3 1 000100 rD rA rB 0 1011101000

Floating-point double-precision multiply efdmul r D,rA,rB rD0:63 ← rA0:63 ×dp rB0:63 rA is multiplied by rB and the result is stored in rD. If rA or rB are zero (or a denormalized number optionally transformed to zero by the implementation), the result is a properly signed zero. Otherwise, if rA or rB are either NaN or infinity, the result is either pmax sign==bsign), or nmax (asign!=bsign). Otherwise, if an overflow occurs, pmax or nmax (as appropriate) is stored in rD. If an underflow occurs, +0 or -0 (as appropriate) is stored in rD. Exceptions: If the contents of rA or rB are Infinity, Denorm, or NaN, SPEFSCR[FINV] is set. If SPEFSCR[FINVE] is set, an interrupt is taken, and the destination register is not updated. Otherwise, if an overflow occurs, SPEFSCR[FOVF] is set, or if an underflow occurs, SPEFSCR[FUNF] is set. If either underflow or overflow exceptions are enabled and the corresponding bit is set, an interrupt is taken. If any of these interrupts are taken, the destination register is not updated. If the result of this instruction is inexact or if an overflow occurs but overflow exceptions are disabled, and no other interrupt is taken, SPEFSCR[FINXS] is set. If the floating-point inexact exception is enabled, an interrupt is taken using the floating-point round interrupt vector. In this case, the destination register is updated with the truncated result, the FG and FX bits are properly updated to allow rounding to be performed in the interrupt handler. FG and FX are cleared if an overflow, underflow, or invalid operation/input error is signaled, regardless of enabled exceptions. Scalar DPFP APU User 0 5 6 1 01 1 1 51 6 2 02 1 3 1 000100 rD rA rB 0 1011101000

Floating-point double-precision negative absolute value efdnabs r D,rA rD0:63 ← 0b1 || rA1:63 The sign bit of rA is set to 1 and the result is placed into rD. Exceptions: Exception detection for embedded floating-point absolute value operations is implementation dependent. An implementation may choose to not detect exceptions and carry out the sign bit operation. If the implementation does not detect exceptions, or if exception detection is disabled, the computation can be carried out in one of two ways, as a sign bit operation ignoring the rest of the contents of the source register, or by examining the input and appropriately saturating the input prior to performing the operation. If an implementation chooses to handle exceptions, the exception is handled as follows: If rA is Infinity, Denorm, or NaN, SPEFSCR[FINV] is set, and FG and FX are cleared. If floating- point invalid input exceptions are enabled, an interrupt is taken and the destination register is not updated. Scalar DPFP APU User 0 5 6 1 01 1 1 51 6 2 02 1 3 1 000100 rD rA 0 00000 1011100101

Floating-point double-precision negate efdneg rD,rA rD0:63 ← ¬rA0 || rA1:63 The sign bit of rA is complemented and the result is placed into rD. Exceptions: Exception detection for embedded floating-point absolute value operations is implementation dependent. An implementation may choose to not detect exceptions and carry out the sign bit operation. If the implementation does not detect exceptions, or if exception detection is disabled, the computation can be carried out in one of two ways, as a sign bit operation ignoring the rest of the contents of the source register, or by examining the input and appropriately saturating the input prior to performing the operation. If an implementation chooses to handle exceptions, the exception is handled as follows: If rA is Infinity, Denorm, or NaN, SPEFSCR[FINV] is set, and FG and FX are cleared. If floating- point invalid input exceptions are enabled, an interrupt is taken and the destination register is not updated. Scalar DPFP APU User 0 5 6 1 01 1 1 51 6 2 02 1 3 1 000100 rD rA 0 00000 1011100110

Floating-point double-precision subtract efdsub r D,rA,rB rD0:63 ← rA0:63 -dp rB0:63 rB is subtracted from rA and the result is stored in rD. If rA is NaN or infinity, the result is either pmax (asign==0), or nmax (asign==1). Otherwise, If rB is NaN or infinity, the result is either nmax (bsign==0), or pmax (bsign==1). Otherwise, if an overflow occurs, pmax or nmax (as appropriate) is stored in rD. If an underflow occurs, +0 (for rounding modes RN, RZ, RP) or -0 (for rounding mode RM) is stored in rD. Exceptions: If the contents of rA or rB are Infinity, Denorm, or NaN, SPEFSCR[FINV] is set. If SPEFSCR[FINVE] is set, an interrupt is taken, and the destination register is not updated. Otherwise, if an overflow occurs, SPEFSCR[FOVF] is set, or if an underflow occurs, SPEFSCR[FUNF] is set. If either underflow or overflow exceptions are enabled and the corresponding bit is set, an interrupt is taken. If any of these interrupts are taken, the destination register is not updated. If the result of this instruction is inexact or if an overflow occurs but overflow exceptions are disabled, and no other interrupt is taken, SPEFSCR[FINXS] is set. If the floating-point inexact exception is enabled, an interrupt is taken using the floating-point round interrupt vector. In this case, the destination register is updated with the truncated result, the FG and FX bits are properly updated to allow rounding to be performed in the interrupt handler. FG and FX are cleared if an overflow, underflow, or invalid operation/input error is signaled, regardless of enabled exceptions. Scalar DPFP APU User 0 5 6 1 01 1 1 51 6 2 02 1 3 1 000100 rD rA rB 0 1011100001

Floating-point double-precision test equal efdtsteq crf D,rA,rB al ← rA0:63 bl ← rB0:63 if (al = bl) then cl ← 1 else cl ← 0 CR4*crD:4*crD+3 ← undefined || cl || undefined || undefined rA is compared against rB. If rA is equal to rB, the bit in the crfD is set, otherwise it is cleared. Comparison ignores the sign of 0 (+0 = -0). The comparison proceeds after treating NaNs, Infinities, and Denorms as normalized numbers, using their values of ‘e’ and ‘f’ directly. No exceptions are generated during the execution of efdtsteq If strict IEEE 754 compliance is required, the program should use efdcmpeq. Implementation note: In an implementation, the execution of efdtsteq is likely to be faster than the execution of efdcmpeq. Scalar DPFP APU User 0 5 6 8 9 1 01 1 1 51 6 2 02 1 3 1 000100 crfD0 0 rA rB 0 1011111110

Floating-point double-precision test greater than efdtstgt crf D,rA,rB al ← rA0:63 bl ← rB0:63 if (al > bl) then cl ← 1 else cl ← 0 CR4*crD:4*crD+3 ← undefined || cl || undefined || undefined rA is compared against rB. If rA is greater than rB, the bit in the crfD is set, otherwise it is cleared. Comparison ignores the sign of 0 (+0 = -0). The comparison proceeds after treating NaNs, Infinities, and Denorms as normalized numbers, using their values of ‘e’ and ‘f’ directly. No exceptions are generated during the execution of efdtstgt. If strict IEEE 754 compliance is required, the program should use efdcmpgt. Note: Implementation note: In an implementation, the execution of efdtstgt is likely to be faster than the execution of efdcmpgt. Scalar DPFP APU User 0 5 6 8 9 1 01 1 1 51 6 2 02 1 3 1 000100 crfD0 0 rA rB 0 1011111100

Floating-point double-precision test less than efdtstlt crf D,rA,rB al ← rA0:63 bl ← rB0:63 if (al < bl) then cl ← 1 else cl ← 0 CR4*crD:4*crD+3 ← undefined || cl || undefined || undefined rA is compared against rB. If rA is less than rB, the bit in the crfD is set, otherwise it is cleared. Comparison ignores the sign of 0 (+0 = -0). The comparison proceeds after treating NaNs, Infinities, and Denorms as normalized numbers, using their values of ‘e’ and ‘f’ directly. No exceptions are generated during the execution of efdtstlt. If strict IEEE 754 compliance is required, the program should use efdcmplt. Implementation note: In an implementation, the execution of efdtstlt is likely to be faster than the execution of efdcmplt. Scalar DPFP APU User 0 5 6 8 9 1 01 1 1 51 6 2 02 1 3 1 000100 crfD0 0 rA rB 0 1011111101

Floating-Point Absolute Value efsabs r D,rA rD32:63 ← 0b0 || rA33:63 The sign bit of rA is cleared and the result is placed into rD. It is implementation dependent if invalid values for rA (NaN, Denorm, Infinity) are detected and exceptions are taken. Scalar SPFP APU User 0 5 6 1 01 1 1 51 6 2 02 1 3 1 000100 rD rA 000000 1011000100

efsadd r D,rA,rB rD32:63 ← rA32:63 +sp rB32:63 The single-precision floating-point value of rA is added to rB and the result is stored in rD. If an overflow condition is detected or the contents of rA or rB are NaN or Infinity, the result is an appropriately signed maximum floating-point value. If an underflow condition is detected, the result is an appropriately signed floating-point 0. The following status bits are set in the SPEFSCR:

  • FINV if the contents of rA or rB are +infinity, –infinity, denorm, or NaN
  • FOFV if an overflow occurs
  • FUNF if an underflow occurs
  • FINXS, FG, FX if the result is inexact or overflow occurred and overflow exceptions are disabled Scalar SPFP APU User 0 5 6 1 01 1 1 51 6 2 02 1 3 1 000100 rD rA rB 0 1011000000

Convert Floating-Point from Signed Fraction efscfsf rD,rB rD32:63 ← CnvtI32ToFP32Sat(rB32:63, SIGN, LOWER, F) The signed fractional value in rB is converted to the nearest single-precision floating-point value using the current rounding mode and placed into rD. The following status bits are set in the SPEFSCR:

  • FINXS, FG, FX if the result is inexact Scalar SPFP APU User 0 5 6 1 01 1 1 51 6 2 02 1 3 1 000100 rD 00000 rB 01011010011

Convert Floating-Point from Signed Integer efscfsi rD,rB rD32:63 ← CnvtSI32ToFP32Sat(rB32:63, SIGN, LOWER, I) The signed integer value in rB is converted to the nearest single-precision floating-point value using the current rounding mode and placed into rD. The following status bits are set in the SPEFSCR:

  • FINXS, FG, FX if the result is inexact Scalar SPFP APU User 0 5 6 1 01 1 1 51 6 2 02 1 3 1 000100 rD 00000 rB 0 1011010001

Convert Floating-Point from Unsigned Fraction efscfuf r D,rB rD32:63 ← CnvtI32ToFP32Sat(rB32:63, UNSIGN, LOWER, F) The unsigned fractional value in rB is converted to the nearest single-precision floating-point value using the current rounding mode and placed into rD. The following status bits are set in the SPEFSCR:

  • FINXS, FG, FX if the result is inexact Scalar SPFP APU User 0 5 6 1 01 1 1 51 6 2 02 1 3 1 000100 rD 00000 rB 0 1011010010

Convert Floating-Point from Unsigned Integer efscfui rD,rB rD32:63 ← CnvtI32ToFP32Sat(rB32:63, UNSIGN, LOWER, I) The unsigned integer value in rB is converted to the nearest single-precision floating-point value using the current rounding mode and placed into rD. The following status bits are set in the SPEFSCR:

  • FINXS, FG, FX if the result is inexact Scalar SPFP APU User 0 5 6 1 01 1 1 51 6 2 02 1 3 1 000100 rD 00000 rB 0 1011010000

Floating-Point Compare Equal efscmpeq cr D,rA,rB al ← rA32:63 bl ← rB32:63 if (al = bl) then cl ← 1 else cl ← 0 CR4*crD:4*crD+3 ← undefined || cl || undefined || undefined The value in rA is compared against rB. If rA equals rB, the crD bit is set, otherwise it is cleared. Comparison ignores the sign of 0 (+0 = –0). If either operand contains a NaN, infinity, or a denorm and floating-point invalid exceptions are enabled in the SPEFSCR, the exception is taken. If the exception is not enabled, the comparison treats NaNs, infinities, and denorms as normalized numbers. The following status bits are set in SPEFSCR:

  • FINV if the contents of rA or rB are +infinity, –infinity, denorm or NaN Scalar SPFP APU User 0 5 6 8 9 1 01 1 1 51 6 2 02 1 3 1 000100 crD0 0 rA rB 0 1011001110

Floating-Point Compare Greater Than efscmpgt cr D,rA,rB al ← rA32:63 bl ← rB32:63 if (al > bl) then cl ← 1 else cl ← 0 CR4*crD:4*crD+3 ← undefined || cl || undefined || undefined The value in rA is compared against rB. If rA is greater than rB, the bit in the crD is set, otherwise it is cleared. Comparison ignores the sign of 0 (+0 = –0). If either operand contains a NaN, infinity, or a denorm and floating-point invalid exceptions are enabled in the SPEFSCR, the exception is taken. If the exception is not enabled, the comparison treats NaNs, infinities, and denorms as normalized numbers. The following status bits are set in SPEFSCR:

  • FINV if the contents of rA or rB are +infinity, –infinity, denorm or NaN Scalar SPFP APU User 0 5 6 8 9 1 01 1 1 51 6 2 02 1 3 1 000100 crD0 0 rA rB 0 1011001100

Floating-Point Compare Less Than efscmplt cr D,rA,rB al ← rA32:63 bl ← rB32:63 if (al < bl) then cl ← 1 else cl ← 0 CR4*crD:4*crD+3 ← undefined || cl || undefined || undefined The value in rA is compared against rB. If rA is less than rB, the bit in the crD is set, otherwise it is cleared. Comparison ignores the sign of 0 (+0 = –0). If either operand contains a NaN, infinity, or a denorm and floating-point invalid exceptions are enabled in the SPEFSCR, the exception is taken. If the exception is not enabled, the comparison treats NaNs, infinities, and denorms as normalized numbers. The following status bits are set in SPEFSCR:

  • FINV if the contents of rA or rB are +infinity, –infinity, denorm or NaN Scalar SPFP APU User 0 5 6 8 9 1 01 1 1 51 6 2 02 1 3 1 000100 crD0 0 rA rB 0 1011001101

Convert Floating-Point to Signed Fraction efsctsf r D,rB rD32:63 ← CnvtFP32ToISat(rB32:63, SIGN, LOWER, ROUND, F) The single-precision floating-point value in rB is converted to a signed fraction using the current rounding mode. The result saturates if it cannot be represented in a 32-bit fraction. NaNs are converted to 0. The following status bits are set in the SPEFSCR:

  • FINV if the contents of rB are +infinity., –infinity, denorm, or NaN, or rB cannot be represented in the target format
  • FINXS, FG, FX if the result is inexact Scalar SPFP APU User 0 5 6 1 01 1 1 51 6 2 02 1 3 1 000100 rD 00000 rB 01011010111

Convert Floating-Point to Signed Integer efsctsi r D,rB rD32:63 ← CnvtFP32ToISat(rB32:63, SIGN, LOWER, ROUND, I) The single-precision floating-point value in rB is converted to a signed integer using the current rounding mode. The result saturates if it cannot be represented in a 32-bit integer. NaNs are converted to 0. The following status bits are set in the SPEFSCR:

  • FINV if the contents of rB are +infinity, –infinity, denorm, or NaN, or rB cannot be represented in the target format
  • FINXS, FG, FX if the result is inexact Scalar SPFP APU User 0 5 6 1 01 1 1 51 6 2 02 1 3 1 000100 rD 00000 rB 01011010101

Convert Floating-Point to Signed Integer with Round toward Zero efsctsiz r D,rB rD32–63 ← CnvtFP32ToISat(rB32:63, SIGN, LOWER, TRUNC, I) The single-precision floating-point value in rB is converted to a signed integer using the rounding mode Round towards Zero. The result saturates if it cannot be represented in a 32- bit integer. NaNs are converted to 0. Scalar SPFP APU User 0 5 6 1 01 1 1 51 6 2 02 1 3 1 000100 rD 00000 rB 01011011010

Convert Floating-Point to Unsigned Fraction efsctuf r D,rB rD32:63 ← CnvtFP32ToISat(rB32:63, UNSIGN, LOWER, ROUND, F) The single-precision floating-point value in rB is converted to an unsigned fraction using the current rounding mode. The result saturates if it cannot be represented in a 32-bit unsigned fraction. NaNs are converted to 0. The following status bits are set in the SPEFSCR:

  • FINV if the contents of rB are +infinity, –infinity, denorm, or NaN, or rB cannot be represented in the target format
  • FINXS, FG, FX if the result is inexact Scalar SPFP APU User 0 5 6 1 01 1 1 51 6 2 02 1 3 1 000100 rD 00000 rB 01011010110

Convert Floating-Point to Unsigned Integer efsctui r D,rB rD32:63 ← CnvtFP32ToISat(rB32:63, UNSIGN, LOWER, ROUND, I) The single-precision floating-point value in rB is converted to an unsigned integer using the current rounding mode. The result saturates if it cannot be represented in a 32-bit unsigned integer. NaNs are converted to 0. The following status bits are set in the SPEFSCR:

  • FINV if the contents of rB are +infinity, –infinity, denorm, or NaN, or rB cannot be represented in the target format
  • FINXS, FG, FX if the result is inexact Scalar SPFP APU User 0 5 6 1 01 1 1 51 6 2 02 1 3 1 000100 rD 00000 rB 01011010100

Convert Floating-Point to Unsigned Integer with Round toward Zero efsctuiz r D,rB rD32:63 ← CnvtFP32ToISat(rB32:63, UNSIGN, LOWER, TRUNC, I) The single-precision floating-point value in rB is converted to an unsigned integer using the rounding mode Round toward Zero. The result saturates if it cannot be represented in a 32- bit unsigned integer. NaNs are converted to 0. The following status bits are set in the SPEFSCR:

  • FINV if the contents of rB are +infinity, –infinity, denorm, or NaN, or rB cannot be represented in the target format
  • FINXS, FG, FX if the result is inexact Scalar SPFP APU User 0 5 6 1 01 1 1 51 6 2 02 1 3 1 000100 rD 00000 rB 01011011000

efsdiv r D,rA,rB rD32:63 ← rA32:63 ÷sp rB32:63 The single-precision floating-point value in rA is divided by rB and the result is stored in rD. If an overflow is detected, or rB is a denorm (or 0 value), or rA is a NaN or Infinity and rB is a normalized number, the result is an appropriately signed maximum floating-point value. If an underflow is detected or rB is a NaN or Infinity, the result is an appropriately signed floating-point 0. The following status bits are set in the SPEFSCR:

  • FINV if the contents of rA or rB are +infinity, –infinity, denorm, or NaN
  • FOFV if an overflow occurs
  • FUNV if an underflow occurs
  • FDBZS, FDBZ if a divide by zero occurs
  • FINXS, FG, FX if the result is inexact or overflow occurred and overflow exceptions are disabled Scalar SPFP APU User 0 5 6 1 01 1 1 51 6 2 02 1 3 1 000100 rD rA rB 01011001001

efsmul r D,rA,rB rD32:63 ← rA32:63 ×sp rB32:63 The single-precision floating-point value in rA is multiplied by rB and the result is stored in rD. If an overflow is detected the result is an appropriately signed maximum floating-point value. If one of rA or rB is a NaN or an Infinity and the other is not a denorm or zero, the result is an appropriately signed maximum floating-point value. If an underflow is detected, or rA or rB is a denorm, the result is an appropriately signed floating-point 0. The following status bits are set in the SPEFSCR:

  • FINV if the contents of rA or rB are +infinity, –infinity, denorm, or NaN
  • FOFV if an overflow occurs
  • FUNV if an underflow occurs
  • FINXS, FG, FX if the result is inexact or overflow occurred and overflow exceptions are disabled Scalar SPFP APU User 0 5 6 1 01 1 1 51 6 2 02 1 3 1 000100 rD rA rB 01011001000

Floating-Point Negative Absolute Value efsnabs r D,rA rD32:63 ← 0b1 || rA33:63 The sign bit of rA is set and the result is stored in rD. It is implementation dependent if invalid values for rA (NaN, Denorm, Infinity) are detected and exceptions are taken. Scalar SPFP APU User 0 5 6 1 01 1 1 51 6 2 02 1 3 1 000100 rD rA 0000001011000101

efsneg r D,rA rD32:63 ← ¬rA32 || rA33:63 The sign bit of rA is complemented and the result is stored in rD. It is implementation dependent if invalid values for rA (NaN, Denorm, Infinity) are detected and exceptions are taken. Scalar SPFP APU User 0 5 6 1 01 1 1 51 6 2 02 1 3 1 000100 rD rA 0000001011000110

efssub r D,rA,rB rD32:63 ← rA32:63 -sp rB32:63 The single-precision floating-point value in rB is subtracted from that in rA and the result is stored in rD. If an overflow condition is detected or the contents of rA or rB are NaN or Infinity, the result is an appropriately signed maximum floating-point value. If an underflow condition is detected, the result is an appropriately signed floating-point 0. The following status bits are set in the SPEFSCR:

  • FINV if the contents of rA or rB are +infinity, –infinity, denorm, or NaN
  • FOFV if an overflow occurs
  • FUNF if an underflow occurs
  • FINXS, FG, FX if the result is inexact or overflow occurred and overflow exceptions are disabled Scalar SPFP APU User 0 5 6 1 01 1 1 51 6 2 02 1 3 1 000100 rD rA rB 01011000001

efststeq cr D,rA,rB al ← rA32:63 bl ← rB32:63 if (al = bl) then cl ← 1 else cl ← 0 CR4*crD:4*crD+3 ← undefined || cl || undefined || undefined The value in rA is compared against rB. If rA equals rB, the bit in crD is set, otherwise it is cleared. Comparison ignores the sign of 0 (+0 = –0). The comparison treats NaNs, infinities, and denorms as normalized numbers. No exceptions are taken during execution of efststeq. If strict IEEE 754 compliance is required, the program should use efscmpeq. Scalar SPFP APU User 0 5 6 8 9 1 01 1 1 51 6 2 02 1 3 1 000100 crD0 0 rA rB 01011011110

Floating-Point Test Greater Than efststgt cr D,rA,rB al ← rA32:63 bl ← rB32:63 if (al > bl) then cl ← 1 else cl ← 0 CR4*crD:4*crD+3 ← undefined || cl || undefined || undefined If rA is greater than rB, the bit in crD is set, otherwise it is cleared. Comparison ignores the sign of 0 (+0 = –0). The comparison treats NaNs, infinities, and denorms as normalized numbers. No exceptions are taken during the execution of efststgt. If strict IEEE 754 compliance is required, the program should use efscmpgt. Scalar SPFP APU User 0 5 6 8 9 1 01 1 1 51 6 2 02 1 3 1 000100 crD0 0 rA rB 01011011100

Floating-Point Test Less Than efststlt cr D,rA,rB al ← rA32:63 bl ← rB32:63 if (al < bl) then cl ← 1 else cl ← 0 CR4*crD:4*crD+3 ← undefined || cl || undefined || undefined If rA is less than rB, the bit in the crD is set, otherwise it is cleared. Comparison ignores the sign of 0 (+0 = –0). The comparison treats NaNs, infinities, and denorms as normalized numbers. No exceptions are taken during the execution of efststlt. If strict IEEE 754 compliance is required, the program should use efscmplt. Scalar SPFP APU User 0 5 6 8 9 1 01 1 1 51 6 2 02 1 3 1 000100 crD0 0 rA rB 01011011101

eqv r A,rS,rB( R c = 0 ) eqv. r A,rS,rB( R c = 1 ) result0:63 ← rS ≡ rB if Rc=1 then do LT ← result32:63 < 0 GT ← result32:63 > 0 EQ ← result32:63 = 0 rA ← result The contents of rS are XORed with the contents of rB and the one’s complement of the result is placed into rA. Other registers altered:

  • CR0 (if Rc=1) Book E User 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1 011111 rS rA rB 0100011100 R c

Figure 24. Vector absolute value (evabs)

Figure 25. Vector add immediate word (evaddiw)

results are placed in rD and into the accumulator. Figure 26. Vector add signed, modulo, integer to accumulator word (evaddsmiaaw)

the SPEFSCR overflow and summary overflow bits. Figure 27. Vector add signed, saturate, integer to accumulator word (evaddssiaaw)

accumulator and the results are placed in rD and the accumulator. Figure 28. Vector add unsigned, modulo, integer to accumulator word

Figure 29. Vector add unsigned, saturate, integer to accumulator word

Figure 30. Vector add word (evaddw)

the corresponding element of rD. Figure 31. Vector AND (evand)

elements of rB. The results are placed in the corresponding element of rD. Figure 32. Vector AND with complement (evandc)

set to the OR and AND of the result of the compare of the high and low elements. Figure 33. Vector Compare Equal (evcmpeq)

crD are set to the OR and AND of the result of the compare of the high and low elements. Figure 34. Vector compare greater than signed (evcmpgts)

crD are set to the OR and AND of the result of the compare of the high and low elements. Figure 35. Vector compare greater than unsigned (evcmpgtu)

are set to the OR and AND of the result of the compare of the high and low elements. Figure 36. Vector compare less than signed (evcmplts)

are set to the OR and AND of the result of the compare of the high and low elements. Figure 37. Vector compare less than unsigned (evcmpltu)

evcntlzw is used for unsigned operands; evcntlsw is used for signed operands. Figure 38. Vector count leading signed bits word (evcntlsw)

Figure 39. Vector count leading zeros word (evcntlzw)

evdivws r D,rA,rB dividendh ← rA0:31 dividendl ← rA32:63 divisorh ← rB0:31 divisorl ← rB32:63 rD0:31 ← dividendh ÷ divisorh rD32:63 ← dividendl ÷ divisorl ovh ← 0 ovl ← 0 if ((dividendh < 0) & (divisorh = 0)) then rD0:31 ← 0x80000000 ovh ← 1 else if ((dividendh >= 0) & (divisorh = 0)) then rD0:31 ← 0x7FFFFFFF ovh ← 1 else if ((dividendh = 0x80000000) & (divisorh = 0xFFFF_FFFF)) then rD0:31 ← 0x7FFFFFFF ovh ← 1 if ((dividendl < 0) & (divisorl = 0)) then rD32:63 ← 0x80000000 ovl ← 1 else if ((dividendl >= 0) & (divisorl = 0)) then rD32:63 ← 0x7FFFFFFF ovl ← 1 else if ((dividendl = 0x80000000) & (divisorl = 0xFFFF_FFFF)) then rD32:63 ← 0x7FFFFFFF ovl ← 1 SPEFSCROVH ← ovh SPEFSCROV ← ovl SPEFSCRSOVH ← SPEFSCRSOVH | ovh SPEFSCRSOV ← SPEFSCRSOV | ovl The two dividends are the two elements of the contents of rA. The two divisors are the two elements of the contents of rB. The resulting two 32-bit quotients on each element are placed into rD. The remainders are not supplied. The operands and quotients are interpreted as signed integers. If overflow, underflow, or divide by zero occurs, the overflow and summary overflow SPEFSCR bits are set. Note that any overflow indication is always set as a side effect of this instruction. No form is defined that disables the setting of the overflow bits. In case of overflow, a saturated value is delivered into the destination register. SPE APU User 0 5 6 1 01 1 1 51 6 2 02 1 3 1 000100 rD rA rB 10011000110

Figure 40. Vector divide word signed (evdivws)

Figure 41. Vector divide word unsigned (evdivwu)

The corresponding elements of rA & rB are XNORed bitwise, & the results are placed in rD. Figure 42. Vector equivalent (eveqv)

Figure 43. Vector extend sign byte (evextsb)

Figure 44. Vector extend sign half word (evextsh)

Vector floating-point single-precision absolute value evfsabs r D,rA rD0:31 ← 0b0 || rA1:31 rD32:63 ← 0b0 || rA33:63 The sign bit of each element in rA is set to 0 and the results are placed into rD. Exceptions: Exception detection for embedded floating-point absolute value operations is implementation dependent. An implementation may choose to not detect exceptions and carry out the computation. If the implementation does not detect exceptions, or if exception detection is disabled, the computation can be carried out in one of two ways, as a sign bit operation ignoring the rest of the contents of the source register, or by examining the input and appropriately saturating the input prior to performing the operation. If an implementation chooses to handle exceptions, the exception is handled as follows: if the contents of either element of rA are Infinity, Denorm, or NaN, SPEFSCR[FINV,FINVH] are set appropriately, and SPEFSCR[FGH,FXH,FG,FX] are cleared appropriately. If floating- point invalid input exceptions are enabled, an interrupt is taken and the destination register is not updated. Vector SPFP APU User 0 5 6 1 01 1 1 51 6 2 02 1 3 1 000100 rD rA 0000001010000100

Vector floating-point single-precision add evfsadd r D,rA,rB rD0:31 ← rA0:31 +sp rB0:31 rD32:63 ← rA32:63 +sp rB32:63 Each single-precision floating-point element of rA is added to the corresponding element of rB and the results are stored in rD. If an element of rA is NaN or infinity, the corresponding result is either pmax (asign==0), or nmax (asign==1). Otherwise, if an element of rB is NaN or infinity, the corresponding result is either pmax (bsign==0), or nmax (bsign==1). Otherwise, if an overflow occurs, pmax or nmax (as appropriate) is stored in the corresponding element of rD. If an underflow occurs, +0 (for rounding modes RN, RZ, RP) or –0 (for rounding mode RM) is stored in the corresponding element of rD. Exceptions: If the contents of either element of rA or rB are Infinity, Denorm, or NaN, SPEFSCR[FINV,FINVH] are set appropriately, and SPEFSCR[FGH,FXH,FG,FX] are cleared appropriately. If SPEFSCR[FINVE] is set, an interrupt is taken and the destination register is not updated. Otherwise, if an overflow occurs, SPEFSCR[FOVF ,FOVFH] are set appropriately, or if an underflow occurs, SPEFSCR[FUNF ,FUNFH] are set appropriately. If either underflow or overflow exceptions are enabled and a corresponding status bit is set, an interrupt is taken. If any of these interrupts are taken, the destination register is not updated. If either result element of this instruction is inexact, or overflows but overflow exceptions are disabled, and no other interrupt is taken, or underflows but underflow exceptions are disabled, and no other interrupt is taken, SPEFSCR[FINXS,FINXSH] is set. If the floating- point inexact exception is enabled, an interrupt is taken using the floating-point round interrupt vector. In this case, the destination register is updated with the truncated result(s). The FG and FX bits are properly updated to allow rounding to be performed in the interrupt handler. FG and FX (FGH and FXH) are cleared if an overflow or underflow interrupt is taken, or if an invalid operation/input error is signaled for the low (high) element (regardless of FINVE). Vector SPFP APU User 0 5 6 1 01 1 1 51 6 2 02 1 3 1 000100 rD rA rB 01010000000

Vector convert floating-point single-precision from signed fraction evfscfsf rD,rB rD0:31 ← CnvtI32ToFP32Sat(rB0:31, SIGN, UPPER, F) rD32:63 ← CnvtI32ToFP32Sat(rB32:63, SIGN, LOWER, F) Each signed fractional element of rB is converted to a single-precision floating-point value using the current rounding mode and the results are placed into the corresponding elements of rD. Exceptions: This instruction can signal an inexact status and set SPEFSCR[FINXS] if the conversions are not exact. If the floating-point inexact exception is enabled, an interrupt is taken using the floating-point round interrupt vector. In this case, the destination register is updated with the truncated result(s). The FGH, FXH, FG and FX bits are properly updated to allow rounding to be performed in the interrupt handler. Vector SPFP APU User 0 5 6 1 01 1 1 51 6 2 02 1 3 1 000100 rD 00000 rB 0 1 0 1 0 0 1 0 0 1 1

Vector convert floating-point single-precision from signed integer evfscfsi rD,rB rD0:31 ← CnvtSI32ToFP32Sat(rB0:31, SIGN, UPPER, I) rD32:63 ← CnvtSI32ToFP32Sat(rB32:63, SIGN, LOWER, I) Each signed integer element of rB is converted to the nearest single-precision floating-point value using the current rounding mode and the results are placed into the corresponding element of rD. Exceptions: This instruction can signal an inexact status and set SPEFSCR[FINXS] if the conversions are not exact. If the floating-point inexact exception is enabled, an interrupt is taken using the floating-point round interrupt vector. In this case, the destination register is updated with the truncated result(s). The FGH, FXH, FG and FX bits are properly updated to allow rounding to be performed in the interrupt handler. Vector SPFP APU User 0 5 6 1 01 1 1 51 6 2 02 1 3 1 000100 rD 00000 rB 0 1 0 1 0 0 1 0 0 0 1

Vector convert floating-point single-precision from unsigned fraction evfscfuf r D,rB rD0:31 ← CnvtI32ToFP32Sat(rB0:31, UNSIGN, UPPER, F) rD32:63 ← CnvtI32ToFP32Sat(rB32:63, UNSIGN, LOWER, F) Each unsigned fractional element of rB is converted to a single-precision floating-point value using the current rounding mode and the results are placed into the corresponding elements of rD. Exceptions: This instruction can signal an inexact status and set SPEFSCR[FINXS] if the conversions are not exact. If the floating-point inexact exception is enabled, an interrupt is taken using the floating-point round interrupt vector. In this case, the destination register is updated with the truncated result(s). The FGH, FXH, FG and FX bits are properly updated to allow rounding to be performed in the interrupt handler. Vector SPFP APU User 0 5 6 1 01 1 1 51 6 2 02 1 3 1 000100 rD 00000 rB 0 1 0 1 0 0 1 0 0 1 0

Vector convert floating-point single-precision from unsigned integer evfscfui rD,rB rD0:31 ← CnvtI32ToFP32Sat(rB031, UNSIGN, UPPER, I) rD32:63 ← CnvtI32ToFP32Sat(rB32:63, UNSIGN, LOWER, I) Each unsigned integer element of rB is converted to the nearest single-precision floating- point value using the current rounding mode and the results are placed into the corresponding elements of rD. Exceptions: This instruction can signal an inexact status and set SPEFSCR[FINXS] if the conversions are not exact. If the floating-point inexact exception is enabled, an interrupt is taken using the floating-point round interrupt vector. In this case, the destination register is updated with the truncated result(s). The FGH, FXH, FG and FX bits are properly updated to allow rounding to be performed in the interrupt handler. Vector SPFP APU User 0 5 6 1 01 1 1 51 6 2 02 1 3 1 000100 rD 00000 rB 0 1 0 1 0 0 1 0 0 0 0

Vector floating-point single-precision compare equal evfscmpeq crf D,rA,rB ah ← rA0:31 al ← rA32:63 bh ← rB0:31 bl ← rB32:63 if (ah = bh) then ch ← 1 else ch ← 0 if (al = bl) then cl ← 1 else cl ← 0 Each element of rA is compared against the corresponding element of rB. If rA equals rB, the crfD bit is set, otherwise it is cleared. Comparison ignores the sign of 0 (+0 = –0). Exceptions: If the contents of either element of rA or rB are Infinity, Denorm, or NaN, SPEFSCR[FINV,FINVH] are set appropriately, and SPEFSCR[FGH,FXH,FG,FX] are cleared appropriately. If floating-point invalid input exceptions are enabled, an interrupt is taken, and the condition register is not updated. Otherwise, the comparison proceeds after treating NaNs, Infinities, and Denorms as normalized numbers, using their values of ‘e’ and ‘f’ directly. Vector SPFP APU User 0 5 6 8 9 1 01 1 1 51 6 2 02 1 3 1 000100 crfD0 0 rA rB 0 1 0 1 0 0 0 1 1 1 0

Vector floating-point single-precision compare greater than evfscmpgt crf D,rA,rB ah ← rA0:31 al ← rA32:63 bh ← rB0:31 bl ← rB32:63 if (ah > bh) then ch ← 1 else ch ← 0 if (al > bl) then cl ← 1 else cl ← 0 Each element of rA is compared against the corresponding element of rB. If rA is greater than rB, the bit in the crfD is set, otherwise it is cleared. Comparison ignores the sign of 0 (+0 = –0). Exceptions: If the contents of either element of rA or rB are Infinity, Denorm, or NaN, SPEFSCR[FINV,FINVH] are set appropriately, and SPEFSCR[FGH,FXH,FG,FX] are cleared appropriately. If floating-point invalid input exceptions are enabled then an interrupt is taken, and the condition register is not updated. Otherwise, the comparison proceeds after treating NaNs, Infinities, and Denorms as normalized numbers, using their values of ‘e’ and ‘f’ directly. Vector SPFP APU User 0 5 6 8 9 1 01 1 1 51 6 2 02 1 3 1 000100 crfD0 0 rA rB 0 1 0 1 0 0 0 1 1 0 0

Vector floating-point single-precision compare less than evfscmplt crf D,rA,rB ah ← rA0:31 al ← rA32:63 bh ← rB0:31 bl ← rB32:63 if (ah < bh) then ch ← 1 else ch ← 0 if (al < bl) then cl ← 1 else cl ← 0 Each element of rA is compared against the corresponding element of rB. If rA is less than rB, the bit in the crfD is set, otherwise it is cleared. Comparison ignores the sign of 0 (+0 = – 0). Exceptions: If the contents of either element of rA or rB are Infinity, Denorm, or NaN, SPEFSCR[FINV,FINVH] are set appropriately, and SPEFSCR[FGH,FXH,FG,FX] are cleared appropriately. If floating-point invalid input exceptions are enabled then an interrupt is taken, and the condition register is not updated. Otherwise, the comparison proceeds after treating NaNs, Infinities, and Denorms as normalized numbers, using their values of ‘e’ and ‘f’ directly. Vector SPFP APU User 0 5 6 8 9 1 01 1 1 51 6 2 02 1 3 1 000100 crfD0 0 rA rB 0 1 0 1 0 0 0 1 1 0 1

Vector convert floating-point single-precision to signed fraction evfsctsf r D,rB rD0:31 ← CnvtFP32ToISat(rB0:31, SIGN, UPPER, ROUND, F) rD32:63 ← CnvtFP32ToISat(rB32:63, SIGN, LOWER, ROUND, F) Each single-precision floating-point element in rB is converted to a signed fraction using the current rounding mode and the result is saturated if it cannot be represented in a 32-bit signed fraction. NaNs are converted as though they were zero. Exceptions: If either element of rB is Infinity, Denorm, or NaN, or if an overflow occurs, SPEFSCR[FINV,FINVH] are set appropriately and SPEFSCR[FGH,FXH,FG,FX] are cleared appropriately. If SPEFSCR[FINVE] is set, an interrupt is taken, the destination register is not updated, and no other status bits are set. If either result element of this instruction is inexact and no other interrupt is taken, SPEFSCR[FINXS] is set. If the floating-point inexact exception is enabled, an interrupt is taken using the floating-point round interrupt vector. In this case, the destination register is updated with the truncated result. The FGH, FXH, FG and FX bits are properly updated to allow rounding to be performed in the interrupt handler. Vector SPFP APU User 0 5 6 1 01 1 1 51 6 2 02 1 3 1 000100 rD 00000 rB 0 1 0 1 0 0 1 0 1 1 1

Vector convert floating-point single-precision to signed integer evfsctsi r D,rB rD0:31 ← CnvtFP32ToISat(rB0:31, SIGN, UPPER, ROUND, I) rD32:63 ← CnvtFP32ToISat(rB32:63, SIGN, LOWER, ROUND, I) Each single-precision floating-point element in rB is converted to a signed integer using the current rounding mode and the result is saturated if it cannot be represented in a 32-bit integer. NaNs are converted as though they were zero. Exceptions: If the contents of either element of rB are Infinity, Denorm, or NaN, or if an overflow occurs on conversion, SPEFSCR[FINV,FINVH] are set appropriately, and SPEFSCR[FGH,FXH,FG,FX] are cleared appropriately. If SPEFSCR[FINVE] is set, an interrupt is taken, the destination register is not updated, and no other status bits are set. If either result element of this instruction is inexact and no other interrupt is taken, SPEFSCR[FINXS] is set. If the floating-point inexact exception is enabled, an interrupt is taken using the floating-point round interrupt vector. In this case, the destination register is updated with the truncated result. The FGH, FXH, FG and FX bits are properly updated to allow rounding to be performed in the interrupt handler. Vector SPFP APU User 0 5 6 1 01 1 1 51 6 2 02 1 3 1 000100 rD 00000 rB 0 1 0 1 0 0 1 0 1 0 1

Vector convert floating-point single-precision to signed integer with round toward zero evfsctsiz r D,rB rD0:31 ← CnvtFP32ToISat(rB0:31, SIGN, UPPER, TRUNC, I) rD32:63 ← CnvtFP32ToISat(rB32:63, SIGN, LOWER, TRUNC, I) Each single-precision floating-point element in rB is converted to a signed integer using the rounding mode Round toward Zero and the result is saturated if it cannot be represented in a 32-bit integer. NaNs are converted as though they were zero. Exceptions: If either element of rB is Infinity, Denorm, or NaN, or if an overflow occurs, SPEFSCR[FINV,FINVH] are set appropriately, and SPEFSCR[FGH,FXH,FG,FX] are cleared appropriately. If SPEFSCR[FINVE] is set, an interrupt is taken, the destination register is not updated, and no other status bits are set. If either result element of this instruction is inexact and no other interrupt is taken, SPEFSCR[FINXS] is set. If the floating-point inexact exception is enabled, an interrupt is taken using the floating-point round interrupt vector. In this case, the destination register is updated with the truncated result. The FGH, FXH, FG and FX bits are properly updated to allow rounding to be performed in the interrupt handler. Vector SPFP APU User 0 5 6 1 01 1 1 51 6 2 02 1 3 1 000100 rD 00000 rB 0 1 0 1 0 0 1 1 0 1 0

Vector convert floating-point single-precision to unsigned fraction evfsctuf r D,rB rD0:31 ← CnvtFP32ToISat(rB0:31, UNSIGN, UPPER, ROUND, F) rD32:63 ← CnvtFP32ToISat(rB32:63, UNSIGN, LOWER, ROUND, F) Each single-precision floating-point element in rB is converted to an unsigned fraction using the current rounding mode and the result is saturated if it cannot be represented in a 32-bit fraction. NaNs are converted as though they were zero. Exceptions: If either element of rB is Infinity, Denorm, or NaN, or if an overflow occurs, SPEFSCR[FINV,FINVH] are set appropriately, and SPEFSCR[FGH,FXH,FG,FX] are cleared appropriately. If SPEFSCR[FINVE] is set, an interrupt is taken, the destination register is not updated, and no other status bits are set. If either result element of this instruction is inexact and no other interrupt is taken, SPEFSCR[FINXS] is set. If the floating-point inexact exception is enabled, an interrupt is taken using the floating-point round interrupt vector. In this case, the destination register is updated with the truncated result. The FGH, FXH, FG and FX bits are properly updated to allow rounding to be performed in the interrupt handler. Vector SPFP APU User 0 5 6 1 01 1 1 51 6 2 02 1 3 1 000100 rD 00000 rB 0 1 0 1 0 0 1 0 1 1 0

Vector convert floating-point single-precision to unsigned integer evfsctui r D,rB rD0:31 ← CnvtFP32ToISat(rB0:31, UNSIGN, UPPER, ROUND, I) rD32:63 ← CnvtFP32ToISat(rB32:63, UNSIGN, LOWER, ROUND, I) Each single-precision floating-point element in rB is converted to an unsigned integer using the current rounding mode and the result is saturated if it cannot be represented in a 32-bit integer. NaNs are converted as though they were zero. Exceptions: If either element of rB is Infinity, Denorm, or NaN, or if an overflow occurs, SPEFSCR[FINV,FINVH] are set appropriately, and SPEFSCR[FGH,FXH,FG,FX] are cleared appropriately. If SPEFSCR[FINVE] is set, an interrupt is taken, the destination register is not updated, and no other status bits are set. If either result element of this instruction is inexact and no other interrupt is taken, SPEFSCR[FINXS] is set. If the floating-point inexact exception is enabled, an interrupt is taken using the floating-point round interrupt vector. In this case, the destination register is updated with the truncated result. The FGH, FXH, FG and FX bits are properly updated to allow rounding to be performed in the interrupt handler. Vector SPFP APU User 0 5 6 1 01 1 1 51 6 2 02 1 3 1 000100 rD 00000 rB 0 1 0 1 0 0 1 0 1 0 0

Vector convert floating-point single-precision to unsigned integer with round toward zero evfsctuiz r D,rB rD0:31 ← CnvtFP32ToISat(rB0:31, UNSIGN, UPPER, TRUNC, I) rD32:63 ← CnvtFP32ToISat(rB32:63, UNSIGN, LOWER, TRUNC, I) Each single-precision floating-point element in rB is converted to an unsigned integer using the rounding mode Round toward Zero and the result is saturated if it cannot be represented in a 32-bit integer. NaNs are converted as though they were zero. Exceptions: If either element of rB is Infinity, Denorm, or NaN, or if an overflow occurs, SPEFSCR[FINV,FINVH] are set appropriately, and SPEFSCR[FGH,FXH,FG,FX] are cleared appropriately. If SPEFSCR[FINVE] is set, an interrupt is taken, the destination register is not updated, and no other status bits are set. If either result element of this instruction is inexact and no other interrupt is taken, SPEFSCR[FINXS] is set. If the floating-point inexact exception is enabled, an interrupt is taken using the floating-point round interrupt vector. In this case, the destination register is updated with the truncated result. The FGH, FXH, FG and FX bits are properly updated to allow rounding to be performed in the interrupt handler. Vector SPFP APU User 0 5 6 1 01 1 1 51 6 2 02 1 3 1 000100 rD 00000 rB 0 1 0 1 0 0 1 1 0 0 0

Vector floating-point single-precision divide evfsdiv r D,rA,rB rD0:31 ← rA0:31 ÷sp rB0:31 rD32:63 ← rA32:63 ÷sp rB32:63 Each single-precision floating-point element of rA is divided by the corresponding element of rB and the result is stored in rD. If an element of rB is a NaN or infinity, the corresponding result is a properly signed zero. Otherwise, if an element of rB is a zero (or a denormalized number optionally transformed to zero by the implementation), or if an element of rA is either NaN or infinity, the corresponding result is either pmax (asign==bsign), or nmax (asign!=bsign). Otherwise, if an overflow occurs, pmax or nmax (as appropriate) is stored in the corresponding element of rD. If an underflow occurs, +0 or –0 (as appropriate) is stored in the corresponding element of rD. Exceptions: If the contents of rA or rB are Infinity, Denorm, or NaN, or if both rA and rB are ±0, SPEFSCR[FINV,FINVH] are set appropriately, and SPEFSCR[FGH,FXH,FG,FX] are cleared appropriately. If SPEFSCR[FINVE] is set, an interrupt is taken and the destination register is not updated. Otherwise, if the content of rB is ±0 and the content of rA is a finite normalized non-zero number, SPEFSCR[FDBZ,FDBZH] are set appropriately. If floating- point divide-by-zero exceptions are enabled, an interrupt is then taken. Otherwise, if an overflow occurs, SPEFSCR[FOVF ,FOVFH] are set appropriately, or if an underflow occurs, SPEFSCR[FUNF ,FUNFH] are set appropriately. If either underflow or overflow exceptions are enabled and a corresponding bit is set, an interrupt is taken. If any of these interrupts are taken, the destination register is not updated. If either result element of this instruction is inexact, or overflows but overflow exceptions are disabled, and no other interrupt is taken, or underflows but underflow exceptions are disabled, and no other interrupt is taken, SPEFSCR[FINXS] is set. If the floating-point inexact exception is enabled, an interrupt is taken using the floating-point round interrupt vector. In this case, the destination register is updated with the truncated result(s). The FG and FX bits are properly updated to allow rounding to be performed in the interrupt handler. FG and FX (FGH and FXH) are cleared if an overflow or underflow interrupt is taken, or if an invalid operation/input error is signaled for the low (high) element (regardless of FINVE). Vector SPFP APU User 0 5 6 1 01 1 1 51 6 2 02 1 3 1 000100 rD rA rB 01010001001

Vector floating-point single-precision multiply evfsmul r D,rA,rB rD0:31 ← rA0:31 ×sp rB0:31 rD32:63 ← rA32:63 ×sp rB32:63 Each single-precision floating-point element of rA is multiplied with the corresponding element of rB and the result is stored in rD. If an element of rA or rB are either zero (or a denormalized number optionally transformed to zero by the implementation), the corresponding result is a properly signed zero. Otherwise, if an element of rA or rB are either NaN or infinity, the corresponding result is either pmax (a sign==bsign), or nmax (asign!=bsign). Otherwise, if an overflow occurs, pmax or nmax (as appropriate) is stored in the corresponding element of rD. If an underflow occurs, +0 or –0 (as appropriate) is stored in the corresponding element of rD. Exceptions: If the contents of either element of rA or rB are Infinity, Denorm, or NaN, SPEFSCR[FINV,FINVH] are set appropriately, and SPEFSCR[FGH,FXH,FG,FX] are cleared appropriately. If SPEFSCR[FINVE] is set, an interrupt is taken and the destination register is not updated. Otherwise, if an overflow occurs, SPEFSCR[FOVF ,FOVFH] are set appropriately, or if an underflow occurs, SPEFSCR[FUNF ,FUNFH] are set appropriately. If either underflow or overflow exceptions are enabled and a corresponding status bit is set, an interrupt is taken. If any of these interrupts are taken, the destination register is not updated. If either result element of this instruction is inexact, or overflows but overflow exceptions are disabled, and no other interrupt is taken, or underflows but underflow exceptions are disabled, and no other interrupt is taken, SPEFSCR[FINXS] is set. If the floating-point inexact exception is enabled, an interrupt is taken using the floating-point round interrupt vector. In this case, the destination register is updated with the truncated result(s). The FG and FX bits are properly updated to allow rounding to be performed in the interrupt handler. FG and FX (FGH and FXH) are cleared if an overflow or underflow exception is taken, or if an invalid operation/input error is signaled for the low (high) element (regardless of FINVE). Vector SPFP APU User 0 5 6 1 01 1 1 51 6 2 02 1 3 1 000100 rD rA rB 01010001000

Vector floating-point single-precision negative absolute value evfsnabs r D,rA rD0:31 ← 0b1 || rA1:31 rD32:63 ← 0b1 || rA33:63 The sign bit of each element in rA is set to 1 and the results are placed into rD. Exceptions: Exception detection for embedded floating-point absolute value operations is implementation dependent. An implementation may choose to not detect exceptions and carry out the sign bit operation. If the implementation does not detect exceptions, or if exception detection is disabled, the computation can be carried out in one of two ways, as a sign bit operation ignoring the rest of the contents of the source register, or by examining the input and appropriately saturating the input prior to performing the operation. If an implementation chooses to handle exceptions, the exception is handled as follows: if the contents of either element of rA are Infinity, Denorm, or NaN, SPEFSCR[FINV,FINVH] are set appropriately, and SPEFSCR[FGH,FXH,FG,FX] are cleared appropriately. If floating- point invalid input exceptions are enabled then an interrupt is taken, and the destination register is not updated. Vector SPFP APU User 0 5 6 1 01 1 1 51 6 2 02 1 3 1 000100 rD rA 0000001010000101

Vector floating-point single-precision negate evfsneg r D,rA rD0:31 ← ¬rA0 || rA1:31 rD32:63 ← ¬rA32 || rA33:63 The sign bit of each element in rA is complemented and the results are placed into rD. Exceptions: Exception detection for embedded floating-point absolute value operations is implementation dependent. An implementation may choose to not detect exceptions and carry out the sign bit operation. If the implementation does not detect exceptions, or if exception detection is disabled, the computation can be carried out in one of two ways, as a sign bit operation ignoring the rest of the contents of the source register, or by examining the input and appropriately saturating the input prior to performing the operation. If an implementation chooses to handle exceptions, the exception is handled as follows: if the contents of either element of rA are Infinity, Denorm, or NaN, SPEFSCR[FINV,FINVH] are set appropriately, and SPEFSCR[FGH,FXH,FG,FX] are cleared appropriately. If floating- point invalid input exceptions are enabled then an interrupt is taken, and the destination register is not updated. Vector SPFP APU User 0 5 6 1 01 1 1 51 6 2 02 1 3 1 000100 rD rA 0000001010000110

Vector floating-point single-precision subtract evfssub r D,rA,rB rD0:31 ← rA0:31 -sp rB0:31 rD32:63 ← rA32:63 -sp rB32:63 Each single-precision floating-point element of rB is subtracted from the corresponding element of rA and the results are stored in rD. If an element of rA is NaN or infinity, the corresponding result is either pmax (asign==0), or nmax (asign==1). Otherwise, if an element of rB is NaN or infinity, the corresponding result is either nmax (bsign==0), or pmax (bsign==1). Otherwise, if an overflow occurs, pmax or nmax (as appropriate) is stored in the corresponding element of rD. If an underflow occurs, +0 (for rounding modes RN, RZ, RP) or –0 (for rounding mode RM) is stored in the corresponding element of rD. Exceptions: If the contents of either element of rA or rB are Infinity, Denorm, or NaN, SPEFSCR[FINV,FINVH] are set appropriately, and SPEFSCR[FGH,FXH,FG,FX] are cleared appropriately. If SPEFSCR[FINVE] is set, an interrupt is taken and the destination register is not updated. Otherwise, if an overflow occurs, SPEFSCR[FOVF ,FOVFH] are set appropriately, or if an underflow occurs, SPEFSCR[FUNF ,FUNFH] are set appropriately. If either underflow or overflow exceptions are enabled and a corresponding status bit is set, an interrupt is taken. If any of these interrupts are taken, the destination register is not updated. If either result element of this instruction is inexact, or overflows but overflow exceptions are disabled, and no other interrupt is taken, or underflows but underflow exceptions are disabled, and no other interrupt is taken, SPEFSCR[FINXS] is set. If the floating-point inexact exception is enabled, an interrupt is taken using the floating-point round interrupt vector. In this case, the destination register is updated with the truncated result(s). The FG and FX bits are properly updated to allow rounding to be performed in the interrupt handler. FG and FX (FGH and FXH) are cleared if an overflow or underflow interrupt is taken, or if an invalid operation/input error is signaled for the low (high) element (regardless of FINVE). Vector SPFP APU User 0 5 6 1 01 1 1 51 6 2 02 1 3 1 000100 rD rA rB 01010000001

Vector floating-point single-precision test equal evfststeq crf D,rA,rB ah ← rA0:31 al ← rA32:63 bh ← rB0:31 bl ← rB32:63 if (ah = bh) then ch ← 1 else ch ← 0 if (al = bl) then cl ← 1 else cl ← 0 Each element of rA is compared against the corresponding element of rB. If rA equals rB, the bit in crfD is set, otherwise it is cleared. Comparison ignores the sign of 0 (+0 = –0). The comparison proceeds after treating NaNs, Infinities, and Denorms as normalized numbers, using their values of ‘e’ and ‘f’ directly. No exceptions are taken during the execution of evfststeq. If strict IEEE 754 compliance is required, the program should use evfscmpeq. Implementation note: In an implementation, the execution of evfststeq is likely to be faster than the execution of evfscmpeq. Vector SPFP APU User 0 5 6 8 9 1 01 1 1 51 6 2 02 1 3 1 000100 crfD0 0 rA rB 0 1 0 1 0 0 1 1 1 1 0

Vector floating-point single-precision test greater than evfststgt crf D,rA,rB ah ← rA0:31 al ← rA32:63 bh ← rB0:31 bl ← rB32:63 if (ah > bh) then ch ← 1 else ch ← 0 if (al > bl) then cl ← 1 else cl ← 0 Each element of rA is compared against the corresponding element of rB. If rA is greater than rB, the bit in crfD is set, otherwise it is cleared. Comparison ignores the sign of 0 (+0 = –0). The comparison proceeds after treating NaNs, Infinities, and Denorms as normalized numbers, using their values of ‘e’ and ‘f’ directly. No exceptions are taken during the execution of evfststgt. If strict IEEE 754 compliance is required, the program should use evfscmpgt. Implementation note: In an implementation, the execution of evfststgt is likely to be faster than the execution of evfscmpgt. Vector SPFP APU User 0 5 6 8 9 1 01 1 1 51 6 2 02 1 3 1 000100 crfD0 0 rA rB 0 1 0 1 0 0 1 1 1 0 0

Vector floating-point single-precision test less than evfststlt crf D,rA,rB ah ← rA0:31 al ← rA32:63 bh ← rB0:31 bl ← rB32:63 if (ah < bh) then ch ← 1 else ch ← 0 if (al < bl) then cl ← 1 else cl ← 0 Each element of rA is compared with the corresponding element of rB. If rA is less than rB, the bit in the crfD is set, otherwise it is cleared. Comparison ignores the sign of 0 (+0 = –0). The comparison proceeds after treating NaNs, Infinities, and Denorms as normalized numbers, using their values of ‘e’ and ‘f’ directly. No exceptions are taken during the execution of evfststlt. If strict IEEE 754 compliance is required, the program should use evfscmplt. Implementation note: In an implementation, the execution of evfststlt is likely to be faster than the execution of evfscmplt. Vector SPFP APU User 0 5 6 8 9 1 01 1 1 51 6 2 02 1 3 1 000100 crfD0 0 rA rB 0 1 0 1 0 0 1 1 1 0 1

Vector load double into four half words evldh r D,d(rA) if (rA = 0) then b ← 0 else b ← (rA) EA ← b + EXTZ(UIMM*8) rD0:15 ← MEM(EA, 2) rD16:31 ← MEM(EA+2,2) rD32:47 ← MEM(EA+4,2) rD48:63 ← MEM(EA+6,2) The double word addressed by EA is loaded from memory and placed in rD. The figure below shows how bytes are loaded into rD as determined by the endian mode. evldh Results in Big- and Little-Endian Modes Implementation note: If the EA is not double-word aligned, an alignment exception occurs. SPE APU User 0 5 6 1 01 1 1 51 6 2 02 1 3 1 000100 rD rAU I M M (1) 01100000101 1. d = UIMM * 8 cdef h 01234567 ab g cdef hab g dc fe gba h Memory GPR in big endian GPR in little endian Byte address

Figure 67. High order element merging (evmergehilo) evmergehilo, and evmergelohi provide a full 32-bit permute of two source operands.

The low-order elements of rA and rB are merged and placed in rD, as shown in Figure 68. Figure 68. Low order element merging (evmergelo) Note: A vector splat low can be performed by specifying the same register in rA and rB.

Figure 69. Low order element merging (evmergelohi) Note: A vector swap can be performed by specifying the same register in rA and rB.

is placed into rD and the accumulator. overflow of the 64-bit sum is not recorded into the SPEFSCR. Figure 70. evmhegsmfaa (even form)

result is placed into rD and the accumulator. overflow of the 64-bit difference is not recorded into the SPEFSCR. Figure 71. evmhegsmfan (even form)

bit accumulator, and the resulting sum is placed into rD and into the accumulator. overflow of the 64-bit sum is not recorded into the SPEFSCR. Figure 72. evmhegsmiaa (even form)

the 64-bit accumulator, and the result is placed into rD and into the accumulator. performed. Any overflow of the 64-bit difference is not recorded into the SPEFSCR. Figure 73. evmhegsmian (even form)

64-bit accumulator. The resulting sum is placed into rD and into the accumulator. overflow of the 64-bit sum is not recorded into the SPEFSCR. Figure 74. evmhegumiaa (even form)

the 64-bit accumulator. The result is placed into rD and into the accumulator. performed. Any overflow of the 64-bit difference is not recorded into the SPEFSCR. Figure 75. evmhegumian (even form)

multiplied then placed into the corresponding words of rD. If A = 1, the result in rD is also placed into the accumulator. Figure 76. Even multiply of two signed modulo fractional elements (to

which are placed into the corresponding rD words and into the accumulator. Figure 77. Even form of vector half-word multiply (evmhesmfaaw)

which are placed into the corresponding rD words and into the accumulator. Figure 78. Even form of vector half-word multiply (evmhesmfanw)

multiplied. The two 32-bit products are placed into the corresponding words of rD. If A = 1, the result in rD is also placed into the accumulator. Figure 79. Even form for vector multiply (to accumulator) (evmhesmi)

placed into the corresponding rD words and into the accumulator. Figure 80. Even form of vector half-word multiply (evmhesmiaaw)

which are placed into the corresponding rD words and into the accumulator. Figure 81. Even form of vector half-word multiply (evmhesmianw)

Vector multiply half words, even, signed, saturate, fractional (to accumulator) evmhessf r D,rA,rB (A = 0) evmhessfa r D,rA,rB (A = 1) // high temp0:31 ← rA0:15 ×sf rB0:15 if (rA0:15 = 0x8000) & (rB0:15 = 0x8000) then rD0:31 ← 0x7FFF_FFFF //saturate movh ← 1 else rD0:31 ← temp0:31 movh ← 0 // low temp0:31 ← rA32:47 ×sf rB32:47 if (rA32:47 = 0x8000) & (rB32:47 = 0x8000) then rD32:63 ← 0x7FFF_FFFF //saturate movl ← 1 else rD32:63 ← temp0:31 movl ← 0 // update accumulator if A = 1 then ACC0:63 ← rD0:63 // update SPEFSCR SPEFSCROVH ← movh SPEFSCROV ← movl SPEFSCRSOVH ← SPEFSCRSOVH | movh SPEFSCRSOV ← SPEFSCRSOV | movl The corresponding even-numbered half-word signed fractional elements in rA and rB are multiplied. The 32 bits of each product are placed into the corresponding words of rD. If both inputs are –1.0, the result saturates to the largest positive signed fraction and the overflow and summary overflow bits are recorded in the SPEFSCR. If A = 1, the result in rD is also placed into the accumulator. Other registers altered: SPEFSCR ACC (If A = 1) SPE APU User 0 5 6 1 01 1 1 51 6 2 02 1 2 52 62 7 3 1 000100 rD rA rB 10000A00011

Figure 82. Even multiply of two signed saturate fractional elements (to

Vector multiply half words, even, signed, saturate, fractional and accumulate into words evmhessfaaw r D,rA,rB // high temp0:31 ← rA0:15 ×sf rB0:15 if (rA0:15 = 0x8000) & (rB0:15 = 0x8000) then temp0:31 ← 0x7FFF_FFFF //saturate movh ← 1 else movh ← 0 temp0:63 ← EXTS(ACC0:31) + EXTS(temp0:31) ovh ← (temp31 ⊕ temp32) rD0:31 ← SATURATE(ovh, temp31, 0x8000_0000, 0x7FFF_FFFF , temp32:63) // low temp0:31 ← rA32:47 ×sf rB32:47 if (rA32:47 = 0x8000) & (rB32:47 = 0x8000) then temp0:31 ← 0x7FFF_FFFF //saturate movl ← 1 else movl ← 0 temp0:63 ← EXTS(ACC32:63) + EXTS(temp0:31) ovl ← (temp31 ⊕ temp32) rD32:63 ← SATURATE(ovl, temp31, 0x8000_0000, 0x7FFF_FFFF , temp32:63) // update accumulator ACC0:63 ← rD0:63 // update SPEFSCR SPEFSCROVH ← movh SPEFSCROV ← movl SPEFSCRSOVH ← SPEFSCRSOVH | ovh | movh SPEFSCRSOV ← SPEFSCRSOV | ovl| movl The corresponding even-numbered half-word signed fractional elements in rA and rB are multiplied producing a 32-bit product. If both inputs are –1.0, the result saturates to 0x7FFF_FFFF . Each 32-bit product is then added to the corresponding word in the accumulator, saturating if overflow or underflow occurs, and the result is placed in rD and the accumulator. If there is an overflow or underflow from either the multiply or the addition, the overflow and summary overflow bits are recorded in the SPEFSCR. Other registers altered: SPEFSCR ACC SPE APU User 0 5 6 1 01 1 1 51 6 2 02 1 3 1 000100 rD rA rB 10100000011

Figure 83. Even form of vector half-word multiply (evmhessfaaw)

summary overflow bits are recorded in the SPEFSCR. Figure 84. Even form of vector half-word multiply (evmhessfanw)

Vector multiply half words, even, signed, saturate, integer and accumulate into words evmhessiaaw r D,rA,rB // high temp0:31 ← rA0:15 ×si rB0:15 temp0:63 ← EXTS(ACC0:31) + EXTS(temp0:31) ovh ← (temp31 ⊕ temp32) rD0:31 ← SATURATE(ovh, temp31, 0x8000_0000, 0x7FFF_FFFF , temp32:63) // low temp0:31 ← rA32:47 ×si rB32:47 temp0:63 ← EXTS(ACC32:63) + EXTS(temp0:31) ovl ← (temp31 ⊕ temp32) rD32:63 ← SATURATE(ovl, temp31, 0x8000_0000, 0x7FFF_FFFF , temp32:63) // update accumulator ACC0:63 ← rD0:63 // update SPEFSCR SPEFSCROVH ← ovh SPEFSCROV ← ovl SPEFSCRSOVH ← SPEFSCRSOVH | ovh SPEFSCRSOV ← SPEFSCRSOV | ovl The corresponding even-numbered half-word signed integer elements in rA and rB are multiplied producing a 32-bit product. Each 32-bit product is then added to the corresponding word in the accumulator, saturating if overflow occurs, and the result is placed in rD and the accumulator. If there is an overflow or underflow from either the multiply or the addition, the overflow and summary overflow bits are recorded in the SPEFSCR. Other registers altered: SPEFSCR ACC SPE APU User 0 5 6 1 01 1 1 51 6 2 02 1 3 1 000100 rD rA rB 10100000001

Figure 85. Even form of vector half-word multiply (evmhessiaaw)

Vector multiply half words, even, signed, saturate, integer and accumulate negative into words evmhessianw r D,rA,rB // high temp0:31 ← rA0:15 ×si rB0:15 temp0:63 ← EXTS(ACC0:31) - EXTS(temp0:31) ovh ← (temp31 ⊕ temp32) rD0:31 ← SATURATE(ovh, temp31, 0x8000_0000, 0x7FFF_FFFF , temp32:63) // low temp0:31 ← rA32:47 ×si rB32:47 temp0:63 ← EXTS(ACC32:63) - EXTS(temp0:31) ovl ← (temp31 ⊕ temp32) rD32:63 ← SATURATE(ovl, temp31, 0x8000_0000, 0x7FFF_FFFF , temp32:63) // update accumulator ACC0:63 ← rD0:63 // update SPEFSCR SPEFSCROVH ← ovh SPEFSCROV ← ovl SPEFSCRSOVH ← SPEFSCRSOVH | ovh SPEFSCRSOV ← SPEFSCRSOV | ovl The corresponding even-numbered half-word signed integer elements in rA and rB are multiplied producing a 32-bit product. Each 32-bit product is then subtracted from the corresponding word in the accumulator, saturating if overflow occurs, and the result is placed in rD and the accumulator. If there is an overflow or underflow from either the multiply or the addition, the overflow and summary overflow bits are recorded in the SPEFSCR. Other registers altered: SPEFSCR ACC SPE APU User 0 5 6 1 01 1 1 51 6 2 02 1 3 1 000100 rD rA rB 10110000001

Figure 86. Even form of vector half-word multiply (evmhessianw)

multiplied. The two 32-bit products are placed into the corresponding words of rD. If A = 1, the result in rD is also placed into the accumulator. Figure 87. Vector multiply half words, even, unsigned, modulo, integer (to

corresponding rD and accumulator words. Figure 88. Even form of vector half-word multiply (evmheumiaaw)

placed into the corresponding rD and accumulator words. Figure 89. Even form of vector half-word multiply (evmheumianw)

Vector multiply half words, even, unsigned, saturate, integer and accumulate into words evmheusiaaw r D,rA,rB // high temp0:31 ← rA0:15 ×ui rB0:15 temp0:63 ← EXTZ(ACC0:31) + EXTZ(temp0:31) ovh ← temp31 rD0:31 ← SATURATE(ovh, 0, 0xFFFF_FFFF , 0xFFFF_FFFF , temp32:63) //low temp0:31 ← rA32:47 ×ui rB32:47 temp0:63 ← EXTZ(ACC32:63) + EXTZ(temp0:31) ovl ← temp31 rD32:63 ← SATURATE(ovl, 0, 0xFFFF_FFFF , 0xFFFF_FFFF , temp32:63) // update accumulator ACC0:63 ← rD0:63 // update SPEFSCR SPEFSCROVH ← ovh SPEFSCROV ← ovl SPEFSCRSOVH ← SPEFSCRSOVH | ovh SPEFSCRSOV ← SPEFSCRSOV | ovl For each word element in the accumulator, corresponding even-numbered half-word unsigned integer elements in rA and rB are multiplied producing a 32-bit product. Each 32- bit product is then added to the corresponding word in the accumulator, saturating if overflow occurs, and the result is placed in rD and the accumulator. If the addition causes overflow, the overflow and summary overflow bits are recorded in the SPEFSCR. Other registers altered: SPEFSCR ACC SPE APU User 0 5 6 1 01 1 1 51 6 2 02 1 3 1 000100 rD rA rB 10100000000

Figure 90. Even form of vector half-word multiply (evmheusiaaw)

Vector multiply half words, even, unsigned, saturate, integer and accumulate negative into words evmheusianw r D,rA,rB // high temp0:31 ← rA0:15 ×ui rB0:15 temp0:63 ← EXTZ(ACC0:31) - EXTZ(temp0:31) ovh ← temp31 rD0:31 ← SATURATE(ovh, 0, 0x0000_0000, 0x0000_0000, temp32:63) //low temp0:31 ← rA32:47 ×ui rB32:47 temp0:63 ← EXTZ(ACC32:63) - EXTZ(temp0:31) ovl ← temp31 rD32:63 ← SATURATE(ovl, 0, 0x0000_0000, 0x0000_0000, temp32:63) // update accumulator ACC0:63 ← rD0:63 // update SPEFSCR SPEFSCROVH ← ovh SPEFSCROV ← ovl SPEFSCRSOVH ← SPEFSCRSOVH | ovh SPEFSCRSOV ← SPEFSCRSOV | ovl For each word element in the accumulator, corresponding even-numbered half-word unsigned integer elements in rA and rB are multiplied producing a 32-bit product. Each 32- bit product is then subtracted from the corresponding word in the accumulator, saturating if underflow occurs, and the result is placed in rD and the accumulator. If there is an underflow from the subtraction, the overflow and summary overflow bits are recorded in the SPEFSCR. Other registers altered: SPEFSCR ACC SPE APU User 0 5 6 1 01 1 1 51 6 2 02 1 3 1 000100 r D r A r B 10110000000

Figure 91. Even form of vector half-word multiply (evmheusianw)

of the 64-bit accumulator, and the result is placed into rD and into the accumulator. Note: This is a modulo sum. There is no check for overflow and no saturation is performed. An overflow from the 64-bit sum, if one occurs, is not recorded into the SPEFSCR. Figure 92. evmhogsmfaa (odd form)

contents of the 64-bit accumulator, and the result is placed into rD and into the accumulator. Note: This is a modulo difference. There is no check for overflow and no saturation is performed. Any overflow of the 64-bit difference is not recorded into the SPEFSCR. Figure 93. evmhogsmfan (odd form)

of the 64-bit accumulator, and the result is placed into rD and into the accumulator. overflow from the 64-bit sum, if one occurs, is not recorded into the SPEFSCR. Figure 94. evmhogsmiaa (odd form)

contents of the 64-bit accumulator, and the result is placed into rD and into the accumulator. performed. Any overflow of the 64-bit difference is not recorded into the SPEFSCR. Figure 95. evmhogsmian (odd form)

of the 64-bit accumulator, and the result is placed into rD and into the accumulator. Note: This is a modulo sum. There is no check for overflow and no saturation is performed. An overflow from the 64-bit sum, if one occurs, is not recorded into the SPEFSCR. Figure 96. evmhogumiaa (odd form)

contents of the 64-bit accumulator, and the result is placed into rD and into the accumulator. performed. Any overflow of the 64-bit difference is not recorded into the SPEFSCR. Figure 97. evmhogumian (odd form)

multiplied. Each product is placed into the corresponding words of rD. If A = 1, the result in rD is also placed into the accumulator. Figure 98. Vector multiply half words, odd, signed, modulo, fractional (to

Figure 99. Odd form of vector half-word multiply (evmhosmfaaw)

results are placed into the corresponding rD words and into the accumulator. Figure 100. Odd form of vector half-word multiply (evmhosmfanw)

multiplied. The two 32-bit products are placed into the corresponding words of rD. If A = 1, the result in rD is also placed into the accumulator. Figure 101. Vector multiply half words, odd, signed, modulo, integer (to

into the corresponding rD words and into the accumulator. Figure 102. Odd form of vector half-word multiply (evmhosmiaaw)

placed into the corresponding rD words and into the accumulator. Figure 103. Odd form of vector half-word multiply (evmhosmianw)

Vector multiply half words, odd, signed, saturate, fractional (to accumulator) evmhossf r D,rA,rB (A = 0) evmhossfa r D,rA,rB (A = 1) // high temp0:31 ← rA16:31 ×sf rB16:31 if (rA16:31 = 0x8000) & (rB16:31 = 0x8000) then rD0:31 ← 0x7FFF_FFFF //saturate movh ← 1 else rD0:31 ← temp0:31 movh ← 0 // low temp0:31 ← rA48:63 ×sf rB48:63 if (rA48:63 = 0x8000) & (rB48:63 = 0x8000) then rD32:63 ← 0x7FFF_FFFF //saturate movl ← 1 else rD32:63 ← temp0:31 movl ← 0 // update accumulator if A = 1 then ACC0:63 ← rD0:63 // update SPEFSCR SPEFSCROVH ← movh SPEFSCROV ← movl SPEFSCRSOVH ← SPEFSCRSOVH | movh SPEFSCRSOV ← SPEFSCRSOV | movl The corresponding odd-numbered half-word signed fractional elements in rA and rB are multiplied. The 32 bits of each product are placed into the corresponding words of rD. If both inputs are –1.0, the result saturates to the largest positive signed fraction and the overflow and summary overflow bits are recorded in the SPEFSCR. If A = 1, the result in rD is also placed into the accumulator. Other registers altered: SPEFSCR ACC (If A = 1) SPE APU User 0 5 6 1 01 1 1 51 6 2 02 1 2 52 62 7 3 1 000100 rD rA rB 10000A00111

Figure 104. Vector multiply half words, odd, signed, saturate, fractional (to

Vector multiply half words, odd, signed, saturate, fractional and accumulate into words evmhossfaaw r D,rA,rB // high temp0:31 ← rA16:31 ×sf rB16:31 if (rA16:31 = 0x8000) & (rB16:31 = 0x8000) then temp0:31 ← 0x7FFF_FFFF //saturate movh ← 1 else movh ← 0 temp0:63 ← EXTS(ACC0:31) + EXTS(temp0:31) ovh ← (temp31 ⊕ temp32) rD0:31 ← SATURATE(ovh, temp31, 0x8000_0000, 0x7FFF_FFFF , temp32:63) // low temp0:31 ← rA48:63 ×sf rB48:63 if (rA48:63 = 0x8000) & (rB48:63 = 0x8000) then temp0:31 ← 0x7FFF_FFFF //saturate movl ← 1 else movl ← 0 temp0:63 ← EXTS(ACC32:63) + EXTS(temp0:31) ovl ← (temp31 ⊕ temp32) rD32:63 ← SATURATE(ovl, temp31, 0x8000_0000, 0x7FFF_FFFF , temp32:63) // update accumulator ACC0:63 ← rD0:63 // update SPEFSCR SPEFSCROVH ← movh SPEFSCROV ← movl SPEFSCRSOVH ← SPEFSCRSOVH | ovh | movh SPEFSCRSOV ← SPEFSCRSOV | ovl| movl The corresponding odd-numbered half-word signed fractional elements in rA and rB are multiplied producing a 32-bit product. If both inputs are –1.0, the result saturates to 0x7FFF_FFFF . Each 32-bit product is then added to the corresponding word in the accumulator, saturating if overflow or underflow occurs, and the result is placed in rD and the accumulator. If there is an overflow or underflow from either the multiply or the addition, the overflow and summary overflow bits are recorded in the SPEFSCR. Other registers altered: SPEFSCR ACC SPE APU User 0 5 6 1 01 1 1 51 6 2 02 1 3 1 000100 rD rA rB 10100000111

Figure 105. Odd form of vector half-word multiply (evmhossfaaw)

Vector multiply half words, odd, signed, saturate, fractional and accumulate negative into words evmhossfanw r D,rA,rB // high temp0:31 ← rA16:31 ×sf rB16:31 if (rA16:31 = 0x8000) & (rB16:31 = 0x8000) then temp0:31 ← 0x7FFF_FFFF //saturate movh ← 1 else movh ← 0 temp0:63 ← EXTS(ACC0:31) - EXTS(temp0:31) ovh ← (temp31 ⊕ temp32) rD0:31 ← SATURATE(ovh, temp31, 0x8000_0000, 0x7FFF_FFFF , temp32:63) // low temp0:31 ← rA48:63 ×sf rB48:63 if (rA48:63 = 0x8000) & (rB48:63 = 0x8000) then temp0:31 ← 0x7FFF_FFFF //saturate movl ← 1 else movl ← 0 temp0:63 ← EXTS(ACC32:63) - EXTS(temp0:31) ovl ← (temp31 ⊕ temp32) rD32:63 ← SATURATE(ovl, temp31, 0x8000_0000, 0x7FFF_FFFF , temp32:63) // update accumulator ACC0:63 ← rD0:63 // update SPEFSCR SPEFSCROVH ← movh SPEFSCROV ← movl SPEFSCRSOVH ← SPEFSCRSOVH | ovh | movh SPEFSCRSOV ← SPEFSCRSOV | ovl| movl The corresponding odd-numbered half-word signed fractional elements in rA and rB are multiplied producing a 32-bit product. If both inputs are –1.0, the result saturates to 0x7FFF_FFFF . Each 32-bit product is then subtracted from the corresponding word in the accumulator, saturating if overflow or underflow occurs, and the result is placed in rD and the accumulator. If there is an overflow or underflow from either the multiply or the subtraction, the overflow and summary overflow bits are recorded in the SPEFSCR. Other registers altered: SPEFSCR ACC SPE APU User 0 5 6 1 01 1 1 51 6 2 02 1 3 1 000100 rD rA rB 10110000111

Figure 106. odd Form of Vector Half-Word Multiply (evmhossfanw)

Vector multiply half words, odd, signed, saturate, integer and accumulate into words evmhossiaaw r D,rA,rB // high temp0:31 ← rA16:31 ×si rB16:31 temp0:63 ← EXTS(ACC0:31) + EXTS(temp0:31) ovh ← (temp31 ⊕ temp32) rD0:31 ← SATURATE(ovh, temp31, 0x8000_0000, 0x7FFF_FFFF , temp32:63) // low temp0:31 ← rA48:63 ×si rB48:63 temp0:63 ← EXTS(ACC32:63) + EXTS(temp0:31) ovl ← (temp31 ⊕ temp32) rD32:63 ← SATURATE(ovl, temp31, 0x8000_0000, 0x7FFF_FFFF , temp32:63) // update accumulator ACC0:63 ← rD0:63 // update SPEFSCR SPEFSCROVH ← ovh SPEFSCROV ← ovl SPEFSCRSOVH ← SPEFSCRSOVH | ovh SPEFSCRSOV ← SPEFSCRSOV | ovl The corresponding odd-numbered half-word signed integer elements in rA and rB are multiplied producing a 32-bit product. Each 32-bit product is then added to the corresponding word in the accumulator, saturating if overflow occurs, and the result is placed in rD and the accumulator. If there is an overflow or underflow from the addition, the overflow and summary overflow bits are recorded in the SPEFSCR. Other registers altered: SPEFSCR ACC SPE APU User 0 5 6 1 01 1 1 51 6 2 02 1 3 1 000100 rD rA rB 10100000101

Figure 107. Odd form of vector half-word multiply (evmhossiaaw)

Vector multiply half words, odd, signed, saturate, integer and accumulate negative into words evmhossianw r D,rA,rB // high temp0:31 ← rA16:31 ×si rB16:31 temp0:63 ← EXTS(ACC0:31) - EXTS(temp0:31) ovh ← (temp31 ⊕ temp32) rD0:31 ← SATURATE(ovh, temp31, 0x8000_0000, 0x7FFF_FFFF , temp32:63) // low temp0:31 ← rA48:63 ×si rB48:63 temp0:63 ← EXTS(ACC32:63) - EXTS(temp0:31) ovl ← (temp31 ⊕ temp32) rD32:63 ← SATURATE(ovl, temp31, 0x8000_0000, 0x7FFF_FFFF , temp32:63) // update accumulator ACC0:63 ← rD0:63 // update SPEFSCR SPEFSCROVH ← ovh SPEFSCROV ← ovl SPEFSCRSOVH ← SPEFSCRSOVH | ovh SPEFSCRSOV ← SPEFSCRSOV | ovl The corresponding odd-numbered half-word signed integer elements in rA and rB are multiplied producing a 32-bit product. Each 32-bit product is then subtracted from the corresponding word in the accumulator, saturating if overflow occurs, and the result is placed in rD and the accumulator. If there is an overflow or underflow from the subtraction, the overflow and summary overflow bits are recorded in the SPEFSCR. Other registers altered: SPEFSCR ACC SPE APU User 0 5 6 1 01 1 1 51 6 2 02 1 3 1 000100 rD rA rB 10110000101

Figure 108. Odd form of vector half-word multiply (evmhossianw)

multiplied. The two 32-bit products are placed into the corresponding words of rD. If A = 1, the result in rD is also placed into the accumulator. Figure 109. Vector multiply half words, odd, unsigned, modulo, integer (to

corresponding rD and accumulator words. Figure 110. Odd form of vector half-word multiply (evmhoumiaaw)

into the corresponding rD and accumulator words. Figure 111. Odd form of vector half-word multiply (evmhoumianw)

Vector multiply half words, odd, unsigned, saturate, integer and accumulate into words evmhousiaaw r D,rA,rB // high temp0:31 ← rA16:31 ×ui rB16:31 temp0:63 ← EXTZ(ACC0:31) + EXTZ(temp0:31) ovh ← temp31 rD0:31 ← SATURATE(ovh, 0, 0xFFFF_FFFF , 0xFFFF_FFFF , temp32:63) //low temp0:31 ← rA48:63 ×ui rB48:63 temp0:63 ← EXTZ(ACC32:63) + EXTZ(temp0:31) ovl ← temp31 rD32:63 ← SATURATE(ovl, 0, 0xFFFF_FFFF , 0xFFFF_FFFF , temp32:63) // update accumulator ACC0:63 ← rD0:63 // update SPEFSCR SPEFSCROVH ← ovh SPEFSCROV ← ovl SPEFSCRSOVH ← SPEFSCRSOVH | ovh SPEFSCRSOV ← SPEFSCRSOV | ovl For each word element in the accumulator, corresponding odd-numbered half-word unsigned integer elements in rA and rB are multiplied producing a 32-bit product. Each 32- bit product is then added to the corresponding word in the accumulator, saturating if overflow occurs, and the result is placed in rD and the accumulator. If the addition causes overflow, the overflow and summary overflow bits are recorded in the SPEFSCR. Other registers altered: SPEFSCR ACC SPE APU User 0 5 6 1 01 1 1 51 6 2 02 1 3 1 000100 rD rA rB 10100000100

Figure 112. Odd form of vector half-word multiply (evmhousiaaw)

Vector multiply half words, odd, unsigned, saturate, integer and accumulate negative into words evmhousianw r D,rA,rB // high temp0:31 ← rA16:31 ×ui rB16:31 temp0:63 ← EXTZ(ACC0:31) - EXTZ(temp0:31) ovh ← temp31 rD0:31 ← SATURATE(ovh, 0, 0xFFFF_FFFF , 0xFFFF_FFFF , temp32:63) //low temp0:31 ← rA48:63 ×ui rB48:63 temp0:63 ← EXTZ(ACC32:63) - EXTZ(temp0:31) ovl ← temp31 rD32:63 ← SATURATE(ovl, 0, 0xFFFF_FFFF , 0xFFFF_FFFF , temp32:63) // update accumulator ACC0:63 ← rD0:63 // update SPEFSCR SPEFSCROVH ← ovh SPEFSCROV ← ovl SPEFSCRSOVH ← SPEFSCRSOVH | ovh SPEFSCRSOV ← SPEFSCRSOV | ovl For each word element in the accumulator, corresponding odd-numbered half-word unsigned integer elements in rA and rB are multiplied producing a 32-bit product. Each 32- bit product is then subtracted from the corresponding word in the accumulator, saturating if overflow occurs, and the result is placed in rD and the accumulator. If subtraction causes overflow, the overflow and summary overflow bits are recorded in the SPEFSCR. Other registers altered: SPEFSCR ACC SPE APU User 0 5 6 1 01 1 1 51 6 2 02 1 3 1 000100 rD rA rB 10110000100

Figure 113. Odd form of vector half-word multiply (evmhousianw)

for initializing the accumulator. Figure 114. Initialize accumulator (evmra)

31 of the two products are placed into the two corresponding words of rD. If A = 1, the result in rD is also placed into the accumulator. Figure 115. Vector multiply word high signed, modulo, fractional (to accumulator)

the two 64-bit products are placed into the two corresponding words of rD. If A = 1,The result in rD is also placed into the accumulator. Figure 116. Vector multiply word high signed, modulo, integer (to accumulator)

Vector multiply word high signed, saturate, fractional (to accumulator) evmwhssf r D,rA,rB (A = 0) evmwhssfa r D,rA,rB (A = 1) // high temp0:63 ← rA0:31 ×sf rB0:31 if (rA0:31 = 0x8000_0000) & (rB0:31 = 0x8000_0000) then rD0:31 ← 0x7FFF_FFFF //saturate movh ← 1 else rD0:31 ← temp0:31 movh ← 0 // low temp0:63 ← rA32:63 ×sf rB32:63 if (rA32:63 = 0x8000_0000) & (rB32:63 = 0x8000_0000) then rD32:63 ← 0x7FFF_FFFF //saturate movl ← 1 else rD32:63 ← temp0:31 movl ← 0 // update accumulator if A = 1 then ACC0:63 ← rD0:63 // update SPEFSCR SPEFSCROVH ← movh SPEFSCROV ← movl SPEFSCRSOVH ← SPEFSCRSOVH | movh SPEFSCRSOV ← SPEFSCRSOV | movl The corresponding word signed fractional elements in rA and rB are multiplied. Bits 0–31 of each product are placed into the corresponding words of rD. If both inputs are –1.0, the result saturates to the largest positive signed fraction and the overflow and summary overflow bits are recorded in the SPEFSCR. Other registers altered: SPEFSCR ACC (If A = 1) SPE APU User 0 5 6 1 01 1 1 51 6 2 02 1 2 52 62 7 3 1 000100 rD rA rB 10001A00111

Figure 117. Vector multiply word high signed, saturate, fractional (to accumulator)

the two products are placed into the two corresponding words of rD. If A = 1, the result in rD is also placed into the accumulator. Figure 118. Vector multiply word high unsigned, modulo, integer (to accumulator)

Figure 119. Vector multiply word low signed, modulo, integer & accumulate in words

Figure 120. Vector multiply word low signed, modulo, integer and accumulate

Vector multiply word low signed, saturate, integer and accumulate in words evmwlssiaaw r D,rA,rB // high temp0:63 ← rA0:31 ×si rB0:31 temp0:63 ← EXTS(ACC0:31) + EXTS(temp32:63) ovh ← (temp31 ⊕ temp32) rD0:31 ← SATURATE(ovh, temp31, 0x8000_0000, 0x7FFF_FFFF , temp32:63) // low temp0:63 ← rA32:63 ×si rB32:63 temp0:63 ← EXTS(ACC32:63) + EXTS(temp32:63) ovl ← (temp31 ⊕ temp32) rD32:63 ← SATURATE(ovl, temp31, 0x8000_0000, 0x7FFF_FFFF , temp32:63) // update accumulator ACC0:63 ← rD0:63 // update SPEFSCR SPEFSCROVH ← ovh SPEFSCROV ← ovl SPEFSCRSOVH ← SPEFSCRSOVH | ovh SPEFSCRSOV ← SPEFSCRSOV | ovl The corresponding word signed integer elements in rA and rB are multiplied producing a 64- bit product. The 32 lsbs of each product is added to the corresponding word in the ACC, saturating if overflow or underflow occurs; the result is placed in rD and the accumulator. If there is overflow or underflow from the addition, overflow and summary overflow bits are recorded in the SPEFSCR. Other registers altered: SPEFSCR ACC SPE APU User 0 5 6 1 01 1 1 51 6 2 02 1 3 1 000100 rD rA rB 10101000001

Figure 121. Vector multiply word low signed, saturate, integer & accumulate in words

Vector multiply word low signed, saturate, integer and accumulate negative in words evmwlssianw r D,rA,rB // high temp0:63 ← rA0:31 ×si rB0:31 temp0:63 ← EXTS(ACC0:31) - EXTS(temp32:63) ovh ← (temp31 ⊕ temp32) rD0:31 ← SATURATE(ovh, temp31, 0x8000_0000, 0x7FFF_FFFF , temp32:63) // low temp0:63 ← rA32:63 ×si rB32:63 temp0:63 ← EXTS(ACC32:63) - EXTS(temp32:63) ovl ← (temp31 ⊕ temp32) rD32:63 ← SATURATE(ovl, temp31, 0x8000_0000, 0x7FFF_FFFF , temp32:63) // update accumulator ACC0:63 ← rD0:63 // update SPEFSCR SPEFSCROVH ← ovh SPEFSCROV ← ovl SPEFSCRSOVH ← SPEFSCRSOVH | ovh SPEFSCRSOV ← SPEFSCRSOV | ovl The corresponding word signed integer elements in rA and rB are multiplied producing a 64- bit product. The 32 lsbs of each product are subtracted from the corresponding ACC word, saturating if overflow or underflow occurs, and the result is placed in rD and the ACC. If addition causes overflow or underflow, overflow and summary overflow SPEFSCR bits are recorded. Other registers altered: SPEFSCR ACC SPE APU User 0 5 6 1 01 1 1 51 6 2 02 1 3 1 000100 rD rA rB 10111000001

Figure 122. Vector multiply word low signed, saturate, integer & accumulate negative

significant 32 bits of each product are placed into the two corresponding words of rD. elements in rA and rB are treated as signed or unsigned 32-bit integers. If A = 1, the result in rD is also placed into the accumulator. Note that evmwlumi and evmwlumia can be used for signed or unsigned integers. Figure 123. Vector multiply word low unsigned, modulo, integer (evmwlumi)

Figure 124. Vector multiply word low unsigned, modulo, integer & accumulate in

into rD and the accumulator. Figure 125. Vector multiply word low unsigned, modulo, integer & accumulate

Vector multiply word low unsigned, saturate, integer and accumulate in words evmwlusiaaw r D,rA,rB // high temp0:63 ← rA0:31 ×ui rB0:31 temp0:63 ← EXTZ(ACC0:31) + EXTZ(temp32:63) ovh ← temp31 rD0:31 ← SATURATE(ovh, 0, 0xFFFF_FFFF , 0xFFFF_FFFF , temp32:63) //low temp0:63 ← rA32:63 ×ui rB32:63 temp0:63 ← EXTZ(ACC32:63) + EXTZ(temp32:63) ovl ← temp31 rD32:63 ← SATURATE(ovl, 0, 0xFFFF_FFFF , 0xFFFF_FFFF , temp32:63) // update accumulator ACC0:63 ← rD0:63 // update SPEFSCR SPEFSCROVH ← ovh SPEFSCROV ← ovl SPEFSCRSOVH ← SPEFSCRSOVH | ovh SPEFSCRSOV ← SPEFSCRSOV | ovl For each word element in the ACC, corresponding word unsigned integer elements in rA and rB are multiplied, producing a 64-bit product. The 32 lsbs of each product are added to the corresponding ACC word, saturating if overflow occurs; the result is placed in rD and the ACC. If the addition causes overflow, the overflow and summary overflow bits are recorded in the SPEFSCR. Other registers altered: SPEFSCR ACC SPE APU User 0 5 6 1 01 1 1 51 6 2 02 1 3 1 000100 rD rA rB 10101000000

Figure 126. Vector multiply word low unsigned, saturate, integer & accumulate in

Vector multiply word low unsigned, saturate, integer and accumulate negative in words evmwlusianw r D,rA,rB // high temp0:63 ← rA0:31 ×ui rB0:31 temp0:63 ← EXTZ(ACC0:31) - EXTZ(temp32:63) ovh ← temp31 rD0:31 ← SATURATE(ovh, 0, 0x0000_0000, 0x0000_0000, temp32:63) //low temp0:63 ← rA32:63 ×ui rB32:63 temp0:63 ← EXTZ(ACC32:63) - EXTZ(temp32:63) ovl ← temp31 rD32:63 ← SATURATE(ovl, 0, 0x0000_0000, 0x0000_0000, temp32:63) // update accumulator ACC 0:63 ← rD0:63 // update SPEFSCR SPEFSCROVH ← ovh SPEFSCROV ← ovl SPEFSCRSOVH ← SPEFSCRSOVH | ovh SPEFSCRSOV ← SPEFSCRSOV | ovl For each ACC word element, corresponding word elements in rA and rB are multiplied producing a 64-bit product. The 32 lsbs of each product are subtracted from corresponding ACC words, saturating if underflow occurs; the result is placed in rD and the ACC. If there is an underflow from the subtraction, the overflow and summary overflow bits are recorded in the SPEFSCR. Other registers altered: SPEFSCR ACC SPE APU User 0 5 6 1 01 1 1 51 6 2 02 1 3 1 000100 rD rA rB 10111000000

Figure 127. Vector multiply word low unsigned, saturate, integer & accumulate

If A = 1, the result in rD is also placed into the accumulator. Figure 128. Vector multiply word signed, modulo, fractional (to accumulator)

placed in rD and the accumulator. Figure 129. Vector multiply word signed, modulo, fractional & accumulate

placed in rD and the accumulator. Figure 130. Vector multiply word signed, modulo, fractional & accumulate negative

If A = 1, the result in rD is also placed into the accumulator. Figure 131. Vector multiply word signed, modulo, integer (to accumulator) (evmwsmi)

Figure 132. Vector multiply word signed, modulo, integer & accumulate (evmwsmiaa)

Figure 133. Vector multiply word signed, modulo, integer & accumulate negative

fraction and the overflow and summary overflow bits are recorded in the SPEFSCR. SPEFSCR[OV] should be set (along with the SOV bit, if it is not already set). If A = 1, the result in rD is also placed into the accumulator. Figure 134. Vector multiply word signed, saturate, fractional (to accumulator)

product. If both inputs are –1.0, the product saturates to the largest positive signed fraction. The 64-bit product is added to the ACC and the result is placed in rD and the ACC. summary overflow bits are recorded. Note: There is no saturation on the addition with the accumulator. Figure 135. Vector multiply word signed, saturate, fractional, & accumulate

Vector multiply word signed, saturate, fractional and accumulate negative evmwssfan r D,rA,rB temp0:63 ← rA32:63 ×sf rB32:63 if (rA32:63 = 0x8000_0000) & (rB32:63 = 0x8000_0000) then temp0:63 ← 0x7FFF_FFFF_FFFF_FFFF //saturate mov ← 1 else mov ← 0 temp0:64 ← EXTS(ACC0:63) - EXTS(temp0:63) ov ← (temp0 ⊕ temp1) rD0:63 ← temp1:64 ) // update accumulator ACC0:63 ← rD0:63 // update SPEFSCR SPEFSCROVH ← 0 SPEFSCROV ← mov SPEFSCRSOV ← SPEFSCRSOV | ov | mov The low word signed fractional elements in rA and rB are multiplied producing a 64-bit product. If both inputs are –1.0, the product saturates to the largest positive signed fraction. The 64-bit product is subtracted from the ACC and the result is placed in rD and the ACC. If there is an overflow from either the multiply or the addition, the SPEFSCR overflow and summary overflow bits are recorded. Note: There is no saturation on the subtraction with the accumulator. Other registers altered: SPEFSCR ACC SPE APU User 0 5 6 1 01 1 1 51 6 2 02 1 3 1 000100 rD rA rB 10111010011

Figure 136. Vector multiply word signed, saturate, fractional & accumulate negative

If A = 1, the result in rD is also placed into the accumulator. Figure 137. Vector multiply word unsigned, modulo, integer (to accumulator)

into the accumulator and into rD. Figure 138. Vector multiply word unsigned, modulo, integer & accumulate

placed into the accumulator and into rD. Figure 139. Vector multiply word unsigned, modulo, integer & accumulate negative

corresponding element of rD. Figure 140. Vector NAND (evnand)

negative number) returns 0x8000_0000. No overflow is detected. Figure 141. Vector negate (evneg)

Note: Use evnand or evnor for evnot. Figure 142. Vector NOR (evnor)

Figure 143. Vector OR (evor) Simplified mnemonic: evmr rD,rA handles moving of the full 64-bit SPE register.

corresponding element of rD. Figure 144. Vector OR with complement (evorc)

Figure 145. Vector rotate left word (evrlw)

Figure 146. Vector rotate left word immediate (evrlwi)

order 16 bits of each element. Figure 147. Vector round word (evrndw)

element of rB is placed into the low-order element of rD. This is shown in Figure 148. Figure 148. Vector select (evsel)

in rB that lie in bit positions 26–31 and 58–63. Shift amounts from 32 to 63 give a zero result. Figure 149. Vector shift left word (evslw)

Figure 150. Vector shift left word immediate (evslwi)

as shown in Figure 151. The SIMM ends up in bit positions rD[0–4] and rD[32–36]. Figure 151. Vector splat fractional immediate (evsplatfi)

significant positions vacated by the shift are filled with a copy of the sign bit. Figure 153. Vector shift right word immediate signed (evsrwis)

Figure 154. Vector shift right word immediate unsigned (evsrwiu)

Shift amounts from 32 to 63 give a result of 32 sign bits. Figure 155. Vector shift right word signed (evsrws)

Shift amounts from 32 to 63 give a zero result. Figure 156. Vector shift right word unsigned (evsrwu)

Figure 165. evstwho Results in big- and little-endian modes

and the difference is placed into the corresponding rD word and into the accumulator. Figure 171. Vector subtract signed, modulo, integer to accumulator word

overflow and summary overflow bits. Figure 172. Vector subtract signed, saturate, integer to accumulator word

the accumulator and the results are placed in rD and into the accumulator. Figure 173. Vector subtract unsigned, modulo, integer to accumulator word

SPEFSCR overflow and summary overflow bits. Figure 174. Vector subtract unsigned, saturate, integer to accumulator word

the results are placed into rD. Figure 175. Vector subtract from word (evsubfw)

the same value is subtracted from both elements of the register. UIMM is 5 bits. Figure 176. Vector subtract immediate from word (evsubifw)

Each element of rA and rB is exclusive-ORed. The results are placed in rD. Figure 177. Vector XOR (evxor)

Extend sign (byte | half word) extsb r A,rS (SZ=0b01, Rc=0) extsb. r A,rS (SZ=0b01, Rc=1) extsh r A,rS (SZ=0b00, Rc=0) extsh. r A,rS (SZ=0b00, Rc=1) if ‘extsb[.]’ then n ← 56 if ‘extsh[.]’ then n ← 48 if ‘extsw’ then n ← 32 if Rc=1 then do LT ← rSn:63 < 0 GT ← rSn:63 > 0 EQ ← rSn:63 = 0 s ← rSn rA ← ns || rSn:63 For extsb[.], the contents of rS[56–63] are placed into rA[56–63]. Bit rS[56] is copied into bits 0–55 of rA. If Rc=1, CR field 0 is set to reflect the result. For extsh[.], the contents of rS[48–63] are placed into rA[48–63]. rS[48] is copied into rA[0– 47]. If Rc=1, CR field 0 is set to reflect the result.

  • Other registers altered: CR0 (if Rc=1) Book E User 0 5 6 1 01 1 1 51 6 2 02 1 2 32 42 52 6 3 03 1 011111 rS rA / / / 111SZ11010 R c

_extsbx _extsbx Extend Sign (Byte | Halfword) se_extsb r X se_extsh r X if se_extsb then n ← 56 if se_extsh then n ← 48 if ‘extsw’ then n ← 32 if Rc=1 then do LT ← GPR(RS)n:63 < 0 GT ← GPR(RS)n:63 > 0 EQ ← GPR(RS)n:63 = 0 s ← GPR(RS or RX)n GPR(RA or RX) ← n-32s || GPR(RS or RX)n:63 For se_extsb, the contents of bits 56–63 of GPR(rX) are placed into bits 56–63 of GPR(rX). Bit 56 of the contents of GPR(rX) is copied into bits 32–55 of GPR( rX). For se_extsh, the contents of bits 48–63 of GPR(rX) are placed into bits 48–63 of GPR(rX). Bit 48 of the contents of GPR(rX) is copied into bits 32–47 of GPR( rX). Special Registers Altered: CR0 (if Rc=1) 05 6 1 1 1 2 1 5

000000001101 R X

000000001111 R X

_extzx _extzx Extend Zero (Byte | Halfword) se_extzb r X se_extzh r X if ‘se_extzb’ then n ← 56 if ‘se_extzh’ then n ← 48 GPR(RX) ← n-320 || GPR(RX)n:63 For se_extzb, the contents of bits 56–63 of GPR(rX) are placed into bits 56–63 of GPR(rX). Bits 32–55 of GPR(rX) are cleared. For se_extzh, the contents of bits 48–63 of GPR(rX) are placed into bits 48–63 of GPR(rX). Bits 32–47 of GPR(rX) are cleared. Special Registers Altered: None 05 6 1 1 1 2 1 5

000000001100 R X

000000001110 R X

fabs fr D,frB( R c = 0 ) fabs. fr D,frB( R c = 1 ) frD) ← 0b0||frB1:63 The contents of frB with bit 0 cleared are placed into frD. If MSR[FP]=0, an attempt to execute fabs[.] causes a floating-point unavailable interrupt. Other registers altered: Book E User 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1 111111 frD /// frB 0100001000 R c

Floating add [single] fadd fr D,frA,frB( P = 1 , R c = 0 ) fadd. fr D,frA,frB( P = 1 , R c = 1 ) fadds fr D,frA,frB( P = 0 , R c = 0 ) fadds. fr D,frA,frB( P = 0 , R c = 1 ) if P=1 then frD ← frA +dp frB else frD ← frA +sp frB The floating-point operand in frA is added to the floating-point operand in frB. If the msb of the resultant significand is not 1, the result is normalized. The result is rounded to the target precision under control of the floating-point rounding control field, FPSCR[RN], and placed into frD. Floating-point addition is based on exponent comparison and addition of the two significands. The exponents of the two operands are compared, and the significand accompanying the smaller exponent is shifted right, with its exponent increased by one for each bit shifted, until the two exponents are equal. The two significands are then added or subtracted as appropriate, depending on the signs of the operands, to form an intermediate sum. All 53 bits of the significand as well as all three guard bits (G, R, and X) enter into the computation. If a carry occurs, the sum’s significand is shifted right one bit position and the exponent is increased by one. FPSCR[FPRF] is set to the class and sign of the result, except for invalid operation exceptions when FPSCR[VE]=1. If MSR[FP]=0, an attempt to execute fadd[s][.] causes a floating-point unavailable interrupt. Other registers altered:

  • FPRF FR FI FX OX UX XX VXSNAN VXISI CR1 ← FX || FEX || VX || OX (if Rc=1) Book E User 0 23456 1 0 1 1 1 5 1 6 2 0 2 1 2 5 2 6 3 0 3 1 111P11 frD frA frB / / / 10101 R c

Floating convert from integer doubleword fcfid fr D,frB sign ← frB0 exp ← 63 frac0:63 ← frB If frac0:63 = 0 then go to Zero Operand If sign = 1 then frac0:63 ← ¬frac0:63 + 1 Do while frac0 = 0 /* do loop 0 times if frB = max negative integer */ frac0:63 ← frac1:63 || 0b0 exp ← exp : 1 End Round Float( sign, exp, frac 0:63, FPSCR[RN] ) If sign = 0 then FPSCR[FPRF] ← ‘+normal number’ If sign = 1 then FPSCR[FPRF] ← ‘:normal number’ frD0 ← sign frD[1-11] ← exp + 1023 /* exp + bias */ frD[12-63] ← frac1:52 Done Zero Operand: FPSCR[FR,FI] ← 0b00 FPSCR[FPRF] ← ‘+zero’ frD ← 0x0000_0000_0000_0000 Done Round Float( sign, exp, frac0:63, round_mode ): inc ← 0 lsb ← frac52 gbit ← frac53 rbit ← frac54 xbit ← frac55:63 > 0 If round_mode = 0b00 then Do /* comparison ignores u bits */ If sign || lsb || gbit || rbit || xbit = 0bu11uu then inc ← 1 If sign || lsb || gbit || rbit || xbit = 0bu011u then inc ← 1 If sign || lsb || gbit || rbit || xbit = 0bu01u1 then inc ← 1 End If round_mode = 0b10 then Do /* comparison ignores u bits */ If sign || lsb || gbit || rbit || xbit = 0b0u1uu then inc ← 1 If sign || lsb || gbit || rbit || xbit = 0b0uu1u then inc ← 1 If sign || lsb || gbit || rbit || xbit = 0b0uuu1 then Book E User 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1 111111 frD /// frB 1101001110 /

inc ← 1 End If round_mode = 0b11 then Do /* comparison ignores u bits */ If sign || lsb || gbit || rbit || xbit = 0b1u1uu then inc ← 1 If sign || lsb || gbit || rbit || xbit = 0b1uu1u then inc ← 1 If sign || lsb || gbit || rbit || xbit = 0b1uuu1 then inc ← 1 End frac0:52 ← frac0:52 + inc If carry_out = 1 then exp ← exp + 1 FPSCR[FR] ← inc FPSCR[FI] ← gbit | rbit | xbit FPSCR[XX] ← FPSCR[XX] | FPSCR[FI] Return The 64-bit signed operand in frB is converted to an infinitely precise floating-point integer. The result of the conversion is rounded to double-precision, as specified by FPSCR[RN], and placed into frD. FPSCR[FPRF] is set to the class and sign of the result. FPSCR[FR] is set if the result is incremented when rounded. FPSCR[FI] is set if the result is inexact. If MSR[FP]=0, an attempt to execute fcfid causes a floating-point unavailable interrupt. Other registers altered: FPRF FR FI FX XX

fcmpu cr D,frA,frB( U = 0 ) fcmpo cr D,frA,frB( U = 1 ) if frA is a NaN or frB is a NaN then c ← 0b0001 else if frA < frB then c ← 0b1000 else if frA > frB then c ← 0b0100 else c ← 0b0010 FPCC ← c CR4×crD:4×crD+3 ← c if ‘fcmpu’ & (frA is a SNaN or frB is a SNaN) then VXSNAN ← 1 if ‘fcmpo’ then do if frA is a SNaN or frB is a SNaN then do if VE=0 then VXVC ← 1 else if frA is a QNaN or frB is a QNaN then VXVC ← 1 The floating-point operand in frA is compared to the floating-point operand in frB. The result of the compare is placed into CR field crD and the FPCC. If either of the operands is a NaN, either quiet or signaling, the CR field crD and the FPCC are set to reflect unordered. If fcmpu, then if either of the operands is a signaling NaN, VXSNAN is set. If fcmpo, then do the following: If either of the operands is a signaling NaN and invalid operation is disabled (VE=0), VXVC is set. If neither operand is a signaling NaN but at least one operand is a quiet NaN, then VXVC is set. If MSR[FP]=0, an attempt to execute fcmpo or fcmpu causes a floating-point unavailable interrupt. Other registers altered:

  • CR field crD FPCC FX VXSNAN VXVC(if fcmpo) Book E User 0 5 6 1 01 1 1 51 6 2 02 1 2 42 52 6 3 03 1 111111 crD/ / frA frB 0000U00000 /

Floating convert to integer doubleword fctid fr D,frB( Z = 0 ) fctidz fr D,frB( Z = 1 ) if ‘fctid[.]’ then round_mode ← FPSCR[RN] if ‘fctidz[.]’ then round_mode ← 0b01 sign ← frB0 If frB[1:11] = 2047 and frB[12:63] = 0 then goto Infinity Operand If frB[1:11] = 2047 and frB 12 = 0 then goto SNaN Operand If frB[1:11] = 2047 and frB12 = 1 then goto QNaN Operand If frB[1:11] > 1086 then goto Large Operand If frB[1:11] > 0 then exp ← frB[1:11] : 1023 /* exp : bias */ If frB[1:11] = 0 then exp ← :1022 /* normal; need leading 0 for later complement */ If frB[1:11] > 0 then frac /* denormal */ If frB[1:11] = 0 then frac0:64 ← 0b00 || frB[12:63] || 110 gbit || rbit || xbit ← 0b000 Do i=1,63:exp /* do the loop 0 times if exp = 63 */ xbit) End Round Integer( sign, frac 0:64, gbit, rbit, xbit, round_mode ) /* needed leading 0 for :264 < frB < :263 */ If sign=1 then frac0:64 ← ¬frac0:64 + 1 If frac0:64 > 263:1 then goto Large Operand If frac0:64 < :263 then goto Large Operand FPSCR[XX] ← FPSCR[XX] | FPSCR[FI] FPSCR[FPRF] ← undefined frD ← frac1:64 Done Round Integer( sign,frac0:64, gbit, rbit, xbit, round_mode ): inc ← 0 If round_mode = 0b00 then /* comparison ignores u bits */ Do If sign || frac64 || gbit || rbit || xbit = 0bu11uu then inc ← 1 If sign || frac64 || gbit || rbit || xbit = 0bu011u then inc ← 1 If sign || frac64 || gbit || rbit || xbit = 0bu01u1 then inc ← 1 End If round_mode = 0b10 then /* comparison ignores u bits */ Do If sign || frac64 || gbit || rbit || xbit = 0b0u1uu then inc ← 1 Book E User 0 5 6 1 01 1 1 51 6 2 02 1 2 93 03 1 111111 frD /// frB 110010111Z/

If sign || frac64 || gbit || rbit || xbit = 0b0uu1u then inc ← 1 If sign || frac64 || gbit || rbit || xbit = 0b0uuu1 then inc ← 1 End If round_mode = 0b11 then /* comparison ignores u bits */ Do If sign || frac64 || gbit || rbit || xbit = 0b1u1uu then inc ← 1 If sign || frac64 || gbit || rbit || xbit = 0b1uu1u then inc ← 1 If sign || frac64 || gbit || rbit || xbit = 0b1uuu1 then inc ← 1 End frac0:64 ← frac0:64 + inc FPSCR[FR] ← inc FPSCR[FI] ← gbit | rbit | xbit Return Infinity Operand: FPSCR[FR,FI,VXCVI] ← 0b001 If FPSCR[VE] = 0 then Do If sign = 0 then frD ← 0x7FFF_FFFF_FFFF_FFFF If sign = 1 then frD ← 0x8000_0000_0000_0000 FPSCR[FPRF] ← undefined End Done SNaN Operand: FPSCR[FR,FI,VXSNAN,VXCVI] ← 0b0011 If FPSCR[VE] = 0 then Do frD ← 0x8000_0000_0000_0000 FPSCR[FPRF] ← undefined End Done QNaN Operand: FPSCR[FR,FI,VXCVI] ← 0b001 If FPSCR[VE] = 0 then Do frD ← 0x8000_0000_0000_0000 FPSCR[FPRF] ← undefined End Done Large Operand: FPSCR[FR,FI,VXCVI] ← 0b001 If FPSCR[VE] = 0 then Do If sign = 0 then frD ← 0x7FFF_FFFF_FFFF_FFFF If sign = 1 then frD ← 0x8000_0000_0000_0000 FPSCR[FPRF] ← undefined End Done For fctid or fctid., the rounding mode is specified by FPSCR[RN]. For fctidz or fctidz., the rounding mode used is round toward zero.

The floating-point operand in frB is converted to a 64-bit signed integer, using the rounding mode specified by the instruction, and placed into frD. If the floating-point operand in frB is greater than 263–1, then 0x7FFF_FFFF_FFFF_FFFF is placed into frD. If the floating-point operand in frB is less than –263, 0x8000_0000_0000_0000 is placed into frD. Except for enabled invalid operation exceptions, FPSCR[FPRF] is undefined. FPSCR[FR] is set if the result is incremented when rounded. FPSCR[FI] is set if the result is inexact. If MSR[FP]=0, an attempt to execute fctid[z] causes a floating-point unavailable interrupt. Other registers altered:

  • FPRF (undefined) FR FI FX XX VXSNAN VXCVI

Floating convert to integer word fctiw fr D,frB( Z = 0 , R c = 0 ) fctiw. fr D,frB( Z = 0 , R c = 1 ) fctiwz fr D,frB( Z = 1 , R c = 0 ) fctiwz. fr D,frB( Z = 1 , R c = 1 ) if ‘fctiw[.]’ then round_mode ← FPSCR[RN] if ‘fctiwz[.]’ then round_mode ← 0b01 sign ← frB0 If frB[1:11] = 2047 and frB[12:63] = 0 then goto Infinity Operand If frB[1:11] = 2047 and frB 12 = 0 then goto SNaN Operand If frB[1:11] = 2047 and frB12 = 1 then goto QNaN Operand If frB[1:11] > 1086 then goto Large Operand If frB[1:11] > 0 then exp ← frB[1:11] : 1023 /* exp : bias */ If frB[1:11] = 0 then exp ← :1022 /* normal; need leading 0 for later complement */ If frB[1:11] > 0 then frac /* denormal */ If frB[1:11] = 0 then frac0:64 ← 0b00 || frB[12:63] || 110 gbit || rbit || xbit ← 0b000 Do i=1,63:exp /* do the loop 0 times if exp = 63 */ xbit) End Round Integer( sign, frac 0:64, gbit, rbit, xbit, round_mode ) /* needed leading 0 for :264 < frB < :263 */ If sign=1 then frac0:64 ← ¬frac0:64 + 1 If frac0:64 > 231:1 then goto Large Operand If frac0:64 < :231 then goto Large Operand FPSCR[XX] ← FPSCR[XX] | FPSCR[FI] frD ← 0xuuuu_uuuu || frac33:64 /* u is undefined hex digit */ FPSCR[FPRF] ← undefined Done Round Integer( sign, frac0:64, gbit, rbit, xbit, round_mode ): inc ← 0 If round_mode = 0b00 then /* comparison ignores u bits */ Do If sign || frac64 || gbit || rbit || xbit = 0bu11uu then inc ← 1 If sign || frac64 || gbit || rbit || xbit = 0bu011u then inc ← 1 If sign || frac64 || gbit || rbit || xbit = 0bu01u1 then inc ← 1 End If round_mode = 0b10 then /* comparison ignores u bits */ Do Book E User 0 5 6 1 01 1 1 51 6 2 02 1 2 93 03 1 111111 frD /// frB 000000111Z R c

If sign || frac64 || gbit || rbit || xbit = 0b0u1uu then inc ← 1 If sign || frac64 || gbit || rbit || xbit = 0b0uu1u then inc ← 1 If sign || frac64 || gbit || rbit || xbit = 0b0uuu1 then inc ← 1 End If round_mode = 0b11 then /* comparison ignores u bits */ Do If sign || frac64 || gbit || rbit || xbit = 0b1u1uu then inc ← 1 If sign || frac64 || gbit || rbit || xbit = 0b1uu1u then inc ← 1 If sign || frac64 || gbit || rbit || xbit = 0b1uuu1 then inc ← 1 End frac0:64 ← frac0:64 + inc FPSCR[FR] ← inc FPSCR[FI] ← gbit | rbit | xbit Return Infinity Operand: FPSCR[FR,FI,VXCVI] ← 0b001 If FPSCR[VE] = 0 then Do /* u is undefined hex digit */ If sign = 0 then frD ← 0xuuuu_uuuu_7FFF_FFFF If sign = 1 then frD ← 0xuuuu_uuuu_8000_0000 FPSCR[FPRF] ← undefined End Done SNaN Operand: FPSCR[FR,FI,VXSNAN,VXCVI] ← 0b0011 If FPSCR[VE] = 0 then Do /* u is undefined hex digit */ frD ← 0xuuuu_uuuu_8000_0000 FPSCR[FPRF] ← undefined End Done QNaN Operand: FPSCR[FR,FI,VXCVI] ← 0b001 If FPSCR[VE] = 0 then Do /* u is undefined hex digit */ frD ← 0xuuuu_uuuu_8000_0000 FPSCR[FPRF] ← undefined End Done Large Operand: FPSCR[FR,FI,VXCVI] ← 0b001 If FPSCR[VE] = 0 then Do /* u is undefined hex digit */ If sign = 0 then frD ← 0xuuuu_uuuu_7FFF_FFFF If sign = 1 then frD ← 0xuuuu_uuuu_8000_0000 FPSCR[FPRF] ← undefined End Done For fctiw or fctiw., the rounding mode is specified by FPSCR[RN].

For fctiwz or fctiwz., the rounding mode used is round toward zero. The floating-point operand in frB is converted to a 32-bit signed integer, using the rounding mode specified by the instruction, and placed into frD[32–63]; frD[0–31] are undefined. If the operand in frB is greater than 231–1, then frD[32–63] are set to 0x7FFF_FFFF . If the operand in frB is less than –231, then frD[32–63] are set to 0x8000_0000. Except for enabled invalid operation exceptions, FPSCR[FPRF] is undefined. FPSCR[FR] is set if the result is incremented when rounded. FPSCR[FI] is set if the result is inexact. If MSR[FP]=0, an attempt to execute fctiw[z][.] causes a floating-point unavailable interrupt. Other registers altered:

  • FPRF (undefined) FR FI FX XX VXSNAN VXCVI CR1 ← FX || FEX || VX || OX (if Rc=1)

Floating divide [single] fdiv fr D,frA,frB( P = 1 , R c = 0 ) fdiv. fr D,frA,frB( P = 1 , R c = 1 ) fdivs fr D,frA,frB( P = 0 , R c = 0 ) fdivs. fr D,frA,frB( P = 0 , R c = 1 ) if P=1 then frD ← frA ÷dp frB else frD ← frA ÷sp frB The floating-point operand in frA is divided by the floating-point operand in frB. The remainder is not supplied as a result. If the msb of the resultant significand is not 1, the result is normalized. The result is rounded to the target precision under control of the floating-point rounding control field, FPSCR[RN], and placed into frD. Floating-point division is based on exponent subtraction and division of the significands. FPSCR[FPRF] is set to the class and sign of the result, except for invalid operation exceptions when FPSCR[VE]=1 and zero divide exceptions when FPSCR[ZE]=1. If MSR[FP]=0, an attempt to execute fdiv[s][.] causes a floating-point unavailable interrupt. Other registers altered:

  • FPRF FR FI FX OX UX ZX XX VXSNAN VXIDI VXZDZ CR1 ← FX || FEX || VX || OX (if Rc=1) Book E User 0 23456 1 0 1 1 1 5 1 6 2 0 2 1 2 5 2 6 3 0 3 1 111P11 frD frA frB / / / 10010 R c

Floating multiply-add [single] fmadd fr D,frA,frC,frB( P = 1 , R c = 0 ) fmadd. fr D,frA,frC,frB( P = 1 , R c = 1 ) fmadds fr D,frA,frC,frB( P = 0 , R c = 0 ) fmadds. fr D,frA,frC,frB( P = 0 , R c = 1 ) if P=1 then frD ← [frA ×fp frC] +dp frB else frD ← [frA ×fp frC] +sp frB The floating-point operand in frA is multiplied by the floating-point operand in frC. The floating-point operand in frB is added to this intermediate result. If the msb of the resultant significand is not 1, the result is normalized. The result is rounded to the target precision under control of the floating-point rounding control field, FPSCR[RN], and placed into frD. FPSCR[FPRF] is set to the class and sign of the result, except for invalid operation exceptions when FPSCR[VE]=1. If MSR[FP]=0, an attempt to execute fmadd[s][.] causes a floating-point unavailable interrupt. Other registers altered:

  • FPRF FR FI FX OX UX XX VXSNAN VXISI VXIMZ CR1 ← FX || FEX || VX || OX (if Rc=1) Book E User 0 23456 1 0 1 1 1 5 1 6 2 0 2 1 2 5 2 6 3 0 3 1 111P11 frD frA frB frC 11101 R c

fmr fr D,frB( R c = 0 ) fmr. fr D,frB( R c = 1 ) frD ← frB The contents of frB are placed into frD. If MSR[FP]=0, an attempt to execute fmr[.] causes a floating-point unavailable interrupt. Other registers altered: Book E User 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1 111111 frD /// frB 0001001000 R c

Floating multiply-subtract [single] fmsub fr D,frA,frC,frB( P = 1 , R c = 0 ) fmsub. fr D,frA,frC,frB( P = 1 , R c = 1 ) fmsubs fr D,frA,frC,frB( P = 0 , R c = 0 ) fmsubs. fr D,frA,frC,frB( P = 0 , R c = 1 ) if P=1 then frD ← [frA ×fp frC] -dp frB else frD ← [frA ×fp frC] -sp frB The floating-point operand in frA is multiplied by the floating-point operand in frC. The floating-point operand in frB is subtracted from this intermediate result. If the msb of the resultant significand is not 1, the result is normalized. The result is rounded to the target precision under control of the floating-point rounding control field, FPSCR[RN], and placed into frD. FPSCR[FPRF] is set to the class and sign of the result, except for invalid operation exceptions when FPSCR[VE]=1. If MSR[FP]=0, an attempt to execute fmsub[s][.] causes a floating-point unavailable interrupt. Other registers altered:

  • FPRF FR FI FX OX UX XX VXSNAN VXISI VXIMZ CR1 ← FX || FEX || VX || OX (if Rc=1) Book E User 0 23456 1 0 1 1 1 5 1 6 2 0 2 1 2 5 2 6 3 0 3 1 111P11 frD frA frB frC 11100 R c

Floating multiply [single] fmul fr D,frA,frC( P = 1 , R c = 0 ) fmul. fr D,frA,frC( P = 1 , R c = 1 ) fmuls fr D,frA,frC( P = 0 , R c = 0 ) fmuls. fr D,frA,frC( P = 0 , R c = 1 ) if P=1 then frD ← frA ×dp frC else frD ← frA ×sp frC The floating-point operand in frA is multiplied by the floating-point operand in frC. If the msb of the resultant significand is not 1, the result is normalized. The result is rounded to the target precision under control of the floating-point rounding control field, FPSCR[RN], and placed into frD. Floating-point multiplication is based on exponent addition and multiplication of the significands. FPSCR[FPRF] is set to the class and sign of the result, except for invalid operation exceptions when FPSCR[VE]=1. If MSR[FP]=0, an attempt to execute fmul[s][.] causes a floating-point unavailable interrupt. Other registers altered:

  • FPRF FR FI FX OX UX XX VXSNAN VXIMZ CR1 ← FX || FEX || VX || OX (if Rc=1) Book E User 0 23456 1 0 1 1 1 5 1 6 2 0 2 1 2 5 2 6 3 0 3 1 111P11 frD frA/ / / frC 11001 R c

Floating negative absolute value fnabs fr D,frB( R c = 0 ) fnabs. fr D,frB( R c = 1 ) frD ← 0b1||frB1:63 The contents of frB with bit 0 set are placed into frD. If MSR[FP]=0, an attempt to execute fnabs[.] causes a floating-point unavailable interrupt. Other registers altered: Book E User 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1 111111 frD /// frB 0010001000 R c

fneg fr D,frB (Rc=0) fneg. fr D,frB( R c = 1 ) frD ← ¬frB0||frB1:63 The contents of frB with bit 0 inverted are placed into frD. If MSR[FP]=0, an attempt to execute fneg[.] causes a floating-point unavailable interrupt. Other registers altered: Book E User 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1 111111 frD /// frB 0000101000 R c

Floating negative multiply-add [single] fnmadd fr D,frA,frC,frB( P = 1 , R c = 0 ) fnmadd. fr D,frA,frC,frB( P = 1 , R c = 1 ) fnmadds fr D,frA,frC,frB( P = 0 , R c = 0 ) fnmadds. fr D,frA,frC,frB( P = 0 , R c = 1 ) if P=1 then frD ← -([frA ×fp frC] +dp frB) else frD ← -([frA ×fp frC] +sp frB) The floating-point operand in frA is multiplied by the floating-point operand in frC. The floating-point operand in frB is added to this intermediate result. If the msb of the resultant significand is not 1, the result is normalized. The result is rounded to the target precision under control of the floating-point rounding control field, FPSCR[RN], then negated and placed into frD. This instruction produces the same result as would be obtained by using the Floating Multiply-Add instruction and then negating the result, with the following exceptions.

  • QNaNs propagate with no effect on their sign bit.
  • QNaNs that are generated as the result of a disabled invalid operation exception have a sign bit of 0.
  • SNaNs that are converted to QNaNs as the result of a disabled invalid operation exception retain the sign bit of the SNaN. FPSCR[FPRF] is set to the class and sign of the result, except for invalid operation exceptions when FPSCR[VE]=1. An attempt to execute fnmadd[s][.] causes a floating-point unavailable interrupt. Other registers altered:
  • FPRF FR FI FX OX UX XX VXSNAN VXISI VXIMZ CR1 ← FX || FEX || VX || OX (if Rc=1) Book E User 0 23456 1 0 1 1 1 5 1 6 2 0 2 1 2 5 2 6 3 0 3 1 111P11 frD frA frB frC 11111 R c

Floating negative multiply-subtract [single] fnmsub fr D,frA,frC,frB( P = 1 , R c = 0 ) fnmsub. fr D,frA,frC,frB( P = 1 , R c = 1 ) fnmsubs fr D,frA,frC,frB( P = 0 , R c = 0 ) fnmsubs. fr D,frA,frC,frB( P = 0 , R c = 1 ) if P=1 then frD ← -([frA ×fp frC] :dp frB) else frD ← -([frA ×fp frC] :sp frB) The floating-point operand in frA is multiplied by the floating-point operand in frC. The floating-point operand in frB is subtracted from this intermediate result. If the msb of the resultant significand is not 1, the result is normalized. The result is rounded to the target precision under control of the floating-point rounding control field, FPSCR[RN], then negated and placed into frD. This instruction produces the same result as would be obtained by using the Floating Multiply-Subtract instruction and then negating the result, with the following exceptions.

  • QNaNs propagate with no effect on their sign bit.
  • QNaNs that are generated as the result of a disabled invalid operation exception have a sign bit of 0.
  • SNaNs that are converted to QNaNs as the result of a disabled invalid operation exception retain the sign bit of the SNaN. FPSCR[FPRF] is set to the class and sign of the result, except for invalid operation exceptions when FPSCR[VE]=1. An attempt to execute fnmsub[s][.] causes a floating-point unavailable interrupt. Other registers altered:
  • FPRF FR FI FX OX UX XX VXSNAN VXISI VXIMZ CR1 ← FX || FEX || VX || OX (if Rc=1) Book E User 0 23456 1 0 1 1 1 5 1 6 2 0 2 1 2 5 2 6 3 0 3 1 111P11 frD frA frB frC 11110 R c

between implementations, and between different executions on the same implementation. Operation with various special values of the operand is summarized in Table 204. exceptions when FPSCR[VE]=1 and zero divide exceptions when FPSCR[ZE]=1. If MSR[FP]=0, an attempt to execute fres[.] causes a floating-point unavailable interrupt.

  • FPRF FR (undefined) FI (undefined) FX OX UX ZX VXSNAN CR1 ← FX || FEX || VX || OX (if Rc=1) Book E User 0 5 6 1 01 1 1 51 6 2 02 1 2 52 6 3 03 1 111011 frD /// frB / / / 11000 R c

Table 204. Operations with special values

Floating round to single-precision frsp fr D,frB( R c = 0 ) frsp. fr D,frB( R c = 1 ) If frB[1:11] < 897 and frB1:63 > 0 then Do If FPSCR[UE] = 0 then goto Disabled Exponent Underflow If FPSCR[UE] = 1 then goto Enabled Exponent Underflow If frB[1:11] > 1150 and frB[1:11] < 2047 then Do If FPSCR[OE] = 0 then goto Disabled Exponent Overflow If FPSCR[OE] = 1 then goto Enabled Exponent Overflow If frB[1:11] > 896 and frB[1:11] < 1151 then goto Normal Operand If frB 1:63 = 0 then goto Zero Operand If frB[1:11] = 2047 then Do If frB[12:63] = 0 then goto Infinity Operand If frB 12 = 1 then goto QNaN Operand If frB12 = 0 and frB[13:63] > 0 then goto SNaN Operand Disabled Exponent Underflow: sign ← frB0 If frB[1:11] = 0 then Do exp ← :1022 frac0:52 ← 0b0 || frB[12:63] If frB[1:11] > 0 then Do exp ← frB[1:11] : 1023 fr ← 0b1 || frB[12:63] Denormalize operand: Do while exp < :126 exp ← exp + 1 FPSCR[UX] ← (frac24:52 || G || R || X) > 0 Round Single(sign,exp,frac0:52,G,R,X) FPSCR[XX] ← FPSCR[XX] | FPSCR[FI] If frac0:52 = 0 then Do frD0 ← sign frD1:63 ← 0 If sign = 0 then FPSCR[FPRF] ← ‘+zero’ If sign = 1 then FPSCR[FPRF] ← ‘:zero’ If frac0:52 > 0 then Do If frac0 = 1 then Do If sign = 0 then FPSCR[FPRF] ← ‘+normal number’ If sign = 1 then FPSCR[FPRF] ← ‘:normal number’ If frac0 = 0 then Do If sign = 0 then FPSCR[FPRF] ← Book E User 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1 111111 frD /// frB 0000001100 R c

‘+denormalized number’ If sign = 1 then FPSCR[FPRF] ← ‘:denormalized number’ Normalize operand: Do while frac0 = 0 exp ← exp:1 frac0:52 ← frac1:52 || 0b0 frD0 ← sign frD[1-11] ← exp + 1023 frD[12-63] ← frac1:52 Done Enabled exponent underflow: FPSCR[UX] ← 1 sign ← frB0 If frB[1:11] = 0 then Do exp ← :1022 frac0:52 ← 0b0 || frB[12:63] If frB[1:11] > 0 then Do exp ← frB[1:11] : 1023 frac0:52 ← 0b1 || frB[12:63] Normalize operand: Do while frac0 = 0 exp ← exp : 1 frac0:52 ← frac1:52 || 0b0 Round Single(sign,exp,frac0:52,0,0,0) FPSCR[XX] ← FPSCR[XX] | FPSCR[FI] exp ← exp + 192 frD0 ← sign frD[1-11] ← exp + 1023 frD[12-63] ← frac1:52 If sign = 0 then FPSCR[FPRF] ← ‘+normal number’ If sign = 1 then FPSCR[FPRF] ← ‘:normal number’ Done Disabled exponent overflow FPSCR[OX] ← 1 If FPSCR[RN] = 0b00 then Do /* Round to Nearest If frB0 = 0 then frD ← 0x7FF0_0000_0000_0000 If frB0 = 1 then frD ← 0xFFF0_0000_0000_0000 If frB0 = 0 then FPSCR[FPRF] ← ‘+infinity’ If frB0 = 1 then FPSCR[FPRF] ← ‘:infinity’ If FPSCR[RN] = 0b01 then Do /* Round toward Zero */ If frB0 = 0 then frD ← 0x47EF_FFFF_E000_0000 If frB0 = 1 then frD ← 0xC7EF_FFFF_E000_0000 If frB0 = 0 then FPSCR[FPRF] ← ‘+normal number’ If frB0 = 1 then FPSCR[FPRF] ← ‘:normal

number’ If FPSCR[RN] = 0b10 then Do /* Round toward +Infinity */ If frB0 = 0 then frD ← 0x7FF0_0000_0000_0000 If frB0 = 1 then frD ← 0xC7EF_FFFF_E000_0000 If frB0 = 0 then FPSCR[FPRF] ← ‘+infinity’ If frB0 = 1 then FPSCR[FPRF] ← ‘:normal number’ If FPSCR[RN] = 0b11 then Do /* Round toward :Infinity */ If frB0 = 0 then frD ← 0x47EF_FFFF_E000_0000 If frB0 = 1 then frD ← 0xFFF0_0000_0000_0000 If frB0 = 0 then FPSCR[FPRF] ← ‘+normal number’ If frB0 = 1 then FPSCR[FPRF] ← ‘:infinity’ FPSCR[FR] ← undefined FPSCR[FI] ← 1 FPSCR[XX] ← 1 Done Enabled Exponent Overflow: sign ← frB0 exp ← frB[1:11] : 1023 frac0:52 ← 0b1 || frB[12:63] Round Single(sign,exp,frac0:52,0,0,0) FPSCR[XX] ← FPSCR[XX] | FPSCR[FI] Enabled Overflow: FPSCR[OX] ← 1 exp ← exp : 192 frD0 ← sign frD[1-11] ← exp + 1023 frD[12-63] ← frac1:52 If sign = 0 then FPSCR[FPRF] ← ‘+normal number’ If sign = 1 then FPSCR[FPRF] ← ‘:normal number’ Done Zero Operand: frD ← frB If frB0 = 0 then FPSCR[FPRF] ← ‘+zero’ If frB0 = 1 then FPSCR[FPRF] ← ‘:zero’ FPSCR[FR,FI] ← 0b00 Done Infinity Operand: frD ← frB If frB0 = 0 then FPSCR[FPRF] ← ‘+infinity’ If frB0 = 1 then FPSCR[FPRF] ← ‘:infinity’ FPSCR[FR,FI] ← 0b00 Done QNaN Operand: frD ← frB0:34 || 290

FPSCR[FPRF] ← ‘QNaN’ FPSCR[FR,FI] ← 0b00 Done SNaN Operand: FPSCR[VXSNAN] ← 1 If FPSCR[VE] = 0 then Do frD[0:11] ← frB[0:11] frD12 ← 1 FPSCR[FPRF] ← ‘QNaN’ FPSCR[FR,FI] ← 0b00 Done Normal Operand: sign ← frB0 exp ← frB[1:11] : 1023 frac0:52 ← 0b1 || frB[12:63] Round Single(sign,exp,frac0:52,0,0,0) FPSCR[XX] ← FPSCR[XX] | FPSCR[FI] If exp > 127 and FPSCR[OE] = 0 then go to Disabled Exponent Overflow If exp > 127 and FPSCR[OE] = 1 then go to Enabled Overflow frD0 ← sign frD[1-11] ← exp + 1023 frD[12-63] ← frac1:52 If sign = 0 then FPSCR[FPRF] ← ‘+normal number’ If sign = 1 then FPSCR[FPRF] ← ‘:normal number’ Done Round Single(sign,exp,frac0:52,G,R,X): inc ← 0 lsb ← frac23 gbit ← frac24 rbit ← frac25 If FPSCR[RN] = 0b00 then Do /* comparison ignores u bits */ If sign || lsb || gbit || rbit || xbit = 0bu11uu then inc ← 1 If sign || lsb || gbit || rbit || xbit = 0bu011u then inc ← 1 If sign || lsb || gbit || rbit || xbit = 0bu01u1 then inc ← 1 If FPSCR[RN] = 0b10 then Do /* comparison ignores u bits */ If sign || lsb || gbit || rbit || xbit = 0b0u1uu then inc ← 1 If sign || lsb || gbit || rbit || xbit = 0b0uu1u then inc ← 1 If sign || lsb || gbit || rbit || xbit = 0b0uuu1 then inc ← 1 If FPSCR[RN] = 0b11 then Do /* comparison ignores u bits */

If sign || lsb || gbit || rbit || xbit = 0b1u1uu then inc ← 1 If sign || lsb || gbit || rbit || xbit = 0b1uu1u then inc ← 1 If sign || lsb || gbit || rbit || xbit = 0b1uuu1 then inc ← 1 frac0:23 ← frac0:23 + inc If carry_out = 1 then Do frac0:23 ← 0b1 || frac0:22 exp ← exp + 1 frac24:52 ← 290 FPSCR[FR] ← inc FPSCR[FI] ← gbit | rbit | xbit Return The floating-point operand in frB is rounded to single-precision, using the rounding mode specified by FPSCR[RN], and placed into frD. FPSCR[FPRF] is set to the class and sign of the result, except for invalid operation exceptions when FPSCR[VE]=1. If MSR[FP]=0, an attempt to execute frsp[.] causes a floating-point unavailable interrupt. Other registers altered:

  • FPRF FR FI FX OX UX XX VXSNAN CR1 ← FX || FEX || VX || OX (if Rc=1)

implementations, and between different executions on the same implementation. exceptions when FPSCR[VE]=1 and zero divide exceptions when FPSCR[ZE]=1. If MSR[FP]=0, attempting to execute frsqrte[.] causes a floating-point unavailable interrupt.

  • FPRF FR (undefined) FI (undefined) FX ZX VXSNAN VXSQRT CR1 ← FX || FEX || VX || OX (if Rc=1) Book E User 0 5 6 1 01 1 1 51 6 2 02 1 2 52 6 3 03 1 111111 frD /// frB / / / 11010 R c

Table 205. Operations with special values

fsel fr D,frA,frC,frB( R c = 0 ) fsel. fr D,frA,frC,frB( R c = 1 ) if frA ≥ 0.0 then frD ← frC else frD ← frB The floating-point operand in frA is compared to the value zero. If the operand is greater than or equal to zero, frD is set to the contents of frC. If the operand is less than zero or is a NaN, frD is set to the contents of frB. The comparison ignores the sign of zero (that is, +0 and –0 are regarded as equal). If MSR[FP]=0, an attempt to execute fsel[.] causes a floating-point unavailable interrupt. Other registers altered: Note: Programming: Examples of uses of this instruction can be found in the appendix Warning: Care must be taken in us ing fsel if IEEE compatibility is required, or if the values being tested can be NaNs or infinities Book E User 0 5 6 1 01 1 1 51 6 2 02 1 2 52 6 3 03 1 111111 frD frA frB frC 10111 R c

The square root of the floating-point operand in frB is placed into frD. exceptions when FPSCR[VE]=1. If MSR[FP]=0, an attempt to execute fsqrt[s][.] causes a floating-point unavailable interrupt.

  • FPRF FR FI FX XX VXSNAN VXSQRT CR1 ← FX || FEX || VX || OX (if Rc=1) Book E User 0 23456 1 0 1 1 1 5 1 6 2 0 2 1 2 5 2 6 3 0 3 1 111P11 frD /// frB / / / 10110 R c

Table 206. Operations with special values

Floating subtract [single] fsub fr D,frA,frB( P = 1 , R c = 0 ) fsub. fr D,frA,frB( P = 1 , R c = 1 ) fsubs fr D,frA,frB( P = 0 , R c = 0 ) fsubs. fr D,frA,frB( P = 0 , R c = 1 ) if P=1 then frD ← frA -dp frB else frD ← frA -sp frB The floating-point operand in frB is subtracted from the floating-point operand in frA. If the msb of the resultant significand is not 1, the result is normalized. The result is rounded to the target precision under control of the floating-point rounding control field, FPSCR[RN]. and placed into frD. The execution of the Floating Subtract instruction is identical to that of Floating Add, except that the contents of frB participate in the operation with the sign bit (bit 0) inverted. FPSCR[FPRF] is set to the class and sign of the result, except for invalid operation exceptions when FPSCR[VE]=1. If MSR[FP]=0, an attempt to execute fsub[s][.] causes a floating-point unavailable interrupt. Other registers altered:

  • FPRF FR FI FX OX UX XX VXSNAN VXISI CR1 ← FX || FEX || VX || OX (if Rc=1) Book E User 0 23456 1 0 1 1 1 5 1 6 2 0 2 1 2 5 2 6 3 0 3 1 111P11 frD frA frB / / / 10100 R c

Instruction cache block invalidate icbi r A,rB if rA=0 then a ← 640 else a ← rA InvalidateInstructionCacheBlock( EA ) EA calculation: Addressing ModeEA for rA=0EA for rA≠0 320 || rB32:63 If the block containing the byte addressed by EA is in memory that is memory-coherence required and a block containing the byte addressed by EA is in the instruction cache of any processors, the block is invalidated in those instruction caches, so that subsequent references cause the block to be fetched from main memory. If the block containing the byte addressed by EA is in memory that is not memory-coherence required and a block containing the byte addressed by EA is in the instruction cache of this processor, the block is invalidated in that instruction cache, so that subsequent references cause the block to be fetched from main memory. The function of this instruction is independent of whether the block containing the byte addressed by EA is in memory that is write-through required or caching-inhibited. This instruction is treated as a load. icbi may cause a cache-locking exception on some implementations. See the implementation documentation. On some implementations, HID1[ABE] must be set to allow management of external L2 caches (for implementations with L2 caches) as well as other L1 caches in the system. Other registers altered: None Book E User 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1 011111 / / / rA rB 1111010110 /

Instruction cache block lock clear icblc CT,rA,rB Form: X if rA = 0 then a ← 640 else a ← GPR(rA) if Mode32 then EA ← 320 || (a + GPR(rB))32:63 if Mode64 then EA ← a + GPR(rB) InstructionCacheBlockClearLock(CT, EA) EA calculation: EA for rA=0EA for rA≠0 320 || GPR(rB)32:63 320 || (GPR(rA)+GPR(rB))32:63 The instruction cache specified by CT has the cache line corresponding to EA unlocked allowing the line to participate in the normal replacement policy. Cache lock clear instructions remove locks previously set by cache lock set instructions. User-level cache instructions on page 180, lists supported CT values. An implementation may use other CT values to enable software to target specific, implementation-dependent portions of its cache hierarchy or structure. The icbtlc instruction requires read (R) or execute (X) permissions with respect to translation and memory protection and can cause DSI and DTLB error interrupts accordingly. An unable-to-unlock condition is said to occur any of the following conditions exist:

  • The target address is marked cache-inhibited, or the storage attributes of the address uses a coherency protocol that does not support locking.
  • The target cache is disabled or not present.
  • The CT field of the instructions contains a value not supported by the implementation.
  • The target address is not in the cache or is present in the cache but is not locked. If an unable-to-unlock condition occurs, no cache operation is performed. EIS specifics Setting L1CSR1[ICLFI] allows system software to clear all L1 instruction cache locking bits without knowing the addresses of the lines locked. Cache locking APU User 0 5 61 0 1 11 5 1 62 0 2 1 3 0 3 1

011111 C T rA rB 0011100110 /

Instruction cache block touch icbt CT,rA,rB if rA=0 then a ← 640 else a ← rA PrefetchInstructionCacheBlock( CT, EA ) EA calculation: Addressing ModeEA for rA=0EA for rA≠0 320 || rB32:63 This instruction is a hint that performance would likely be improved if the block containing the byte addressed by EA is fetched into the instruction cache, because the program will probably soon execute code from the addressed location. User-level cache instructions on page 180,” lists supported CT values. An implementation may use other CT values to enable software to target specific, implementation-dependent portions of its cache hierarchy or structure. Implementations should perform no operation when CT specifies a value not supported by the implementation. The hint is ignored if the block is caching-inhibited. This instruction treated as a load (see the discussion of cache and MMU operation in the user’s manual), except that an interrupt is not taken for a translation or protection violation. Other registers altered: None Book E User 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1

011111 C T rA rB 0000010110 /

Instruction cache block touch and lock set icbtls CT,rA,rB Form: X if rA = 0 then a ← 640 else a ← GPR(rA) if Mode32 then EA ← 320 || (a + GPR(rB))32:63 if Mode64 then EA ← a + GPR(rB) PrefetchInstructionCacheBlockLockSet(CT, EA) EA calculation: EA for rA=0EA for rA≠0 320 || GPR(rB)32:63 320 || (GPR(rA)+GPR(rB))32:63 The instruction cache specified by CT has the line corresponding to EA loaded and locked. If the line exists in the cache, it is locked without refetching from memory. Cache touch and lock set instructions allow software to lock lines into the cache to shorten latency for critical cache accesses and more deterministic behavior. Lines locked in the cache do not participate in the normal replacement policy when a line must be victimized for replacement. User-level cache instructions on page 180,” lists supported CT values. An implementation may use other CT values to enable software to target specific, implementation-dependent portions of its cache hierarchy or structure. The icbtls requires read (R) or execute (X) permissions for translation and memory protection and can cause DSI and DTLB error interrupts accordingly. For unable-to-lock conditions, described in Unable-to-lock conditions on page 849,” no cache operation is performed and LICSR0[ICUL] is set. An overlocking condition is said to exist is all the available ways for a given cache index are already locked. If an overlocking condition occurs for a icbtls instruction and if the lock was targeted for the primary cache or secondary cache (CT = 0 or CT = 2), the requested line is not locked into the cache. When an overlock condition occurs, L1CSR1[ICLO] (L2CSR[L2CLO] for CT = 2) is set. If L1CSR1[ICLOA] is set (or L2CSR[L2CLOA] for CT = 2), the requested line is locked into the cache and implementation dependent line currently locked in the cache is evicted. Results of overlocking and unable-to-lock conditions for caches other than the primary and secondary cache are defined as part of the architecture for the cache hierarchy designated by CT. If a unified primary cache is implemented and L1CSR1 is not implemented, L1CSR0[DCUL] and L1CSR0[DCLO] are updated instead of the corresponding L1CSR1 bits. Other registers altered:

  • L1CSR1[ICUL] if unable to lock occurs
  • L1CSRI[ICLO] (L2CSR[L2CLO]) if lock overflow occurs Cache locking APU User 0 5 61 0 1 11 5 1 62 0 2 1 3 0 3 1

011111 C T rA rB 0111100110 /

_illegal _illegal Illegal se_illegal SRR1 ← MSR SRR0 ← CIA NIA ← IVPR32:47 || IVOR648:59 || 0b0000 MSRWE,EE,PR,IS,DS,FP,FE0,FE1 ← 0b0000_0000 se_illegal is used to request an illegal instruction exception. A program interrupt is generated. The contents of the MSR are copied into SRR1 and the address of the se_illegal instruction is placed into SRR0. MSR[WE,EE,PR,IS,DS,FP ,FE0,FE1] are cleared. The interrupt causes the next instruction to be fetched from address IVPR[32– This instruction is context synchronizing. Special Registers Altered: SRR0 SR R1 MSR[WE,EE,PR,IS,DS,FP ,FE0,FE1] 01 5 0000000000000000 VLE User

isel r D, rA, rB, crb if (rA = 0) then a ← 640 else a ← GPR(rA) c ← crcrb + 32 if c then rD ← a else rD ← GPR(rB) If CR[crb + 32] is set, the contents of rA|0 are copied into rD. If CR[crb + 32] is clear, the contents of rB are copied into rD. Integer Select APU User 0 5 6 1 0 1 11 5 1 62 0 2 12 5 2 63 0 3 1 011111 rD rA rB crb 011110

isync provides an ordering function for the effects of all instructions executed by the processor executing the isync instruction. Executing an isync ensures that all instructions preceding the isync have completed before isync completes, and that no subsequent instructions are initiated until after isync completes. It also causes any prefetched instructions to be discarded, with the effect that subsequent instructions are fetched and executed in the context established by the instructions preceding isync. isync may complete before memory accesses associated with instructions preceding isync have been performed. isync is context synchronizing. See Context synchronization on page 144.” Other registers altered: None Book E User 0 5 6 2 02 1 3 03 1 010011 / / / 0010010110 /

_isync _isync Instruction Synchronize se_isync The se_isync instruction provides an ordering function for the effects of all instructions executed by the processor executing the se_isync instruction. Executing an se_isync instruction ensures that all instructions preceding the se_isync instruction have completed before the se_isync instruction completes, and that no subsequent instructions are initiated until after the se_isync instruction completes. It also causes any prefetched instructions to be discarded, with the effect that subsequent instructions are fetched and executed in the context established by the instructions preceding the se_isync instruction. The se_isync instruction may complete before memory accesses associated with instructions preceding the se_isync instruction have been performed. This instruction is context synchronizing (see Book E). It has identical semantics to Book E isync, just a different encoding. Special Registers Altered: None 01 5 0000000000000001 VLE User

Load byte and zero [with update] [indexed] lbz r D,D(rA) (D-mode, U=0) lbzu r D,D(rA) (D-mode, U=1) lbzx r D,rA,rB (X-mode, U=0) lbzux r D,rA,rB (X-mode, U=1) if rA=0 then a ← 640 else a ← rA if D-mode then EA ← 320 || (a + EXTS(D))32:63 if X-mode then EA ← 320 || (a + rB)32:63 rD ← 560 || MEM(EA,1) if U=1 then rA ← EA The EA is calculated as follows:

  • For lbz and lbzu, EA is bits 32–63 of the sum of the contents of rA, or 64 zeros if rA=0, and the sign-extended value of the D field.
  • For lbzx and lbzux, EA is bits 32–63 of the sum of the contents of rA, or 64 zeros if rA=0, and the contents of rB. The byte in memory addressed by EA is loaded into rD[56–63]; rD[0–55] are cleared. If U=1 (with update), EA is placed into rA. If U=1 (with update), and rA=0 or rA=rD, the instruction form is invalid. Other registers altered: None Book E User 05 6 1 0 1 1 1 5 1 6 3 1 10001U rD rAD 0 5 6 1 01 1 1 51 6 2 02 1 2 42 52 6 3 03 1 011111 rD rA rB 0001U10111 /

_lbzx _lbzx Load Byte and Zero [with Update] [Indexed] e_lbz r D,D(rA) (D-mode) se_lbz r Z,SD4(rX) (SD4-mode) e_lbzu r D,D8(rA) (D8-mode) if (RA=0 & !se_lbz) then a ← 320 else a ← GPR(RA or RX) if D-mode then EA ← (a + EXTS(D))32:63 if D8-mode then EA ← (a + EXTS(D8))32:63 if SD4-mode then EA ← (a + (280 || SD4))32:63 GPR(RD or RZ) ← 240 || MEM(EA,1) if e_lbzu then GPR(RA) ← EA Let the EA be calculated as follows:

  • For e_lbz and e_lbzu, let EA be the sum of the contents of GPR(rA), or 32 0s if rA=0 , and the sign-extended value of the D or D8 instruction field.
  • For se_lbz, let EA be the sum of the contents of GPR(rX) and the zero-extended value of the SD4 instruction field. The byte in memory addressed by EA is loaded into bits 56–63 of GPR(rD or rZ). Bits 32–55 of GPR(rD or rZ) are cleared. If e_lbzu, EA is placed into GPR(rA). If e_lbzu and rA = 0 or rA= rD, the instruction form is invalid. Special Registers Altered: None VLE User 0 5 6 1 01 1 1 51 6 3 1

001100 R D R A D

0 5 6 1 01 1 1 51 6 2 32 4 3 1

000110 R D R A 00000000 D 8

Load floating-point double lfd fr D,D(rA) (D-mode, U=0) lfdu fr D,D(rA) (D-mode, U=1) lfdx fr D,rA,rB (X-mode, U=0) lfdux fr D,rA,rB (X-mode, U=1) if rA=0 then a ← 640 else a ← rA if D-mode then EA ← 320 || (a + EXTS(D))32:63 if X-mode then EA ← 320 || (a + rB)32:63 frD ← MEM(EA,8) if U=1 then rA ← EA The EA is calculated as follows:

  • For lfd and lfdu, EA is 32 zeros concatenated with bits 32–63 of the sum of the contents of rA, or 64 zeros if rA=0, and the sign-extended value of the D field.
  • For lfdx and lfdux, EA is 32 zeros concatenated with bits 32–63 of the sum of the contents of rA, or 64 zeros if rA=0, and the contents of rB. The double word addressed by EA is placed into frD. If U=1 (with update), EA is placed into register rA. If U=1 (with update) and rA=0, the instruction form is invalid. If MSR[FP]=0, an attempt to execute lfd[u][x] causes a floating-point unavailable interrupt. Other registers altered: None Book E User 05 6 1 0 1 1 1 5 1 6 3 1 11001U frD rAD 0 5 6 1 01 1 1 51 6 2 02 1 2 42 52 6 3 03 1 011111 frD rA rB 1001U10111 /

Load floating-point single lfs fr D,D(rA) (D-mode, U=0) lfsu fr D,D(rA) (D-mode, U=1) lfsx fr D,rA,rB (X-mode, U=0) lfsux fr D,rA,rB (X-mode, U=1) if rA=0 then a ← 640 else a ← rA if D-mode then EA ← 320 || (a + EXTS(D))32:63 if X-mode then EA ← 320 || (a + rB)32:63 frD ← DOUBLE(MEM(EA,4)) if U=1 then rA ← EA The EA is calculated as follows:

  • For lfs and lfsu, EA is 32 zeros concatenated with bits 32–63 of the sum of the contents of rA, or 64 zeros if rA=0, and the sign-extended value of the D field.
  • For lfsx and lfsux, EA is 32 zeros concatenated with bits 32–63 of the sum of the contents of rA, or 64 zeros if rA=0, and the contents of rB. The word addressed by EA is interpreted as a single-precision operand, converted to floating-point double format, and placed into frD. If U=1 (with update), EA is placed into register rA. If U=1 (with update) and rA=0, the instruction form is invalid. If MSR[FP]=0, an attempt to execute lfs[u][x] causes a floating-point unavailable interrupt. Other registers altered: None Book E User 0 4 5 6 10 11 15 16 31 11000U frD rAD 0 5 6 1 01 1 1 51 6 2 02 1 2 42 52 6 3 03 1 011111 frD rA rB 1000U10111 /

Load half word algebraic [with update] [indexed] lha r D,D(rA) (D-mode, U=0) lhau r D,D(rA) (D-mode, U=1) lhax r D,rA,rB (X-mode, U=0) lhaux r D,rA,rB (X-mode, U=1) if rA=0 then a ← 640 else a ← rA if D-mode then EA ← 320 || (a + EXTS(D))32:63 if X-mode then EA ← 320 || (a + rB)32:63 rD ← 320 || EXTS(MEM(EA,2))32:63 if U=1 then rA ← EA The EA is calculated as follows:

  • For lha and lhau, EA is bits 32–63 of the sum of the contents of rA, or 64 zeros if rA=0, and the sign-extended value of the D field.
  • For lhax and lhaux, EA is bits 32–63 of the sum of the contents of rA, or 64 zeros if rA=0, and the contents of rB. The half word addressed by EA is loaded into rD[48–63]. rD[32–47] are filled with a copy of bit 0 of the loaded half word. Bits rD[0–31] are cleared. If U=1 (with update), EA is placed into rA. If U=1 (with update), and rA=0 or rA=rD, the instruction form is invalid. Other registers altered: None Book E User 0 4 5 6 10 11 15 16 31 10101U rD rAD 0 5 6 1 01 1 1 51 6 2 02 1 2 42 52 6 3 03 1 011111 rD rA rB 0101U10111 /

_lhax _lhax Load Halfword Algebraic [with Update] [Indexed] e_lha r D,D(rA) (D-mode) e_lhau r D,D8(rA) (D8-mode) if RA=0 then a ← 320 else a ← GPR(RA) if D-mode then EA ← (a + EXTS(D))32:63 if D8-mode then EA ← (a + EXTS(D8))32:63 GPR(RD) ← EXTS(MEM(EA,2))32:63 if e_lhau then GPR(RA) ← EA Let the EA be calculated as follows:

  • For e_lha and e_lhau, let EA be the sum of the contents of GPR(rA), or 32 0s if rA=0 , and the sign-extended value of the D or D8 instruction field. The half word in memory addressed by EA is loaded into bits 48–63 of GPR(rD). Bits 32–47 of GPR(rD) are filled with a copy of bit 0 of the loaded half word. If e_lhau, EA is placed into GPR(rA). If e_lhau and rA = 0 or rA= rD, the instruction form is invalid. Special Registers Altered: None VLE User 0 5 6 1 01 1 1 51 6 3 1

001110 R D R A D

0 5 6 1 01 1 1 51 6 2 32 4 3 1

000110 R D R A 00000011 D 8

Load half word byte-reverse indexed lhbrx r D,rA,rB if rA=0 then a ← 640 else a ← rA data0:15 ← MEM(EA,2) rD ← 480 || data8:15 || data0:7 The EA is calculated as follows:

  • For lhbrx, EA is bits 32–63 of the sum of the contents of rA, or 64 zeros if rA=0, and the contents of rB. Bits 0–7 of the half word addressed by EA are loaded into rD[56–63]. Bits 8–15 of the half word addressed by EA are loaded into rD[48–55]; rD[0–47] are cleared. Other registers altered: None Programming notes:
  • When EA references big-endian memory, these instructions have the effect of loading data in little-endian byte order. Likewise, when EA references little-endian memory, these instructions have the effect of loading data in big-endian byte order.
  • In some implementations, the Load Half Word Byte-Reverse Indexed instructions may have greater latency than other load instructions. Book E User 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1 011111 rD rA rB 1100010110 /

Load half word and zero [with update] [indexed] lhz r D,D(rA) (D-mode, U=0) lhzu r D,D(rA) (D-mode, U=1) lhzx r D,rA,rB (X-mode, U=0) lhzux r D,rA,rB (X-mode, U=1) if rA=0 then a ← 640 else a ← rA if D-mode then EA ← 320 || (a + EXTS(D))32:63 if X-mode then EA ← 320 || (a + rB)32:63 rD ← 480 || MEM(EA,2) if U=1 then rA ← EA The EA is calculated as follows:

  • For lhz and lhzu, EA is bits 32–63 of the sum of the contents of rA, or 64 zeros if rA=0, and the sign-extended value of the D field.
  • For lhzx and lhzux, EA is bits 32–63 of the sum of the contents of rA, or 64 zeros if rA=0, and the contents of rB. The half word addressed by EA is loaded into rD[48–63]; rD[0–47] are cleared. If U=1 (with update), EA is placed into rA. If U=1 (with update), and rA=0 or rA=rD, the instruction form is invalid. Other registers altered: None Book E User 0 4 5 6 10 11 15 16 31 10100U rD rAD 0 5 6 1 01 1 1 51 6 2 02 1 2 42 52 6 3 03 1 011111 rD rA rB 0100U10111 /

_lhzx _lhzx Load Halfword and Zero [with Update] [Indexed] e_lhz r D,D(rA) (D-mode) se_lhz r Z,SD4(rX) (SD4-mode) e_lhzu r D,D8(rA) (D8-mode) if (RA=0 & !se_lhz) then a ← 320 else a ← GPR(RA or RX) if D-mode then EA ← (a + EXTS(D))32:63 if D8-mode then EA ← (a + EXTS(D8))32:63 if SD4-mode then EA ← (a + (270 || SD4 || 0))32:63 GPR(RD or RZ) ← 160 || MEM(EA,2) if e_lhzu then GPR(RA) ← EA Let the EA be calculated as follows:

  • For e_lhz and e_lhzu, let EA be the sum of the contents of GPR(rA), or 32 0s if rA=0 , and the sign-extended value of the D or D8 instruction field.
  • For se_lhz let EA be the sum of the contents of GPR(rX) and the zero-extended value of the SD4 instruction field shifted left by 1 bit. The half word in memory addressed by EA is loaded into bits 48–63 of GPR(rD). Bits 32–47 of GPR(rD) are cleared. If e_lhzu, EA is placed into GPR(rA). If e_lhzu and rA = 0 or rA= rD, the instruction form is invalid. Special Registers Altered: None VLE User 0 5 6 1 01 1 1 51 6 3 1

010110 R D R A D

0 5 6 1 01 1 1 51 6 2 32 4 3 1

000110 R D R A 00000001 D 8

_lix _lix Load Immediate [Shifted] e_li r D,LI20 (LI20-mode) LI20 ← LI200:3 || LI204:8 || LI209:19 GPR(RD) ← EXTS(LI20) For e_li, the sign-extended LI20 field is placed into GPR(rD). Special Registers Altered: None e_lis r D,UI UI ← UI 0:4 || UI5:15 GPR(RD) ← UI || 160 For e_lis, the UI field is concatenated on the right with 16 0’s and placed into GPR(rD). Special Registers Altered: None se_li r X,UI7 GPR(RX) ← 250 || UI7 For se_li, the zero-extended UI7 field is placed into GPR(rX). Special Registers Altered: None VLE User 0 5 6 1 01 1 1 51 61 7 2 02 1 3 1

011100 R D LI204:8 0 LI200:3 LI209:19

0 5 6 1 01 1 1 51 6 2 02 1 3 1

011100 R D UI0:4 11100 UI5:15

01001 U I 7 R X

lmw r D,D(rA) if rA=0 then EA ← 320 || EXTS(D)32:63 else EA ← 320 || (rA+EXTS(D))32:63 r ← rD do while r ≤ 31 GPR(r) ← 320 || MEM(EA,4) r ← r + 1 The EA is bits 32–63 of the sum of the contents of rA, or 64 zeros if rA=0, and the sign- extended value of the D instruction field. Here n=(32–rD). n consecutive words starting at EA are loaded into bits 32–63 of registers rD through GPR31. Bits 0–31 of these GPRs are cleared. EA must be a multiple of 4. If it is not, either an alignment interrupt is invoked or the results are boundedly undefined. If rA is in the range of registers to be loaded, including the case in which rA=0, the instruction form is invalid. Other registers altered: None Book E User 05 6 1 0 1 1 1 5 1 6 3 1 101110 rD rAD

_lmw _lmw Load Multiple Word e_lmw r D,D8(rA) if RA=0 then EA ← EXTS(D8)32:63 else EA ← (GPR(RA)+EXTS(D8))32:63 r ← RD do while r ≤ 31 GPR(r) ← MEM(EA,4) r ← r + 1 EA ← (EA+4)32:63 Let the EA be the sum of the contents of GPR(rA), or 32 0s if rA = 0, and the sign-extended value of the D8 instruction field. Let n = (32-rD). n consecutive words starting at EA are loaded into bits 32–63 of registers GPR(rD) through GPR(31). EA must be a multiple of 4. If it is not, either an alignment interrupt is invoked or the results are boundedly undefined. If rA is in the range of registers to be loaded, including the case in which rA = 0, the instruction form is invalid. Special Registers Altered: None VLE User 0 5 6 1 01 1 1 51 6 2 32 4 3 1

000110 R D R A 0 0 0 0 1 0 0 0 D 8

Load string word (immediate | indexed) lswi r D,rA,NB lswx r D,rA,rB if rA=0 then a ← 640 else a ← rA if ‘lswi’ then EA ← 320 || a32:63 if ‘lswx’ then EA ← 320 || (a + rB)32:63 if ‘lswi’ & NB=0 then n ← 32 if ‘lswi’ & NB≠0 then n ← NB if ‘lswx’ then n ← XER57:63 r ← rD : 1 i ← 32 rD ← undefined do while n > 0 if i = 32 then r ← r + 1 (mod 32) GPR(r) ← 0 GPR(r)i:i+7 ← MEM(EA,1) i ← i + 8 if i = 64 then i ← 32 n ← n : 1 The EA is calculated as follows:

  • For lswi, EA is 32 zeros concatenated with rA[32–63], or 32 zeros if rA=0.
  • For lwsx, EA is 32 zeros concatenated with bits 32–63 of the sum of the contents of rA, or 64 zeros if rA=0, and the contents of rB. If lswi, n = NB if NB ≠ 0, n = 32 if NB=0. If lswx, n=XER[57–63]. n is the number of bytes to load. Here nr=CEIL(n÷4): nr is the number of registers to receive data. If n>0, n consecutive bytes starting at EA are loaded into registers rD through (rD+nr–1). Data is loaded into the low-order 4 bytes of each GPR; the high-order 4 bytes are cleared. Bytes are loaded left to right in each GPR. The sequence wraps to GPR0 if required. If the 4 LSBs of GPR(rD+nr–1) are partially fill ed, the unfilled LSBs of that GPR are cleared. If lswx and n=0, the contents of rD are undefined. If rA, or rB for lswx, is in the range of registers to be loaded, including where rA=0, an illegal instruction type program interrupt is invoked or results are boundedly undefined. If rD=rA, or rD=rB for lswx, the instruction form is invalid. Other registers altered: None Note: Programming: String instructions move data without concern for alignment. They can perform short moves between arbitrary locations or long moves between misaligned memory fields. Book E User 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1 011111 rD rA N B 1001010101 / 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1 011111 rD rA rB 1000010101 /

Load word and reserve indexed lwarx r D,rA,rB if rA=0 then a ← 640 else a ← rA RESERVE ← 1 RESERVE_ADDR ← real_addr(EA) rD ← 320 || MEM(EA,4) EA is bits 32–63 of the sum of the contents of rA (32 zeros if rA=0), and the contents of rB. The word addressed by EA is loaded into rD[32–63]; rD[0–31] are cleared. lwarx creates a reservation for use by a stwcx. instruction. An address computed from the EA is associated with the reservation and replaces any previously associated address. See Atomic update primitives using lwarx and stwcx. on page 176.” If EA is not a multiple of 4, an alignment interrupt occurs or results are boundedly undefined. Other registers altered: None Programming notes:

  • lwarx, and stwcx. permit programmers to write an instruction sequence that appears to perform an atomic update operation on a memory location. This operation depends on a single reservation resource in each processor. At most one reservation exists on any given processor.
  • Because lwarx instructions have implementation dependencies (such as the granularity at which reservations are managed), they must be used with care. System library programs should use these instructions to implement high-level synchronization functions (such as test and set, compare and swap) needed by application programs. Application programs should use these library programs, rather than use lwarx directly The granularity with which reservations are managed is implementation-dependent. Therefore the location to be accessed by lwarx should be allocated by a system library program. See Atomic update primitives using lwarx and stwcx. on page 176.” Book E User 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1 011111 rD rA rB 0000010100 /

Load word byte-reverse indexed lwbrx r D,rA,rB if rA=0 then a ← 640 else a ← rA data0:31 ← MEM(EA,4) rD ← 320 || data24:31 || data16:23 || data8:15 || data0:7 The EA is calculated as follows:

  • For lwbrx, EA is bits 32–63 of the sum of the contents of rA, or 64 zeros if rA=0, and the contents of rB. Bits 0–7 of the word addressed by EA are loaded into rD[56–63]. Bits 8–15 of the word addressed by EA are loaded into rD[48–55]. Bits 16–23 of the word addressed by EA are loaded into rD[40–47]. Bits 24–31 of the word addressed by EA are loaded into rD[32–39]. Bits rD[0–31] are cleared. Other registers altered: None Programming notes:
  • When EA references big-endian memory, these instructions have the effect of loading data in little-endian byte order. Likewise, when EA references little-endian memory, these instructions have the effect of loading data in big-endian byte order.
  • In some implementations, the load word byte-reverse instructions may have greater latency than other load instructions. Book E User 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1 011111 rD rA rB 1000010110 /

Load word and zero [with update] [indexed] lwz r D,D(rA) (D-mode, U=0) lwzu r D,D(rA) (D-mode, U=1) lwzx r D,rA,rB (X-mode, U=0) lwzux r D,rA,rB (X-mode, U=1) if rA=0 then a ← 640 else a ← rA if D-mode then EA ← 320 || (a + EXTS(D))32:63 if X-mode then EA ← 320 || (a + rB)32:63 rD ← 320 || MEM(EA,4) if U=1 then rA ← EA The EA is calculated as follows:

  • For lwz and lwzu, EA is bits 32–63 of the sum of the contents of rA, or 64 zeros if rA=0, and the sign-extended value of the D field.
  • For lwzx and lwzux, EA is bits 32–63 of the sum of the contents of rA, or 64 zeros if rA=0, and the contents of rB. The word addressed by the EA is loaded into rD[32–63]; rD[0–31] are cleared. If U=1 (with update), EA is placed into rA. If U=1 (with update), and rA=0 or rA=rD, the instruction form is invalid. Other registers altered: None Book E User 05 6 1 0 1 1 1 5 1 6 3 1 10000U rD rAD 0 5 6 1 01 1 1 51 6 2 02 1 2 42 52 6 3 03 1 011111 rD rA rB 0000U10111 /

_lwz _lwz Load Word and Zero [with Update] [Indexed] e_lwz r D,D(rA) (D-mode) se_lwz r Z,SD4(rX) (SD4-mode) e_lwzu r D,D8(rA) (D8-mode) if (RA=0 & !se_lwz) then a ← 320 else a ← GPR(RA or RX) if D-mode then EA ← (a + EXTS(D))32:63 if D8-mode then EA ← (a + EXTS(D8))32:63 if SD4-mode then EA ← (a + (260 || SD4 || 20))32:63 GPR(RD or RZ) ← MEM(EA,4) if e_lwzu then GPR(RA) ← EA Let the EA be calculated as follows:

  • For e_lwz and e_lwzu, let EA be the sum of the contents of GPR(rA), or 32 0s if rA=0 , and the sign-extended value of the D or D8 instruction field.
  • For se_lwz let EA be the sum of the contents of GPR(rX) and the zero-extended value of the SD4 instruction field shifted left by 2 bits. The word in memory addressed by the EA is loaded into bits 32–63 of GPR( rD). If e_lwzu, EA is placed into GPR(rA). If e_lwzu and rA = 0 or rA= rD, the instruction form is invalid. Special Registers Altered: None VLE User 0 5 6 1 01 1 1 51 6 3 1

010100 R D R A D

0 5 6 1 01 1 1 51 6 2 32 4 3 1

000110 R D R A 00000010 D 8

When MO=0, mbar provides a memory ordering function for all memory access instructions executed by the processor executing the mbar instruction. Executing an mbar instruction ensures that all data memory accesses caused by instructions preceding the mbar have completed before any data memory accesses caused by any instructions after the mbar. This order is seen by all mechanisms. When mbar (MO = 1), as defined by the EIS, mbar functions like eieio as it is defined by the Classic PowerPC architecture. It provides ordering for the effects of load and store instructions. These instructions consist of two sets, which are ordered separately. Memory accesses caused by a dcbz or a dcba are ordered like a store. The two sets follow:

  • Caching-inhibited, guarded loads and stores to memory and write-through-required stores to memory. mbar (MO=1) controls the order in which accesses are performed in main memory. It ensures that all applicable memory accesses caused by instructions preceding the mbar have completed with respect to main memory before any applicable memory accesses caused by instructions following mbar access main memory. It acts like a barrier that flows through the memory queues and to main memory, preventing the reordering of memory accesses across the barrier. No ordering is performed for dcbz if the instruction causes the system alignment error handler to be invoked. All accesses in this set are ordered as one set; there is not one order for guarded, caching-inhibited loads and stores and another for write-through-required stores.
  • Stores to memory that are caching-allowed, write-through not required, and memory- coherency required. mbar (MO=1) controls the order in which accesses are performed with respect to coherent memory. It ensures that, with respect to coherent memory, applicable stores caused by instructions before the mbar complete before any applicable stores caused by instructions after it. Except for dcbz and dcba, mbar (MO=1) does not affect the order of cache operations (whether caused explicitly by a cache management instruction or implicitly by the cache coherency mechanism). Also. mbar does not affect the order of accesses in one set with respect to accesses in the other. mbar (MO=1) may complete before memory accesses caused by instructions preceding it have been performed with respect to main memory or coherent memory as appropriate. mbar (MO=1) is intended for use in managing shared data structures, in accessing memory-mapped I/O, and in preventing load/store combining operations in main memory. For the first use, the shared data structure and the lock that protects it must be altered only by stores that are in the same set (for both cases described above). For the second use, mbar (MO=1) can be thought of as placing a barrier into the stream of memory accesses issued by a core, such that any given memory access appears to be on the same side of the barrier to both the core and the I/O device. Because the core performs store operations in order to memory that is designated as both caching-inhibited and guarded, mbar (MO=1) is needed for such memory only when loads must be ordered with respect to stores or with respect to other loads. Book E User 0 5 6 1 0 1 12 0 2 13 0 3 1

011111 M O / / / 1101010110 /

Note that mbar (MO=1) does not connect hardware considerations to it such as multiprocessor implementations that send an mbar (MO=1) address-only broadcast (useful in some designs). For example, if a design has an external buffer that re-orders loads and stores for better bus efficiency, mbar (MO=1) broadcasts signals to that buffer that previous loads/stores (marked caching-inhibited, guarded, or write-through required) must complete before any following loads/stores (marked caching-inhibited, guarded, or write-through required). If MO is not 0 or 1, an implementation may support the mbar instruction ordering a particular subset of memory accesses. An implementation may also support multiple, non- zero values of MO that each specify a different subset of memory accesses that are ordered by the mbar instruction. Which subsets of memory accesses are ordered and which values of MO specify these subsets is implementation-dependent. See the user’s manual for the implementation. On some implementations, HID1[ABE] must be set to allow management of external L2 caches (for implementations with L2 caches) as well as other L1 caches in the system. Other registers altered: None Programming note: mbar is provided to implement a pipelined memory barrier. The following sequence shows one use of mbar in supporting shared data, ensuring the action is completed before releasing the lock. P1 P2 lock . . . read & write . . . mbar . . . free lock . . . ... l o c k . . . read & write ... m b a r ... f r e e l o c k

Move condition register field mcrf cr D,crS CR4xBF+32:4xBF+35 ← CR4xcrS+32:4xcrS+35 The contents of field crS (bits 4×crS+32–4×crS+35) of CR are copied to field crD (bits 4×crD+32–4×crD+35) of CR. Other registers altered: CR Book E User 0 5 6 8 9 1 01 1 1 31 4 2 02 1 3 03 1 010011 crD/ / crS / / / 0000000000 /

_mcrf _mcrf Move CR Field e_mcrf cr D,crS CR4xCRD+32:4xCRD+35 ← CR4xCRS+32:4xCRS+35 The contents of field crS (bits 4×CRS+32 through 4×CRS+35) of the CR are copied to field crD (bits 4×CRD+32 through 4×CRD+35) of the CR. Special Registers Altered: CR VLE User 0 5 6 8 9 1 01 1 1 31 4 2 02 1 3 03 1

011111 C R D / / C R S / / / 0000010000 /

Move to condition register from FPSCR mcrfs cr D,crS CRBF×4:crD×4+3 ← FPSCRcrS×4:crS×4+3 FPSCRcrS×4:crS×4+3 ← 0b0000 The contents of FPSCR[crS] are copied to CR field crD. All exception bits copied are cleared in the FPSCR. If the FX bit is copied, it is cleared in the FPSCR. If MSR[FP]=0, an attempt to execute mcrfs causes a floating-point unavailable interrupt. Other registers altered:

  • CR field crD FX OX(if crS=0) UX ZX XX VXSNAN(if crS=1) VXISI VXIDI VXZDZ VXIMZ(if crS=2) VXVC(if crS=3) VXSOFT VXSQRT VXCVI(if crS=5) Book E User 0 5 6 8 9 1 01 1 1 31 4 2 02 1 3 03 1 111111 crD/ / crS / / / 0001000000 /

Move to condition register from integer exception register mcrxr cr D CR4×crD+32:4×crD+35 ← XER32:35 XER32:35 ← 0b0000 The contents of XER[32–35] are copied to CR field crD. XER[32–35] are cleared. Other registers altered: CR XER[32–35] Book E User 0 5 6 8 9 2 02 1 3 03 1 011111 crD / / / 1000000000 /

mfapidi r D,rA rD ¨ implementation-dependent value based on rA The contents of rA are provided to any auxiliary processing extensions that may be present. A value, that is implementation-dependent and extension-dependent, is placed in rD. Other registers altered: None Programming note: This instruction is provided as a mechanism for software to query the presence and configuration of one or more auxiliary processing extensions. See user’s manual for the implementation for details on the behavior of this instruction. Book E User 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1 011111 rD rA / / / 0100010011 /

_mfar _mfar Move from Alternate Register se_mfar r X,arY GPR(RX) ← GPR(ARY) For se_mfar, the contents of GPR(arY) are placed into GPR(rX). arY specifies a GPR in the range R8–R23. The encoding 0000 specifies R8, 0001 specifies R9,…, 1111 specifies R23. Special Registers Altered: None 05 6 7 8 1 1 1 2 1 5

00000011 A R Y R X

Move from condition register mfcr r D rD ← 320 || CR The contents of the CR are placed into rD[32–63]. Bits rD[0–31] are cleared. Other registers altered: None Book E User 0 5 6 1 0 1 12 0 2 13 0 3 1 011111 rD / / / 0000010011 /

_mfctr _mfctr Move From Count Register se_mfctr r X GPR(RX) ← CTR The CTR contents are placed into bits 32–63 of GPR( rX). Special Registers Altered: None 05 6 1 1 1 2 1 5 000000 0 0 1 0 1 0 R X VLE User

Move from device control register mfdcr r D,DCRN rD ← DCREG(DCRN) DCRN identifies the DCR (see the user’s manual for a list of DCRs supported by the implementation). The contents of the designated DCR are placed into rD. For 32-bit DCRs, the contents of the DCR are placed into rD[32–63]. Bits rD[0–31] are cleared. Execution of this instruction is restricted to supervisor mode. Other registers altered: None Book E User 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1 011111 rD DCRN 5–9 DCRN0–4 0101000011 /

mffs fr D( R c = 0 ) mffs. fr D( R c = 1 ) frD ← FPSCR The contents of the FPSCR are placed into frD[32–63]; frD[0–31] are undefined. If MSR[FP]=0, an attempt to execute mffs[.] causes a floating-point unavailable interrupt. Other registers altered: Book E User 0 5 6 1 0 1 12 0 2 13 0 3 1 111111 frD / / / 1001000111 R c

_mflr _mflr Move From Link Register se_mflr r X GPR(RX) ← LR The LR contents are placed into bits 32–63 of GPR( rX). Special Registers Altered: None 05 6 1 1 1 2 1 5 000000 0 0 1 0 0 0 R X VLE User

Move from machine state register mfmsr r D rD ← 320 || MSR The contents of the MSR are placed into rD[32–63]. Bits rD[0–31] are cleared. Execution of this instruction is restricted to supervisor mode. Other registers altered: None Book E User 0 5 6 1 0 1 12 0 2 13 0 3 1 011111 rD / / / 0001010011 /

Move from Performance Monitor Register mfpmr r D,PMRN GPR(rD) ← PMREG(PMRN) PMRN denotes a performance monitor register. Section 2.16: Performance monitor registers (PMRs),” lists supported performance monitor registers. The contents of the designated performance monitor register are placed into GPR[rD]. When MSR[PR] = 1, specifying a performance monitor register that is not implemented and is not privileged (PMRN[5] = 0) results in an illegal instruction exception-type program interrupt. When MSR[PR] = 1, specifying a performance monitor register that is privileged (PMRN[5] = 1) results in a privileged instruction exception-type program interrupt. When MSR[PR] = 0, specifying an unimplemented performance monitor register is boundedly undefined. Other registers altered: None Book E User 05 6 1 0 1 1 1 5 1 6 2 0 2 1 3 1 011111 rD PMRN5–9 PMRN0–4 01010011100

SPRN denotes an SPR (see Chapter 2.18: Book E SPR model on page 130”). SPR are placed into rD[32–63]. Bits rD[0–31] are cleared. MSR[PR]=1 results in a privileged instruction exception-type program interrupt. Table 207. Effect of SPRN[5] and MSR[PR]

01 D e f i n e d If not implemented, illegal instruction exception

_mr _mr Move Register se_mr r X,rY GPR(RX) ← GPR(RY) For se_mr, the contents of GPR(rY) are placed into GPR(rX). Special Registers Altered: None 05 6 7 8 1 1 1 2 1 5

00000001 R Y R X

The msync instruction provides an ordering function for the effects of all instructions executed by the processor executing the msync. Executing msync ensures that all instructions preceding the msync have completed before msync completes and that no subsequent instructions are initiated until after the msync completes. It also creates a memory barrier (see Atomic update primitives using lwarx and stwcx. on page 176”), which orders the memory accesses associated with these instructions. The msync may not complete before memory accesses associated with instructions preceding msync have been performed. On some implementations, HID1[ABE] must be set to allow management of external L2 caches (for implementations with L2 caches) as well as other L1 caches in the system. msync is execution synchronizing. (See Execution synchronization on page 145.”) Other registers altered: None Programming notes:

  • msync can be used to ensure that all stores into a data structure, caused by store instructions executed in a critical section of a program, are performed with respect to another processor before the store that releases the lock is performed with respect to that processor. The functions performed by the msync may take a significant amount of time to complete, so indiscriminate use of this instruction may adversely affect performance. The Memory Barrier (mbar) instruction may be more appropriate than msync for many cases.
  • msync replaces the sync instruction; it uses the same opcode as sync such that PowerPC applications calling for sync invoke the msync when executed on an Book E implementation. The functionality of msync is identical to sync except that msync also does not complete until all previous memory accesses complete. mbar is provided in the Book E for those occasions when only ordering of memory accesses is required without execution synchronization. Book E User 0 5 6 2 02 1 3 03 1 011111 / / / 1001010110 /

_mtar _mtar Move to Alternate Register se_mtar ar X,rY GPR(ARX) ← GPR(RY) For se_mtar, the contents of GPR(rY) are placed into GPR(arX). arX specifies a GPR in the range R8–R23. The encoding 0000 specifies R8, 0001 specifies R9,…, 1111 specifies R23. Special Registers Altered: None 05 6 7 8 1 1 1 2 1 5

00000010 R Y A R X

Move to condition register fields mtcrf CRM,rS i ← 0 do while i < 8 if CRMi=1 then CR4×i+32:4×i+35 ← rS4×i+32:4×i+35 i ← i+1 The contents of rS[32–63] are placed into the CR under control of the field mask specified by CRM. The field mask identifies the 4-bit fields affected. Let i be an integer in the range 0– 7. If CRM i = 1, CR field i (CR bits 4×i+32 through 4×i+35) is set to the contents of the corresponding field of rS[32–63]. Other registers altered: CR fields selected by mask Book E User 0 5 6 1 01 11 2 1 92 02 1 3 03 1 011111 rS / C R M / 0010010000 /

_mtctr _mtctr Move To Count Register se_mtctr r X CTR ← GPR(RX) The contents of bits 32–63 of GPR( rX) are placed into the CTR. Special Registers Altered: CTR 05 6 1 1 1 2 1 5 000000 0 0 1 0 1 1 R X VLE User

Move to device control register mtdcr DCRN,rS DCREG(DCRN) ← rS DCRN identifies the DCR (see user’s manual for a list of DCRs supported by the implementation). The contents of rS are placed into the designated DCR. For 32-bit DCRs, rS[32–63] are placed into the DCR. Execution of this instruction is restricted to supervisor mode. Other registers altered: See the user’s manual for the implementation Book E User 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1 011111 rS DCRN 5–9 DCRN0–4 0111000011 /

mtfsb0 crb D( R c = 0 ) mtfsb0. crb D( R c = 1 ) FPSCR[BT]← 0b0 FPSCR[BT] is cleared. If MSR[FP]=0, an attempt to execute mtfsb0[.] causes a floating-point unavailable interrupt. Other registers altered:

  • FPSCR[BT] CR1 ← FX || FEX || VX || OX (if Rc=1) Programming note: Bits 1 and 2 (FEX and VX) cannot be explicitly reset. Book E User 0 5 6 1 0 1 12 0 2 13 0 3 1 111111 crbD / / / 0001000110 R c

mtfsb1 crb D( R c = 0 ) mtfsb1. crb D( R c = 1 ) FPSCR[BT] ← 0b1 FPSCR[BT] is set. If MSR[FP]=0, an attempt to execute mtfsb1[.] causes a floating-point unavailable interrupt. Other registers altered:

  • FPSClR[BT,FX] CR1 ← FX || FEX || VX || OX (if Rc=1) Programming note: Bits 1 and 2 (FEX and VX) cannot be explicitly set. Book E User 0 5 6 1 0 1 12 0 2 13 0 3 1 111111 crbD / / / 0000100110 R c

mtfsf FM,frB( R c = 0 ) mtfsf. FM,frB( R c = 1 ) i ← 0 do while i<8 if FMi=1 then FPSCR4×i:4×i+3 ← frB4×i:4×i+3 i ← i+1 The contents of frB[32–63] are placed into the FPSCR under control of the field mask specified by FM. The field mask identifies the 4-bit fields affected. Let i be an integer in the range 0–7. If FM i=1, FPSCR field i (FPSCR bits 4×i through 4×i+3) is set to the contents of the corresponding field of the low-order 32 bits of frB. FPSCR[FX] is altered only if FM0 = 1. If MSR[FP]=0, an attempt to execute mtfsf[.] causes a floating-point unavailable interrupt. Other registers altered:

  • FPSCR fields selected by mask CR1 ← FX || FEX || VX || OX (if Rc=1) Programming notes:
  • Updating fewer than all eight fields of the FPSCR may have substantially poorer performance on some implementations than updating all the fields.
  • When FPSCR[0–3] is specified, bits 0 (FX) and 3 (OX) are set to the values of (frB)32 and (frB)35 (that is, even if this instruction causes OX to change from 0 to 1, FX is set from (frB)32 and not by the usual rule that FX is set when an exception bit changes from 0 to 1). Bits 1 and 2 (FEX and VX) are set according to the usual rule (see Table 10: FPSCR field descriptions on page 59) and not from (frB) 33–34. Book E User 0 5 6 7 14 15 16 20 21 30 31 111111 / F M / frB 1011000111 R c

Move to FPSCR field immediate mtfsfi cr D,UIMM (Rc=0) mtfsfi. cr D,UIMM (Rc=1) FPSCRBF×4:crD×4+3 ← UIMM The value of the UIMM field is placed into FPSCR[crD]. FPSCR[FX] is altered only if crD = 0. If MSR[FP]=0, an attempt to execute mtfsfi[.] causes a floating-point unavailable interrupt. Other registers altered:

  • FPSCR[crD] CR1 ← FX || FEX || VX || OX (if Rc=1) Programming note: When FPSCR[0–3] is specified, bits 0 (FX) and 3 (OX) are set to the values of U0 and U3 (that is, even if this instruction causes OX to change from 0 to 1, FX is set from U0 and not by the usual rule that FX is set when an exception bit changes from 0 to 1). Bits 1 and 2 (FEX and VX) are set according to the usual rule (see Table 10: FPSCR field descriptions on page 59), and not from U1–2. Book E User 0 5 6 8 9 1 51 6 1 92 02 1 3 03 1 111111 crD / / / U I M M / 0010000110 R c

_mtlr _mtlr Move To Link Register se_mtlr r X LR ← GPR(RX) The contents of bits 32–63 of GPR( rX) are placed into the LR. Special Registers Altered: LR 05 6 1 1 1 2 1 5 000000 0 0 1 0 0 1 R X VLE User

Move to machine state register mtmsr r S MSR ← rS32:63 The contents of rS[32–63] are placed into the MSR. Execution of this instruction is restricted to supervisor mode. Execution of this instruction is execution synchronizing. See Execution synchronization on page 145.” In addition, changes to the EE or CE bits are effective as soon as the instruction completes. Thus if MSR[EE]=0 and an external interrupt is pending, executing an mtmsr that sets MSR[EE] causes the external interrupt to be taken before the next instruction is executed, if no higher priority exception exists. Likewise, if MSR[CE]=0 and a critical input interrupt is pending, executing an mtmsr that sets MSR[CE] causes the critical input interrupt to be taken before the next instruction is executed if no higher priority exception exists. Other registers altered: MSR Programming note: For a discussion of software synchronization requirements when altering certain MSR bits, refer to Chapter 2.18.2: Synchronization requirements for SPRs on page 130.” Book E Supervisor 0 5 6 1 0 1 12 0 2 13 0 3 1 011111 rS / / / 0010010010 /

Move To Performance Monitor Register mtpmr PMRN,rS PMREG(PMRN) ← GPR(RS) PMRN denotes a performance monitor register. Section 2.16: Performance monitor registers (PMRs),” lists supported performance monitor registers). The contents of GPR[rS] are placed into the designated performance monitor register. When MSR[PR] = 1, specifying a performance monitor register that is not implemented and is not privileged (PMRN[5] = 0) results in an illegal instruction exception-type program interrupt. When MSR[PR] = 1, specifying a performance monitor register that is privileged (PMRN[5] = 1) results in a privileged instruction exception-type program interrupt. When MSR[PR] = 0, specifying a unimplemented performance monitor register is boundedly undefined. Other registers altered: None Performance Monitor APU User/Supervisor 05 6 1 0 1 1 1 5 1 6 2 0 2 1 3 1 011111 rS PMRN5–9 PMRN0–4 01110011100

Move to special purpose register mtspr SPRN,rS SPREG(SPRN) ← rS SPRN denotes an SPR (see Chapter 2.18: Book E SPR model on page 130,” and the user’s manual of the implementation for a list of all SPRs that are implemented). The contents of rS are placed into the designated SPR. For 32-bit SPRs, the contents of rS[32–63] are placed into the SPR. When MSR[PR]=1, specifying an SPR that is not implemented and is not privileged (SPRN[5]=0) results in an illegal instruction exception-type program interrupt. When MSR[PR]=1, specifying an SPR that is privileged (SPRN[5]=1) results in a privileged instruction exception-type program interrupt. When MSR[PR]=0, specifying an SPR that is not implemented is boundedly undefined. Other registers altered: See Chapter 2.18: Book E SPR model on page 130,” or the user’s manual for the implementation. Programming note: For a discussion of software synchronization requirements when altering certain SPRs, please refer to Chapter 2.18.2: Synchronization requirements for SPRs on page 130.” Book E User/Supervisor 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1 011111 rS S P R N [ 5 – 9 ] S P R N [ 0 – 4 ] 0111010011 /

mulhw r D,rA,rB( R c = 0 ) mulhw. r D,rA,rB( R c = 1 ) prod0:63 ← rA32:63 × rB32:63 if Rc=1 then do LT ← prod0:31 < 0 GT ← prod0:31 > 0 EQ ← prod0:31 = 0 rD32:63 ← prod0:31 rD0:31 ← undefined Bits 0–31 of the 64-bit product of the contents of rA[32–63] and the contents of rB[32–63] are placed into rD[32–63]. Bits rD[0–31] are undefined. Both operands and the product are interpreted as signed integers. Other registers altered: CR0 (if Rc=1) Book E User 0 5 6 1 01 1 1 51 6 2 02 12 2 3 03 1 011111 rD rA rB / 001001011 R c

Multiply high word unsigned mulhwu r D,rA,rB( R c = 0 ) mulhwu. r D,rA,rB( R c = 1 ) prod0:63 ← rA32:63 × rB32:63 if Rc=1 then do LT ← prod0:31 < 0 GT ← prod0:31 > 0 EQ ← prod0:31 = 0 rD32:63 ← prod0:31 rD0:31 ← undefined Bits 0–31 of the 64-bit product the contents of rA[32–63] and the contents of rB[32–63] are placed into rD[32–63]. Bits rD[0–31] are undefined. Both operands and the product are interpreted as unsigned integers, except that if Rc=1 the first three bits of CR field 0 are set by signed comparison of the result to zero. Other registers altered: CR0 (if Rc=1) Book E User 0 5 6 1 01 1 1 51 6 2 02 12 2 3 03 1 011111 rD rA rB / 000001011 R c

mulli r D,rA,SIMM prod0:127 ← rA × EXTS(SIMM) rD ← prod64:127 Bits 64–127 of the 128-bit product of the contents of rA and the sign-extended value of the SIMM field are placed into rD. Both operands and the product are interpreted as signed integers. Other registers altered: None Programming notes:

  • For mulli, the low-order 64 bits of the product are independent of whether the operands are regarded as signed or unsigned 64-bit integers.
  • For mulli and mullw, bits 32–63 of the product are independent of whether the operands are regarded as signed or unsigned 32-bit integers. Book E User 05 6 1 0 1 1 1 5 1 6 3 1 000111 rD rAS I M M

_mullix _mullix Multiply Low [2 operand] Immediate e_mulli r D,rA,SCI8 imm ← SCI8(F ,SCL,UI8) prod0:63 ← GPR(RA) × imm GPR(RD) ← prod32:63 Bits 32–63 of the 64-bit product of the contents of GPR(rA) and the value of SCI8 are placed into GPR(rD). Both operands and the product are interpreted as signed integers. Special Registers Altered: None e_mull2i r A,SI prod 0:63 ← GPR(RA) × EXTS(SI0:4 || SI5:15) GPR(RA) ← prod32:63 Bits 32–63 of the 64-bit product of the contents of GPR( rA) and the sign-extended value of the SI field are placed into GPR(rA). Both operands and the product are interpreted as signed integers. Special Registers Altered: None VLE User 0 5 6 1 01 1 1 51 6 2 02 12 22 32 4 3 1

000110 R D R A 10100F S C L U I 8

0 5 6 1 01 1 1 51 6 2 02 1 3 1

011100 SI0:4 R A 1 0100 SI5:15

mullw r D,rA,rB( O E = 0 , R c = 0 ) mullw. r D,rA,rB( O E = 0 , R c = 1 ) mullwo r D,rA,rB( O E = 1 , R c = 0 ) mullwo. r D,rA,rB( O E = 1 , R c = 1 ) prod0:63 ← rA32:63 × rB32:63 if OE=1 then do OV ← (prod0:31 ≠ 320) & (prod0:31 ≠ 321) SO ← SO | OV if Rc=1 then do LT ← prod32:63 < 0 GT ← prod32:63 > 0 EQ ← prod32:63 = 0 rD ← prod0:63 The 64-bit product of the contents of rA[32–63] and the contents of rB[32–63] is placed into rD. If OE=1, OV is set if the product cannot be represented in 32 bits. Both operands and the product are interpreted as signed integers. Other registers altered:

  • CR0 (if Rc=1) SO OV (if OE=1) Programming notes:
  • Bits 32–63 of the product are independent of whether the operands are regarded as signed or unsigned 32-bit integers. Book E User 0 5 6 1 01 1 1 51 6 2 02 12 2 3 03 1 011111 rD rA rB O E 011101011 R c

_mullwx _mullwx Multiply Low Word se_mullw r X,rY prod0:63 ← GPR(RX)32:63 × GPR(RY)32:63 GPR(RX) ← prod32:63 Bits 32–63 of the 64-bit product of the contents of bits 32–63 of GPR(rX) and the contents of bits 32–63 of GPR(rY) is placed into GPR(rX). Special Registers Altered: None 05 6 7 8 1 1 1 2 1 5

00000101 R Y R X

nand r A,rS,rB( R c = 0 ) nand. r A,rS,rB( R c = 1 ) result0:63 ← ¬(rS & rB) if Rc=1 then do LT ← result32:63 < 0 GT ← result32:63 > 0 EQ ← result32:63 = 0 rA ← result The contents of rS are ANDed with the contents of rB and the one’s complement of the result is placed into rA. Other registers altered: CR0 (if Rc=1) Book E User 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1 011111 rS rA rB 0111011100 R c

neg r D,rA( O E = 0 , R c = 0 ) neg. r D,rA( O E = 0 , R c = 1 ) nego r D,rA( O E = 1 , R c = 0 ) nego. r D,rA( O E = 1 , R c = 1 ) carry0:63 ← Carry(¬rA + 1) sum0:63 ← ¬rA + 1 if OE=1 then do OV ← carry32 ⊕ carry33 SO ← SO | (carry32 ⊕ carry33) if Rc=1 then do LT ← sum32:63 < 0 GT ← sum32:63 > 0 EQ ← sum32:63 = 0 rD ← sum The sum of the one’s complement of the contents of rA and 1 is placed into rD. If rA contains the most negative 64-bit number (0x8000_0000_0000_0000), the result is the most negative number. Similarly, if rA[32–63] contain the most negative 32-bit number (0x8000_0000), bits 32–63 of the result contain the most negative 32-bit number and, if OE=1, OV is set. Other registers altered:

  • CR0 (if Rc=1) SO OV (if OE=1) Book E User 0 5 6 1 01 1 1 51 6 2 02 12 2 3 03 1 011111 rD rA / / / O E 001101000 R c

_negx _negx Negate se_neg r X result32:63 ← ¬GPR(RX)+ 1 GPR(RX) ← result32:63 The sum of the one’s complement of the contents of GPR(rX) and 1 is placed into GPR(rX). If bits 32–63 of GPR(rX) contain the most negative 32-bit number (0x8000_0000), bits 32– 63 of the result contain the most negative 32-bit number Special Registers Altered: None 05 6 1 1 1 2 1 5

000000000011 R X

nor r A,rS,rB( R c = 0 ) nor. r A,rS,rB( R c = 1 ) result0:63 ← ¬(rS | rB) if Rc=1 then do LT ← result32:63 < 0 GT ← result32:63 > 0 EQ ← result32:63 = 0 rA ← result The contents of rS are ORed with the contents of rB and the one’s complement of the result is placed into rA. Other registers altered: CR0 (if Rc=1) Book E User 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1 011111 rS rA rB 0001111100 R c

_notx _notx NOT se_not r X result32:63 ← ¬GPR(RX) GPR(RX) ← result32:63 The contents of GPR(rX) are inverted. Special Registers Altered: None 05 6 1 1 2 1 1 5

000000000010 R X

OR [Immediate [shifted] | with complement] or r A,rS,rB( R c = 0 ) or. r A,rS,rB( R c = 1 ) ori r A,rS,UIMM (S=0, Rc=0) oris r A,rS,UIMM (S=1, Rc=0) orc r A,rS,rB( R c = 0 ) orc. r A,rS,rB( R c = 1 ) if ‘ori’ then b ← 480 || UIMM if ‘oris’ then b ← 320 || UIMM || 160 if ‘or[.]’ then b ← rB if ‘orc[.]’ then b ← ¬rB result0:63 ← rS | b if Rc=1 then do LT ← result32:63 < 0 GT ← result32:63 > 0 EQ ← result32:63 = 0 rA ← result For ori, the contents of rS are ORed with 480 || UIMM. For oris, the contents of rS are ORed with 320 || UIMM || 160. For or[.], the contents of rS are ORed with the contents of rB. For orc[.], the contents of rS are ORed with the one’s complement of the contents of rB. The result is placed into rA. The preferred no-op is ori 0,0,0 Other registers altered: CR0 (if Rc=1) Book E User 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1 011111 rS rA rB 0110111100 R c 05 6 1 0 1 1 1 5 1 6 3 1 01100S rS rAU I M M 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1 011111 rS rA rB 0110011100 R c

_orx _orx OR [2 operand] [Immediate | with Complement] [Shifted][and Record] se_or r X,rY e_or2i r D,UI e_or2is r D,UI e_ori r A,rS,SCI8 (Rc = 0) e_ori. r A,rS,SCI8 (Rc = 1) if ‘e_ori[.]’ then b ← SCI8(F ,SCL,UI8) if ‘e_or2i’ then b ← 160 || UI0:4 || UI5:15 if ‘e_or2is’ then b ← UI0:4 || UI5:15 || 160 if ‘se_or’ then b ← GPR(RB) result0:63 ← GPR(RS or RD or RX) | b if Rc=1 then do LT ← result32:63 < 0 GT ← result32:63 > 0 EQ ← result32:63 = 0 GPR(RA or RD or RX) ← result For e_ori[.], the contents of GPR(rS) are ORed with the value of SCI8. For e_or2i, the contents of GPR(rD) are ORed with 160 || UI. For e_or2is, the contents of GPR(rD) are ORed with UI || 160. For se_or, the contents of GPR(rX) are ORed with the contents of GPR(rY). The result is placed into GPR(rA or rX). The preferred ‘no-op’ (an instruction that does nothing) is: e_ori 0,0,0 Special Registers Altered: CR0 (if Rc = 1) 05 6 7 8 1 1 1 2 1 5

01000100 R Y R X

0 5 6 1 01 1 1 51 6 2 02 1 3 1

011100 R D UI0:4 11000 UI5:15

0 5 6 1 01 1 1 51 6 2 02 1 3 1

011100 R D UI0:4 11010 UI5:15

0 5 6 1 01 1 1 51 6 1 92 02 12 22 32 4 3 1

000110 R S R A 1101 R c F S C L U I 8

Return from critical interrupt rfci MSR ← CSRR1 NIA ← CSRR0[0:61] || 0b00 The rfci instruction is used to return from a critical class interrupt, or as a means of establishing a new context and synchronizing on that new context simultaneously. The contents of CSRR1 are placed into the MSR. If the new MSR value does not enable any pending exceptions, then the next instruction is fetched, under control of the new MSR value, from the address CSRR0[0–61]||0b00. If the new MSR value enables one or more pending exceptions, the interrupt associated with the highest priority pending exception is generated; in this case, the value placed into SRR0 or CSRR0 by the interrupt processing mechanism is the address of the instruction that would have been executed next had the interrupt not occurred (that is, the address in CSRR0 at the time of the execution of the rfci). Execution of this instruction is restricted to supervisor mode. Execution of this instruction is context synchronizing. See Context synchronization on page 144.” Other registers altered: MSR Programming note: In addition to Branch to LR (bclr[l]) and Branch to CTR (bcctr[l]) instructions, rfi and rfci allow software to branch to any valid 64-bit address by using the respective 64-bit SRR0 and CSRR0. Book E Supervisor 0 5 6 2 02 1 3 03 1 010011 / / / 0000110011 /

_rfci _rfci Return From Critical Interrupt se_rfci MSR ← CSRR1 NIA ← CSRR00:62 || 0b0 The se_rfci instruction is used to return from a critical class interrupt, or as a means of establishing a new context and synchronizing on that new context simultaneously. The contents of CSRR1 are placed into the MSR. If the new MSR value does not enable any pending exceptions, then the next instruction is fetched, under control of the new MSR value, from the address CSRR0[32–62]||0b0. If the new MSR value enables one or more pending exceptions, the interrupt associated with the highest priority pending exception is generated; in this case the value placed into SRR0 or CSRR0 by the interrupt processing mechanism (see Book E) is the address of the instruction that would have been executed next had the interrupt not occurred (that is, the address in CSRR0 at the time of the execution of the se_rfci). Execution of this instruction is privileged and restricted to supervisor mode. Execution of this instruction is context synchronizing. Special Registers Altered: MSR 01 5 0000000000001001 VLE Supervisor

Return from debug interrupt rfdi if Mode32 then m ← 32 if Mode64 then m ← 0 MSR ← DSRR1 NIA ← m0 || DSRR0m:61 || 0b00 The rfdi instruction is used to return from a debug interrupt, or as a means of establishing a new context and synchronizing on that new context simultaneously. The contents of DSRR1 are placed into the MSR. If the new MSR value does not enable any pending exceptions, then the next instruction is fetched, under control of the new MSR value, from the address DSRR0[0–61]||0b00. If the new MSR value enables one or more pending exceptions, the interrupt associated with the highest priority pending exception is generated; in this case the value placed into SRR0, CSRR0, or DSRR0 by the interrupt processing mechanism is the address of the instruction that would have been executed next had the interrupt not occurred (that is, the address in DSRR0 at the time of the execution of the rfdi). Execution of this instruction is privileged and restricted to supervisor mode. Execution of this instruction is context synchronizing. Other registers altered:

  • MSR set as described above. Debug APU Supervisor 0 5 61 0 1 11 5 1 62 0 2 1 3 0 3 1 010011 /// 0000100111 /

MSR ← SRR1 NIA ← SRR0[0:61] || 0b00 The rfi instruction is used to return from a non-critical class interrupt, or as a means of simultaneously establishing a new context and synchronizing on that new context. The contents of SRR1 are placed into the MSR. If the new MSR value does not enable any pending exceptions, then the next instruction is fetched, under control of the new MSR value, from the address SRR0[0–61]||0b00. If the new MSR value enables one or more pending exceptions, the interrupt associated with the highest priority pending exception is generated; in this case the value placed into SRR0 or CSRR0 by the interrupt processing mechanism is the address of the instruction that would have been executed next had the interrupt not occurred (that is, the address in SRR0 at the time of the execution of the rfi). Execution of this instruction is restricted to supervisor mode. Execution of this instruction is context synchronizing. See Context synchronization on page 144.” Other registers altered: MSR Book E Supervisor 0 5 6 2 02 1 3 03 1 010011 / / / 0000110010 /

_rfi _rfi Return From Interrupt se_rfi MSR ← SRR1 NIA ← SRR00:62 || 0b0 The se_rfi instruction is used to return from a non-critical class interrupt, or as a means of simultaneously establishing a new context and synchronizing on that new context. The contents of SRR1 are placed into the MSR. If the new MSR value does not enable any pending exceptions, then the next instruction is fetched under control of the new MSR value from the address SRR0[32–62]||0b0. If the new MSR value enables one or more pending exceptions, the interrupt associated with the highest priority pending exception is generated; in this case the value placed into SRR0 or CSRR0 by the interrupt processing mechanism (see Book E) is the address of the instruction that would have been executed next had the interrupt not occurred (that is, the address in SRR0 at the time of the execution of the se_rfi). Execution of this instruction is privileged and restricted to supervisor mode. Execution of this instruction is context synchronizing. Special Registers Altered: MSR 01 5 0000000000001000 VLE Supervisor

Return from Machine Check Interrupt rfmci MSR ← MCSRR1 NIA ← MCSRR00:61 || 0b00 The rfmci instruction is used to return from a machine check interrupt, or as a means of simultaneously establishing a new context and synchronizing on that new context. The contents of machine check save/restore register 1 (MCSRR1) are placed into the MSR. If the new MSR value does not enable any pending exceptions, the next instruction is fetched, under control of the new MSR value from the address MCSRR0[32-61]|| 0b00. If the new MSR value enables one or more pending exceptions, the interrupt associated with the highest priority pending exception is generated; in this case the value placed into SRR0 or CSRR0 by the interrupt processing mechanism is the address of the instruction that would have been executed next had the interrupt not occurred (that is, the address in MCSRR0 at the time of the execution of rfi or rfci). Execution of this instruction is privileged and context synchronizing. Special registers altered: MSR Machine Check APU Supervisor 0 5 6 2 02 1 3 03 1 010011 / / / 0 0001001100

_rlw _rlw Rotate Left Word [Immediate] e_rlw r A,rS,rB( R c = 0 ) e_rlw. r A,rS,rB( R c = 1 ) e_rlwi r A,rS,SH (Rc = 0) e_rlwi. r A,rS,SH (Rc = 1) if ‘e_rlw[.]’ then n ← GPR(RB)59:63 else n ← SH result32:63 ← ROTL32(GPR(RS)32:63,n) if Rc=1 then do LT ← result32:63 < 0 GT ← result32:63 > 0 EQ ← result32:63 = 0 GPR(RA) ← result32:63 If e_rlw[.], let the shift count n be the contents of bits 59–63 of GPR( rB). If e_rlwi[.], let the shift count n be SH. The contents of GPR(rS) are rotated32 left n bits. The rotated data is placed into GPR(rA). Special Registers Altered: CR0 (if Rc = 1) VLE User 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1

011111 R S R A R B 0100011000 R c

0 5 6 1 01 1 1 51 6 2 02 1 3 03 1

011111 R S R A S H 0100111000 R c

Rotate left word immediate then mask insert rlwimi r A,rS,SH,MB,ME (Rc=0) rlwimi. r A,rS,SH,MB,ME (Rc=1) n ← SH b ← MB+32 e ← ME+32 r ← (rS32:63,n) m ← MASK(b,e) if Rc=1 then do LT ← result32:63 < 0 GT ← result32:63 > 0 EQ ← result32:63 = 0 rA ← result0:63 The shift count n is the value SH. The contents of rS are rotated32 left n bits. A mask is generated having 1 bits from bit MB+32 through bit ME+32 and 0 bits elsewhere. The rotated data is inserted into rA under control of the generated mask. (If a mask bit is 1 the associated bit of the rotated data is placed into the target register, and if the mask bit is 0 the associated bit in the target register remains unchanged.) Other registers altered: CR0 (if Rc=1) Programming note: Uses for rlwimi[.]:

  • To insert a k-bit field that is left-justified in rS[32–63], into rA[32–63] starting at bit position j, by setting SH=64-j, MB=j-32, and ME=(j+k)-33.
  • To insert an k-bit field that is right-justified in rS[32–63], into rA[32–63] starting at bit position j, by setting SH=64-(j+k), MB=j-32, and ME=(j+k)-33. Book E User 0 5 6 1 01 1 1 51 6 2 02 1 2 52 6 3 03 1 010100 rS rAS H M B M E R c

_rlwimi _rlwimi Rotate Left Word Immediate then Mask Insert e_rlwimi r A,rS,SH,MB,ME n ← SH b ← MB+32 e ← ME+32 r ← ROTL32(GPR(RS)32:63,n) m ← MASK(b,e) result32:63 ← r&m | GPR(RA)&¬m GPR(RA) ← result32:63 Let the shift count n be the value SH. The contents of GPR(rS) are rotated32 left n bits. A mask is generated having 1 bits from bit MB+32 through bit ME+32 and 0 bits elsewhere. The rotated data are inserted into GPR(rA) under control of the generated mask (if a mask bit is 1 the associated bit of the rotated data is placed into the target register, and if the mask bit is 0 the associated bit in the target register remains unchanged). Special Registers Altered: None VLE User 0 5 6 1 01 1 1 51 6 2 02 1 2 52 6 3 03 1

011101 R S R A S H M B M E 0

_rlwinm _rlwinm Rotate Left Word Immediate then AND with Mask e_rlwinm r A,rS,SH,MB,ME n ← SH b ← MB+32 e ← ME+32 r ← ROTL32(GPR(RS)32:63,n) m ← MASK(b,e) result32:63 ← r & m GPR(RA) ← result32:63 Let the shift count n be SH. The contents of GPR(rS) are rotated32 left n bits. A mask is generated having 1 bits from bit MB+32 through bit ME+32 and 0 bits elsewhere. The rotated data are ANDed with the generated mask and the result is placed into GPR(rA). Special Registers Altered: None VLE User 0 5 6 1 01 1 1 51 6 2 02 1 2 52 6 3 03 1

011101 R S R A S H M B M E 1

Rotate left word [immediate] then AND with mask rlwnm r A,rS,rB,MB,ME (Rc=0) rlwnm. r A,rS,rB,MB,ME (Rc=1) rlwinm r A,rS,SH,MB,ME (Rc=0) rlwinm. r A,rS,SH,MB,ME (Rc=1) if ‘rlwnm[.]’ then n ← rB59:63 else n ← SH b ← MB+32 e ← ME+32 r ← (rS32–63,n) m ← MASK(b,e) result0:63 ← r & m if Rc=1 then do LT ← result32:63 < 0 GT ← result32:63 > 0 EQ ← result32:63 = 0 rA ← result0:63 If rlwnm[.], the shift count, n, is the contents of rB[59–63]. If rlwinm[.], n is SH. The rS contents are rotated32 left n bits. The mask has 1s from bit MB+32 through bit ME+32 and 0s elsewhere. The rotated data is ANDed with the mask and the result is placed into rA. Other registers altered: CR0 (if Rc=1) Book E User 0 5 6 1 01 1 1 51 6 2 02 1 2 52 6 3 03 1 010111 rS rA rBM B M E R c 0 5 6 1 01 1 1 51 6 2 02 1 2 52 6 3 03 1 010101 rS rAS H M B M E R c Uses for rlwnm[.] Uses for rlwinm[.] To extract a k-bit field starting at bit position j in rS[32–63], right-justified into rA[32–63] (clearing the remaining 32– k bits of rA[32–63])… …by setting rB[59–63]=j+k-32, MB=32–k, and ME=31. …by setting SH= j+k-32, MB=32–k, and ME=31. To extract a k-bit field that starts at bit position j in rS[32–63], left-justified into rA[32–63] (clearing the remaining 32– k bits of rA[32–63])… …by setting rB[59–63]=j-32, MB=0, and ME=k–1. …by setting SH= j-32, MB=0, and ME=k–1. To rotate the contents of bits 32–63 of a register left by k bits… …setting rB[59–63]=k, MB=0, and ME=31. …setting SH= k, MB=0, and ME=31. To rotate the contents of bits 32–63 of a register right by k bits… …by setting rB[59–63] =32– k, MB=0, and ME=31. …by setting SH=32– k, MB=0, and ME=31.

To shift the contents of bits 32–63 of a register right by k bits, by setting SH=32–k, MB=k, and ME=31. To clear the high-order j bits of the contents of bits 32–63 of a register and then shift the result left by k bits, by setting SH=k, MB=j–k and ME=31–k. To clear the low-order k bits of bits 32–63 of a register, by setting SH=0, MB=0, and ME=31–k. For the uses given above, bits rA[0–31] are cleared. Uses for rlwnm[.] Uses for rlwinm[.]

SRR1 ← MSR SRR0 ← CIA+4 NIA ← EVPR[0:47] || IVOR8[48-59] || 0b0000 MSR[WE,EE,PR,IS,DS,FP ,FE0,FE1] ← 0b0000_0000 sc is used to request a system service. A system call interrupt is generated. The MSR contents are copied into SRR1 and the address of the instruction after the sc instruction is placed into SRR0. MSR[WE,EE,PR,IS,DS,FP ,FE0,FE1] are cleared. The interrupt causes the next instruction to be fetched from the address IVPR[0–47]||IVOR8[48-59]||0b0000. sc is context synchronizing. See Context synchronization on page 144.” Other registers altered: SRR0 SRR1 MSR[WE,EE,PR,IS,DS,FP ,FE0,FE1] Book E Supervisor 05 6 29 30 31 010001 / / / 1 /

slw r A,rS,rB( R c = 0 ) slw. r A,rS,rB( R c = 1 ) n ← rB59:63 r ← ROTL32(rS32:63,n) if rB58=0 then m ← MASK(32,63-n) else m ← 640 result0:63 ← r & m if Rc=1 then do LT ← result32:63 < 0 GT ← result32:63 > 0 EQ ← result32:63 = 0 rA ← result0:63 The shift count n is the value specified by the contents of rB[58–63]. The contents of rS[32–63] are shifted left n bits. Bits shifted out of position 32 are lost. Zeros are supplied to the vacated positions on the right. The 32-bit result is placed into rA[32–63]. Bits rA[0–31] are cleared. Shift amounts from 32 to 63 give a zero result. Other registers altered: CR0 (if Rc=1) Book E User 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1 011111 rS rA rB 0000011000 R c

_slwx _slwx Shift Left Word [Immediate] [and Record] e_slwi r A,rS,SH (Rc = 0) e_slwi. r A,rS,SH (Rc = 1) se_slw r X,rY se_slwi r X,UI5 if ‘e_slwi[.]’ then n ← SH if se_slw then n ← GPR(RY)58:63 if se_slwi then n ← UI5 r ← ROTL32(GPR(RS or RX)32:63,n) if n<32 then m ← MASK(32,63-n) else m ← 320 result32:63 ← r & m if Rc=1 then do LT ← result32:63 < 0 GT ← result32:63 > 0 EQ ← result32:63 = 0 GPR(RA or RX) ← result32:63 Let the shift count n be the value specified by the contents of bits 58–63 of GPR(rB or rY), or by the value of the SH or UI5 field. The contents of bits 32–63 of GPR(rS or rX) are shifted left n bits. Bits shifted out of position 32 are lost. Zeros are supplied to the vacated positions on the right. The 32-bit result is placed into bits 32–63 of GPR( rA or rX). Shift amounts from 32 to 63 give a zero result. Special Registers Altered: CR0 (if Rc = 1) VLE User 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1

011111 R S R A S H 0000111000 R c

01000010 R Y R X

0110110 U I 5 R X

_sc _sc System Call se_sc SRR1 ← MSR SRR0 ← CIA+2 NIA ← IVPR32:47 || IVOR848:59 || 0b0000 MSRWE,EE,PR,IS,DS,FP,FE0,FE1 ← 0b0000_0000 se_sc is used to request a system service. A system call interrupt is generated. The contents of the MSR are copied into SRR1 and the address of the instruction after the se_sc instruction is placed into SRR0. MSR[WE,EE,PR,IS,DS,FP ,FE0,FE1] are cleared. The interrupt causes the next instruction to be fetched from the address IVPR[32–47]||IVOR8[48–59]||0b0000 This instruction is context synchronizing. Special Registers Altered: SRR0 SR R1 MSR[WE,EE,PR,IS,DS,FP ,FE0,FE1] 01 5 0000000000000010 VLE User

Shift right algebraic word [immediate] sraw r A,rS,rB( R c = 0 ) sraw. r A,rS,rB( R c = 1 ) srawi r A,rS,SH (Rc=0) srawi. r A,rS,SH (Rc=1) if ‘sraw[.]’ then n ← rB59:63 else n ← SH r ← ROTL64(rS[32:63],64-n) if ‘sraw[.]’ & rB58=1 then m ← 640 else m ← MASK(n+32,63) s ← rS32 result0:63 ← r&m | (64s)&¬m if Rc=1 then do LT ← result32:63 < 0 GT ← result32:63 > 0 EQ ← result32:63 = 0 rA ← result0:63 If sraw[.], the shift count n is the contents of rB[58–63]. If srawi[.], the shift count n is the value of the SH field. The contents of rS[32–63] are shifted right n bits. Bits shifted out of position 63 are lost. Bit 32 of rS is replicated to fill the vacated positions on the left. The 32-bit result is placed into rA[32–63]. rS[32] is replicated to fill bits rA[0–31]. CA is set if rS[32–63] contain a negative value and any 1 bits are shifted out of bit position 63; otherwise CA is cleared. A shift amount of zero causes rA to receive EXTS(rS[32–63]), and CA to be cleared. For sraw[.] shift amounts from 32 to 63 give a result of 64 signed bits, and cause CA to receive rS[32] (that is, sign bit of rS[32–63]). Other registers altered: CA CR0 (if Rc=1) Book E User 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1 011111 rS rA rB 1100011000 R c 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1 011111 rS rA S H 1100111000 R c

_srawx _srawx Shift Right Algebraic Word [Immediate] [and Record] se_sraw r X,rY se_srawi r X,UI5 if ‘se_sraw’ then n ← GPR(RY)59:63 if ‘se_srawi’ then n ← UI5 r ← ROTL32(GPR(RS or RX)32:63,32-n) if ((se_sraw & GPR(RY)58=1) then m ← 320 else m ← MASK(n+32,63) s ← GPR(RS or RX)32 result0:63 ← r&m | (32s)&¬m if Rc=1 then do LT ← result32:63 < 0 GT ← result32:63 > 0 EQ ← result32:63 = 0 GPR(RA or RX) ← result32:63 If se_sraw, let the shift count n be the contents of bits 58–63 of GPR( rY). If se_srawi, let the shift count n be the value of the UI5 field. The contents of bits 32–63 of GPR( rS or rX) are shifted right n bits. Bits shifted out of position 63 are lost. Bit 32 of rS or rX is replicated to fill vacated positions on the left. The 32-bit result is placed into bits 32–63 of GPR( rA or rX). CA is set if bits 32–63 of GPR(rS or rX) contain a negative value and any 1 bits are shifted out of bit position 63; otherwise CA is cleared. A shift amount of zero causes GPR(rA or rX) to receive EXTS(GPR(rS or rX)32:63), and CA to be cleared. For se_sraw, shift amounts from 32 to 63 give a result of 64 sign bits, and cause CA to receive bit 32 of the contents of GPR(rS or rX) (that is, sign bit of GPR(rS or rX)32:63). Special Registers Altered: CA CR0 (if Rc = 1) 05 6 7 8 1 1 1 2 1 5

01000001 R Y R X

0110101 U I 5 R X

srw r A,rS,rB( R c = 0 ) srw. r A,rS,rB( R c = 1 ) n ← rB59–63 r ← ROTL64(rS32–63,64-n) if rB58=0 then m ← MASK(n+32,63) else m ← 640 result0:63 ← r & m if Rc=1 then do LT ← result32:63 < 0 GT ← result32:63 > 0 EQ ← result32:63 = 0 rA ← result0:63 The shift count n is the value specified by the contents of rB[58–63]. The contents of rS[32–63] are shifted right n bits. Bits shifted out of position 63 are lost. Zeros are supplied to the vacated positions on the left. The 32-bit result is placed into rA[32–63]. Bits rA[0–31] are cleared. Shift amounts from 32 to 63 give a zero result. Other registers altered: CR0 (if Rc=1) Book E User 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1 011111 rS rA rB 1000011000 R c

_srwx _srwx Shift Right Word [Immediate] [and Record] e_srwi r A,rS,SH (Rc = 0) e_srwi. r A,rS,SH (Rc = 1) se_srw r X,rY se_srwi r X,UI5 n ← GPR(RB)59:63 if ‘e_srwi[.]’ then n ← SH if ‘se_srw’ then n ← GPR(RY)59:63 if ‘se_srwi’ then n ← UI5 r ← ROTL32(GPR(RS or RX)32:63,32-n) if ((se_srw & GPR(RY)58=1) then m ← 320 else m ← MASK(n+32,63) result32:63 ← r & m if Rc=1 then do LT ← result32:63 < 0 GT ← result32:63 > 0 EQ ← result32:63 = 0 GPR(RA or RX) ← result32:63 If e_srwi, let the shift count n be the value of the SH field. If se_srw, let the shift count n be the contents of bits 58–63 of GPR( rY). If se_srwi, let the shift count n be the value of the UI5 field. The contents of bits 32–63 of GPR( rS or rX) are shifted right n bits. Bits shifted out of position 63 are lost. Zeros are supplied to the vacated positions on the left. The 32-bit result is placed into bits 32–63 of GPR( rA or rX). Shift amounts from 32 to 63 give a zero result. Special Registers Altered: CR0 (if Rc = 1) VLE User 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1

011111 R S R A S H 1000111000 R c

01000000 R Y R X

0110100 U I 5 R X

Store byte [with update] [indexed] stb r S,D(rA) (D-mode, I=0) stbu r S,D(rA) (D-mode, I=1) stbx r S,rA,rB (X-mode, U=0) stbux r S,rA,rB (X-mode, U=1) if rA=0 then a ← 640 else a ← rA if D-mode then EA ← 320 || (a + EXTS(D))32:63 if X-mode then EA ← 320 || (a + rB)32:63 MEM(EA,1) ← rS56:63 if U=1 then rA ← EA The EA is calculated as follows:

  • For stb and stbu, EA is bits 32–63 of the sum of the contents of rA, or 64 zeros if rA=0, and the sign-extended value of the D field.
  • For stbx and stbux, EA is bits 32–63 of the sum of the contents of rA, or 64 zeros if rA=0, and the contents of rB. The contents of rS[56–63] are stored into the byte addressed by EA. If U=1 (with update), EA is placed into rA. If U=1 (with update) and rA=0, the instruction form is invalid. Other registers altered: None Book E User 05 6 1 0 1 1 1 5 1 6 3 1 10011U rS rAD 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1 011111 rS rA rB 0011U10111 /

_stbx _stbx Store Byte [with Update] [Indexed] e_stb r S,D(rA) (D-mode) se_stb r Z,SD4(rX) (SD4-mode) e_stbu r S,D8(rA) (D8-mode) if (RA=0 & !se_stb) then a ← 320 else a ← GPR(RA or RX) if D-mode then EA ← (a + EXTS(D))32:63 if D8-mode then EA ← (a + EXTS(D8))32:63 if SD4-mode then EA ← (a + (280 || SD4))32:63 MEM(EA,1) ← GPR(RS or RZ)56:63 if e_stbu then GPR(RA) ← EA Let the EA be calculated as follows:

  • For e_stb and e_stbu, let EA be the sum of the contents of GPR(rA), or 32 0s if rA=0 , and the sign-extended value of the D or D8 instruction field.
  • For se_stb, let EA be the sum of the contents of GPR(rX) and the zero-extended value of the SD4 instruction field. The contents of bits 56–63 of GPR(rS) are stored into the byte in memory addressed by EA.
  • If e_stbu, EA is placed into GPR(rA).
  • If e_stbu and rA = 0, the instruction form is invalid.
  • None VLE User 0 5 6 1 01 1 1 51 6 3 1

001101 R S R A D

0 5 6 1 01 1 1 51 6 2 32 4 3 1

000110 R S R A 00000100 D 8

Store floating-point double [with update] [indexed] stfd fr S,D(rA) (D-mode, U=0) stfdu fr S,D(rA) (D-mode, U=1) stfdx fr S,rA,rB (X-mode, U=0) stfdux fr S,rA,rB (X-mode, U=1) if rA=0 then a ← 640 else a ← rA if D-mode then EA ← 320 || (a + EXTS(D))32:63 if X-mode then EA ← 320 || (a + rB)32:63 MEM(EA,8) ← frS if U=1 then rA ← EA The EA is calculated as follows:

  • For stfd and stfdu, EA is 32 zeros concatenated with bits 32–63 of the sum of the contents of rA, or 64 zeros if rA=0, and the sign-extended value of the D instruction field.
  • For stfdx and stfdux, EA is 32 zeros concatenated with bits 32–63 of the sum of the contents of rA, or 64 zeros if rA=0, and the contents of rB. The contents of frS are stored into the double word addressed by EA. If U=1 (with update), EA is placed into rA. If U=1 (with update) and rA=0, the instruction form is invalid. If MSR[FP]=0, stfd[u][x] causes a floating-point unavailable interrupt. Other registers altered: None Book E User 05 6 1 0 1 1 1 5 1 6 3 1 11011U frS rAD 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1 011111 frS rA rB 1011U10111 /

Store floating-point as integer word indexed stfiwx fr S,rA,rB if rA=0 then a ← 640 else a ← rA MEM(EA,4) ← frS[32:63] The EA is calculated as follows:

  • For stfiwx, EA is 32 zeros concatenated with bits 32–63 of the sum of the contents of rA, or 64 zeros if rA=0, and the contents of rB. The contents of frS[32–63] are stored, without conversion, into the word addressed by EA. If the contents of frS were produced, either directly or indirectly, by a load floating-point single instruction, a single-precision arithmetic instruction, or frsp, the value stored is undefined. (The contents of frS are produced directly by such an instruction if frS is the target register for the instruction. The contents of frS are produced indirectly by such an instruction if frS is the final target register of a sequence of one or more floating-point move instructions, with the input to the sequence having been produced directly by such an instruction.) If MSR[FP]=0, an attempt to execute stfiwx causes a floating-point unavailable interrupt. Other registers altered: None Book E User 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1 011111 frS rA rB 1111010111 /

Store floating-point single [with update] [indexed] stfs fr S,D(rA) (D-mode, U=0) stfsu fr S,D(rA) (D-mode, U=1) stfsx fr S,rA,rB (X-mode, U=0) stfsux fr S,rA,rB (X-mode, U=1) if rA=0 then a ← 640 else a ← rA if D-mode then EA ← 320 || (a + EXTS(D))32:63 if X-mode then EA ← 320 || (a + rB)32:63 MEM(EA,4) ← SINGLE(frS) if U=1 then rA ← EA The EA is calculated as follows:

  • For stfs and stfsu, EA is 32 zeros concatenated with bits 32–63 of the sum of the contents of rA, or 64 zeros if rA=0, and the sign-extended value of the D field.
  • For stfsx and stfsux, EA is 32 zeros concatenated with bits 32–63 of the sum of the contents of rA, or 64 zeros if rA=0, and the contents of rB. The contents of frS are converted to single format and stored into the word addressed by EA. If U=1 (with update), EA is placed into rA. If U=1 (with update) and rA=0, the instruction form is invalid. If MSR[FP]=0, stfs[u][x] causes a floating-point unavailable interrupt. Other registers altered: None Book E User 0 4 5 6 10 11 15 16 31 11010U frS rAD 0 5 6 1 01 1 1 51 6 2 02 1 2 42 52 6 3 03 1 011111 frS rA rB 1010U10111 /

Store half word [with update] [indexed] sth r S,D(rA) (D-mode, U=0) sthu r S,D(rA) (D-mode, U=1) sthx r S,rA,rB (X-mode, U=0) sthux r S,rA,rB (X-mode, U=1) if rA=0 then a ← 640 else a ← rA if D-mode then EA ← 320 || (a + EXTS(D))32:63 if X-mode then EA ← 320 || (a + rB)32:63 MEM(EA,2) ← rS48:63 if U=1 then rA ← EA The EA is calculated as follows:

  • For sth and sthu, EA is bits 32–63 of the sum of the contents of rA, or 64 zeros if rA=0, and the sign-extended value of the D field.
  • For sthx and sthux, EA is bits 32–63 of the sum of the contents of rA, or 64 zeros if rA=0, and the contents of rB. The contents of rS[48–63] are stored into the half word addressed by EA. If U=1 (with update), EA is placed into rA. If U=1 (with update) and rA=0, the instruction form is invalid. Other registers altered: None Book E User 0 4 5 6 10 11 15 16 31 10110U rS rAD 0 5 6 1 01 1 1 51 6 2 02 1 2 42 52 6 3 03 1 011111 rS rA rB 0110U10111 /

_sthx _sthx Store Halfword [with Update] [Indexed] e_sth r S,D(rA) (D-mode) se_sth r Z,SD4(rX) (SD4-mode) e_sthu r S,D8(rA) (D8-mode) if (RA=0 & !se_sth) then a ← 320 else a ← GPR(RA or RX) if D-mode then EA ← (a + EXTS(D))32:63 if D8-mode then EA ← (a + EXTS(D8))32:63 if SD4-mode then EA ← (a + (270 || SD4 || 0))32:63 MEM(EA,2) ← GPR(RS or RZ)48:63 if e_sthu then GPR(RA) ← EA Let the EA be calculated as follows:

  • For e_sth and e_sthu, let EA be the sum of the contents of GPR(rA), or 32 0s if rA=0 , and the sign-extended value of the D or D8 instruction field.
  • For se_sth let EA be the sum of the contents of GPR(rX) and the zero-extended value of the SD4 instruction field shifted left by 1 bit. The contents of bits 48–63 of GPR( rS) are stored into the half word in memory addressed by EA. If e_sthu, EA is placed into GPR(rA). If e_sthu and rA = 0, the instruction form is invalid. Special Registers Altered: None VLE User 0 5 6 1 01 1 1 51 6 3 1

010111 R S R A D

0 5 6 1 01 1 1 51 6 2 32 4 3 1

000110 R S R A 0 0 0 0 0 1 0 1 D 8

Store half word byte-reverse sthbrx r S,rA,rB if rA=0 then a ← 640 else a ← rA MEM(EA,2) ← rS56:63 || rS48:55 The EA is calculated as follows:

  • For sthbrx, EA is bits 32–63 of the sum of the contents of rA, or 64 zeros if rA=0, and the contents of rB. rS[56–63] are stored into bits 0–7 of the half word addressed by EA. Bits 48–55 of rS are stored into bits 8–15 of the half word addressed by EA. Other registers altered: None Programming note: When EA references big-endian memory, these instructions have the effect of storing data in little-endian byte order. Likewise, when EA references little-endian memory, these instructions have the effect of storing data in big-endian byte order. Book E User 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1 011111 rS rA rB 1110010110 /

stmw r S,D(rA) if rA=0 then EA ← 320 || EXTS(D)32:63 else EA ← 320 || (rA+EXTS(D))32:63 r ← rS do while r ≤ 31 MEM(EA,4) ← GPR(r)32:63 r ← r + 1 The EA is bits 32–63 of the sum of the contents of rA, or 64 zeros if rA=0, and the sign- extended value of the D instruction field. EA must be a multiple of 4. If it is not, either an alignment interrupt is invoked or the results are boundedly undefined. Other registers altered: None Book E User 05 6 1 0 1 1 1 5 1 6 3 1 101111 rS rAD

_stmw _stmw Store Multiple Word e_stmw r S,D8(rA) (D8-mode) if RA=0 then EA ← EXTS(D8)32:63 else EA ← (GPR(RA)+EXTS(D8))32:63 r ← RS do while r ≤ 31 MEM(EA,4) ← GPR(r)32:63 r ← r + 1 EA ← (EA+4)32:63 Let the EA be the sum of the contents of GPR(rA), or 32 0s if rA = 0, and the sign-extended value of the D8 instruction field. Let n = (32 - rS). Bits 32–63 of registers GPR(rS) through GPR(31) are stored in n consecutive words in memory starting at address EA. EA must be a multiple of 4. If it is not, either an alignment interrupt is invoked or the results are boundedly undefined. Special Registers Altered: None VLE User 0 5 6 1 01 1 1 51 6 2 32 4 3 1

000110 R S R A 00001001 D 8

Store string word (immediate | indexed) stswi r S,rA,NB stswx r S,rA,rB if rA=0 then a ← 640 else a ← rA if ‘stswi’ then EA ← 320 || a32:63 if ‘stswx’ then EA ← 320 || (a + rB)32:63 if ‘stswi’ & NB=0 then n ← 32 if ‘stswi’ & NB≠0 then n ← NB if ‘stswx’ then n ← XER57:63 r ← rS - 1 i ← 32 do while n > 0 if i=32 then r ← r + 1 (mod 32) MEM(EA,1) ← GPR(r)i:i+7 i ← i + 8 if i = 64 then i ← 32 n ← n - 1 The EA is calculated as follows:

  • For stswi, EA is 32 zeros concatenated with bits 32–63 of the contents of rA, or 32 zeros if rA=0.
  • For stswx, EA is 32 zeros concatenated with bits 32–63 of the sum of the contents of rA, or 64 zeros if rA=0, and the contents of rB. If stswi, let n=NB if NB≠0, n=32 if NB=0. If stswx, let n=XER[57–63]. n is the number of bytes to store. Let nr=CEIL(n÷4): nr is the number of registers to supply data. n consecutive bytes starting at EA are stored from registers rS through GPR(rS+nr–1). Data is stored from the low-order 4 bytes of each GPR. Bytes are stored left to right from each GPR. The register sequence can wrap to GPR0. If stswx and n=0, no bytes are stored. Other registers altered: None Programming note: Store string word and load string word instructions allow movement of data between memory and registers without concern for alignment. They can be used for a short move between arbitrary locations or long moves between misaligned memory fields. Book E User 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1 011111 rS rA N B 1011010101 / 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1 011111 rS rA rB 1010010101 /

Store word [with update] [indexed] stw r S,D(rA) (D-mode, U=0) stwu r S,D(rA) (D-mode, U=1) stwx r S,rA,rB (X-mode, U=0) stwux r S,rA,rB (X-mode, U=1) if rA=0 then a ← 640 else a ← rA if D-mode then EA ← 320 || (a + EXTS(D))32:63 if X-mode then EA ← 320 || (a + rB)32:63 MEM(EA,4) ← rS[32:63] if U=1 then rA ← EA The EA is calculated as follows:

  • For stw and stwu, EA is bits 32–63 of the sum of the contents of rA, or 64 zeros if rA=0, and the sign-extended value of the D field.
  • For stwx and stwux, EA is bits 32–63 of the sum of the contents of rA, or 64 zeros if rA=0, and the contents of rB. The contents of rS[32–63] are stored into the word addressed by EA. If U=1 (with update), EA is placed into rA. If U=1 (with update) and rA=0, the instruction form is invalid. Other registers altered: None Book E User 05 6 1 0 1 1 1 5 1 6 3 1 10010U rS rAD 0 5 6 1 01 1 1 51 6 2 02 1 2 42 52 6 3 03 1 011111 rS rA rB 0010U10111 /

_stwx _stwx Store Word [with Update] [Indexed] e_stw r S,D(rA) (D-mode) se_stw r Z,SD4(rX) (SD4-mode) e_stwu r S,D8(rA) (D8-mode) if (RA=0 & !se_stw) then a ← 320 else a ← GPR(RA or RX) if D-mode then EA ← (a + EXTS(D))32:63 if D8-mode then EA ← (a + EXTS(D8))32:63 if SD4-mode then EA ← (a + (260 || SD4 || 20))32:63 MEM(EA,4) ← GPR(RS or RZ)32:63 Let the EA be calculated as follows:

  • For e_stw and e_stwu, let EA be the sum of the contents of GPR(rA), or 32 0s if rA = 0, and the sign-extended value of the D or D8 instruction field.
  • For se_stw, let EA be the sum of the contents of GPR(rX) and the zero-extended value of the SD4 instruction field shifted left by 2 bits. The contents of bits 32–63 of GPR( rS) are stored into the word in memory addressed by EA. If e_stwu, EA is placed into GPR(rA). If e_stwu and rA = 0, the instruction form is invalid. Special Registers Altered: None VLE User 0 5 6 1 01 1 1 51 6 3 1

010101 R S R A D

0 5 6 1 01 1 1 51 6 2 32 4 3 1

000110 R S R A 00000110 D 8

stwbrx r S,rA,rB if rA=0 then a ← 640 else a ← rA MEM(EA,4) ← rS56:63 || rS48:55 || rS40:47 || rS32:39 The EA is calculated as follows:

  • For stwbrx, EA is bits 32–63 of the sum of the contents of rA, or 64 zeros if rA=0, and the contents of rB. Bits 56–63 of rS are stored into bits 0–7 of the word addressed by EA. Bits 48–55 of rS are stored into bits 8–15 of the word addressed by EA. Bits 40–47 of rS are stored into bits 16– 23 of the word addressed by EA. Bits 32–39 of rS are stored into bits 24–31 of the word addressed by EA. Other registers altered: None Programming note: When EA references big-endian memory, these instructions have the effect of storing data in little-endian byte order. Likewise, when EA references little-endian memory, these instructions have the effect of storing data in big-endian byte order. Book E User 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1 011111 rS rA rB 1010010110 /

stwcx. stwcx. Store word conditional indexed stwcx. rS,rA,rB if rA=0 then a ← 640 else a ← rA if RESERVE then if RESERVE_ADDR = real_addr(EA) then MEM(EA,4) ← rS[32:63] CR0 ← 0b00 || 0b1 || XERSO else u ← undefined 1-bit value if u then MEM(EA,4) ← rS[32:63] RESERVE ← 0 else CR0 ← 0b00 || 0b0 || XERSO The EA is calculated as follows:

  • For stwcx., EA is bits 32–63 of the sum of the contents of rA, or 64 zeros if rA=0, and the contents of rB. If a reservation exists and the address specified by the stwcx. is the same as that specified by the lwarx instruction that established the reservation, the contents of rS[32–63] are stored into the word addressed by EA and the reservation is cleared. If a reservation exists but the address specified by stwcx. is not the same as that specified by the load and reserve instruction that established the reservation, the reservation is cleared, and it is undefined whether the instruction completes without altering memory. If a reservation does not exist, the instruction completes without altering memory. CR field 0 is set to reflect whether the store operation was performed, as follows: CR0[LT,GT,EQ,SO] = 0b00 || store_performed || XER[SO] EA must be a multiple of 4. If it is not, either an alignment interrupt is invoked or the results are boundedly undefined. Other registers altered: CR0 Programming notes:
  • stwcx., in combination with lwarx, permits the programmer to write a sequence of instructions that appear to perform an atomic update operation on a memory location. This operation depends on a single reservation resource in each processor. At most one reservation exists on any given processor: there are not separate reservations for words and for double words.
  • Because stwcx. instructions have implementation dependencies (such as the granularity at which reservations are managed), they must be used with care. The operating system should provide system library programs that use these instructions to implement the high-level synchronization functions (such as, test and set, and compare Book E User 0 5 6 1 01 1 1 51 6 2 02 1 3 1 011111 rS rA rB 00100101101

and swap) needed by application programs. Application programs should use these library programs, rather than use stwcx. directly.

  • The granularity with which reservations are managed is implementation-dependent. Therefore, the memory to be accessed by stwcx. should be allocated by a system library program. Additional information can be found in Atomic update primitives using lwarx and stwcx. on page 176.”
  • When correctly used, the load and reserve and store conditional instructions can provide an atomic update function for a single aligned word (lwarx and stwcx.) of memory. In general, correct use requires that lwarx be paired with stwcx. with the same address specified by both instructions of the pair. The only exception is that an unpaired stwcx. to any (scratch) effective address can be used to clear any reservation held by the processor. Examples of correct uses of these instructions to emulate primitives such as fetch and add, test and set, and compare and swap can be found in Appendix C: Programming examples on page 1143. A reservation is cleared if any of the following events occur: – The processor holding the reservation executes another load and reserve instruction; this clears the first reservation and establishes a new one. – The processor holding the reservation executes a store conditional instruction to any address. – Another processor executes any store instruction to the address associated with the reservation. – Any mechanism, other than the processor holding the reservation, stores to the address associated with the reservation. See Atomic update primitives using lwarx and stwcx. on page 176,” for additional information.

_sub _sub Subtract se_sub r X,rY sum32:63 ← GPR(RX) + ¬GPR(RY) + 1 GPR(RX) ← sum32:63 The sum of the contents of GPR(rX), the one’s complement of contents of GPR(rY), and 1 is placed into GPR(rX). Special Registers Altered: None 05 6 7 8 1 1 1 2 1 5

00000110 R Y R X

subf r D,rA,rB( O E = 0 , R c = 0 ) subf. r D,rA,rB( O E = 0 , R c = 1 ) subfo r D,rA,rB( O E = 1 , R c = 0 ) subfo. r D,rA,rB( O E = 1 , R c = 1 ) carry0:63 ← Carry(¬rA + rB + 1) sum0:63 ← ¬rA + rB + 1 if OE=1 then do OV ← carry32 ⊕ carry33 SO ← SO | (carry32 ⊕ carry33) if Rc=1 then do LT ← sum32:63 < 0 GT ← sum32:63 > 0 EQ ← sum32:63 = 0 rD ← sum The sum of the one’s complement of the contents of rA, the contents of rB, and 1 is placed into rD. Other registers altered:

  • CR0 (if Rc=1) SO OV (if OE=1) Book E User 0 5 6 1 01 1 1 51 6 2 02 12 2 3 03 1 011111 rD rA rB O E 000101000 R c

_subfx _subfx Subtract From se_subf r X,rY sum32:63 ← ¬GPR(RX) + GPR(RY) + 1 GPR(RX) ← sum32:63 The sum of the one’s complement of the contents of GPR(rX), the contents of GPR(rY), and 1 is placed into GPR(rX). Special Registers Altered: None 05 6 1 0 1 1 1 5 VLE User

subfc r D,rA,rB( O E = 0 , R c = 0 ) subfc. r D,rA,rB( O E = 0 , R c = 1 ) subfco r D,rA,rB( O E = 1 , R c = 0 ) subfco. r D,rA,rB( O E = 1 , R c = 1 ) carry0:63 ← Carry(¬rA + rB + 1) sum0:63 ← ¬rA + rB + 1 if OE=1 then do OV ← carry32 ⊕ carry33 SO ← SO | (carry32 ⊕ carry33) if Rc=1 then do LT ← sum32:63 < 0 GT ← sum32:63 > 0 EQ ← sum32:63 = 0 rD ← sum CA ← carry32 The sum of the one’s complement of the contents of rA, the contents of rB, and 1 is placed into rD. Other registers altered:

  • CA CR0 (if Rc=1) SO OV (if OE=1) Book E User 0 5 6 1 01 1 1 51 6 2 02 12 2 3 03 1 011111 rD rA rB O E 000001000 R c

subfe r D,rA,rB( O E = 0 , R c = 0 ) subfe. r D,rA,rB( O E = 0 , R c = 1 ) subfeo r D,rA,rB( O E = 1 , R c = 0 ) subfeo. r D,rA,rB( O E = 1 , R c = 1 ) if E=0 then Cin ← CA carry0:63 ← Carry(¬rA + rB + Cin) sum0:63 ← ¬rA + rB + Cin if OE=1 then do OV ← carry32 ⊕ carry33 SO ← SO | (carry32 ⊕ carry33) if Rc=1 then do LT ← sum32:63 < 0 GT ← sum32:63 > 0 EQ ← sum32:63 = 0 rD ← sum CA ← carry32 For subfe[o][.], the sum of the one’s complement of the contents of rA, the contents of rB, and CA is placed into rD. Other registers altered:

  • CA CR0 (if Rc=1) SO OV (if OE=1) Book E User 0 5 6 1 01 1 1 51 6 2 02 12 2 3 03 1 011111 rD rA rB O E 010001000 R c

Subtract from immediate carrying subfic r D,rA,SIMM carry0:63 ← Carry(¬rA + EXTS(SIMM) + 1) sum0:63 ← ¬rA + EXTS(SIMM) + 1 rD ← sum CA ← carry32 The sum of the one’s complement of the contents of rA, the sign-extended value of the SIMM field, and 1 is placed into rD. Other registers altered: CA Book E User 05 6 1 0 1 1 1 5 1 6 3 1 001000 rD rAS I M M

_subficx _subficx Subtract From Immediate Carrying [and Record] e_subfic r D,rA,SCI8 (Rc = 0) e_subfic. r D,rA,SCI8 (Rc = 1) imm ← SCI8(F ,SCL,UI8) carry32:63 ← Carry(¬GPR(RA) + imm + 1) sum32:63 ← ¬GPR(RA) + imm + 1 if Rc=1 then do LT ← sum 32:63 < 0 GT ← sum32:63 > 0 EQ ← sum32:63 = 0 GPR(RD) ← sum32:63 CA ← carry32 The sum of the one’s complement of the contents of GPR(rA), the value of SCI8, and 1 is placed into GPR(rD). Special Registers Altered: CA CR0 (if Rc=1) VLE User 0 5 6 1 01 1 1 51 6 2 02 12 22 32 4 3 1

000110 R D R A 1011 R c F S C L U I 8

Subtract from minus one extended subfme r D,rA( O E = 0 , R c = 0 ) subfme. r D,rA( O E = 0 , R c = 1 ) subfmeo r D,rA( O E = 1 , R c = 0 ) subfmeo. r D,rA( O E = 1 , R c = 1 ) if E=0 then Cin ← CA carry0:63 ← Carry(¬rA + Cin + 0xFFFF_FFFF_FFFF_FFFF) sum0:63 ← ¬rA + Cin + 0xFFFF_FFFF_FFFF_FFFF if OE=1 then do OV ← carry32 ⊕ carry33 SO ← SO | (carry32 ⊕ carry33) if Rc=1 then do LT ← sum32:63 < 0 GT ← sum32:63 > 0 EQ ← sum32:63 = 0 rD ← sum CA ← carry32 For subfme[o][.], the sum of CA, 641, and the one’s complement of the contents of rA is placed into rD. Other registers altered:

  • CA CR0 (if Rc=1) SO OV (if OE=1) Book E User 0 5 6 1 01 1 1 51 6 2 02 12 2 3 03 1 011111 rD rA / / / O E 011101000 R c

Subtract from zero extended subfze r D,rA( O E = 0 , R c = 0 ) subfze. r D,rA( O E = 0 , R c = 1 ) subfzeo r D,rA( O E = 1 , R c = 0 ) subfzeo. r D,rA( O E = 1 , R c = 1 ) if E=0 then Cin ← CA carry0:63 ← Carry(¬rA + Cin) sum0:63 ← ¬rA + Cin if OE=1 then do OV ← carry32 ⊕ carry33 SO ← SO | (carry32 ⊕ carry33) if Rc=1 then do LT ← sum32:63 < 0 GT ← sum32:63 > 0 EQ ← sum32:63 = 0 rD ← sum CA ← carry32 For subfze[o][.], the sum of the one’s complement of the contents of rA and CA is placed into rD. Other registers altered:

  • CA CR0 (if Rc=1) SO OV (if OE=1) Book E User 0 5 6 1 01 1 1 51 6 2 02 12 2 3 03 1 011111 rD rA / / / O E 011001000 R c

_subix _subix Subtract Immediate [and Record] se_subi r X,OIMM (Rc = 0) se_subi. r X,OIMM (Rc = 1) sum32:63 ← GPR(RX) + ¬(270 || OFFSET(OIM5)) + 1 if Rc=1 then do LT ← sum32:63 < 0 GT ← sum32:63 > 0 EQ ← sum32:63 = 0 GPR(RX) ← sum32:63 The sum of the contents of GPR(rX), the one’s complement of the zero-extended value of the offseted OIM5 field (a final value in the range 1–32), and 1 is placed into GPR( rX). Special Registers Altered: CR0 (if Rc = 1) 0 5 6 7 11 12 15

001001 R c OIM5(1)

  1. OIMM = OIM5 +1 RX VLE User

TLB Invalidate virtual address indexed tlbivax r A,rB if rA=0 then a ← 640 else a ← rA AS ← implementation-dependent value ProcessID ← implementation-dependent value VA ← AS || ProcessID || EA InvalidateTLB(VA) EIS note: Executing tlbivax invalidates any TLB entry that corresponds to a virtual address calculated by this instruction if IPROT is not set; this includes invalidating TLB entries on other devices as well as on the processor executing tlbivax. Thus an invalidate operation is broadcast throughout the coherent domain of the processor executing tlbivax. On some implementations, HID1[ABE] must be set to allow management of external L2 caches (for implementations with L2 caches) as well as other L1 caches in the system. EA calculation: Addressing ModeEA for rA=0EA for rA≠0 320 || rB32:63 Address space (AS) is defined as implementation-dependent (for example, it could be MSR[DS] or a bit from an implementation-dependent SPR). ProcessID is implementation-dependent (for example, it could be from the PID or from an implementation-dependent SPR). The EIS implements the architected PID and additional implementation-specific PIDs. See Section 2.12.1: Process ID registers (PID0–PIDn).” The virtual address (VA) is the value AS || ProcessID || EA. A TLB entry corresponding to VA is made invalid (that is, removed from the TLB). This instruction causes the target TLB entry to be invalidated in all processors. The operation performed by this instruction is ordered by mbar (or msync) with respect to a subsequent tlbsync executed by the processor executing tlbivax. Operations caused by tlbivax and tlbsync are ordered by mbar as a set of operations independent of the other sets that mbar orders. Other registers altered: None Programming notes:

  • The effects of the invalidation are not guaranteed to be visible to the programming model until the completion of a context synchronizing operation. See Context synchronization on page 144.”
  • Care must be taken not to invalidate TLB entries that contain interrupt vector mappings. Book E Supervisor 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1 011111 / / / rA rB 1100010010 /

The RTL for the EIS definition of tlbre is as follows: tlb_entry_id = MAS0(TLBSEL, ESEL | MAS2(EPN) result = MMU(tlb_entry_id) MAS0, MAS1, MAS2, MAS3, (and MAS7 if HID0[EN_MAS7_UPDATE] = 1) = result Bits 6–20 of the encoding are allocated for implementation-dependent use and may be used to specify the source TLB entry, the source portion of the source TLB entry, and the target resource into which the result is placed. The EIS makes no use of these bits. The implementation-defined TLB entry is read, and the implementation-defined portion of the TLB entry is extracted and placed into an implementation-defined target resource. If the instruction specifies a TLB entry that does not exist, the results are undefined. EIS implementation note: tlbre causes the contents of a single TLB entry to be extracted from the MMU and be placed in the corresponding fields of the MMU assist (MAS) registers. The entry extracted is specified by the TLBSEL, ESEL and EPN fields of MAS0 and MAS2. The contents extracted from the MMU are placed in MAS0–MAS3. See the user’s manual for the implementation. Execution of this instruction is restricted to supervisor mode. Other registers altered: MAS0, MAS1, MAS2, and MAS3, as defined by the EIS Book E Supervisor 0 5 6 2 02 1 3 03 1 011111 / / / (1) 1110110010/ 1 1. This field is defined as allocated by the Book E architecture, for possible use in an implementation. These bits are not implemented by the EIS.

tlbsx r A,rB if RA!=0 then generate exception EA = 320 || GPR(RB)32:63 ProcessID = MAS6(SPID) AS = MAS6(SAS) VA0 = AS || (MMUCFG[PIDSIZE] + 1)0 || EA VA1 = AS || ProcessID || EA if Valid_TLB_matching_entry_exists (VA0) or Valid_TLB_matching_entry_exists (VA1) MAS0, MAS1, MAS2, MAS3 = result EA calculation: Addressing ModeEA for rA=0EA for rA≠0 320 || rB32:63 Note that rA = 0 is a preferred form for tlbsx and that some ST implementations take an illegal instruction exception program interrupt if rA != 0. Virtual address 0 (VA0) is the value AS || (MMUCFG[PIDSIZE] + 1)0 || EA Virtual address 1 (VA1) is the value AS || ProcessID || EA If the TLB contains an entry corresponding to VA, an implementation-dependent value is placed into an implementation-dependent-specified target. Otherwise the contents of the implementation-dependent-specified target are left undefined. Other registers altered: implementation-dependent. See Supervisor-level tlb management instructions on page 183. Book E Supervisor 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1 011111 / / / (1) rA rB 1110010010 / 1 1. This field is defined as allocated by the Book E architecture, for possible use in an implementation. These bits are not implemented by the EIS.

tlbsync provides an ordering function for the effects of all tlbivax instructions executed by the processor executing tlbsync, with respect to the memory barrier created by a subsequent msync instruction executed by the same processor. Executing tlbsync ensures that all of the following occur.

  • All TLB invalidations caused by tlbivax instructions preceding the tlbsync instruction will have completed on any other processor before any memory accesses associated with data accesses caused by instructions following the msync instruction are performed with respect to that processor.
  • All memory accesses by other processors for which the address was translated using the translations being invalidated, will have been performed with respect to the processor executing the msync instruction, to the extent required by the associated memory-coherence required attributes, before the mbar or msync instruction’s memory barrier is created. The operation performed by this instruction is ordered by the mbar and msync instructions with respect to preceding tlbivax instructions executed by the processor executing the tlbsync instruction. The operations caused by tlbivax and tlbsync are ordered by mbar as a set of operations that is independent of the other sets that mbar orders. The tlbsync instruction may complete before operations caused by tlbivax instructions preceding the tlbsync instruction have been performed. Execution of this instruction is restricted to supervisor mode. Other registers altered: None Book E Supervisor 0 5 6 2 02 1 3 03 1 011111 / / / 1000110110 /

Bits 6–20 of the instruction encoding are allocated for implementation-dependent use, and may be used to specify the target TLB entry, the target portion of the target TLB entry, and the source of the value that is to be written into the TLB. The EIS does not make use of these bits. The contents of the implementation-dependent–specified source are written into the implementation-dependent–specified portion of the implementation-dependent–specified TLB entry. If the instruction specifies a TLB entry that does not exist, the results are undefined. Execution of this instruction may cause other implementation-dependent effects. See the user’s manual for the implementation. Execution of this instruction is restricted to supervisor mode. Other registers altered: None Programming notes:

  • The effects of the update are not guaranteed to be visible to the programming model until the completion of a context synchronizing operation. See Context synchronization on page 144.”
  • Care must be taken not to invalidate any TLB entry that contains the mapping for any interrupt vector. Book E Supervisor 0 5 6 2 02 1 3 03 1 011111 / / / (1) 1111010010 / 1. This field is defined as allocated by the Book E architecture, for possible use in an implementation. These bits are not implemented by the EIS.

Trap word [immediate] tw TO,rA,rB twi TO,rA,SIMM a ← EXTS(rA32:63) if ‘tw’ then b ← EXTS(rB32:63) if ‘twi’ then b ← EXTS(SIMM) if (a < b) & TO0 then TRAP if (a > b) & TO1 then TRAP if (a = b) & TO2 then TRAP if (a <u b) & TO3 then TRAP if (a >u b) & TO4 then TRAP For tw, the contents of rA[32–63] are compared with the contents of rB[32–63]. For twi, the contents of rA[32–63] are compared with the sign-extended value of the SIMM field. If any bit in the TO field is set and its corresponding condition is met by the result of the comparison, then the system trap handler is invoked. Other registers altered: None Book E User 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1

011111 T O rA rB 0000000100 /

000011 T O rAS I M M

Write MSR external enable [immediate] wrtee r S wrteei E if ‘wrtee’ then MSR[EE]← rS48 if ‘wrteei’ then MSR[EE]← E For wrtee, rS[48] is placed into MSR[EE]. For wrteei, the value specified in the E field is placed into MSR[EE]. Execution of this instruction is restricted to supervisor mode. In addition, changes to MSR[EE] are effective as soon as the instruction completes. Thus if MSR[EE]=0 and an external interrupt is pending, executing a wrtee or wrteei that sets MSR[EE] causes the external interrupt to be taken before the next instruction is executed, if no higher priority exception exists. Other registers altered: MSR Programming note: wrtee and wrteei are used to update of MSR[EE] without affecting other MSR bits. Typical usage is as follows: mfmsr Rn #save EE in GPR(Rn) wrteei 0 #turn off EE :: : : : #code with EE disabled :: : wrtee Rn #restore EE without altering other MSR bits that may have changed Book E Supervisor 0 5 6 1 0 1 12 0 2 13 0 3 1 011111 rS / / / 0010000011 / 0 5 6 1 51 61 7 2 02 1 3 03 1 011111 / / / E / / / 0010100011 /

XOR [Immediate [shifted]] xor r A,rS,rB( R c = 0 ) xor. r A,rS,rB( R c = 1 ) xori r A,rS,UIMM (S=0, Rc=0) xoris r A,rS,UIMM (S=1, Rc=0) if ‘xori’ then b ← 480 || UIMM if ‘xoris’ then b ← 320 || UIMM || 160 if ‘xor[.]’ then b ← rB result0:63 ← rS ⊕ b if Rc=1 then do LT ← result32:63 < 0 GT ← result32:63 > 0 EQ ← result32:63 = 0 rA ← result For xori, the contents of rS are XORed with 480 || UIMM. For xoris, the contents of rS are XORed with 320 || UIMM || 160. For xor[.], the contents of rS are XORed with the contents of rB. The result is placed into rA. Other registers altered: CR0 (if Rc=1) Book E User 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1 011111 rS rA rB 0100111100 R c 0 4 5 6 10 11 15 16 31 01101S rS rAU I M M

_xorx _xorx XOR [Immediate] [and Record] e_xori r A,rS,SCI8 (Rc = 0) e_xori. r A,rS,SCI8 (Rc = 1) if ‘e_xori[.]’ then b ← SCI8(F ,SCL,UI8) result32:63 ← GPR(RS) ⊕ b if Rc=1 then do LT ← result32:63 < 0 GT ← result32:63 > 0 EQ ← result32:63 = 0 GPR(RA) ← result For e_xori[.], the contents of GPR(rS) are XORed with SCI8. The result is placed into GPR(rA). Special Registers Altered: CR0 (if Rc = 1) VLE User 0 5 6 1 01 1 1 51 6 1 92 02 12 22 32 4 3 1

000110 R S R A 1110 R c F S C L U I 8

RM0004 Part II: EIS-defined extensions to the Book E architecture Part II: EIS-defined extensions to the Book E architecture This part describes the extensions defined by the Book E Implementation Standards (EIS). It consists of the following:

  • Chapter 7: Auxiliary processing units (APUs) on page 823,” describes APUs such as the isel instruction, performance monitor, signal processing engine (SPE), locking, and machine check APUs.
  • Chapter 8: Storage-related APUs on page 848,” describes the following APUs defined by the storage architecture: – Chapter 8.1: Cache line locking APU on page 848” – Chapter 8.2: Direct cache flush APU on page 850” – Chapter 8.3: Cache way partitioning APU on page 851”
  • Subsequent chapters describe the VLE extension – Chapter 9: VLE introduction on page 852” – Chapter 10: VLE storage addressing on page 759 – Chapter 11: VLE compatibility with the EIS on page 856” – Chapter 12: VLE instruction classes on page 860” – Chapter 13: VLE instruction set on page 891” – Chapter 14: VLE instruction index on page 967”

Auxiliary processing units (APUs) RM0004

7 Auxiliary processing units (APUs)

This chapter describes the APUs defined by the EIS, which are as follows:

  • Chapter 7.1: Integer select APU”
  • Chapter 7.2: Performance monitor APU”
  • Chapter 7.3: Signal processing engine APU (SPE APU)”
  • Chapter 7.4: Embedded vector and scalar single-precision floating-point APUs (SPFP APUs)”
  • Chapter 7.5: Machine check APU”
  • Chapter 7.6: Debug APU”
  • Chapter 7.7: Alternate time base” Note that individual processors may implement APUs that are not defined by the EIS. Individual processors may either further extend these APUs or may implement a subset of the resources described here. See the documentation for the individual implementation.

7.1 Integer select APU

Control code, which is characterized by unpredictable short branches, is common in embedded applications. When mispredicted, these branches cause long pipeline delays. The integer select (isel) APU consists of a single instruction (isel), a conditional register move that helps eliminate some of these branches. The isel instruction works as follows: if crB then rD = rA else rD = rB The isel instruction allows more efficient implementation of a condition sequence such as the one in the following generic example: int16 global1,…, global37,...; .... void procedure17(int16 parm) { if (global1 == 27) { global37 = parm + 17; else { global37 = parm - 17;

7.1.1 Integer select APU programming model

The integer select APU includes only the isel instruction, described in Chapter 6: Instruction set on page 330.” It accesses the GPRs and the CR and does not implement additional registers or interrupt resources.

7.1.2 Using isel to Improve conditional branch performance

Table 208 shows a coding example with and without the isel instruction. resulting throughput is typically low.

  • Sets a condition code according to the results of a comparison
  • Has code that executes both the IF and the ELSE segments
  • Has a final statement that copies the results of one of the segments to the desired destination register
  • Works well for small code segments and for unpredictable branches
  • Can reduce code size

7.2 Performance monitor APU

performance monitor interrupt.

7.2.1 Performance monito r APU programming model

Table 208. Recoding with isel

while in user-mode causes a privilege exception. an illegal instruction exception. Table 209. Performance monitor registers—supervisor level Table 210. Performance monitor registers—user level (read-only)

The APU also defines the instructions in Table 211 to move to and move from these PMRs. an enabled condition or event.

7.3 Signal processing engine APU (SPE APU)

This section describes the SPE APU programming model, exceptions, and functions.

7.3.1 Overview

back operations without loop unrolling. Table 211. Performance monitor apu instructions Table 210. Performance monitor registers—user level (read-only) (continued)

Auxiliary processing units (APUs) RM0004

7.3.2 Nomenclature and conventions

Several conventions regarding nomenclature are used in this document:

  • The signal processing engine APU is abbreviated as SPE.
  • All register bit numbering is 64-bit, with bit 0 being the most significant bit. Registers that are only 32-bit define bit 32 as the most significant bit. For both 32- and 64-bit registers, bit 63 is the least significant bit.
  • Bits 0 to 31 of a 64-bit register are referenced as upper word, even word or high word element of the register. Bits 32–63 are referred to as lower word, odd word, or low word element of the register. Each half is an element of a 64-bit GPR.
  • Bits 0 to 15 and bits 32 to 47 are referenced as even half words. Bits 16 to 31 and bits 48 to 63 are referenced as odd half words.
  • Mnemonics for SPE instructions generally begin with the letters ‘ev’ (embedded vector).

7.3.3 Programming model

This section describes SPE registers, instructions, and interrupts. General operation SPE instructions generally take elements from each source register and operate on them with the corresponding elements of a second source register (and/or the accumulator) to produce results. Results are placed in the destination register and/or the accumulator. Instructions that are vector in nature (that is, they produce results of more than one element) provide results for each element that are independent of the computation of the other elements. These instructions can also be used to perform scalar DSP operations by ignoring the results of the upper 32-bit half of the register file. There are no record forms of SPE instructions. SPE compare instructions store the compare result into the condition register (CR). The meaning of the CR bits is now overloaded for SPE operations. SPE compare instructions specify a CR field, two source registers, and the type of compare: greater than, less than, or, equal. Two bits of the CR field are written with the result of the vector compare, one for each element. The remaining two bits reflect the ANDing and ORing of the vector compare results. GPR registers The SPE APU requires a GPR register file with thirty-two 64-bit registers. For 32-bit implementations, PowerPC Book E instructions that normally operate on a 32-bit register file access and change only the least significant 32 bits of the GPRs, leaving the most significant 32 bits unchanged. For 64-bit implementations, operation of these instructions is unchanged, that is, those instructions continue to operate on the 64-bit registers as they would if the SPE APU was not implemented. SPE APU instructions view the 64-bit register as being composed of a vector of two elements, each of which is 32 bits wide. (Some instructions read or write 16-bit elements.) The most significant 32 bits are called the upper word, high word or even word. The least significant 32 bits are called the lower word, low word or odd word. Unless otherwise specified, SPE instructions write all 64 bits of the destination register. 03 1 3 26 3 GPR Upper word Lower word

RM0004 Auxiliary processing units (APUs) Accumulator register A partially visible accumulator register (ACC) is provided for the integer/fractional multiply accumulate (MAC) forms of instructions. The accumulator is a 64-bit register that holds the results of the multiply accumulate forms of SPE fixed-point instructions. The accumulator allows the back-to-back execution of dependent MAC instructions, something that is found in the inner loops of DSP code such as FIR and FFT filters. The accumulator is partially visible to the programmer in the sense that its results do not have to be explicitly read to use them. Instead they are always copied into a 64-bit destination GPR, which is specified as part of the instruction. Based upon the type of instruction, the accumulator can hold either a single 64-bit value or a vector of two 32-bit elements. Signal processing embedded floating-point status and control register (SPEFSCR) Status and control for SPE uses the SPEFSCR, described in Chapter 2.14.1: Signal processing, embedded floating-point status, control register (SPEFSCR) on page 119.” The embedded floating-point APUs also use SPEFSCR. Status and control bits are shared for embedded floating-point operations and SPE vector operations. The SPEFSCR is implemented as SPR number 512 and is read and written by the mfspr and mtspr instructions in both user and supervisor mode. SPE exception bit in ESR ESR[SPE] is defined as the SPE exception bit. This bit is set whenever the processor takes an interrupt related to the execution of SPE instructions. (Note that the same bit is used for embedded floating-point APU exceptions. Thus, SPE and embedded floating-point exceptions are indistinguishable in the ESR.) SPE available bit in MSR MSR[SPE] is defined as the SPE available bit. If this bit is not set and software attempts to execute an SPE instruction, the SPE APU unavailable interrupt is taken. Software note: This bit can be used by software to detect when a process uses the upper 32 bits of a 64-bit register on a 32-bit implementation and thus save them on context switch. Data formats The SPE APU provides two different data formats, integer and fractional. Both data formats can be treated as signed or unsigned quantities. Integer format Integer data format is the same as what is conventionally used in computing. Unsigned integers consist of 16-, 32-, or 64-bit binary integer values. The largest representable value is 2n – 1, where n represents the number of bits in the value. The smallest representable value is 0. Computations that produce values larger than 2n –1 o r smaller than 0 set OV or OVH in SPEFSCR. Signed integers consist of 16-, 32-, or 64-bit binary values in two’s-complement form. The largest representable value is 2n–1 – 1, where n represents the number of bits in the value. 03 1 3 26 3 ACC Upper word Lower word

2n–1 – 1 or smaller than –2 n–1 set OV or OVH in SPEFSCR. Fractional data format is the same that is conventionally used for DSP fractional arithmetic. Fractional data is useful for representing data converted from analog devices. n, where n represents the number of bits in the value. The smallest representable value is 0. would; therefore, unsigned fractional instruction forms are not defined for SPE. represents the number of bits in the value. The smallest representable value is –1. bit is concatenated as the least-significant bit (lsb) of the shifted result.

  • Simple vector instructions. These instructions use the corresponding low- and high- word elements of the operands to produce a vector result that is placed in the destination register, the accumulator, or both. Figure 178 shows how operations are typically performed in vector operations.

Figure 178. Two-element vector operations

  • Multiply and accumulate instructions. These instructions perform multiply operations, add the result to the accumulator and place the result into the destination register and the accumulator. These instructions are composed of different multiply forms, data 03 1 3 2 6 3 rA rB operation operation rD

their various characteristics. These are shown in Table 212.

  • Load and store instructions. These instructions provide load and store capabilities for moving data to and from memory. A variety of forms are provided that position data for efficient computation.
  • Compare and miscellaneous instructions. These instructions perform miscellaneous functions such as field manipulation, bit reversed incrementing, and vector compares.

Table 212. Mnemonic extensions for multiply accumulate instructions

Auxiliary processing units (APUs) RM0004 SPE exceptions and interrupts The APU defines the following SPE exceptions:

  • SPE/embedded floating-point unavailable exception (causes the SPE/embedded floating point unavailable interrupt)
  • SPE vector alignment exception (causes the alignment interrupt) Interrupt vector offset registers (IVORs) IVOR32 (SPE/embedded floating-point unavailable interrupt) and IVOR5 (alignment interrupt) are used by the interrupt model. The SPR number for IVOR32 is 528; IVOR5 is defined by Book E. These registers are privileged. SPE/Embedded floating point unavailable exception The SPE/embedded floating point unavailable exception occurs when execution of an SPE instruction (except brinc) is attempted and bit 38 (SPE available, MSR[SPE]) is not set. If the SPE/embedded floating point unavailable exception occurs, a SPE/embedded floating point unavailable exception interrupt is taken and the processor suppresses execution of the instruction causing the exception. SRR0, SRR1, MSR, and ESR are modified as follows:
  • SRR0 is set to the EA of the instruction causing the interrupt.
  • SRR1 is set to the contents of the MSR at the time of the interrupt.
  • MSR bits CE, ME, and DE are unchanged. All other bits are cleared.
  • ESR[36] bit is set. All other ESR bits are cleared. Instruction execution resumes at address IVPR[0–47]||IVOR32[48–59]||0b0000. Software note: This exception is also used by the embedded floating-point APUs in the same manner. It should be used by software to determine if the application is using the upper 32 bits of the GPRs and thus is required to save and restore them on a context switch. SPE vector alignment exception The SPE vector alignment exception is taken if the EA of any of the following instructions in not aligned to a 64-bit boundary: evldd, evlddx, evldw, evldwx, evldh, evldhx, evstdd, evstddx, evstdw, evstdwx, evstdh, or evstdhx. When an SPE vector alignment exception occurs, an alignment interrupt is taken and the processor suppresses execution of the instruction causing the exception. SRR0, SRR1, MSR, ESR, and DEAR are modified as follows:
  • SRR0 is set to the EA of the instruction causing the interrupt.
  • SRR1 is set to the contents of the MSR at the time of the interrupt.
  • MSR bits CE, ME, and DE are unchanged. All other bits are cleared.
  • ESR[56] bit is set. ESR[ST] is set if the instruction causing the interrupt is a store. All other ESR bits are cleared.
  • DEAR is updated with the EA used in the load or the store. Instruction execution resumes at address IVPR[0–47]||IVOR32[48–59]||0b0000. Interrupt priorities The following list shows the priority order in which SPE APU and SPFP APU interrupts are taken (see Embedded floating-point interrupts on page 837”):

RM0004 Auxiliary processing units (APUs) 1. SPE APU unavailable interrupt 2. SPE vector alignment interrupt 3. Embedded floating-point data interrupt 4. Embedded floating-point round interrupt

7.3.4 Instruction definitions

Chapter 6: Instruction set on page 330,” gives complete descriptions of SPE and embedded floating-point instructions. Chapter 6.3.1 on page 336,” provides pseudo RTL for saturation and bit reversal to more accurately describe those functions that are referenced in the instruction pseudo RTL.

7.4 Embedded vector and scalar si ngle-precision floating-point

APUs (SPFP APUs) This section describes the instruction set architecture of the embedded floating-point APUs. The EIS defines the following APUs:

  • Embedded vector single-precision floating-point APU
  • Embedded scalar single-precision floating-point APU
  • Embedded scalar double-precision floating-point APU Each of these APUs may be implemented independently of the other. In addition, there is a strong relationship with the SPE APU in that each of the embedded floating-point APUs shares a common status register with the SPE.

7.4.1 Nomenclature and conventions

Several conventions regarding nomenclature are used in this document:

  • The embedded vector single-precision floating-point APU operations are abbreviated as vector floating-point or vector SPFP .
  • The embedded scalar single-precision floating-point APU operations are abbreviated as scalar SPFP .
  • The embedded scalar double-precision floating-point APU operations are abbreviated as scalar DPFP .
  • Bits 0 to 31 of a 64-bit register are referenced as field 0, upper half, upper word, or high-word element of the register. Bits 32–63 are referred to as field 1, lower half, or lower-word element of the register. Each half is an element of a 64-bit GPR.
  • Mnemonics for vector floating-point instructions generally begin with the letters ‘evf’ (embedded vector float).
  • Mnemonics for single-precision floating-point instructions generally begin with the letters ‘efs’ (embedded floating single).
  • References to ‘floating-point’ or ‘embedded SPFP’ refer to both APUs.

7.4.2 Embedded floating-poi nt APUs programming model

The embedded floating-point APUs use the GPRs as source and destination operands; however, double precision and vector instruction require 64-bit GPRs as described in Embedded floating-point APUs GPR implementations on page 836.”

  • Opcodes for embedded vector floating-point instructions on page 833”
  • Opcodes for embedded scalar single-precision floating-point instructions on page 833”
  • Opcodes for embedded scalar double-precision floating-point instructions on page 834” Opcodes for embedded vector floating-point instructions Table 213 lists the embedded vector floating-point opcodes. Opcodes for embedded scalar single-precision floating-point instructions Table 214 lists the embedded scalar single-precision floating-point opcodes.

Table 213. Embedded vector floating-point instruction opcodes

Table 215 lists the embedded scalar double-precision floating-point opcodes. Table 214. Embedded scalar single-precision floating-point instruction opcodes Table 215. Embedded scalar double-precision floating-point instruction opcodes

All embedded floating-point APUs use GPRs to hold and operate on floating-point values. implement the following load/store instructions from the SPE APU.

RM0004 Auxiliary processing units (APUs) For scalar double-precision:

  • evldd—Vector load doubleword into doubleword
  • evlddx—Vector load doubleword into doubleword indexed
  • evstdd—Vector store doubleword of doubleword
  • evstddx—Vector store doubleword of doubleword
  • evmergehi—Vector merge high
  • evmergelo—Vector merge low For vector single-precision, all of the vector load/store word and doubleword instructions, merge instructions, and word forms of splat instructions may be implemented. Because the vector single-precision embedded floating-point APU uses a significant set of the SPE vector load/store/merge instructions, it is strongly recommended that the SPE APU be present when implementing the vector single-precision embedded floating-point APU. Floating-point conversion models Each APU contains floating-point conversion to and from integer and fractional type instructions. The floating-point to and from non–floating-point conversion model pseudo RTL is provided in Chapter 6.3.2: Embedded floating-point conversion models on page 337,” as a group of functions that is called from the individual instruction pseudo-RTL descriptions included in the instruction descriptions in Chapter 6: Instruction set on page 330.” Embedded floating-point registers The embedded floating-point APUs share register resources with the SPE APU, as described in the following sections. Embedded floating-point APUs GPR implementations Embedded floating-point operations are performed in the GPRs of the processor. The vector floating-point and double-precision floating-point require a GPR register file with thirty-two 64-bit registers. This is consistent with the SPE APU. Thus, these can coexist with the SPE APU. Single-precision floating-point requires a GPR register file with thirty-two 32-bit or 64-bit registers. When implemented with a 64-bit register file on a 32-bit implementation, single- precision floating-point operations only use and modify bits 32–63 of the GPR. In this case, bits 0–31 of the GPR are left unchanged by a si ngle-precision floating-point operation. For 64-bit implementations, bits 0–31 are undefined after a single-precision floating-point operation. Floating-point double-precision instructions operate on the entire 64 bits of the GPRs where a floating-point data item consists of 64 bits. Vector floating-point instructions operate on the entire 64 bits of the GPRs as well, but contain two 32-bit data items that are operated on independently of each other in a SIMD fashion. The format of both data items is the same as a single-precision floating-point value. The data item contained in bits 0–31 is called the ‘high word’. The data item contained in bits 32–63 is called the low word There are no record forms of embedded floating-point instructions. Floating-point compare instructions treat NaNs, Infinity and Denorm as normalized numbers for the comparison calculation when default results are provided.

Auxiliary processing units (APUs) RM0004 Signal processing embedded floating-point status and control register (SPEFSCR) The embedded floating-point APUs use the SPEFSCR, which is described in Chapter 2.14.1: Signal processing, embedded floating-point status, control register (SPEFSCR) on page 119.” The SPE APU also uses SPEFSCR. Status and control bits are shared for vector floating-point operations, single-precision floating-point operations and SPE vector operations. The SPEFSCR is implemented as SPR number 512 and is read and written by mfspr and mtspr in both user and supervisor mode. Vector floating-point instructions affect both the high- and low-element floating-point status flags (bits 34–39 and 50–55). Scalar SPFP instructions affect only the low-element flags and leave the high element flags undefined. Embedded floating-point exception bit—ESR[SPE] ESR[SPE] is defined as the embedded floating-point exception bit. This bit is set whenever the processor takes an interrupt related to the execution of the embedded floating-point instructions. (Note that the same bit is used for SPE APU exceptions. Thus, SPE and embedded floating-point interrupts are indistinguishable in the ESR.) Embedded floating-point interrupts The following sections describe the embedded floating-point APU interrupts:

  • SPE/embedded floating-point unavailable interrupt on page 837”
  • Embedded floating-point data interrupt on page 837”
  • Embedded floating-point round interrupt on page 838” SPE/embedded floating-point unavailable interrupt The SPE/embedded floating-point unavailable interrupt vector is used by the embedded scalar double-precision floating-point APU and the embedded vector single-precision floating-point APU. It is not used by the embedded scalar single-precision floating-point APU. The SPE/embedded floating-point unavailable interrupt occurs when an embedded vector floating-point or an embedded scalar double-precision floating-point instruction is executed and bit 38 of the MSR is not set. If the SPE/embedded floating-point unavailable interrupt occurs, the processor suppresses execution of the instruction causing the exception. The SRR0, SRR1, MSR, and ESR registers are modified as follows:
  • SRR0 is set to the EA of the instruction causing the interrupt.
  • SRR1 is set to the contents of the MSR at the time of the interrupt.
  • MSR bits CE, ME, and DE are unchanged. All other bits are cleared.
  • ESR[24] is set. All other ESR bits are cleared. Instruction execution resumes at address IVPR[0–47]||IVOR32[48–59]||0b0000. This interrupt is also used by the SPE APU in the same manner. It should be used by software to determine if the application is using the upper 32 bits of the GPRs and thus is required to save and restore them on a context switch. Embedded floating-point data interrupt The embedded floating-point data interrupt vector is used for enabled floating-point invalid operation/input error, underflow, overflow, and divide-by-zero exceptions (collectively called floating-point data exceptions). When one of these enabled exceptions occurs, the

RM0004 Auxiliary processing units (APUs) processor suppresses execution of the instruction causing the exception. The SRR0, SRR1, MSR, ESR, and SPEFSCR are modified as follows:

  • SRR0 is set to the EA of the instruction causing the interrupt.
  • SRR1 is set to the contents of the MSR at the time of the interrupt.
  • MSR bits CE, ME and DE are unchanged. All other bits are cleared.
  • ESR[SPE] is set. All other ESR bits are cleared.
  • One or more SPEFSCR status bits are set to indicate the type of exception. The affected bits are FINVH, FINV, FDBZH, FDBZ, FOVFH, FOVF , FUNFH, and FUNF . SPEFSCR[FG,FGH, FX, FXH] are cleared. Instruction execution resumes at address IVPR[0–47]||IVOR32[48–59]||0b0000. Embedded floating-point round interrupt The embedded floating-point round interrupt occurs if no other floating-point data interrupt is taken and one of the following conditions is met:
  • SPEFSCR[FINXE] is set and the unrounded result of an operation is not exact
  • SPEFSCR[FINXE] is set, an overflow occurs, and overflow exceptions are disabled (FOVF or FOVFH set with FOVFE cleared)
  • An underflow occurs and underflow exceptions are disabled (FUNF set with FUNFE cleared) The embedded floating-point round interrupt does not occur if an enabled embedded floating-point data interrupt occurs. If an implementation does not support ±infinity rounding modes and the rounding mode is set to be +infinity or –infinity, an embedded floating-point round interrupt occurs after every floating-point instruction for which rounding might occur regardless of the value of FINXE unless an embedded floating-point data interrupt also occurs and is taken. When the embedded floating-point round interrupt occurs, the unrounded (truncated) result of an inexact high or low element is placed in the target register. If only a single element is inexact, the other exact element is updated with the correctly rounded result, and the FG and FX bits corresponding to the other exact element are both zero. The FG and FX bits are provided so that an interrupt handler can round the result as it desires. FG (the guard bit) is the value of the bit immediately to the right of the least significant bit of the destination format mantissa from the infinitely precise intermediate calculation before rounding. FX (the sticky bit) is the value of the OR of all bits to the right of the guard bit (FG) of the destination format mantissa from the infinitely precise intermediate calculation before rounding. The SRR0, SRR1, MSR, ESR, and SPEFSCR are modified as follows:
  • SRR0 is set to the EA of the instruction following the instruction causing the interrupt.
  • SRR1 is set to the contents of the MSR at the time of the interrupt.
  • MSR bits CE, ME, and DE are unchanged. All other bits are cleared.
  • ESR[SPE] is set. All other ESR bits are cleared.
  • SPEFSCR FGH, FG, FXH, and FX are set appropriately. SPEFSCR[FINXS] is set. Instruction execution resumes at address IVPR[0–47]||IVOR32[48–59]||0b0000.
  1. SPE/embedded floating-poin t unavailable interrupt
  2. SPE vector alignment interrupt
  3. Embedded floating-point data interrupt
  4. Embedded floating-point round interrupt

floating-point data exception.

7.4.3 Embedded floating- point APU operations

underflow and overflow handling, IEEE 754 compliance, and conversion models. saturates results. Other modes are currently not defined. biased exponent (exp) and 23 bits of fraction. Figure 179. Floating-point data formats

RM0004 Auxiliary processing units (APUs) For single-precision normalized numbers, the biased exponent value, e, lies in the range of 1 to 254 corresponding to an actual exponent value E in the range –126 to +127. With the hidden bit implied to be 1 (for normalized numbers), the value of the number is interpreted as follows: where E is the unbiased exponent and 1.fraction is the mantissa (or significand) consisting of a leading 1 (the hidden bit) and a fractional part (fraction field). For the single-precision format, the maximum positive normalized number (pmax) is represented by the encoding 0x7F7F_FFFF , which is approximately 3.4E+38 (2 128), and the minimum positive normalized value (pmin) is represented by the encoding 0x0080_0000, which is approximately 1.2E–38 (2 –126). Two specific values of the biased exponent are reserved (0 and 255 for single-precision) for encoding special values of +0, –0, +infinity, –infinity, and NaNs. Zeros of both positive and negative sign are represented by a biased exponent value (e) of zero and a fraction that is zero. Infinities of both positive and negative sign are represented by a maximum exponent field value (255 for single-precision) and a fraction that is zero. Denormalized numbers of both positive and negative sign are represented by a biased exponent value of 0 and a non-zero fraction. For these numbers, the hidden bit is defined by the IEEE 754 standard to be zero. This number type is not directly supported in hardware. Instead, either a software interrupt handler is invoked or a default value is defined. Not-a-Numbers (NaNs) are represented by a maximum exponent field value (255 for single- precision) and a fraction that is non-zero. Overflow and underflow Defining pmax to be the most positive normalized value (farthest from zero), pmin the smallest positive normalized value (closest to zero), nmax the most negative normalized value (farthest from zero) and nmin the smallest normalized negative value (closest to zero), an overflow is said to have occurred if the numerically correct result of an instruction is such that r > pmax or r < nmax. Additionally, an implementation may also signal overflow by comparing the exponents of the operands. In this case, the hardware examines both exponents ignoring the fractional values. If it is determined that the operation to be performed may overflow (ignoring the fractional values), an overflow may be said to occur. For addition and subtraction this can occur if the larger exponent of both operands is 254. For multiplication this can occur if the sum of the exponents of the operands less the bias is 254. Thus: single-precision addition: if A exp >= 254 | Bexp >= 254 then overflow double-precision addition: if Aexp >= 2046 | Bexp >= 2046 then overflow single-precision multiplication: if Aexp + Bexp - 127 >= 254 then overflow double-precision multiplication: if Aexp + Bexp - 1023 >= 2046 then overflow 1–() s 2E× 1.fraction()×

Auxiliary processing units (APUs) RM0004 An underflow is said to have occurred if the numerically correct result of an instruction is such that 0<r<pmin or nmin<r<0. In this case, r may be denormalized, or may be smaller than the smallest denormalized number. As with overflow detection, an implementation may also signal underflow by comparing the exponents of the operands. In this case, the hardware examines both exponents regardless of the fractional values. If it is determined that the operation to be performed may underflow (ignoring the fractional values), an underflow may be said to occur. For division this can occur if the difference of the exponent of the A operand less the exponent of the B operand less the bias is 1. Thus: single-precision division: if A exp - Bexp - 127 <= 1 then underflow double-precision multiplication: if Aexp - Bexp - 1023 <= 1 then underflow The embedded floating-point APUs will not produce +Inf, –Inf, NaN, or a Denormalized number. If the result of an instruction overflows and floating-point overflow exceptions are disabled (SPEFSCR[FOVFE] is cleared), pmax or nmax is generated as the result of that instruction depending upon the sign of the result. If the result of an instruction underflows and floating-point underflow exceptions are disabled (SPEFSCR[FUNFE] is cleared), +0 or - 0 is generated as the result of that instruction based upon the sign of the result. IEEE 754 compliance The embedded floating-point APU implements a floating-point system as defined in ANSI/IEEE Standard 754-1985 but may rely on software support in order to conform fully with the standard. Thus, whenever an input operand of a floating-point instruction has data values that are +infinity, –infinity, denorm, or NaN, or when the result of an operation produces an overflow or an underflow, an interrupt may be taken and the interrupt handler is responsible for delivering IEEE 754–compliant behavior if desired. When floating-point invalid input exceptions are disabled (SPEFSCR[FINVE] is cleared), default results are provided by the hardware when an infinity, denorm, or NaN input is received, or for the operation 0/0. When floating-point underflow exceptions are disabled (SPEFSCR[FUNFE] is cleared) and the result of a floating-point operation underflows, a signed zero result is produced. The inexact exception is also signaled for this condition. When floating-point overflow exceptions are disabled (EFSCR[FOVFE] is cleared) and the result of a floating-point operation overflows, a pmax or nmax result is produced. The inexact exception is also signaled for this condition. An exception enable flag (SPEFSCR[FINXE]) is also provided for generating an interrupt when an inexact result is produced, to allow a software handler to conform to the IEEE 754 standard. A divide-by-zero exception enable flag (SPEFSCR[FDBZE]) is provided for generating an interrupt when a divide-by-zero operation is attempted to allow a software handler to conform to the IEEE 754 standard. All of these exceptions may be disabled, and the hardware then delivers an appropriate default result. The sign of the result of an addition operation is the sign of the source operand having the larger absolute value. If both operands have the same sign, the sign of the result is the same as the sign of the operands. This includes subtraction, which is addition with the negation of the sign of the second operand. The sign of the result of an addition operation with operands of differing signs for which the result is zero is positive except when rounding to –infinity. Thus, –0 + –0 = –0 is the only case in which the result is a –0; all other cases that result in a zero value give +0 unless the rounding mode is round to –infinity. Note that when exceptions are disabled and default results computed, operations having input values that are denormalized may provide different bit-exact results on different

RM0004 Auxiliary processing units (APUs) implementations. An implementation may choose to use the denormalized value or a zero value for any computation. Thus a computational operation involving a denormalized value and a normal value may return different results on other implementations. Sticky bit handling for exception conditions The SPEFSCR defines sticky bits for retaining information about exception conditions that are detected. These sticky bits (FINXS, FINVS, FDBZS, FUNFS, and FOVFS) can be used to help provide IEEE 754 compliance. The sticky bits represent the combined OR of all previous status bits produced from any embedded floating-point operation before the last time software zeroed the sticky bit. Only software can zero a sticky bit; hardware can only set sticky bits. Not all sticky bits are required to be updated by an implementation. Only the FINXS and FDBZS sticky bits are required to be set by hardware. Thus for FINVS, FUNFS and FOVFS, software is required to perform sticky bit setting unless software knows that a given implementation updates them in hardware. This can be achieved by enabling the appropriate exceptions and performing the sticky bit updating in the software interrupt handler. If an implementation provides sticky bit handling for any sticky bits other than FINXS and FDBZS, it must provide it for all sticky bits.

7.4.4 Implementation options summary

There are several options that may be chosen for a given implementation. This section summarizes all the items that are implementation dependent and should be used to help decide which implementation dependent features are chosen.

  • APUs. Each of the APUs can be implemented independently of one another. The vector single-precision floating-point APU should be implemented only if the SPE APU is implemented; however, this is not required.
  • Both the vector single-precision floating-point APU and the scalar double-precision floating-point APU allow the optional implementation of 64-bit load and store instructions as well as merge upper and lower instructions from the SPE APU. This allows data to be moved in and out of the upper half of a register for 32-bit implementations with 64-bit registers.
  • Overflow and underflow conditions may be signaled by doing exponent evaluation of the operation. If by examining the exponents, an overflow or underflow could occur, the implementation may choose to signal an overflow or underflow. It is recommended that future implementations do not use this estimation and signal overflow or underflow when they actually occur.
  • If an operand for a calculation or conversion is denormalized, the implementation may choose to use a same-signed zero value in place of the denormalized operand.
  • The rounding modes of +Infinity and -Infinity are not required to handled by an implementation. If an implementation does not support ±Infinity rounding modes and the rounding mode is set to be +Infinity or -Infinity, an embedded floating-point round interrupt occurs after every floating-point instruction for which rounding may occur

Auxiliary processing units (APUs) RM0004 regardless of the value of FINXE unless an embedded floating-point data interrupt also occurs and is taken.

  • For absolute value, negate, negative absolute value operations, an implementation may choose to either simply perform the sign bit operation ignoring exceptions, or to compute the operation and handle exceptions and saturation where appropriate.
  • The FGH and FXH bits of the SPEFSCR are undefined upon the completion of a scalar floating-point operation. An implementation may choose to zero them or leave them unchanged.
  • An implementation may choose to only implement sticky bit setting by hardware for FDBZS and FINXS allowing software to manage the other sticky bits. It is recommended that all future implementations implement all sticky bit setting in hardware.
  • For 64-bit implementations, the upper 32 bits of the destination register are undefined when the result of a scalar floating-point operation is a 32-bit result. It is recommended that future 64-bit implementations produce 64-bit results for the results of 64-bit conversions to integer values.

7.5 Machine check APU

The machine check APU defines features for the machine check interrupt in addition to those defined by the PowerPC architecture and the Book E version of the PowerPC architecture. The machine check APU includes an enhanced definition of the machine check interrupt type similar to the Book E–defined critical interrupt.

7.5.1 Machine check APU programming model

The APU defines dedicated save and restore SPRs, MSRR0 and MSRR1, so a machine check interrupt does not affect the CSRR0, CSRR1, or ESR registers as defined by the Book E architecture. The APU also defines a separate Return from Machine Check Interrupt instruction, rfmci, that restores context from MSRR0 and MSRR1 when the machine check interrupt handler completes. Machine check APU register model The machine check APU defines different register for the machine check interrupt resources than the Book E definition. These are as follows:

  • Machine-check save/restore register 0 (MCSRR0)—SPR 570. Holds the instruction where fetching begins after rfmci executes, typically at the end of the machine check interrupt handler. See Machine check save/restore register 0 (MCSRR0) on page 87.”
  • Machine-check save/restore register 1 (MCSRR1)—SPR 571. Holds the machine state copied to the MSR when a machine check interrupt occurs. The MCSRR1 value is restored to the MSR when rfmci executes, typically at the end of the machine check interrupt handler. See Machine check save/restore register 1 (MCSRR1) on page 87.”
  • Machine check syndrome register (MCSR)—SPR 572. MCSR has fields that identify causes for a machine check interrupt along with an indication of whether the processor can recover from the machine check interrupt. See Machine check syndrome register (MCSR) on page 88.”

RM0004 Auxiliary processing units (APUs) Note, however, that the MSR[ME] bit, defined by the original PowerPC architecture, is also used in Book E and in the machine check APU to enable the machine check interrupt. Machine check APU instruction model The Return from Machine Check Interrupt instruction, rfmci, is context-synchronizing; it works its way to the final execute stage, updates architected registers, and redirects instruction flow. When rfmci executes, data is restored from MCSRR0 and MCSRR1. The rfi and rfci instructions do not affect MCSRR0 and MCSRR1. This instruction is described in Chapter 3: Instruction model on page 133.” Machine check interrupt The machine check APU is consistent with the machine check exception as defined in Book E with the following differences:

  • Machine check is no longer a critical interrupt but uses MCSRR0 and MCSRR1 for saving the return address and the MSR in case the machine check is recoverable.
  • The Return from Machine Check Interrupt instruction (rfmci) is implemented to support the return to the address saved in MCSRR0.
  • The machine check syndrome register, MCSR, is used (instead of ESR) to log the cause of the machine check.

7.6 Debug APU

This section describes the instruction set architecture of software accessible debug related items for Book E Implementations (EIS). The debug APU defines an additional interrupt class for debug interrupts. This allows the debug features to be used in the software that is providing service for critical class interrupts. This is accomplished by providing specific save and restore registers for debug interrupts and providing a new return from interrupt instruction (return from debug interrupt). The debug APU reassigns debug interrupts into its own interrupt class, adding a new set of registers used to save the machine context upon the occurrence of a debug interrupt, and adds a new instruction, Return From Debug Interrupt (rfdi), to return from a debug interrupt and restore the machine state from the new set of registers. This APU redefines PowerPC Book E debug interrupt behavior. An implementation may choose to provide the debug APU and also provide a method to disable the debug APU, reverting to using the critical interrupt as defined in Book E. If such a capability is provided, HID0[DAPUEN] should be implemented.

7.6.1 Debug APU programming model

The following sections described the debug APU’s extensions to the Book E interrupt, register, and interrupt models.

7.6.2 Debug APU register model

  • Debug save/restore register 0 (DSRR0). When a debug interrupt is taken, DSRR0 is set to the current or next instruction address. When rfdi is executed, instruction execution continues at the address in DSRR0.
  • Debug save/restore register 1 (DSRR1), When a debug interrupt is taken, the contents of the MSR are placed into DSRR1. When rfdi is executed, the contents of DSRR1 are placed into the MSR. Bits of DSRR1 that correspond to reserved bits in the MSR are also reserved. This instruction is fully described in Chapter 6: Instruction set on page 330.” The debug APU defines fields in the following Book E–defined registers:
  • Debug status register (DBSR). New event fields, described in Table 216, have been added to DBSR to record critical interrupt taken events and critical interrupt return events.
  • The debug control register 0 (DBCR0), The debug APU adds event enable bits to DBCR0, described in Table 217, to control critical interrupt taken events, and critical interrupt return events.

Table 216. EIS-defined DB SR field descriptions class, that is, uses CSRR0 and CSRR1) occurs. 0No critical interrupt taken debug event has occurred. 1A critical interrupt taken debug event occurred. 0No critical interrupt return debug event has occurred. 1A critical interrupt return debug event occurred. Table 217. DBCR0 field descriptions critical class, that is, uses CSRR0 and CSRR1) occurs. 0 Critical interrupt taken debug events are disabled. 1 Critical interrupt taken debug events are enabled. instruction is executed) occurs. 0 Critical interrupt return debug events are disabled. 1 Critical interrupt return debug events are enabled.

RM0004 Auxiliary processing units (APUs)

7.6.3 Debug APU instruction model

The debug APU defines the supervisor-level rfdi instruction to restore state after a debug interrupt. The contents of DSRR1 are placed into the MSR. If the new MSR value does not enable any pending exceptions, then the next instruction is fetched, under control of the new MSR value, from the address DSRR0[0–61]||0b00. If the new MSR value enables one or more pending exceptions, the interrupt associated with the highest priority pending exception is generated; in this case the value placed into SRR0, CSRR0, or DSRR0 by the interrupt processing mechanism is the address of the instruction that would have been executed next had the interrupt not occurred (that is, the address in DSRR0 at the time of the execution of the rfdi). This instruction is fully described in Chapter 6.” Debug APU interrupt model A debug interrupt occurs when no higher priority exception exists, a debug exception is presented to the interrupt mechanism, and MSR[DE] = 1. The specific cause or causes of debug exceptions are unchanged from Book E. DSRR0, DSRR1, MSR, debug address register, and debug status register are updated as follows: Debug save/restore register 0 (DSRR0) is set to an instruction address. DSRR0 is set to the EA of an instruction that was executing or just completed execution when the debug exception occurred. DSRR0 is set the same as CSRR0 is defined to be set in Book E on a debug interrupt. CSRR0 is not changed as the result of a debug interrupt. Debug save/restore register 1 (DSRR1) is set to the contents of the MSR at the time of the interrupt. CSRR1 is not changed as the result of a debug interrupt. MSR[CM] is set to the value of MSR[ICM]. MSR[ICM] and MSR[ME] are unchanged and all other defined MSR bits are cleared. The DBSR and the debug control registers (DBCR0–DBCR2) operate as described in Book E with the addition of a critical interrupt taken debug event and a critical return debug event. Instruction execution resumes at address IVPR[0–47]||IVOR15[48–59]||0b0000.

7.7 Alternate time base

The alternate time base APU defines a time base counter similar to the time base defined in the PowerPC architecture. It is intended to be used for measuring time in implementation defined intervals. It differs from the time base defined by the PowerPC architecture in that it is not writable and always counts up, wrapping when the 64-bit count overflows.

7.7.1 Programming model

The alternate time base is simply a 64-bit counter that counts up at some implementation dependent rate. Although not required, it is recommended that the rate be at the core clock frequency or as small a multiple of the frequency as practical by the implementation. Consult the user documentation for devices that support this feature. The counter can be read by executing an mfspr instruction specifying the ATB (or ATBL) register, but cannot be written. In 32-bit mode, reading the ATB (or ATBL) register will place the lower 32 bits of the counter into the target register. In 64-bit mode all 64 bits of the counter are placed in the target register. A second SPR register ATBU, is defined that

Auxiliary processing units (APUs) RM0004 accesses only the upper 32 bits of the counter. Thus the upper 32 bits of the counter may be read into a register by reading the ATBU register regardless of computation mode. The alternate time base is analogous to the time base in the PowerPC architecture except that it counts at a different frequency and is not writable. The effect of power savings mode or core frequency changes on counting in the alternate time base is implementation dependent. See the user document for details. Implementation Note: An implementation may choose to directly alias the alternate time base to the time base counter if the granularity of time base counting is acceptable. Registers The programming model consists of two SPRs, alternate time base lower and upper (ATBL and ATBU). Alternate time base registers (ATBL and ATBU) The ATBL and ATBU registers are described in Chapter 2.15: Alternate time base registers (ATBL and ATBU) on page 123.” The alternate time base counter (ATB) is formed by concatenating the upper and lower alternate time base registers (ATBU and ATBL). ATBL (SPR 526) provides read-only access to the 64-bit alternate time base counter, which is incremented at an implementation-defined frequency. ATB registers are accessible in both user and supervisor mode. Like the TB implementation, the ATBL register is an aliased name for ATB.

RM0004 Storage-related APUs

8 Storage-related APUs

This chapter describes the following APUs that are defined as part of the EIS storage architecture:

  • Chapter 8.1: Cache line locking APU”
  • Chapter 8.2: Direct cache flush APU”
  • Chapter 8.3: Cache way partitioning APU”

8.1 Cache line locking APU

The cache line locking APU defines instructions and methods for locking frequently used instructions and data into their cache lines. Cache locking allows software to mark individual cache lines (blocks) as locked, instructing the cache to keep latency-sensitive data available for fast access. Unlike normal cache lines, locked cache lines do not participate in the normal replacement policy.

8.1.1 Programming model

This section gives a general description of the instructions defined by the cache line locking APU. Full descriptions are provided in Chapter 6: Instruction set on page 330.” Lock setting and clearing Lines are locked into the cache by software using a series of touch and lock set instructions. The following instructions are provided to lock data items into the data and instruction cache:

  • dcbtls—Data Cache Block Touch and Lock Set
  • dcbtstls—Data Cache Block Touch for Store and Lock Set
  • icbtls—Instruction Cache Block Touch and Lock Set The rA and rB operands to these instructions form a effective address identifying the line to be locked. The CT field indicates which cache in the cache hierarchy should be targeted. These instructions are similar to the dcbt, dcbtst, and icbt instructions, but locking instructions can not execute speculatively and may cause additional exceptions. For unified caches, both the instruction lock set and the data lock set target the same cache. Similarly, lines are unlocked from the cache by software using a series of lock-clear instructions. The following instructions are provided to lock instructions into the instruction cache:
  • dcblc—Data Cache Block Lock Clear
  • icblc—Instruction Cache Block Lock Clear The rA and rB operands to these instructions form an EA identifying the line to be unlocked. The CT field indicates which cache in the cache hierarchy should be targeted. Additionally, software may clear all the locks in the cache. For the primary cache, this is accomplished by setting the CLFC (DCLFC, ICLFC) bit in L1CSR0 (L1CSR1).

Storage-related APUs RM0004 Cache lines can also be implicitly unlocked in the following ways:

  • A locked line is invalidated if it is targeted by a dcbi, dcbf, or icbi instruction.
  • A snoop hit on a locked line that requires the line to be invalidated. This can occur because the data the line contains has been modified external to the processor, or another processor has explicitly invalidated the line.
  • The entire cache containing the locked line is flash invalidated. An implementation is not required to unlock lines if data is invalidated in the cache. Although the data may be invalidated (and thus not in the cache), the line can remain locked and be filled from the memory subsystem when the next access occurs. This method of not clearing locks when the associated line is invalidated, is called persistent locking. An implementation may choose to implement locks as persistent or not persistent; the preferred method is persistent. Error conditions Setting locks in the cache can fail for several reasons. An address specified with a lock set instruction that does not have the proper permission causes a data storage interrupt (DSI). Cache locking addresses are always translated as data references, therefore icbtls instructions that fail to translate or fail permissions cause DTLB and DSI errors respectively. Additionally, cache locking and clearing operations can fail due to restricted user mode access. See Cache locking (user mode) exceptions on page 850.” Overlocking If no exceptions occur for the execution of an dcbtls, dcbtstls, or icbtls instruction an attempt is made to lock the corresponding line in the cache. If all of the available ways are already locked in the given cache set, the requested line is not locked. This is considered an overlocking situation and if the lock was targeted for the primary cache (CT = 0) then L1CSR0[DCLO] (or L1CSR1[ICLO] if icbtls) is set appropriately. A processor may optionally allow victimizing a locked line in an overlocking situation. If L1CSR0[DCLOA] (L1CSR0[ICLOA] for the primary instruction cache,) is set, an overlocking condition causes the replacement of an existing locked line with the requested line. The selection of the line to replace in an overlocking situation is implementation dependent. The overlocking condition is still said to exist and is appropriatly reflected in the status bits for lock overflow. An attempt to lock a line that is present and valid in the cache does not cause an overlocking condition. A non–lock-setting cache-line fill or line replacement request to a cache that has all ways locked for a given set does not cause a lock to be cleared. Unable-to-lock conditions If no exceptions occur and no overlocking condition exists, an attempt to set a lock can fail if any of the following is true:
  • The target address is marked cache-inhibited or the storage attributes of the address uses a coherency protocol that does not support locking.
  • The target cache is disabled or not present.
  • The CT field specifies a value not supported by the implementation.
  • Any other implementation-specific error condition.

RM0004 Storage-related APUs If an unable-to-lock condition occurs, the lock set instruction is treated as a NOP . If the lock targeted the data cache (dcbtls, dcbtstls), L1CSR0[DCUL] is set to indicate the unable-to- lock condition; if the lock targeted the instruction cache (icbtls), L1CSR1[ICUL] is set. L1CSR0[DCUL] or L1CSR0[ICUL] is set regardless of the CT value in the lock-setting instruction. Cache locking (user mode) exceptions Setting and clearing cache locks can be restricted to supervisor mode only access. If set, MSR[UCLE] allows cache locking operations to be performed in user mode. If MSR[UCLE] = 0 and MSR[PR] = 1 and execution of a cache lock or cache clear instruction occurs, a cache locking exception occurs. In this case the processor suppresses execution of the instruction causing the exception. A DSI interrupt is taken and SRR0, SRR1, MSR, and ESR are modified as follows:

  • SRR0 is set to the EA of the instruction causing the interrupt.
  • SRR1 is set to the contents of the MSR at the time of the interrupt.
  • MSR[CE,ME,DE] are unchanged. All other bits are cleared.
  • ESR[DLK] is set if the instruction was a dcbtls, dcbtstls, or a dcblc.
  • ESR[ILK] is set if the instruction was a icbtls or a icblc.
  • All other ESR bits are cleared. Instruction execution resumes at address IVPR[0–47]||IVOR2[48–59]||0b0000.

8.2 Direct cache flush APU

8.2.1 Overview

To assist in software flush of the L1 cache, the direct cache flush APU allows the programmer to flush and/or invalidate the cache by specifying the cache set and cache way. Without such a feature, the programmer must either:

  • Know the virtual addresses of the lines that need to be flushed and issue dcbst or dcbf instructions to those addresses.
  • Flush the entire cache by causing all the lines to be replaced. This requires a virtual address range that is mapped as a contiguous physical address range, that the programmer knows and can manipulate the replacement policy of the cache, and the size and organization of the cache. With the direct cache flush APU the program needs only specify the way and set of the cache to flush. The direct cache flush APU available bit, L1CFG0[CFISWA], is set for implementations that contain the direct cache flush APU.

8.2.2 Programming model

To address a specific physical block of the cache, the L1 flush and invalidate control register 0 (L1FINV0) is written with the cache set (L1FINV0[CSET]) and cache way (L1FINV0[CWAY]) of the line that is to be flushed. L1FINV0 is written using a mtspr instruction specifying the L1FINV0 register. No tag match in the cache is required. An additional field, L1FINV0[CCMD], is used to specify the type of flush to be performed on the line addressed by L1FINV0[CWAY] and L1FINV0[CSET].

Storage-related APUs RM0004 The available L1FINV0[CCMD] encodings are described in Table 33 on page 96. Only the L1 data cache (or unified cache) is manipulated by the direct cache flush APU. The L1 instruction cache or any other caches in the cache hierarchy are not explicitly targeted by this APU. Register model The direct cache flush APU defined one register, the L1 flush and invalidate control register 0, described in Chapter 2.11.5 on page 96.” L1FINV0 contains fields to provide the way and set selection of a cache line to flush and or invalidate.

8.3 Cache way partitioning APU

The cache way partitioning APU allows ways in a unified L1 cache to be configured to accept either data or instruction miss line-fill replacements.

8.3.1 Programming model

The cache way partitioning APU is comprised of bits in L1CSR0 and L1CFG0, as follows:

  • Way instruction disable field (L1CSR0[WID]) is a 4-bit field that that determines which of ways 0–3 are available for replacement by instruction miss line refills.
  • The additional ways instruction disable bit (L1CSR0[AWID]) determines whether ways 4 and above are available for replacement by instruction miss line refills.
  • Way data disable field (L1CSR0[WDD]) is a 4-bit field that that determines which of ways 0–3 are available for replacement by data miss line refills.
  • The additional ways data disable bit (L1CSR0[AWDD]) determines whether ways 4 and above are available for replacement by instruction miss line refills.
  • See Chapter 2.11.1: L1 cache control and status register 0 (L1CSR0) on page 90.”
  • Way access mode bit, L1CSR0[WAM], Determines whether all ways are available for access or only ways partitioned for the specific type of access are used for a fetch or read operation. See Chapter 2.11.1 on page 90.”
  • Cache way partitioning APU available bit, L1CFG0[CWPA], indicates whether the cache way partitioning APU is available. See Chapter 2.11.3 on page 94.” These fields are described in detail in Chapter 2.11.3 on page 94,” and in Chapter 2.11.1: L1 cache control and status register 0 (L1CSR0) on page 90.”

8.3.2 Interaction with the cache locking APU

Note that the cache way partitioning APU can affect the cache line locking APU’s ability to control replacement of lines. If any cache line locking instruction (icbtls, dcbtls, dcbtstls) is allowed to execute and finds a matching line in the cache, the line’s lock bit is set regardless of the L1CSR0[WID,AWID,WDD,AWDD] settings. In this case, no replacement has been made. However, for cache misses that occur while executing a cache line lock set instruction, the only candidate lines available for locking are those that correspond to ways of the cache that have not been disabled for the particular type of line locking instruction (controlled by WDD and AWDD for dcbtls and dcbtstls, controlled by WID and AWID for icbtls). Thus, an overlocking condition may result even though fewer than eight lines with the same index are locked.y

9 VLE introduction

This body of this document describes the VLE (variable length encoding) extension to the Book E architecture. The VLE extension offers more efficient binary representations of applications for the embedded processor spaces where code density plays a major role in affecting overall system cost, and to a somewhat lesser extent, performance. The intent of the VLE extension is not to define an entirely different ISA nor to supplant the PowerPC ISA; instead the VLE extension can be viewed as a supplement that is can be applied to an application or to part of an application to improve code density. Chapter 11: VLE compatibility with the EIS on page 856,” describes additional VLE extensions to the EIS. The major objectives of the VLE extension are as follows:

  • Coexistence and consistency with the Book E ISA and general architecture
  • Maintain a common programming model and instruction operation model in the VLE extension
  • Reduce overall code size by ~30% over existing PowerPC text segments
  • Limit the increase in execution path length to under 10% for most important

applications

  • Limit the increase in hardware complexity for implementations containing the VLE extension

9.1 Compatibility with PowerPC Book E

VLE provides an extension to Book E. There are additional operations defined using an alternate instruction encoding to enable reduced code footprint. This alternate encoding set is selected on an instruction page basis. A single page attribute bit selects between standard Book E instruction encodings and VLE instructions for that page of memory. This attribute is an extension to the Book E page attributes. Pages can be freely intermixed, allowing for a mixture of both types of encodings. Instruction encodings in pages marked as using the VLE extension are either 16 or 32 bits long, and are aligned on 16-bit boundaries. Because of this, all instruction pages marked as VLE are required to use big-endian byte ordering. The programmer’s model uses the same register set with both instruction encodings, although certain registers are not accessible by VLE instructions using the 16-bit formats and not all condition register (CR) fields are used by condition setting or conditional branch instructions executing from a VLE instruction page. In addition, immediate fields and displacements differ in size and use, due to the more restrictive encodings imposed by VLE instructions. The VLE extension defines additional fields in registers defined by Book E and the EIS. These are described in Chapter 11.2: VLE extension processor and storage control extensions on page 856.” Other than the requirement of big-endian byte ordering for instruction pages and the additional page attribute to identify whether the instruction page corresponds to a VLE section of code, VLE complies with the memory model defined in Book E and the Book E Implementation Specifications (EIS). Likewise, the VLE extension complies with the Book E

and EIS definitions of the exception and interrupt model, the timer facilities, the debug facilities and the special-purpose registers (SPRs).

9.2 Instruction mnem onics and operands

The description of each instruction includes the mnemonic and a formatted list of operands. VLE instruction semantics are either identical or similar to Book E instruction semantics. Where the semantics, side-effects, and binary encodings are identical, Book E mnemonics and formats are used. Where the semantics are similar but the binary encodings differ, the Book E mnemonic is typically preceded with an e_. To distinguish similar instructions available in both 16- and 32-bit forms under VLE and standard Book E instructions, VLE instructions encoded with 16 bits have an se_ prefix. Those VLE instructions encoded with 32 bits that have different binary encodings or semantics than the equivalent Book E instruction have an e_ prefix. The following are examples: stw rS,D(rA) // standard Book E instruction e_stw rS,D(rA) // 32-bit VLE instruction se_stw rZ,SD4(rX) // 16-bit VLE instruction

10 VLE storage addressing

instruction, or when it fetches the next sequential instruction.

10.1 Data memory addressing modes

Table 218 lists data memory addressing modes supported by the VLE extension.

10.2 Instruction memory addressing modes

Table 219 lists instruction memory addressing modes supported by the VLE extension. Table 218. Data storage addressing modes rX = 0 is not a special case). Table 219. Instruction storage addressing modes extended, and then added to the address of the branch instruction. form the EA of the next instruction. form the EA of the next instruction. the next sequential instruction is undefined.

Table 219. Instruction storage addressing modes (continued)

RM0004 VLE compatibility with the EIS

11 VLE compatibility with the EIS

The body of this document addresses the relationship between VLE and Book E. It does not explicitly address EIS-defined features, such as the APUs or the use of MAS registers. However, the information in the previous chapters provides a model for how the VLE extension is integrated with features defined by the layer of architecture defined by the EIS.

11.1 Overview

The VLE extension uses the same semantics as the Book E architecture. Due to the limited instruction encoding formats, VLE instructions typically support reduced immediate fields and displacements, and not all Book E operations are encoded in the VLE extension. The basic philosophy is to capture all useful operations, with most frequent operations given priority. Immediate fields and displacements are provided to cover the majority of ranges encountered in embedded control code. Instructions are encoded in either a 16- or 32-bit format, and these may be freely intermixed. Book E floating-point registers (FPRs) are not accessible by VLE instructions. VLE instructions use Book E GPR and SPR registers with the following limitations:

  • VLE instructions using the 16-bit formats are limited to addressing GPR0–GPR7, and GPR24–GPR31 in most instructions. Move instructions are provided to transfer register contents between these registers and GPR8–GPR23.
  • VLE instructions using the 16-bit formats are limited to addressing CR0
  • VLE instructions using the 32-bit formats are limited to addressing CR0–CR3 VLE instruction encodings are generally different than Book E instructions, except that most Book E instructions falling within Book E major opcode 31 are encoded identically in 32-bit VLE instructions and have identical semantics unless they affect or access a resource not supported by the VLE extension. Also, major opcode 4 is available to support additional APUs using identical encodings for both Book E and the VLE extension. This allows an implementation of the VLE extension to include additional APUs, such as the cache-line locking, single-precision floating-point, and SPE APUs, and to use the exact encodings. Because future compatibility is desired, and to avoid confusion with Book E, register bit numbering remains the same as in Book E.

11.2 VLE extension processor and storage control extensions

This section describes additional functionality and extensions to the EIS to support the VLE extension.

11.2.1 EIS instruction extensions

This section describes extensions to EIS instructions to support VLE operations. Because instructions may reside on a half-word boundary, bit 62 is not masked by instructions that cause fetching from a register, such as the LR, CTR, or a save/restore register 0, that holds an instruction address:

  • Return from interrupt instructions, such as rfdi (defined as part of the debug APU) and rfmci (defined as part of the machine check APU) no longer mask bit 62 of the respective save/restore register 0. The destination address is xSRR0[32–62] || 1’b0.

11.2.2 Book E instruction extensions

  • rfci, rfdi, and rfi no longer mask bit 62 of CSRR0, DSRR0, or SRR0. The destination address is xSRR0[32–62] || 1’b0.
  • bclr, bclrl, bcctr, and bcctrl no longer mask bit 62 of the LR or CTR. The destination address is [LR,CTR][32–62] || 1’b0.

11.2.3 EIS MMU extensions

indicates the corresponding page of memory is a VLE page. additional value shown in Table 220. described in the following sections. the VLE extension is not present, this bit is always read as zero and writes are ignored. Table 220. TLB Entry 0 reset value

MAS2[VLE] is defined in Table 221. read as zero and writes are ignored. MAS4 is shown below. MAS4[VLED]is described in Table 222. Table 221. MAS2 field descriptions Table 222. MAS4 field descriptions

58 VLED

Default VLE value. Defined by the EIS.

0 This page is a standard Book E page

1 This page is a VLE page

VLE compatibility with the EIS RM0004

11.2.4 EIS debug APU extensions

The se_rfdi instruction is provided to support the EIS debug interrupt APU. rfdi rfdi Return from debug interrupt se_rfdi MSR ← DSRR1 NIA ← DSRR032:62 || 0b0 The se_rfdi instruction is used to return from a debug class interrupt, or as a means of establishing a new context and synchronizing on that new context simultaneously. The contents of DSRR1 are placed into the MSR. If the new MSR value does not enable any pending exceptions, then the next instruction is fetched, under control of the new MSR value, from the address DSRR0[32–62]||0b0. If the new MSR value enables one or more pending exceptions, the interrupt associated with the highest priority pending exception is generated; in this case the value placed into SRR0 or CSRR0 by the interrupt processing mechanism (see Book E) is the address of the instruction that would have been executed next had the interrupt not occurred (that is, the address in DSRR0 at the time of the execution of se_rfdi). Execution of this instruction is privileged and restricted to supervisor mode only. Execution of this instruction is context synchronizing. When the debug APU is disabled, this instruction is treated as an illegal instruction. Special Registers Altered: MSR 0 15 0000000000001010

12 VLE instruction classes

12.1 Processor control instructions

  • Chapter 12.1.1: System linkage instructions on page 860”
  • Chapter 12.1.2: Processor control register manipulation instructions on page 860”
  • Chapter 12.1.3: Instruction synchronization instruction on page 861”

12.1.1 System linkage instructions

which the system can return from performing a service or from processing an interrupt. Table 223 lists system linkage instructions.

12.1.2 Processor control register manipulation instructions

lists the processor control register manipulation instructions. Table 223. System linkage instruction set index Table 224. System register manipulation instruction set index

12.1.3 Instruction synch ronization instruction

Table 225 lists the VLE-defined se_isync instruction.

12.2 Branch operation instructions

and the registers that support them.

12.2.1 Registers for branch operations

  • Chapter 2.5.1: Condition register (CR) on page 61”
  • Chapter 2.5.2: Link register (LR) on page 66”
  • Chapter 2.5.3: Count register (CTR) on page 67” Condition register (CR) The condition register (CR) is a 32-bit register. CR bits are numbered 32 (most-significant bit) to 63 (least-significant bit). The CR reflects the result of certain operations, and provides a mechanism for testing (and branching). The VLE extension implements the entire CR, but some comparison operations and all branch instructions are limited to using CR0–CR3. The full Book E condition register field and logical operations are provided however. mfspr rD,SPRN Move From Special Purpose Register Book E se_mtctr rX Move To Count Register Page -942 mtdcr DCRN,rS Move To Device Control Register Book E se_mtlr rX Move To Link Register Page -943 mtmsr rS Move To Machine State Register Book E mtspr SPRN,rS Move To Special Purpose Register Book E wrtee rA Write MSR External Enable Book E wrteei E Write MSR External Enable Immediate Book E

Table 224. System register manipulation instruction set index (continued) Table 225. Instruction Synchronization Instruction Set Index

  • Specified fields of the condition register can be set by a move to the CR from a GPR (mtcrf).
  • A specified CR field can be set by a move to the CR from another CR field (e_mcrf).
  • CR field 0 can be set as the implicit result of an integer instruction.
  • A specified condition register field can be set as the result of an integer compare instruction.
  • CR field 0 can be set as the result of an integer bit test instruction. Instructions are provided to perform logical operations on individual CR bits and to test individual condition register bits (see Book E). Condition register settings for integer instructions For all integer word instructions in which the Rc bit is defined and set, and for addic., the first three bits of CR field 0 (CR[32–34]) are se t by signed comparison of bits 32–63 of the result to zero, and the fourth bit of CR field 0 (CR[35]) is copied from the final state of XER[SO]. if (target_register) 32:63 < 0 then c ← 0b100 else if (target_register)32:63 > 0 then c ← 0b010 else c ← 0b001 CR0 ← c || XERSO If any portion of the result is undefined, the value placed into the first three bits of CR field 0 is undefined. The bits of CR field 0 are interpreted as shown in Table 226. Condition register setting for compare instructions For compare instructions, a CR field specified by the crD operand in for the e_cmph, e_cmphl, e_cmpi, and e_cmpli instructions, or CR0 for the e_cmp16i, e_cmph16i, e_cmphl16i, e_cmpl16i, se_cmp, se_cmph, se_cmphl, se_cmpi, and se_cmpli instructions, is set to reflect the result of the comparison. The CR field bits are interpreted as shown in Table 227. A complete description of how the bits are set is given in the instruction descriptions and Chapter 12.4.5: Integer compare and bit test instructions on page 872.”

Table 226. CR0 encodings 32 Negative (LT). Bit 32 of the result is equal to 1. 34 Zero (EQ). Bits 32–63 of the result are equal to 0.

Table 227. Condition register setting for compare instructions For signed-integer compare, GPR(rA or rX) < SCI8 or SI or GPR(rB or rY). For signed-integer compare, GPR(rA or rX) > SCI8 or SI or UI5 or GPR(rB or rY). For integer compare, GPR(rA or rX) = SCI8 or UI5 or SI or UI or GPR(rB or rY). Table 228. Branch to link register instruction comparison

12.2.2 Branch instructions

generated branch target address is forced to 0 by the processor in performing the branch.

  1. Adding a displacement to the address of the branch instruction.
  2. Using the address contained in the LR (Branch to Link Register [and Link]).
  3. Using the address contained in the CTR (Branch to Count Register [and Link]).

computed: this is done whether or not the branch is taken. which the branch is taken and how the branch is affected by or affects the CR and CTR.

  1. BI refers to the BI field in the branch instruction encoding. For example, specifying BI = 2

Encodings for the BO32 field for the VLE extension are shown in Table 230. The encoding for the BO16 field for the VLE extension is shown in Table 231. Table 229. Branch to count register instruction comparison Table 230. VLE extension BO32 encodings 00 Branch if the condition is FALSE. 01 Branch if the condition is TRUE. 10 Decrement CTR[32–63] , then branch if the decremented CTR[32–63]≠0. 11 Decrement CTR[32–63], then branch if the decremented CTR[32–63] = 0.

The various branch instructions supported by the VLE extension are shown in Table 232.

12.3 Condition register instructions

Table 231. VLE extension BO16 encodings 0 Branch if the condition is FALSE. 1 Branch if the condition is TRUE. Table 232. Branch instruction set index

12.4 Integer instructions

This section lists the integer instructions supported by the VLE extension.

12.4.1 Integer load instructions

The VLE extension supports both big- and little-endian byte ordering for data accesses. load with update instructions in Book E. Basic integer load instructions are listed in Table 234. Table 233. Condition register instruction set index Table 234. Basic integer load instruction set index

Integer load byte-reversed instructions are listed in Table 235. The VLE-defined integer load multiple instruction is listed in Table 236. The VLE-defined integer load and reserve instruction is listed in Table 237.

12.4.2 Integer store instructions

The VLE extension supports both big- and little-endian byte ordering for data accesses. Table 235. Integer Load Byte-Reverse Instruction Set Index Table 236. Integer load multiple instruction set index Table 237. Integer load and reserve instruction set index Table 234. Basic integer load instruction set index (continued)

EA. For these forms, the following rules (from Book E) apply.

  • If rA ≠ 0, the EA is placed into GPR(rA).
  • If rS = rA, the contents of GPR(rS) are copied to the target memory element and then EA is placed into GPR(rA). The basic integer store instructions are listed in Table 238. The integer store byte-reverse instructions are listed in Table 239. The integer store multiple instruction is listed in Table 240. The integer store conditional instruction is listed in Table 241.

Table 238. Basic integer store instruction set index Table 239. Integer store byte-reverse instruction set index Table 240. Integer store multiple instruction set index Table 241. Integer store conditional instruction set index

12.4.3 Integer arithmetic instructions

place results into GPRs, into status bits in the XER and into CR0. integers unless the instruction is explicitly identified as performing an unsigned operation. 32–63 of the result to zero. e_addic[.] and e_subfic[.] always set CA to reflect the carry out of bit 32. The integer arithmetic instructions are listed in Table 242. Table 242. Integer arithmetic instruction set index

12.4.4 Integer logical and move instructions

another GPR, or an immediate value. instructions on page 869.” The logical instructions do not change XER[SO,OV,CA]. The integer logical instructions are listed in Table 243. Table 242. Integer arithmetic instruction set index (continued)

Table 243. Integer logical instruction set index

12.4.5 Integer compare and bit test instructions

  • The value of the SCI8 field
  • The zero-extended value of the UI field
  • The zero-extended value of the UI5 field
  • The sign-extended value of the SI field
  • The contents of GPR(rB) or GPR(rY). The following comparisons are signed: e_cmph, e_cmpi, e_cmp16i, e_cmph16i, se_cmp, se_cmph, and se_cmpi. The following comparisons are unsigned: e_cmphl, e_cmpli, e_cmphl16i, e_cmpl16i, se_cmpli, se_cmpl, and se_cmphl. When operands are treated as 32-bit signed quantities, GPRn[32] is the sign bit. When operands are treated as 16-bit signed quantities, GPRn[48] is the sign bit. For 32-bit implementations, the L field must be zero. Compare instructions set one of the left-most three bits of the designated CR field and clears the other two. XER[SO] is copied to bit 3 of the designated CR field. The CR field is set as shown in Table 244. The integer bit test instruction tests the bit specified by the UI5 instruction field and sets the CR0 field as shown in Table 245. orc rA,rS,rB orc. rA,rS,rB OR with Complement Book E e_ori[.] rA,rS,SCI8 e_or2i rD,UI OR Immediate Page -966 e_or2is rD,UI OR Immediate Shifted Page -966 xor rA,rS,rB xor. rA,rS,rB XOR Book E e_xori[.] rA,rS,SCI8 XOR Immediate Page -966

Table 243. Integer logical instruction set index (continued) Table 244. CR settings for compare instructions

Table 246 is an index for integer compare and bit test operations.

12.4.6 Integer select instruction

destination register under the control of a predicate value supplied by a CR bit. The integer select instruction is listed in Table 247.

12.4.7 Integer trap instructions

Table 245. CR settings for integer bit test instructions Table 246. Integer compare and bit test instruction set index Table 247. Integer select instruction set index

instruction execution continues normally. the contents of bits 32–63 of rA (and rB) participate in the comparison. The integer trap instruction is listed in Table 249.

12.4.8 Integer rotate and shift instructions

result, or a portion of the result, to a GPR. The rotation operations rotate a 32-bit quantity left by a specified number of bit positions. Bits that exit from position 32 enter at position 63. The rotate32 operation is used to rotate a given 32-bit quantity. There is no way to specify an all-zero mask. Table 248. Integer trap conditions Table 249. Integer trap instruction set index

The use of the mask is described in following sections. and shift instructions, except algebraic right shifts, do not change the CA bit. type, the amount of the rotation is either specified as an immediate, or contained in a GPR. ANDed with a mask before being placed into the target register. (in concept) by a left-rotation of 32-n, where n is the number of bits by which to rotate right. concept) by a left-rotation of 32-n, where n is the number of bits by which to rotate right. The integer shift instructions are listed in <Cross Refs>Table 252. Table 250. Integer rotate instruction set index Table 251. Integer rotate with mask instruction set index Table 252. Integer shift instruction set index

12.5 Storage control instructions

  • Chapter 12.5.1: Storage synchronization instructions on page 876”
  • Chapter 12.5.2: Cache management instructions on page 876”
  • Chapter 12.5.3: TLB management instructions on page 877”

12.5.1 Storage synchronization instructions

The storage synchronization instructions are listed in Table 253.

12.5.2 Cache management instructions

The cache management instructions are listed in Table 254. Table 252. Integer shift instruction set index (continued) Table 253. Storage synchronization instruction set index

12.5.3 TLB management instructions

defined in Book E and in the EIS. The TLB management instructions are listed in Table 255.

12.5.4 Instruction alig nment and byte ordering

generate an instruction storage interrupt byte-ordering exception.

12.6 Instruction listings

This section lists instructions either defined or supported by the VLE extension. Table 256 lists instructions by instruction name. Table 254. Cache management instruction set index Table 255. TLB management instruction set index

Table 256. Instructions listed by name

Table 256. Instructions listed by name (continued)

Table 257 lists instructions by mnemonic. Table 257. Instructions listed by mnemonic

Table 257. Instructions listed by mnemonic (continued)

13 VLE instruction set

instruction descriptions generally assume the appropriate calculation has been performed. e_cmpi and se_cmpi are both listed under cmpi.

13.1 Book E– and EIS-defined instructions

the EIS. Full descriptions of those instructions can be found in the EREF . compared to their Book E and EIS equivalents. Table 258. Book E– and EIS-defined instructions listed by mnemonic

Table 258. Book E– and EIS-defined instructions listed by mnemonic (continued)

13.2 Immediate field and displacement field encodings

Table 259 shows encodings for immediate and displacement fields. Table 259. Immediate field and displacement field encodings to the current instruction address to form the branch target address. instruction address to form the branch target address. added to the current instruction address to form the branch target address. extended to 32 bits for the e_li instruction. value, not the OIM5 instruction field binary encoding. immediate is shifted left three bits (concatenated with 0b000).

instruction. The instruction encoding differs between the I16A and I16L instruction formats. select a register bit in the range 0–31. zero-extended to 32 bits and used as the operand of the instruction. Table 259. Immediate field and displacement field encodings (continued)

VLE instruction set RM0004 _addx _addx Add se_add r X,rY sum32:63 ← GPR(RX) + GPR(RY) GPR(RX) ← sum32:63 The sum of the contents of GPR(rX) and the contents of GPR(rY) is placed into GPR(rX). Special registers altered: None 05 6 1 0 1 1 1 5 0000010 0 R Y R X VLE User

RM0004 VLE instruction set _addix _addix Add [2 operand] Immediate [Shifted] [and Record] e_add16i r D,rA,SI a ← GPR(RA) b ← EXTS(SI) GPR(RD) ← a + b The sum of the contents of GPR(rA) and the sign-extended value of field SI is placed into GPR(rD). Special Registers Altered: None e_add2i. r A,SI SI ← SI 0:4 || SI5:15 sum32:63 ← GPR(RA) + EXTS(SI) LT ← sum32:63 < 0 GT ← sum32:63 > 0 EQ ← sum32:63 = 0 GPR(RA) ← sum32:63 The sum of the contents of GPR(rA) and the sign-extended value of SI is placed into GPR(rA). Special Registers Altered: CR0 e_add2is r A,SI SI ← SI 0:4 || SI5:15 sum32:63 ← GPR(RD) + (SI || 160) GPR(RA) ← sum32:63 The sum of the contents of GPR(rA) and the value of SI concatenated with 16 zeros is placed into GPR(rA). Special Registers Altered: None e_addi r D,rA,SCI8 (Rc = 0) e_addi. r D,rA,SCI8 (Rc = 1) VLE User 05 6 1 0 1 1 1 5 1 6 3 1 0 5 6 1 01 1 1 51 6 2 02 1 3 1

011100 S I 0:4 R A 1 0001 S I 5:15

0 5 6 1 01 1 1 51 6 2 02 1 3 1

011100 S I 0:4 R A 1 0010 S I 5:15

0 5 6 1 01 1 1 51 6 2 02 12 22 32 4 3 1

000110 R D R A 1000 R c FS C L U I 8

VLE instruction set RM0004 imm ← SCI8(F ,SCL,UI8) sum32:63 ← GPR(RA) + imm if Rc=1 then do LT ← sum 32:63 < 0 GT ← sum32:63 > 0 EQ ← sum32:63 = 0 GPR(RD) ← sum32:63 The sum of the contents of GPR(rA) and the value of SCI8 is placed into GPR(rD). Special Registers Altered: CR0 (if Rc = 1) se_addi r X,OIMM GPR(RX) ← GPR(RX) + ( 270 || OFFSET(OIM5)) The sum of the contents of GPR(rX) and the zero-extended offset value of OIM5 (a final value in the range 1–32), is placed into GPR(rX). Special Registers Altered: None 0 5 6 7 11 12 15

0010000 O I M 5 (1)

  1. OIMM = OIM5 +1 RX

RM0004 VLE instruction set _addicx _addicx Add Immediate Carrying [and Record] e_addic r D,rA,SCI8 (Rc = 0) e_addic. r D,rA,SCI8 (Rc = 1) imm ← SCI8(F ,SCL,UI8) carry32:63 ← Carry(GPR(RA) + imm) sum32:63 ← GPR(RA) + imm if Rc=1 then do LT ← sum 32:63 < 0 GT ← sum32:63 > 0 EQ ← sum32:63 = 0 GPR(RD) ← sum32:63 CA ← carry32 The sum of the contents of GPR(rA) and the value of SCI8 is placed into GPR(rD). Special Registers Altered: CA, CR0 (if Rc=1) VLE User 0 5 6 1 01 1 1 51 6 1 92 02 12 22 32 4 3 1

000110 R D R A 1001 R c FS C L U I 8

VLE instruction set RM0004 _andx _andx AND [2 operand] [Immediate | with Complement] [and Record] se_and r X,rY( R c = 0 ) se_and. r X,rY( R c = 1 ) e_and2i. r D,UI e_and2is. r D,UI e_andi r A,rS,SCI8 (Rc = 0) e_andi. r A,rS,SCI8 (Rc = 1) se_andi r X,UI5 se_andc r X,rY if ‘e_andi[.]’ then b ← SCI8(F ,SCL,UI8) if ‘se_andi’ then b ← UI5 if ‘se_and[.]’ then b ← GPR(RY) if ‘se_andc’ then b ← ¬GPR(RY) if ‘e_and2i.’ then b ← 160 || UI0:4 || UI5:15 if ‘e_and2is.’ then b ← UI0:4 || UI5:15 || 160 result32:63 ← GPR(RS or RD or RX) & b if Rc=1 then do LT ← result32:63 < 0 GT ← result32:63 > 0 EQ ← result32:63 = 0 if ‘se_and[ci]’ then GPR(RX) ← result32:63 else GPR(RA or RD) ← result32:63 For e_andi[.], the contents of GPR(rS) are ANDed with the value of SCI8. For e_and2i., the contents of GPR(rD) are ANDed with 160 || UI. For e_and2is., the contents of GPR(rD) are ANDed with UI || 160. For se_andi, the contents of GPR(rX) are ANDed with the value of UI5. For se_and[.], the contents of GPR(rX) are ANDed with the contents of GPR(rY). 0 5678 1 1 1 2 1 5 0 5 6 1 01 1 1 51 6 2 02 1 3 1

011100 R D U I 0:4 11001 U I 5:15

0 5 6 1 01 1 1 51 6 2 02 1 3 1 0 5 6 1 01 1 1 51 6 2 02 12 22 32 4 3 1

000110 R S R A 1100 R c FS C L U I 8

RM0004 VLE instruction set For se_andc, the contents of GPR(rX) are ANDed with the one’s complement of the contents of GPR(rY). The result is placed into GPR(rA) or GPR(rX) (se_and[ic][.]) Special Registers Altered: CR0 (if Rc = 1)

VLE instruction set RM0004 _bx _bx Branch [and Link] e_b BD24 (LK = 0) e_bl BD24 (LK = 1) a ← CIA NIA ← (a + EXTS(BD24||0b0))32:63 if LK=1 then LR ← CIA + 4 Let the BTEA be calculated as follows:

  • For e_b[l], let BTEA be the sum of the CIA and the sign-extended value of the BD24 instruction field concatenated with 0b0. The BTEA is the address of the next instruction to be executed. If LK = 1, the sum CIA+4 is placed into the LR. Special Registers Altered: LR (if LK = 1) se_b BD8 (LK = 0) se_bl BD8 (LK = 1) a ← CIA NIA ← (a + EXTS(BD8||0b0)) 32:63 if LK=1 then LR ← CIA + 2 Let the BTEA be calculated as follows:
  • For se_b[l], let BTEA be the sum of the CIA and the sign-extended value of the BD8 instruction field concatenated with 0b0. The BTEA is the address of the next instruction to be executed. If LK = 1, the sum CIA+2 is placed into the LR. Special Registers Altered: LR (if LK = 1) VLE User 0 567 30 31

RM0004 VLE instruction set _bcx _bcx Branch Conditional [and Link] e_bc BO32,BI32,BD15 (LK = 0) e_bcl BO32,BI32,BD15 (LK = 1) if BO320 then CTR32:63 ← CTR32:63 – 1 ctr_ok ← ¬BO320 | ((CTR32:63 ≠ 0) ⊕ BO321) cond_ok ← BO320 | (CRBI32+32 ≡ BO321) if ctr_ok & cond_ok then NIA ← (CIA + EXTS(BD15 || 0b0)) 32:63 else NIA ← CIA + 4 if LK=1 then LR ← CIA + 4 Let the BTEA be calculated as follows:

  • For e_bc[l], let BTEA be the sum of the CIA and the sign-extended value of the BD15 instruction field concatenated with 0b0. BO32 specifies any conditions that must be met for the branch to be taken, as defined in Chapter 12.2.2: Branch instructions on page 864.” The sum BI32+32 specifies the CR bit. Only CR[32–47] may be specified. If the branch conditions are met, the BTEA is the address of the next instruction to be executed. If LK = 1, the sum CIA + 4 is placed into the LR. Special Registers Altered: CTR (if BO32 0 =1 ) LR (if LK = 1) se_bc BO16,BI16,BD8 cond_ok ← (CRBI16+32 ≡ BO16) if cond_ok then NIA ← (CIA + EXTS(BD8 || 0b0)) 32:63 else NIA ← CIA + 2 Let the BTEA be calculated as follows:
  • For se_bc, BTEA is the sum of the CIA and the sign-extended value of the BD8 instruction field concatenated with 0b0. BO16 specifies any conditions that must be met for the branch to be taken, as defined in Chapter 12.2.2: Branch instructions on page 864.” The sum BI16+32 specifies CR bit; only CR[32–35] may be specified. If the branch conditions are met, the BTEA is the address of the next instruction to be executed. Special Registers Altered: None VLE User 0 5 6 9 10 11 12 15 16 30 31 011110 1 0 0 0B O 3 2 B I 3 2 B D 1 5 L K 0 4 5 678 1 5

11100 B O 1 6 B I 1 6 B D 8

VLE instruction set RM0004 _bclri _bclri Bit Clear Immediate se_bclri r X,UI5 a ← UI5 result32:63 ← GPR(RX) & b GPR(RX) ← result32:63 For se_bclri, the bit of GPR(rX) specified by the value of UI5 is cleared and all other bits in GPR(rX) remain unaffected. Special Registers Altered: None 0 5 6 7 11 12 15

RM0004 VLE instruction set _bctrx _bctrx Branch to Count Register [and Link] se_bctr (LK = 0) se_bctrl (LK = 1) NIA ← CTR32:62 || 0b0 if LK=1 then LR ← CIA + 2 Let the BTEA be calculated as follows:

  • For se_bctr[l], let BTEA be bits 32–62 of the contents of the CTR concatenated with 0b0. The BTEA is the address of the next instruction to be executed. If LK = 1, the sum CIA + 2 is placed into the LR. Special Registers Altered: LR (if LK = 1) 01 4 1 5

VLE instruction set RM0004 _bgeni _bgeni Bit Generate Immediate se_bgeni r X,UI5 a ← UI5 GPR(RX) ← b For se_bgeni, a constant value consisting of a single ’1’ bit surrounded by ’0’s is generated and the value is placed into GPR(rX). The position of the ’1’ bit is specified by the UI5 field. Special Registers Altered: None 0 5 6 7 11 12 15

RM0004 VLE instruction set _blrx _blrx Branch to Link Register [and Link] se_blr (LK = 0) se_blrl (LK = 1) NIA ← LR32:62 || 0b0 if LK=1 then LR ← CIA + 2 Let the BTEA be calculated as follows:

  • For se_blr[l], let BTEA be bits 32–62 of the contents of the LR concatenated with 0b0. The BTEA is the address of the next instruction to be executed. If LK = 1, the sum CIA + 2 is placed into the LR. Special Registers Altered: LR (if LK = 1) 01 4 1 5

VLE instruction set RM0004 _bmaski _bmaski Bit Mask Generate Immediate se_bmaski r X,UI5 a ← UI5 if a = 0 then b ← 321 else b ← 32-a0 || a1 GPR(RX) ← b For se_bmaski, a constant value consisting of a mask of low-order ’1’ bits that is zero- extended to 32 bits is generated, and the value is placed into GPR(rX). The number of low- order ’1’ bits is specified by the UI5 field. If UI5 is 0b00000, a value of all ’1’s is generated Special Registers Altered: None 0 5 6 7 11 12 15

RM0004 VLE instruction set _bseti _bseti Bit Set Immediate se_bseti r X,UI5 a ← UI5 result32:63 ← GPR(RX) | b GPR(RX) ← result32:63 For se_bseti, the bit of GPR(rX) specified by the value of UI5 is set, and all other bits in GPR(rX) remain unaffected. Special Registers Altered: None 0 5 6 7 11 12 15

VLE instruction set RM0004 _btsti _btsti Bit Test Immediate se_btsti r X,UI5 a ← UI5 c ← GPR(RX) & b if c = 320 then d ← 0b001 else d ← 0b010 CR0:3 ← d || XERSO For se_btsti, the bit of GPR(rX) specified by the value of UI5 is tested for equality to ’1’. The result of the test is recorded in the CR. EQ is set if the tested bit is clear, LT is cleared, and GT is set to the inverse value of EQ. Special Registers Altered: CR[0–3] 0 5 6 7 11 12 15

RM0004 VLE instruction set _cmp _cmp Compare [Immediate] e_cmp16i r A,SI e_cmpi cr D32,rA,SCI8 a ← GPR(RA)32:63 if ‘e_cmpi’ then b ← SCI8(F ,SCL,UI8) if ‘e_cmp16i’ then b ← EXTS(SI0:4 || SI5:15) if a < b then c ← 0b100 if a > b then c ← 0b010 if a = b then c ← 0b001 if ‘e_cmpi’ then CR4×CRD32+32:4×CRD32+35 ← c || XERSO // only CR0-CR3 if ‘e_cmp16i’ then CR32:35 ← c || XERSO // only CR0 If e_cmpi, GPR(rA) contents are compared with the value of SCI8, treating operands as signed integers. If e_cmp16i, GPR(rA) contents are compared with the sign-extended value of the SI field, treating operands as signed integers. The result of the comparison is placed into CR field crD (crD32). For e_cmpi, only CR0– CR3 may be specified. For e_cmp16i, only CR0 may be specified. Special Registers Altered: CR field crD (crD32) (CR0 for e_cmp16i) se_cmp r X,rY se_cmpi r X,UI5 a ← GPR(RX)32:63 if ‘se_cmpi’ then b ← 270 || UI5 if ‘se_cmp’ then b ← GPR(RY)32:63 if a < b then c ← 0b100 if a > b then c ← 0b010 if a = b then c ← 0b001 CR0:3 ← c || XERSO VLE User 0 5 6 1 01 1 1 51 6 2 02 1 3 1

011100 S I 0:4 R A 1 0011 S I 5:15

0 5 6 8 9 1 01 1 1 51 6 2 02 12 22 32 4 3 1 0001100 00 CRD32 RA 10101FS C L U I 8 0 5678 1 1 1 2 1 5

VLE instruction set RM0004 If se_cmp, the contents of GPR(rX) are compared with the contents of GPR(rY), treating the operands as signed integers. The result of the comparison is placed into CR field 0. If se_cmpi, the contents of GPR(rX) are compared with the value of the zero-extended UI5 field, treating the operands as signed integers. The result of the comparison is placed into CR field 0. Special Registers Altered: CR[0–3]

RM0004 VLE instruction set _cmph _cmph Compare Halfword [Immediate] e_cmph cr D,rA,rB a ← EXTS(GPR(RA)48:63) b ← EXTS(GPR(RB)48:63) if a < b then c ← 0b100 if a > b then c ← 0b010 if a = b then c ← 0b001 CR4×CRD+32:4×CRD+35 ← c || XERSO For e_cmph, the contents of the low-order 16 bits of GPR(rA) and GPR(rB) are compared, treating the operands as signed integers. The result of the comparison is placed into CR field CRD. Special Registers Altered: CR field CRD se_cmph r X,rY a ← EXTS(GPR(RX) 48:63) b ← EXTS(GPR(RY)48:63) if a < b then c ← 0b100 if a > b then c ← 0b010 if a = b then c ← 0b001 CR0:3 ← c || XERSO For se_cmph, the contents of the low-order 16 bits of GPR(rX) and GPR(rY) are compared, treating the operands as signed integers. The result of the comparison is placed into CR field 0. Special Registers Altered: CR[0–3] e_cmph16i r A,SI a ← EXTS(GPR(RA) 48:63) b ← EXTS(SI0:4 || SI5:15) if a < b then c ← 0b100 if a > b then c ← 0b010 if a = b then c ← 0b001 CR32:35 ← c || XERSO // only CR0 VLE User 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1 0 5 6 1 01 1 1 51 6 2 02 1 3 1

011100 S I 0:4 R A 1 0110 S I 5:15

VLE instruction set RM0004 The contents of the lower 16-bits of GPR(rA) are sign-extended and compared with the sign-extended value of the SI field, treating the operands as signed integers. The result of the comparison is placed into CR0. Special Registers Altered: CR0

RM0004 VLE instruction set _cmphl __cmphl _ Compare Halfword Logical [Immediate] e_cmphl cr D,rA,rB a ← EXTZ(GPR(RA)48:63) b ← EXTZ(GPR(RB)48:63) if a < b then c ← 0b100 if a > b then c ← 0b010 if a = b then c ← 0b001 CR4×CRD+32:4×CRD+35 ← c || XERSO For e_cmphl, the contents of the low-order 16 bits of GPR(rA) and GPR(rB) are compared, treating the operands as unsigned integers. The result of the comparison is placed into CR field CRD. Special Registers Altered: CR field CRD se_cmphl r X,rY a ← GPR(RX) 48:63 b ← GPR(RY)48:63 if a < b then c ← 0b100 if a > b then c ← 0b010 if a = b then c ← 0b001 CR0:3 ← c || XERSO For se_cmphl, the contents of the low-order 16 bits of GPR(rX) and GPR(rY) are compared, treating the operands as unsigned integers. The result of the comparison is placed into CR field 0. Special Registers Altered: CR[0–3] e_cmphl16i r A,UI a ← 160 || GPR(RA)48:63) if a < b then c ← 0b100 if a > b then c ← 0b010 if a = b then c ← 0b001 CR32:35 ← c || XERSO // only CR0 VLE User 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1 0 5 6 1 01 1 1 51 6 2 02 1 3 1

011100 U I 0:4 R A 1 0111 U I 5:15

VLE instruction set RM0004 The contents of the lower 16-bits of GPR(rA) are zero-extended and compared with the zero-extended value of the UI field, treating the operands as unsigned integers. The result of the comparison is placed into CR0. Special Registers Altered: CR0

RM0004 VLE instruction set _cmpl __cmpl _ Compare Logical [Immediate] e_cmpl16i r A,UI e_cmpli cr D32,rA,SCI8 a ← GPR(RA)32:63 if ‘e_cmpli’ then b ← SCI8(F ,SCL,UI8) if ‘e_cmpl16i’ then b ← 160 || UI0:4 || UI5:15 if a <u b then c ← 0b100 if a >u b then c ← 0b010 if a = b then c ← 0b001 if ‘e_cmpli’ then CR4×CRD32+32:4×CRD32+35 ← c || XERSO // only CR0-CR3 if ‘e_cmp16i’ then CR32:35 ← c || XERSO // only CR0 If e_cmpi, the contents of bits 32–63 of GPR( rA) are compared with the value of SCI8, treating the operands as unsigned integers. L must be 0 for 32-bit implementations If e_cmpl16i, the contents of GPR(rA) are compared with the zero-extended value of the UI field, treating the operands as unsigned integers. The result of the comparison is placed into CR field CRD (CRD32). For e_cmpli, only CR0– CR3 may be specified. For e_cmpl16i, only CR0 may be specified. Special Registers Altered: CR field CRD (CRD32) (CR0 for e_cmpl16i) se_cmpl r X,rY se_cmpli r X,OIMM a ← GPR(RX) 32:63 if ‘se_cmpli’ then b ← 270 || OFFSET(OIM5) if ‘se_cmpl’ then b ← GPR(RY)32:63 if a <u b then c ← 0b100 if a >u b then c ← 0b010 if a = b then c ← 0b001 VLE User 0 5 6 1 01 1 1 51 6 2 02 1 3 1

011100 U I 0:4 R A 1 0101 U I 5:15

0 5 6 8 9 1 01 1 1 51 6 2 02 12 22 32 4 3 1 0001100 01 CRD32 RA 10101FS C L U I 8 0 5678 1 1 1 2 1 5

0010001 O I M 5 (1)

  1. OIMM = OIM5 +1 RX

VLE instruction set RM0004 CR0:3 ← c || XERSO If se_cmpl, the contents of GPR(rX) are compared with the contents of GPR(rY), treating the operands as unsigned integers. The result of the comparison is placed into CR field 0. If se_cmpli, the contents of GPR(rX) are compared with the value of the zero-extended offset value of the OIM5 field (a final value in the range 1–32), treating the operands as unsigned integers. The result of the comparison is placed into CR field 0. Special Registers Altered: CR[0–3]

RM0004 VLE instruction set _crand __crand _ Condition Register AND e_crand crb D,crbA,crbB CRBT+32 ← CRBA+32 & CRBB+32 The content of bit CRBA+32 of the CR is ANDed with the content of bit CRBB+32 of the CR, and the result is placed into bit CRBD+32 of the CR. Special Registers Altered: CR Condition Register AND with Complement e_crandc crb D,crbA,crbB CR BT+32 ← CRBA+32 & ¬CRBB+32 The content of bit CRBA+32 of the CR is ANDed with the one’s complement of the content of bit CRBB+32 of the CR, and the result is placed into bit CRBD+32 of the CR. Special Registers Altered: CR CR Equivalent e_creqv crb D,crbA,crbB CR BT+32 ← CRBA+32 ≡ CRBB+32 The content of bit CRBA+32 of the CR is XORed with the content of bit CRBB+32 of the CR, and the one’s complement of result is placed into bit CRBD+32 of the CR. Special Registers Altered: CR VLE User 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1

VLE instruction set RM0004 _crnand __crnand _ Condition Register NAND e_crnand crb D,crbA,crbB CRBT+32 ← ¬(CRBA+32 & CRBB+32) The content of bit CRBA+32 of the CR is ANDed with the content of bit CRBB+32 of the CR, and the one’s complement of the result is placed into bit CRBD+32 of the CR. Special Registers Altered: CR VLE User 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1

RM0004 VLE instruction set _crnor _ crnor Condition Register NOR e_crnor crb D,crbA,crbB CRBT+32 ← ¬(CRBA+32 | CRBB+32) The content of bit CRBA+32 of the CR is ORed with the content of bit CRBB+32 of the CR, and the one’s complement of the result is placed into bit CRBD+32 of the CR. Special Registers Altered: CR VLE User 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1

VLE instruction set RM0004 _cror _cror Condition Register OR e_cror crb D,crbA,crbB CRBT+32 ← CRBA+32 | CRBB+32 The content of bit CRBA+32 of the CR is ORed with the content of bit CRBB+32 of the CR, and the result is placed into bit CRBD+32 of the CR. Special Registers Altered: CR VLE User 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1

RM0004 VLE instruction set _cror __cror _ Condition Register OR with Complement e_crorc crb D,crbA,crbB CRBT+32 ← CRBA+32 | ¬CRBB+32 The content of bit CRBA+32 of the CR is ORed with the one’s complement of the content of bit CRBB+32 of the CR, and the result is placed into bit CRBD+32 of the CR. Special Registers Altered: CR VLE User 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1

VLE instruction set RM0004 _crxor __crxor _ Condition Register XOR e_crxor crb D,crbA,crbB CRcrbD+32 ← CRBA+32 ⊕ CRBB+32 The content of bit CRBA+32 of the CR is XORed with the content of bit CRBB+32 of the CR, and the result is placed into bit CRBD+32 of the CR. Special Registers Altered: CR VLE User 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1

RM0004 VLE instruction set _extsbx _extsbx Extend Sign (Byte | Halfword) se_extsb r X se_extsh0. r X if se_extsb then n ← 56 if se_extsh then n ← 48 if ‘extsw’ then n ← 32 if Rc=1 then do LT ← GPR(RS)n:63 < 0 GT ← GPR(RS)n:63 > 0 EQ ← GPR(RS)n:63 = 0 s ← GPR(RS or RX)n GPR(RA or RX) ← n-32s || GPR(RS or RX)n:63 For se_extsb, the contents of bits 56–63 of GPR(rX) are placed into bits 56–63 of GPR(rX). Bit 56 of the contents of GPR(rX) is copied into bits 32–55 of GPR( rX). For se_extsh, the contents of bits 48–63 of GPR(rX) are placed into bits 48–63 of GPR(rX). Bit 48 of the contents of GPR(rX) is copied into bits 32–47 of GPR( rX). Special Registers Altered: CR0 (if Rc=1) 05 6 1 1 1 2 1 5

VLE instruction set RM0004 _extzx extzx Extend Zero (Byte | Halfword) se_extzb r X se_extzh r X if ‘se_extzb’ then n ← 56 if ‘se_extzh’ then n ← 48 GPR(RX) ← n-320 || GPR(RX)n:63 For se_extzb, the contents of bits 56–63 of GPR(rX) are placed into bits 56–63 of GPR(rX). Bits 32–55 of GPR(rX) are cleared. For se_extzh, the contents of bits 48–63 of GPR(rX) are placed into bits 48–63 of GPR(rX). Bits 32–47 of GPR(rX) are cleared. Special Registers Altered: None 05 6 1 1 1 2 1 5

RM0004 VLE instruction set _illegal _illegal Illegal se_illegal SRR1 ← MSR SRR0 ← CIA NIA ← IVPR32:47 || IVOR648:59 || 0b0000 MSRWE,EE,PR,IS,DS,FP,FE0,FE1 ← 0b0000_0000 se_illegal is used to request an illegal instruction exception. A program interrupt is generated. The contents of the MSR are copied into SRR1 and the address of the se_illegal instruction is placed into SRR0. MSR[WE,EE,PR,IS,DS,FP ,FE0,FE1] are cleared. The interrupt causes the next instruction to be fetched from address IVPR[32– This instruction is context synchronizing. Special Registers Altered: SRR0 SR R1 MSR[WE,EE,PR,IS,DS,FP ,FE0,FE1] 01 5 0000000000000000 VLE User

VLE instruction set RM0004 _isync _isync Instruction Synchronize se_isync The se_isync instruction provides an ordering function for the effects of all instructions executed by the processor executing the se_isync instruction. Executing an se_isync instruction ensures that all instructions preceding the se_isync instruction have completed before the se_isync instruction completes, and that no subsequent instructions are initiated until after the se_isync instruction completes. It also causes any prefetched instructions to be discarded, with the effect that subsequent instructions are fetched and executed in the context established by the instructions preceding the se_isync instruction. The se_isync instruction may complete before memory accesses associated with instructions preceding the se_isync instruction have been performed. This instruction is context synchronizing (see Book E). It has identical semantics to Book E isync, just a different encoding. Special Registers Altered: None 01 5 0000000000000001 VLE User

RM0004 VLE instruction set _lbzx _lbzx Load Byte and Zero [with Update] [Indexed] e_lbz r D,D(rA) (D-mode) se_lbz r Z,SD4(rX) (SD4-mode) e_lbzu r D,D8(rA) (D8-mode) if (RA=0 & !se_lbz) then a ← 320 else a ← GPR(RA or RX) if D-mode then EA ← (a + EXTS(D))32:63 if D8-mode then EA ← (a + EXTS(D8))32:63 if SD4-mode then EA ← (a + (280 || SD4))32:63 GPR(RD or RZ) ← 240 || MEM(EA,1) if e_lbzu then GPR(RA) ← EA Let the EA be calculated as follows:

  • For e_lbz and e_lbzu, let EA be the sum of the contents of GPR(rA), or 32 0s if rA=0 , and the sign-extended value of the D or D8 instruction field.
  • For se_lbz, let EA be the sum of the contents of GPR(rX) and the zero-extended value of the SD4 instruction field. The byte in memory addressed by EA is loaded into bits 56–63 of GPR(rD or rZ). Bits 32–55 of GPR(rD or rZ) are cleared. If e_lbzu, EA is placed into GPR(rA). If e_lbzu and rA = 0 or rA= rD, the instruction form is invalid. Special Registers Altered: None VLE User 05 6 1 0 1 1 1 5 1 6 3 1

0 5 6 1 01 1 1 51 6 2 32 4 3 1

VLE instruction set RM0004 _lhax _lhax Load Halfword Algebraic [with Update] [Indexed] e_lha r D,D(rA) (D-mode) e_lhau r D,D8(rA) (D8-mode) if RA=0 then a ← 320 else a ← GPR(RA) if D-mode then EA ← (a + EXTS(D))32:63 if D8-mode then EA ← (a + EXTS(D8))32:63 GPR(RD) ← EXTS(MEM(EA,2))32:63 if e_lhau then GPR(RA) ← EA Let the EA be calculated as follows:

  • For e_lha and e_lhau, let EA be the sum of the contents of GPR(rA), or 32 0s if rA=0 , and the sign-extended value of the D or D8 instruction field. The half word in memory addressed by EA is loaded into bits 48–63 of GPR(rD). Bits 32–47 of GPR(rD) are filled with a copy of bit 0 of the loaded half word. If e_lhau, EA is placed into GPR(rA). If e_lhau and rA = 0 or rA= rD, the instruction form is invalid. Special Registers Altered: None VLE User 05 6 1 0 1 1 1 5 1 6 3 1

0 5 6 1 01 1 1 51 6 2 32 4 3 1

RM0004 VLE instruction set _lhzx _lhzx Load Halfword and Zero [with Update] [Indexed] e_lhz r D,D(rA) (D-mode) se_lhz r Z,SD4(rX) (SD4-mode) e_lhzu r D,D8(rA) (D8-mode) if (RA=0 & !se_lhz) then a ← 320 else a ← GPR(RA or RX) if D-mode then EA ← (a + EXTS(D))32:63 if D8-mode then EA ← (a + EXTS(D8))32:63 if SD4-mode then EA ← (a + (270 || SD4 || 0))32:63 GPR(RD or RZ) ← 160 || MEM(EA,2) if e_lhzu then GPR(RA) ← EA Let the EA be calculated as follows:

  • For e_lhz and e_lhzu, let EA be the sum of the contents of GPR(rA), or 32 0s if rA=0 , and the sign-extended value of the D or D8 instruction field.
  • For se_lhz let EA be the sum of the contents of GPR(rX) and the zero-extended value of the SD4 instruction field shifted left by 1 bit. The half word in memory addressed by EA is loaded into bits 48–63 of GPR(rD). Bits 32–47 of GPR(rD) are cleared. If e_lhzu, EA is placed into GPR(rA). If e_lhzu and rA = 0 or rA= rD, the instruction form is invalid. Special Registers Altered: None VLE User 05 6 1 0 1 1 1 5 1 6 3 1

0 5 6 1 01 1 1 51 6 2 32 4 3 1

VLE instruction set RM0004 _lix _lix Load Immediate [Shifted] e_li r D,LI20 (LI20-mode) LI20 ← LI200:3 || LI204:8 || LI209:19 GPR(RD) ← EXTS(LI20) For e_li, the sign-extended LI20 field is placed into GPR(rD). Special Registers Altered: None e_lis r D,UI UI ← UI 0:4 || UI5:15 GPR(RD) ← UI || 160 For e_lis, the UI field is concatenated on the right with 16 0’s and placed into GPR(rD). Special Registers Altered: None se_li r X,UI7 GPR(RX) ← 250 || UI7 For se_li, the zero-extended UI7 field is placed into GPR(rX). Special Registers Altered: None VLE User 0 5 6 1 01 1 1 51 61 7 2 02 1 3 1

011100 R D L I 2 0 4:8 0L I 2 0 0:3 LI209:19

0 5 6 1 01 1 1 51 6 2 02 1 3 1

011100 R D U I 0:4 11100 U I 5:15

RM0004 VLE instruction set _lmw _lmw Load Multiple Word e_lmw r D,D8(rA) if RA=0 then EA ← EXTS(D8)32:63 else EA ← (GPR(RA)+EXTS(D8))32:63 r ← RD do while r ≤ 31 GPR(r) ← MEM(EA,4) r ← r + 1 EA ← (EA+4)32:63 Let the EA be the sum of the contents of GPR(rA), or 32 0s if rA = 0, and the sign-extended value of the D8 instruction field. Let n = (32-rD). n consecutive words starting at EA are loaded into bits 32–63 of registers GPR(rD) through GPR(31). EA must be a multiple of 4. If it is not, either an alignment interrupt is invoked or the results are boundedly undefined. If rA is in the range of registers to be loaded, including the case in which rA = 0, the instruction form is invalid. Special Registers Altered: None VLE User 0 5 6 1 01 1 1 51 6 2 32 4 3 1

VLE instruction set RM0004 _lwz _lwz Load Word and Zero [with Update] [Indexed] e_lwz r D,D(rA) (D-mode) se_lwz r Z,SD4(rX) (SD4-mode) e_lwzu r D,D8(rA) (D8-mode) if (RA=0 & !se_lwz) then a ← 320 else a ← GPR(RA or RX) if D-mode then EA ← (a + EXTS(D))32:63 if D8-mode then EA ← (a + EXTS(D8))32:63 if SD4-mode then EA ← (a + (260 || SD4 || 20))32:63 GPR(RD or RZ) ← MEM(EA,4) if e_lwzu then GPR(RA) ← EA Let the EA be calculated as follows:

  • For e_lwz and e_lwzu, let EA be the sum of the contents of GPR(rA), or 32 0s if rA=0 , and the sign-extended value of the D or D8 instruction field.
  • For se_lwz let EA be the sum of the contents of GPR(rX) and the zero-extended value of the SD4 instruction field shifted left by 2 bits. The word in memory addressed by the EA is loaded into bits 32–63 of GPR( rD). If e_lwzu, EA is placed into GPR(rA). If e_lwzu and rA = 0 or rA= rD, the instruction form is invalid. Special Registers Altered: None VLE User 05 6 1 0 1 1 1 5 1 6 3 1

0 5 6 1 01 1 1 51 6 2 32 4 3 1

RM0004 VLE instruction set _mcrf _mcrf Move CR Field e_mcrf cr D,crS CR4xCRD+32:4xCRD+35 ← CR4xCRS+32:4xCRS+35 The contents of field crS (bits 4×CRS+32 through 4×CRS+35) of the CR are copied to field crD (bits 4×CRD+32 through 4×CRD+35) of the CR. Special Registers Altered: CR VLE User 0 5 6 8 9 1 01 1 1 31 4 2 02 1 3 03 1

VLE instruction set RM0004 _mfar _mfar Move from Alternate Register se_mfar r X,arY GPR(RX) ← GPR(ARY) For se_mfar, the contents of GPR(arY) are placed into GPR(rX). arY specifies a GPR in the range R8–R23. The encoding 0000 specifies R8, 0001 specifies R9,…, 1111 specifies R23. Special Registers Altered: None 0 5678 1 1 1 2 1 5

RM0004 VLE instruction set _mfctr _mfctr Move From Count Register se_mfctr r X GPR(RX) ← CTR The CTR contents are placed into bits 32–63 of GPR( rX). Special Registers Altered: None 05 6 1 1 1 2 1 5 000000 0 0 1 0 1 0 R X VLE User

VLE instruction set RM0004 _mflr _mflr Move From Link Register se_mflr r X GPR(RX) ← LR The LR contents are placed into bits 32–63 of GPR( rX). Special Registers Altered: None 05 6 1 1 1 2 1 5 000000 0 0 1 0 0 0 R X VLE User

RM0004 VLE instruction set _mr _mr Move Register se_mr r X,rY GPR(RX) ← GPR(RY) For se_mr, the contents of GPR(rY) are placed into GPR(rX). Special Registers Altered: None 0 5678 1 1 1 2 1 5

VLE instruction set RM0004 _mtar _mtar Move to Alternate Register se_mtar ar X,rY GPR(ARX) ← GPR(RY) For se_mtar, the contents of GPR(rY) are placed into GPR(arX). arX specifies a GPR in the range R8–R23. The encoding 0000 specifies R8, 0001 specifies R9,…, 1111 specifies R23. Special Registers Altered: None 0 5678 1 1 1 2 1 5

RM0004 VLE instruction set _mtctr _mtctr Move To Count Register se_mtctr r X CTR ← GPR(RX) The contents of bits 32–63 of GPR( rX) are placed into the CTR. Special Registers Altered: CTR 05 6 1 1 1 2 1 5 000000 0 0 1 0 1 1 R X VLE User

VLE instruction set RM0004 _mtlr _mtlr Move To Link Register se_mtlr r X LR ← GPR(RX) The contents of bits 32–63 of GPR( rX) are placed into the LR. Special Registers Altered: LR 05 6 1 1 1 2 1 5 000000 0 0 1 0 0 1 R X VLE User

RM0004 VLE instruction set _mullix _mullix Multiply Low [2 operand] Immediate e_mulli r D,rA,SCI8 imm ← SCI8(F ,SCL,UI8) prod0:63 ← GPR(RA) × imm GPR(RD) ← prod32:63 Bits 32–63 of the 64-bit product of the contents of GPR(rA) and the value of SCI8 are placed into GPR(rD). Both operands and the product are interpreted as signed integers. Special Registers Altered: None e_mull2i r A,SI prod 0:63 ← GPR(RA) × EXTS(SI0:4 || SI5:15) GPR(RA) ← prod32:63 Bits 32–63 of the 64-bit product of the contents of GPR( rA) and the sign-extended value of the SI field are placed into GPR(rA). Both operands and the product are interpreted as signed integers. Special Registers Altered: None VLE User 0 5 6 1 01 1 1 51 6 2 02 12 22 32 4 3 1

000110 R D R A 10100FS C L U I 8

0 5 6 1 01 1 1 51 6 2 02 1 3 1

011100 S I 0:4 R A 1 0100 S I 5:15

VLE instruction set RM0004 _mullwx _mullwx Multiply Low Word se_mullw r X,rY prod0:63 ← GPR(RX)32:63 × GPR(RY)32:63 GPR(RX) ← prod32:63 Bits 32–63 of the 64-bit product of the contents of bits 32–63 of GPR(rX) and the contents of bits 32–63 of GPR(rY) is placed into GPR(rX). Special Registers Altered: None 0 5678 1 1 1 2 1 5

RM0004 VLE instruction set _negx _negx Negate se_neg r X result32:63 ← ¬GPR(RX)+ 1 GPR(RX) ← result32:63 The sum of the one’s complement of the contents of GPR(rX) and 1 is placed into GPR(rX). If bits 32–63 of GPR(rX) contain the most negative 32-bit number (0x8000_0000), bits 32– 63 of the result contain the most negative 32-bit number Special Registers Altered: None 05 6 1 1 1 2 1 5

VLE instruction set RM0004 _notx _notx NOT se_not r X result32:63 ← ¬GPR(RX) GPR(RX) ← result32:63 The contents of GPR(rX) are inverted. Special Registers Altered: None 05 6 1 1 2 1 1 5

RM0004 VLE instruction set _orx _orx OR [2 operand] [Immediate | with Complement] [Shifted][and Record] se_or r X,rY e_or2i r D,UI e_or2is r D,UI e_ori r A,rS,SCI8 (Rc = 0) e_ori. r A,rS,SCI8 (Rc = 1) if ‘e_ori[.]’ then b ← SCI8(F ,SCL,UI8) if ‘e_or2i’ then b ← 160 || UI0:4 || UI5:15 if ‘e_or2is’ then b ← UI0:4 || UI5:15 || 160 if ‘se_or’ then b ← GPR(RB) result0:63 ← GPR(RS or RD or RX) | b if Rc=1 then do LT ← result32:63 < 0 GT ← result32:63 > 0 EQ ← result32:63 = 0 GPR(RA or RD or RX) ← result For e_ori[.], the contents of GPR(rS) are ORed with the value of SCI8. For e_or2i, the contents of GPR(rD) are ORed with 160 || UI. For e_or2is, the contents of GPR(rD) are ORed with UI || 160. For se_or, the contents of GPR(rX) are ORed with the contents of GPR(rY). The result is placed into GPR(rA or rX). The preferred ‘no-op’ (an instruction that does nothing) is: e_ori 0,0,0 Special Registers Altered: CR0 (if Rc = 1) 0 5678 1 1 1 2 1 5 0 5 6 1 01 1 1 51 6 2 02 1 3 1

011100 R D U I 0:4 11000 U I 5:15

0 5 6 1 01 1 1 51 6 2 02 1 3 1

011100 R D U I 0:4 11010 U I 5:15

0 5 6 1 01 1 1 51 6 1 92 02 12 22 32 4 3 1

000110 R S R A 1101 R c FS C L U I 8

VLE instruction set RM0004 _rfci _rfci Return From Critical Interrupt se_rfci MSR ← CSRR1 NIA ← CSRR00:62 || 0b0 The se_rfci instruction is used to return from a critical class interrupt, or as a means of establishing a new context and synchronizing on that new context simultaneously. The contents of CSRR1 are placed into the MSR. If the new MSR value does not enable any pending exceptions, then the next instruction is fetched, under control of the new MSR value, from the address CSRR0[32–62]||0b0. If the new MSR value enables one or more pending exceptions, the interrupt associated with the highest priority pending exception is generated; in this case the value placed into SRR0 or CSRR0 by the interrupt processing mechanism (see Book E) is the address of the instruction that would have been executed next had the interrupt not occurred (that is, the address in CSRR0 at the time of the execution of the se_rfci). Execution of this instruction is privileged and restricted to supervisor mode. Execution of this instruction is context synchronizing. Special Registers Altered: MSR 01 5 0000000000001001 VLE Supervisor

RM0004 VLE instruction set _rfi _rfi Return From Interrupt se_rfi MSR ← SRR1 NIA ← SRR00:62 || 0b0 The se_rfi instruction is used to return from a non-critical class interrupt, or as a means of simultaneously establishing a new context and synchronizing on that new context. The contents of SRR1 are placed into the MSR. If the new MSR value does not enable any pending exceptions, then the next instruction is fetched under control of the new MSR value from the address SRR0[32–62]||0b0. If the new MSR value enables one or more pending exceptions, the interrupt associated with the highest priority pending exception is generated; in this case the value placed into SRR0 or CSRR0 by the interrupt processing mechanism (see Book E) is the address of the instruction that would have been executed next had the interrupt not occurred (that is, the address in SRR0 at the time of the execution of the se_rfi). Execution of this instruction is privileged and restricted to supervisor mode. Execution of this instruction is context synchronizing. Special Registers Altered: MSR 01 5 0000000000001000 VLE Supervisor

VLE instruction set RM0004 _rlw _rlw Rotate Left Word [Immediate] e_rlw r A,rS,rB( R c = 0 ) e_rlw. r A,rS,rB( R c = 1 ) e_rlwi r A,rS,SH (Rc = 0) e_rlwi. r A,rS,SH (Rc = 1) if ‘e_rlw[.]’ then n ← GPR(RB)59:63 else n ← SH result32:63 ← ROTL32(GPR(RS)32:63,n) if Rc=1 then do LT ← result32:63 < 0 GT ← result32:63 > 0 EQ ← result32:63 = 0 GPR(RA) ← result32:63 If e_rlw[.], let the shift count n be the contents of bits 59–63 of GPR( rB). If e_rlwi[.], let the shift count n be SH. The contents of GPR(rS) are rotated32 left n bits. The rotated data is placed into GPR(rA). Special Registers Altered: CR0 (if Rc = 1) VLE User 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1

RM0004 VLE instruction set _rlwimi _rlwimi Rotate Left Word Immediate then Mask Insert e_rlwimi r A,rS,SH,MB,ME n ← SH b ← MB+32 e ← ME+32 r ← ROTL32(GPR(RS)32:63,n) m ← MASK(b,e) result32:63 ← r&m | GPR(RA)&¬m GPR(RA) ← result32:63 Let the shift count n be the value SH. The contents of GPR(rS) are rotated32 left n bits. A mask is generated having 1 bits from bit MB+32 through bit ME+32 and 0 bits elsewhere. The rotated data are inserted into GPR(rA) under control of the generated mask (if a mask bit is 1 the associated bit of the rotated data is placed into the target register, and if the mask bit is 0 the associated bit in the target register remains unchanged). Special Registers Altered: None VLE User 0 5 6 1 01 1 1 51 6 2 02 1 2 52 6 3 03 1

VLE instruction set RM0004 _rlwinm _rlwinm Rotate Left Word Immediate then AND with Mask e_rlwinm r A,rS,SH,MB,ME n ← SH b ← MB+32 e ← ME+32 r ← ROTL32(GPR(RS)32:63,n) m ← MASK(b,e) result32:63 ← r & m GPR(RA) ← result32:63 Let the shift count n be SH. The contents of GPR(rS) are rotated32 left n bits. A mask is generated having 1 bits from bit MB+32 through bit ME+32 and 0 bits elsewhere. The rotated data are ANDed with the generated mask and the result is placed into GPR(rA). Special Registers Altered: None VLE User 0 5 6 1 01 1 1 51 6 2 02 1 2 52 6 3 03 1

RM0004 VLE instruction set _sc _sc System Call se_sc SRR1 ← MSR SRR0 ← CIA+2 NIA ← IVPR32:47 || IVOR848:59 || 0b0000 MSRWE,EE,PR,IS,DS,FP,FE0,FE1 ← 0b0000_0000 se_sc is used to request a system service. A system call interrupt is generated. The contents of the MSR are copied into SRR1 and the address of the instruction after the se_sc instruction is placed into SRR0. MSR[WE,EE,PR,IS,DS,FP ,FE0,FE1] are cleared. The interrupt causes the next instruction to be fetched from the address IVPR[32–47]||IVOR8[48–59]||0b0000 This instruction is context synchronizing. Special Registers Altered: SRR0 SR R1 MSR[WE,EE,PR,IS,DS,FP ,FE0,FE1] 01 5 0000000000000010 VLE User

VLE instruction set RM0004 _slwx _slwx Shift Left Word [Immediate] [and Record] e_slwi r A,rS,SH (Rc = 0) e_slwi. r A,rS,SH (Rc = 1) se_slw r X,rY se_slwi r X,UI5 if ‘e_slwi[.]’ then n ← SH if se_slw then n ← GPR(RY)58:63 if se_slwi then n ← UI5 r ← ROTL32(GPR(RS or RX)32:63,n) if n<32 then m ← MASK(32,63-n) else m ← 320 result32:63 ← r & m if Rc=1 then do LT ← result32:63 < 0 GT ← result32:63 > 0 EQ ← result32:63 = 0 GPR(RA or RX) ← result32:63 Let the shift count n be the value specified by the contents of bits 58–63 of GPR(rB or rY), or by the value of the SH or UI5 field. The contents of bits 32–63 of GPR(rS or rX) are shifted left n bits. Bits shifted out of position 32 are lost. Zeros are supplied to the vacated positions on the right. The 32-bit result is placed into bits 32–63 of GPR( rA or rX). Shift amounts from 32 to 63 give a zero result. Special Registers Altered: CR0 (if Rc = 1) VLE User 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1

RM0004 VLE instruction set _srawx _srawx Shift Right Algebraic Word [Immediate] [and Record] se_sraw r X,rY se_srawi r X,UI5 if ‘se_sraw’ then n ← GPR(RY)59:63 if ‘se_srawi’ then n ← UI5 r ← ROTL32(GPR(RS or RX)32:63,32-n) if ((se_sraw & GPR(RY)58=1) then m ← 320 else m ← MASK(n+32,63) s ← GPR(RS or RX)32 result0:63 ← r&m | (32s)&¬m if Rc=1 then do LT ← result32:63 < 0 GT ← result32:63 > 0 EQ ← result32:63 = 0 GPR(RA or RX) ← result32:63 If se_sraw, let the shift count n be the contents of bits 58–63 of GPR( rY). If se_srawi, let the shift count n be the value of the UI5 field. The contents of bits 32–63 of GPR( rS or rX) are shifted right n bits. Bits shifted out of position 63 are lost. Bit 32 of rS or rX is replicated to fill vacated positions on the left. The 32-bit result is placed into bits 32–63 of GPR( rA or rX). CA is set if bits 32–63 of GPR(rS or rX) contain a negative value and any 1 bits are shifted out of bit position 63; otherwise CA is cleared. A shift amount of zero causes GPR(rA or rX) to receive EXTS(GPR(rS or rX)32:63), and CA to be cleared. For se_sraw, shift amounts from 32 to 63 give a result of 64 sign bits, and cause CA to receive bit 32 of the contents of GPR(rS or rX) (that is, sign bit of GPR(rS or rX)32:63). Special Registers Altered: CA CR0 (if Rc = 1) 0 5678 1 1 1 2 1 5

VLE instruction set RM0004 _srwx _srwx Shift Right Word [Immediate] [and Record] e_srwi r A,rS,SH (Rc = 0) e_srwi. r A,rS,SH (Rc = 1) se_srw r X,rY se_srwi r X,UI5 n ← GPR(RB)59:63 if ‘e_srwi[.]’ then n ← SH if ‘se_srw’ then n ← GPR(RY)59:63 if ‘se_srwi’ then n ← UI5 r ← ROTL32(GPR(RS or RX)32:63,32-n) if ((se_srw & GPR(RY)58=1) then m ← 320 else m ← MASK(n+32,63) result32:63 ← r & m if Rc=1 then do LT ← result32:63 < 0 GT ← result32:63 > 0 EQ ← result32:63 = 0 GPR(RA or RX) ← result32:63 If e_srwi, let the shift count n be the value of the SH field. If se_srw, let the shift count n be the contents of bits 58–63 of GPR( rY). If se_srwi, let the shift count n be the value of the UI5 field. The contents of bits 32–63 of GPR( rS or rX) are shifted right n bits. Bits shifted out of position 63 are lost. Zeros are supplied to the vacated positions on the left. The 32-bit result is placed into bits 32–63 of GPR( rA or rX). Shift amounts from 32 to 63 give a zero result. Special Registers Altered: CR0 (if Rc = 1) VLE User 0 5 6 1 01 1 1 51 6 2 02 1 3 03 1

RM0004 VLE instruction set _stbx _stbx Store Byte [with Update] [Indexed] e_stb r S,D(rA) (D-mode) se_stb r Z,SD4(rX) (SD4-mode) e_stbu r S,D8(rA) (D8-mode) if (RA=0 & !se_stb) then a ← 320 else a ← GPR(RA or RX) if D-mode then EA ← (a + EXTS(D))32:63 if D8-mode then EA ← (a + EXTS(D8))32:63 if SD4-mode then EA ← (a + (280 || SD4))32:63 MEM(EA,1) ← GPR(RS or RZ)56:63 if e_stbu then GPR(RA) ← EA Let the EA be calculated as follows:

  • For e_stb and e_stbu, let EA be the sum of the contents of GPR(rA), or 32 0s if rA=0 , and the sign-extended value of the D or D8 instruction field.
  • For se_stb, let EA be the sum of the contents of GPR(rX) and the zero-extended value of the SD4 instruction field. The contents of bits 56–63 of GPR(rS) are stored into the byte in memory addressed by EA.
  • If e_stbu, EA is placed into GPR(rA).
  • If e_stbu and rA = 0, the instruction form is invalid.
  • None VLE User 05 6 1 0 1 1 1 5 1 6 3 1

0 5 6 1 01 1 1 51 6 2 32 4 3 1

VLE instruction set RM0004 _sthx _sthx Store Halfword [with Update] [Indexed] e_sth r S,D(rA) (D-mode) se_sth r Z,SD4(rX) (SD4-mode) e_sthu r S,D8(rA) (D8-mode) if (RA=0 & !se_sth) then a ← 320 else a ← GPR(RA or RX) if D-mode then EA ← (a + EXTS(D))32:63 if D8-mode then EA ← (a + EXTS(D8))32:63 if SD4-mode then EA ← (a + (270 || SD4 || 0))32:63 MEM(EA,2) ← GPR(RS or RZ)48:63 if e_sthu then GPR(RA) ← EA Let the EA be calculated as follows:

  • For e_sth and e_sthu, let EA be the sum of the contents of GPR(rA), or 32 0s if rA=0 , and the sign-extended value of the D or D8 instruction field.
  • For se_sth let EA be the sum of the contents of GPR(rX) and the zero-extended value of the SD4 instruction field shifted left by 1 bit. The contents of bits 48–63 of GPR( rS) are stored into the half word in memory addressed by EA. If e_sthu, EA is placed into GPR(rA). If e_sthu and rA = 0, the instruction form is invalid. Special Registers Altered: None VLE User 05 6 1 0 1 1 1 5 1 6 3 1

0 5 6 1 01 1 1 51 6 2 32 4 3 1

RM0004 VLE instruction set _stmw _stmw Store Multiple Word e_stmw r S,D8(rA) (D8-mode) if RA=0 then EA ← EXTS(D8)32:63 else EA ← (GPR(RA)+EXTS(D8))32:63 r ← RS do while r ≤ 31 MEM(EA,4) ← GPR(r)32:63 r ← r + 1 EA ← (EA+4)32:63 Let the EA be the sum of the contents of GPR(rA), or 32 0s if rA = 0, and the sign-extended value of the D8 instruction field. Let n = (32 - rS). Bits 32–63 of registers GPR(rS) through GPR(31) are stored in n consecutive words in memory starting at address EA. EA must be a multiple of 4. If it is not, either an alignment interrupt is invoked or the results are boundedly undefined. Special Registers Altered: None VLE User 0 5 6 1 01 1 1 51 6 2 32 4 3 1

VLE instruction set RM0004 _stwx _stwx Store Word [with Update] [Indexed] e_stw r S,D(rA) (D-mode) se_stw r Z,SD4(rX) (SD4-mode) e_stwu r S,D8(rA) (D8-mode) if (RA=0 & !se_stw) then a ← 320 else a ← GPR(RA or RX) if D-mode then EA ← (a + EXTS(D))32:63 if D8-mode then EA ← (a + EXTS(D8))32:63 if SD4-mode then EA ← (a + (260 || SD4 || 20))32:63 MEM(EA,4) ← GPR(RS or RZ)32:63 Let the EA be calculated as follows:

  • For e_stw and e_stwu, let EA be the sum of the contents of GPR(rA), or 32 0s if rA = 0, and the sign-extended value of the D or D8 instruction field.
  • For se_stw, let EA be the sum of the contents of GPR(rX) and the zero-extended value of the SD4 instruction field shifted left by 2 bits. The contents of bits 32–63 of GPR( rS) are stored into the word in memory addressed by EA. If e_stwu, EA is placed into GPR(rA). If e_stwu and rA = 0, the instruction form is invalid. Special Registers Altered: None VLE User 05 6 1 0 1 1 1 5 1 6 3 1

0 5 6 1 01 1 1 51 6 2 32 4 3 1

RM0004 VLE instruction set _sub _sub Subtract se_sub r X,rY sum32:63 ← GPR(RX) + ¬GPR(RY) + 1 GPR(RX) ← sum32:63 The sum of the contents of GPR(rX), the one’s complement of contents of GPR(rY), and 1 is placed into GPR(rX). Special Registers Altered: None 0 5678 1 1 1 2 1 5

VLE instruction set RM0004 _subfx _subfx Subtract From se_subf r X,rY sum32:63 ← ¬GPR(RX) + GPR(RY) + 1 GPR(RX) ← sum32:63 The sum of the one’s complement of the contents of GPR(rX), the contents of GPR(rY), and 1 is placed into GPR(rX). Special Registers Altered: None 05 6 1 0 1 1 1 5 VLE User

RM0004 VLE instruction set _subficx _subficx Subtract From Immediate Carrying [and Record] e_subfic r D,rA,SCI8 (Rc = 0) e_subfic. r D,rA,SCI8 (Rc = 1) imm ← SCI8(F ,SCL,UI8) carry32:63 ← Carry(¬GPR(RA) + imm + 1) sum32:63 ← ¬GPR(RA) + imm + 1 if Rc=1 then do LT ← sum 32:63 < 0 GT ← sum32:63 > 0 EQ ← sum32:63 = 0 GPR(RD) ← sum32:63 CA ← carry32 The sum of the one’s complement of the contents of GPR(rA), the value of SCI8, and 1 is placed into GPR(rD). Special Registers Altered: CA CR0 (if Rc=1) VLE User 0 5 6 1 01 1 1 51 6 2 02 12 22 32 4 3 1

000110 R D R A 1011 R c FS C L U I 8

VLE instruction set RM0004 _subix _subix Subtract Immediate [and Record] se_subi r X,OIMM (Rc = 0) se_subi. r X,OIMM (Rc = 1) sum32:63 ← GPR(RX) + ¬(270 || OFFSET(OIM5)) + 1 if Rc=1 then do LT ← sum32:63 < 0 GT ← sum32:63 > 0 EQ ← sum32:63 = 0 GPR(RX) ← sum32:63 The sum of the contents of GPR(rX), the one’s complement of the zero-extended value of the offseted OIM5 field (a final value in the range 1–32), and 1 is placed into GPR( rX). Special Registers Altered: CR0 (if Rc = 1) 0 5 6 7 11 12 15

001001 R c O I M 5 (1)

  1. OIMM = OIM5 +1 RX VLE User

RM0004 VLE instruction set _xorx _xorx XOR [Immediate] [and Record] e_xori r A,rS,SCI8 (Rc = 0) e_xori. r A,rS,SCI8 (Rc = 1) if ‘e_xori[.]’ then b ← SCI8(F ,SCL,UI8) result32:63 ← GPR(RS) ⊕ b if Rc=1 then do LT ← result32:63 < 0 GT ← result32:63 > 0 EQ ← result32:63 = 0 GPR(RA) ← result For e_xori[.], the contents of GPR(rS) are XORed with SCI8. The result is placed into GPR(rA). Special Registers Altered: CR0 (if Rc = 1) VLE User 0 5 6 1 01 1 1 51 6 1 92 02 12 22 32 4 3 1

000110 R S R A 1110 R c FS C L U I 8

14 VLE instruction index

14.1 Instruction index sorted by opcode

Table 261 lists the 16-bit VLE instructions, sorted by opcode. Table 260. Notation conventions ? Allocated for implementatio n-dependent use. See the implementation documentation. Table 261. Instruction index sorted by opcode

Table 261. Instruction index sorted by opcode (continued)

Table 262 shows 32-bit instruction encodings. Table 262. 32-bit instruction encodings

S C I 8 000110t t t t t aaaaa 10011FSSi i i i i i i i e_addic. Table 262. 32-bit instruction encodings (continued)

S C I 8 000110t t t t t aaaaa 10111FSSi i i i i i i i e_subfic.

14.2 Instruction index sorted by mnemonic

Table 263 lists all of the 16-bit VLE instructions, sorted by mnemonic. Table 263. 16-Bit VLE instructions sorted by mnemonic

Table 263. 16-Bit VLE instructions sorted by mnemonic (continued)

Table 264 outlines the 32-bit instruction encodings. Table 264. 32-bit instruction encodings (by mnemonic) I 1 6 A 011100i i i i i aaaaa 10001 i i i i i iiiii i e_add2i.

S C I 8 000110t t t t t aaaaa 10011FSSi i iiiii i e_addic. Table 264. 32-bit instruction encodings (by mnemonic) (continued)

S C I 8 000110t t t t t aaaaa 10111FSSi i iiiii i e_subfic.

14.3 Instruction index sorted by opcode

Table 265 lists all the 16-bit Power*Embedded instructions, sorted by opcode. Table 265. Instruction index sorted by opcode

Table 265. Instruction index sorted by opcode (continued)

Table 266 outlines the 32-bit instruction encodings. Table 266. 32-bit instruction encodings

Table 266. 32-bit instruction encodings (continued)

14.4 Instruction index sorted by mnemonic

Table 267 lists all the 16-bit Power*Embedded instructions, sorted by mnemonic. Table 267. Instruction index sorted by mnemonic

Table 267. Instruction index sorted by mnemonic (continued)

Table 268 sorts 32-bit instructions by mnemonic, ignoring the e_ prefix. Table 268. 32-bit instructions by mnemonic (ignoring the e_ prefix) S C I 8 001110t t t t t aaaaa 11001 FSSi i i i i i i i e_addic.

Table 268. 32-bit instructions by mnemonic (ignoring the e_ prefix) (continued)

S C I 8 001110t t t t t aaaaa 11110 FSSi i i i i i i i e_mulli.

011101SSSSS aaaaa hhhhh bbbbb eeeee 1 e_rlwinm. 011101SSSSS aaaaa hhhhh bbbbb eeeee 0 e_rlwimi.

S C I 8 001110t t t t t aaaaa 11100 FSSi i i i i i i i e_subfic.

  • Chapter A.1: Instructions sorted by mnemonic (decimal and hexadecimal)”
  • Chapter A.2: Instructions sorted by primary opcodes (decimal and hexadecimal) ”
  • Chapter A.3: Instructions sorted by mnemonic (binary) ”
  • Chapter A.4: Instructions sorted by opcode (binary) ”
  • Chapter A.5: Instruction set legend” Note that this appendix does not include instructions defined by the VLE extension. These instructions are listed in Chapter 14: VLE instruction index on page 862. A.1 Instructions sorted by mnem onic (decimal and hexadecimal) Table 269 lists instructions in alphabetical order by mnemonic, showing decimal and hexadecimal values of the primary opcode (0–5) and binary values of the secondary opcode (21–31). This list also includes simplified mn emonics and their equivalents using standard mnemonics.

Table 269. Instructions sorted by mnemonic (decimal and hexadecimal) addic. 13 (0x0D) rD rA SIMM D addic.

andi. 28 (0x1C) rS rA UIMM D andi. andis. 29 (0x1D) rS rA UIMM D andis. Table 269. Instructions sorted by mnemonic (decimal and hexadecimal) (continued)

rlwimi. 20 (0x14) rS rA SH MB ME 1 M rlwimi. rlwinm. 21 (0x15) rS rA SH MB ME 1 M rlwinm. rlwnm. 23 (0x17) rS rA rB MB ME 1 M rlwnm.

subic. subic. r D,rA,value equivalent to addic. rD,rA,–value subic.

  1. Simplified mnemonics for branch instructions that do not test a CR bit should not specify one; a programming error may occur.
  2. The value in the BI operand selects CR n[2], the EQ bit.
  3. The value in the BI operand selects CR n[0], the LT bit.
  4. The value in the BI operand selects CR n[1], the GT bit.
  5. The value in the BI operand selects CR n[3], the SO bit.
  6. Optional to the PowerPC classic architecture.
  7. Supervisor-level instruction
  8. Access level is detemined by whether the SPR is defined as a user- or supervisor-level SPR.

Table 270. Instructions sorted by primary opcodes (decimal and hexadecimal)

Table 270. Instructions sorted by primary opcodes (decimal and hexadecimal) (continued)

addic. 13 (0x0D) rD rA SIMM D addic. rlwimi. 20 (0x14) rS rA SH MB ME 1 M rlwimi. rlwinm. 21 (0x15) rS rA SH MB ME 1 M rlwinm. rlwnm. 23 (0x17) rS rA rB MB ME 1 M rlwnm.

andi. 28 (0x1C) rS rA UIMM D andi. andis. 29 (0x1D) rS rA UIMM D andis.

  1. Supervisor-level instruction

also includes simplified mnemonics and their equivalents using standard mnemonics. Table 271. Instructions sorted by mnemonic (binary) addic. 001101 r D r A S I M M D addic. andi. 011100 r S r A U I M M D andi. andis. 011101 r S r A U I M M D andis.

Table 271. Instructions sorted by mnemonic (binary) (continued)

rlwimi. 010100 r S r A S H M B M E 1 M rlwimi. rlwinm. 010101 r S r A S H M B M E 1 M rlwinm. rlwnm. 010111 r S r A r B M B M E 1 M rlwnm.

subic. subic. r D,rA,valueequivalent to addic. rD,rA,–value subic.

  1. Simplified mnemonics for branch instructions that do not test a CR bit should not specify one; a programming error may occur.
  2. The value in the BI operand selects CRn[2], the EQ bit.
  3. The value in the BI operand selects CRn[0], the LT bit.
  4. The value in the BI operand selects CRn[1], the GT bit.
  5. The value in the BI operand selects CRn[3], the SO bit.
  6. Optional to the PowerPC classic architecture.
  7. Supervisor-level instruction.
  8. Access level is detemined by whether the SPR is defined as a user or supervisor level SPR.

Table 272 lists instructions by opcode, shown in binary. Table 272. Instructions sorted by opcode (binary)

addic. 001101 r D r A S I M M D addic. Table 272. Instructions sorted by opcode (binary) (continued)

rlwimi. 010100 r S r A S H M B M E 1 M rlwimi. rlwinm. 010101 r S r A S H M B M E 1 M rlwinm. rlwnm. 010111 r S r A r B M B M E 1 M rlwnm. andi. 011100 r S r A U I M M D andi. andis. 011101 r S r A U I M M D andis.

  1. Supervisor-level instruction

Table 273. PowerPC instruction set legend

Table 273. PowerPC instruction set legend (continued)

  1. Load/Store string or multiple.
  2. Load/Store string or multiple.
  3. Supervisor and user level instruction.
  4. Supervisor and user level instruction.

Table 274. PowerPC instruction set legend

Table 274. PowerPC instruction set legend (continued)

RM0004 Simplified mnemonics for PowerPC instructions Appendix B Simplified mnemonics for PowerPC instructions This chapter describes simplified mnemonics, which are provided for easier coding of assembly language programs. Simplified mnemonics are defined for the most frequently used forms of branch conditional, compare, trap, rotate and shift, and certain other instructions defined by the PowerPC™ architecture and by implementations of and extensions to the PowerPC architecture. Chapter B.11: Comprehensive list of simplified mnemonics on page 1133,” provides an alphabetical listing of simplified mnemonics. Some assemblers may define additional simplified mnemonics not included here. The simplified mnemonics listed here should be supported by all compilers. B.1 Overview Simplified (or extended) mnemonics allow an assembly-language programmer to program using more intuitive mnemonics and symbols than the instructions and syntax defined by the instruction set architecture. For example, to code the conditional call “branch to an absolute target if CR4 specifies a greater than condition, setting the LR without simplified mnemonics, the programmer would write the branch conditional instruction, bc 12,17,target. The simplified mnemonic, branch if greater than, bgt cr4, target, incorporates the conditions. Not only is it easier to remember the symbols than the numbers when programming, it is also easier to interpret simplified mnemonics when reading existing code. Although the original PowerPC architecture documents include a set of simplified mnemonics, these are not a formal part of the architecture, but rather a recommendation for assemblers that support the instruction set. Many simplified mnemonics have been added to those originally included in the architecture documentation. Some assemblers created their own, and others have been added to support extensions to the instruction set (for example, AltiVec instructions and Book E auxiliary processing units (APUs)). Simplified mnemonics have been added for new architecturally defined and new implementation-specific special-purpose registers (SPRs). These simplified mnemonics are described only in a very general way. B.2 Subtract simplified mnemonics This section describes simplified mnemonics for subtract instructions. B.2.1 Subtract immediate There is no subtract immediate instruction, however, its effect is achieved by negating the immediate operand of an Add Immediate instruction, addi. Simplified mnemonics include this negation, making the intent of the computation more clear. These are listed in Table 275.

  • Extract—Select a field of n bits starting at bit position b in the source register; left or right justify this field in the target register; clear all other bits of the target register.
  • Insert—Select a left- or right-justified field of n bits in the source register; insert this field starting at bit position b of the target register; leave other bits of the target register unchanged.
  • Rotate—Rotate the contents of a register right or left n bits without masking.
  • Shift—Shift the contents of a register right or left n bits, clearing vacated bits (logical shift).
  • Clear—Clear the leftmost or rightmost n bits of a register.
  • Clear left and shift left—Clear the leftmost b bits of a register, then shift the register left by n bits. This operation can be used to scale a (known non-negative) array index by the width of an element. B.3.1 Operations on words The simplified mnemonics in Table 277 can be coded with a dot (.) suffix to cause the Rc bit to be set in the underlying instruction.

Table 275. Subtract immediate simplified mnemonics Table 276. Subtract simplified mnemonics

  1. rD,rB,rA is not the standard order for the operands. The order of rB and rA is reversed to show the

equivalent behavior of the simplified mnemonic.

  1. Extract the sign bit (bit 0) of rS and place the result right-justified into rA.
  2. Insert the bit extracted in (1) into the sign bit (bit 0) of rB.
  3. Shift the contents of rA left 8 bits.
  4. Clear the high-order 16 bits of rS and place the result into rA.

and a prediction, as part of the mnemonic rather than as numeric BO and BI operands. simplify unconditional branch mnemonics. Table 277. Word rotate and shift simplified mnemonics Table 278. Branch instructions

Simplified mnemonics for PowerPC instructions RM0004 The BO and BI operands correspond to two fields in the instruction opcode, as figure below shows for Branch Conditional (bc, bca, bcl, and bcla) instructions. The BO operand specifies branch operations that involve decrementing CTR. It is also used to determine whether testing a CR bit causes a branch to occur if the condition is true or false. The BI operand identifies a CR bit to test (whether a comparison is less than or greater than, for example). The simplified mnemonics avoid the need to memorize the numerical values for BO and BI. For example, bc 16,0,target is a conditional branch that, as a BO value of 16 (0b1_0000) indicates, decrements CTR, then branches if the decremented CTR is not zero. The operation specified by BO is abbreviated as d (for decrement) and nz (for not zero), which replace the c in the original mnemonic; so the simplified mnemonic for bc becomes bdnz. The branch does not depend on a condition in the CR, so BI can be eliminated, reducing the expression to bdnz target. In addition to CTR operations, the BO operand provides an optional prediction bit and a true or false indicator can be added. For example, if the previous instruction should branch only on an equal condition in CR0, the instruction becomes bc 8,2,target. To incorporate a true condition, the BO value becomes 8 (as shown in Table 280); the CR0 equal field is indicated by a BI value of 2 (as shown in Table 281). Incorporating the branch-if-true condition adds a ‘t’ to the simplified mnemonic, bdnzt. The equal condition, that is specified by a BI value of 2 (indicating the EQ bit in CR0) is replaced by the eq symbol. Using the simplified mnemonic and the eq operand, the expression becomes bdnzt eq,target. This example tests CR0[EQ]; however, to test the equal condition in CR5 (CR bit 22), the expression becomes bc 8,22,target. The BI operand of 22 indicates CR[22] (CR5[2], or BI field 0b10110), as shown in Table 281. This can be expressed as the simplified mnemonic. bdnzt 4 * cr5 + eq,target. The notation, 4 * cr5 + eq may at first seem awkward, but it eliminates computing the value of the CR bit. It can be seen that (4 * 5) + 2 = 22. Note that although 32-bit registers in Book E processors are numbered 32–63, only values 0–31 are valid (or possible) for BI operands. As shown in Table 282, a Book E–compliant processor automatically translates the bit values; specifying a BI value of 22 selects bit 55 on a Book E processor, or CR5[2] = CR5[EQ]. 0 5 6 1 01 1 1 51 6 2 93 03 1

001000 B O B I B D A A L K

RM0004 Simplified mnemonics for PowerPC instructions B.4.1 Key facts about si mplified branch mnemonics The following key points are helpful in understanding how to use simplified branch mnemonics:

  • All simplified branch mnemonics eliminate the BO operand, so if any operand is present in a branch simplified mnemonic, it is the BI operand (or a reduced form of it).
  • If the CR is not involved in the branch, the BI operand can be deleted.
  • If the CR is involved in the branch, the BI operand can be treated in the following ways: – It can be specified as a numeric value, just as it is in the architecturally defined instruction, or it can be indicated with an easier to remember formula, 4 * crn + [test bit symbol], where n indicates the CR field number. – The condition of the test bit (eq, lt, gt, and so) can be incorporated into the mnemonic, leaving the need for an operand that defines only the CR field. - If the test bit is in CR0, no operand is needed. - If the test bit is in CR1–CR7, the BI operand can be replaced with a crS operand (that is, cr1, cr2, cr3, and so forth). B.4.2 Eliminating the BO operand The 5-bit BO field, shown below, encodes the following operations in conditional branch instructions:
  • Decrement count register (CTR) – And test if result is equal to zero – And test if result is not equal to zero
  • Test condition register (CR) – Test condition true – Test condition false
  • Branch prediction (taken, fall through). If the prediction bit, y, is needed, it is signified by appending a plus or minus sign as described in Chapter B.4.3: Incorporating the BO branch prediction on page 1116.”

BO bits can be interpreted individually as described in Table 279. be cleared, as they may be assigned a meaning in a future version of the architecture. Table 279. BO bit encodings 0 If set, ignore the CR bit comparison.

1 If set, the CR bit comparison is against true, if not set the CR bit comparison is against false

2 If set, the CTR is not decremented.

3 If BO[2] is set, this bit determines whether the CTR comparison is for equal to zero or not

Table 280. BO operand encodings

  1. Assumes y = z = 0. Chapter B.4.3: Incorporating the BO branch prediction,” describes how to use simplified mnemonics to

program the y bit for static prediction.

  1. Instructions for which B0 is 12 (branc h if condition true) or 4 (branch if condition false) do not depend on the CTR value and
  2. A z bit indicates a bit that is ignored. However, these bits should be cleared, as they may be assigned a meaning in a

future version of the architecture.

RM0004 Simplified mnemonics for PowerPC instructions B.4.3 Incorporating the BO branch prediction As shown in Table 280, the low-order bit (y bit) of the BO field provides a hint about whether the branch is likely to be taken (static branch prediction). Assemblers should clear this bit unless otherwise directed. This default action indicates the following:

  • A branch conditional with a negative displacement field is predicted to be taken.
  • A branch conditional with a non-negative displacement field is predicted not to be taken (fall through).
  • A branch conditional to an address in the LR or CTR is predicted not to be taken (fall through). If the likely outcome (branch or fall through) of a given branch conditional instruction is known, a suffix can be added to the mnemonic that tells the assembler how to set the y bit. That is, ‘+’ indicates that the branch is to be taken and ‘–’ indicates that the branch is not to be taken. This suffix can be added to any branch conditional mnemonic, standard or simplified. For relative and absolute branches (bc[l][a]), the setting of the y bit depends on whether the displacement field is negative or non-negative. For negative displacement fields, coding the suffix ‘+’ causes the bit to be cleared, and coding the suffix ‘–’ causes the bit to be set. For non-negative displacement fields, coding the suffix ‘+’ causes the bit to be set, and coding the suffix ‘–’ causes the bit to be cleared. For branches to an address in the LR or CTR (bclr[l] or bcctr[l]), coding the suffix ‘+’ causes the y bit to be set, and coding the suffix ‘–’ causes the bit to be cleared. Examples of branch prediction follow: 1. Branch if CR0 reflects less than condition, specifying that the branch should be predicted as taken. blt+ target 2. Same as (1), but target address is in the LR and the branch should be predicted as not taken. bltlr– 4. Simplified mnemonics for branch instruct ions that do not test CR bits (BO = 16, 18, and 20) should specify only a target. Otherwise a programming error may occur. 5. Notice that these instructions do not use the branch if condition true or false operations. For that reason, simplified mnemonics for these should not specify a BI operand.

Simplified mnemonics for PowerPC instructions RM0004 B.4.4 The BI opera nd—CR bit and field representations With standard branch mnemonics, the BI operand is used when it is necessary to test a CR bit, as shown in the example in Chapter B.4: Branch instruction simplified mnemonics.” With simplified mnemonics, the BI operand is handled differently depending on whether the simplified mnemonic incorporates a CR condition to test, as follows:

  • Some branch simplified mnemonics incorporate only the BO operand. These simplified mnemonics can use the architecturally defined BI operand to specify the CR bit, as follows: – The BI operand can be presented exactly as it is with standard mnemonics—as a decimal number, 0–31. – Symbols can be used to replace the decimal operand, as shown in the example in Chapter B.4: Branch instruction simplified mnemonics,” where bdnzt 4 * cr5 + eq,target could be used instead of bdnzt 22,target. This is described in Specifying a CR bit on page 1118.” The simplified mnemonics in Chapter B.4.5: Simplified mnemonics that incorporate the BO operand,” use one of these two methods to specify a CR bit. – Additional simplified mnemonics are specified that incorporate CR conditions that would otherwise be specified by the BI operand, so the BI operand is replaced by the crS operand to specify the CR field, CR0–CR7. See BI operand instruction encoding on page 1117.” These mnemonics are described in Chapter B.4.6: Simplified mnemonics that incorporate CR conditions (eliminates BO and replaces BI with crS).” BI operand instruction encoding The entire 5-bit BI field, shown in Figure 180, represents the bit number for the CR bit to be tested. For standard branch mnemonics and for branch simplified mnemonics that do not incorporate a CR condition, the BI operand provides all 5 bits. For simplified branch mnemonics described in Chapter B.4.6,” the BI operand is replaced by a crS operand. To understand this, it is useful to view the BI operand as comprised of two parts. As Figure 180 shows, BI[0–2] indicates the CR field and BI[3–4] represents the condition to test.

RM0004 Simplified mnemonics for PowerPC instructions Figure 180. BI field (Bits 11–14 of the instruction encoding) Integer record-form instructions update CR0 as described in Table 281. 32 is automatically added to the BI value, as shown in Table 281 and Table 282. Table 281. CR0 and CR1 fields as updated by integer instructions

Description

E 0–2 3–4 CR0[0] 0 32 000 00 Negative (LT)—Set when the result is negative. CR0[1] 1 33 000 01 Positive (GT)—Set when t he result is positive (and not zero). CR0[2] 2 34 000 10 Zero (EQ)—Set when the result is zero. CR0[3] 3 35 000 11 Summary overflow (SO). Copy of XER[SO] at the instruction’s completion. 01234 BI[0–2] specifies CR field, CR0–CR7. Simplified mnemonics based on CR conditions but not CTR values—BO = 12 (branch if true) and BO = 4 branch if false) Specified by a separate, reduced BI operand (crS) Incorporated into the simplified mnemonic. Standard branch mnemonics and simplified mnemonics based on CTR values The BI operand specifies the entire 5-bit field. If CR0 is used, the bit can be identified by LT, GT, EQ, or SO. If CR1–CR7 are used, the form 4 * crS + LT|GT|EQ|SO can be used. BI Opcode Field BI[3–4] specifies one of the 4 bits in a CR field. (LT, GT, EQ,SO)

coded using a standard branch conditional syntax. can be used in the syntax used with the simplified mnemonic. Table 282. BI operand settings for CR fields for branch comparisons

added to a bit-number-within-CR-field symbol can be used, (for example, cr0 * 4 + eq). the basic mnemonics b, ba, bl, and bla are used. Table 283. CR field identification symbols Table 284. Branch simplified mnemonics

  1. Simplified mnemonics for branch instru ctions that do not test CR bits should specify only a target. Otherwise a

programming error may occur.

substituted for numeric values.

  1. Decrement CTR and branch if it is still nonz ero (closure of a loop controlled by a count

loaded into CTR) (note that no CR bits are tested). may be considered a programming error. Subsequent examples test conditions).

  1. Same as (1) but branch only if CTR is nonzero and equal condition in CR0.
  2. Same as (2), but equ al condition is in CR5.
  3. Branch if bit 59 of CR is false.
  4. Same as (4), but set the link regist er. This is a form of conditional call.

Table 286 lists simplified mnemonics and syntax for bc and bca without LR updating. Table 285. Branch instructions

  1. x stands for one of the symbols in Table 280, where applicable.
  2. BI can be a numeric value or an expression as shown in Table 283.

Table 286. Simplified mnemonics for bc and bca without LR update

Table 287 lists simplified mnemonics and syntax for bclr and bcctr without LR updating. Table 288 provides simplified mnemonics and syntax for bcl and bcla.

  1. Instructions for which B0 is either 12 (branch if conditi on true) or 4 (branch if condition false) do not depend on the CTR
  2. Simplified mnemonics for branch instru ctions that do not test CR bits should specify only a target. Otherwise a

programming error may occur. Table 286. Simplified mnemonics for bc and bca without LR update (continued) Table 287. Simplified mnemonics for bclr and bcctr without LR update

  1. Simplified mnemonics for branch instru ctions that do not test a CR bit should not specify one; a programming error may
  2. Instructions for which B0 is 12 (branc h if condition true) or 4 (branch if condition false) do not depend on a CTR value and

Table 288. Simplified mnemonics for bcl and bcla with LR update

Table 289 provides simplified mnemonics and syntax for bclrl and bcctrl with LR updating.

  1. Instructions for which B0 is either 12 (branch if conditi on true) or 4 (branch if condition false) do not depend on the CTR
  2. Simplified mnemonics for branch instruct ions that do not test CR bits should specify only a target. A programming error

Table 288. Simplified mnemonics for bcl and bcla with LR update (continued) Table 289. Simplified mnemonics for bclrl and bcctrl with LR update

  1. Simplified mnemonics for branch instru ctions that do not test a CR bit should not specify one. A programming error may

the test bit falls, so the BI operand is replaced by a crS operand. example, less than or equal (le) and not greater than (ng) achieve the same result. CR0 is used, no crS is necessary. no crS is specified, CR0 is used. Table 292 shows the simplified branch mnemonics incorporating conditions. Table 290. Standard coding for branch conditions Table 291. Branch instructions and simplified mnemonics that incorporate CR conditions

  1. x stands for one of the symbols in Table 290, where applicable.
  2. BI can be a numeric value or an expression as shown in Table 283.

cr7) are used for this operand, as shown in examples 2–4 below.

  1. Branch if CR0 reflects not-equal condition.
  2. Same as (1) but condition is in CR3.
  3. Branch to an absolute target if CR4 specifies greater than condition, setting the LR.

This is a form of conditional call.

  1. Same as (3), but target address is in the CTR.

Table 292. Simplified mnemonics with comparison conditions

Table 293. Simplified mnemonics for bc and bca without comparison conditions or LR Update

  1. The value in the BI operand selects CR n[0], the LT bit.
  2. The value in the BI operand selects CR n[1], the GT bit.
  3. The value in the BI operand selects CR n[2], the EQ bit.
  4. The value in the BI operand selects CR n[3], the SO bit.

Table 294. Simplified mnemonics for bclr and bcctr without comparison conditions or LR

  1. The value in the BI operand selects CR n[0], the LT bit.
  2. The value in the BI operand selects CR n[1], the GT bit.
  3. The value in the BI operand selects CR n[2], the EQ bit.
  4. The value in the BI operand selects CR n[3], the SO bit.

Table 295 shows simplified branch mnemonics and syntax for bcl and bcla. Table 295. Simplified mnemonics for bcl and bcla with comparison conditions, LR update

  1. The value in the BI operand selects CR n[0], the LT bit.
  2. The value in the BI operand selects CR n[1], the GT bit.
  3. The value in the BI operand selects CR n[2], the EQ bit.
  4. The value in the BI operand selects CR n[3], the SO bit.

Table 296. Simplified mnemonics for bclrl and bcctrl with comparison conditions, LR update

  1. The value in the BI operand selects CR n[0], the LT bit.
  2. The value in the BI operand selects CR n[1], the GT bit.
  3. The value in the BI operand selects CR n[2], the EQ bit.
  4. The value in the BI operand selects CR n[3], the SO bit.

(L = 1). Simplified mnemonics in Table 297 eliminate the L operand for word comparisons.

  1. Compare rA with immediate value 100 as signed 32-bit integers and place result in
  2. Same as (1), but place results in CR4.
  3. Compare rA and rB as unsigned 32-bit integers and place result in CR0.

symbols defined in Table 282 can be used to identify the CR bit. Table 297. Word compare simplified mnemonics Table 298. Condition register logical simplified mnemonics

  1. Same as (2), but clear CR3[SO].
  2. Same as (4), but CR4[EQ] is inverted and the result is placed into CR5[EQ].

The codes in Table 299 are for the most common combinations of trap conditions. values represented in the mnemonic rather than specified as a numeric operand. Table 299. Standard codes for trap instructions

  1. The symbol ‘<U’ indicates an unsi gned less-than evaluation is performed.
  2. The symbol ‘>U’ indicates an unsi gned greater-than evaluation is performed.
  1. Trap if rA is not equal to rB.
  2. Trap if rA is logically greater than 0x7FF .

either the sign-extended SIMM field or the contents of rB, depending on the trap instruction. not 0, the trap exception handler is invoked. See Table 301 for these conditions. Table 300. Trap simplified mnemonics Table 301. TO operand bit encoding

leaving the source or destination GPR operand, rS or rD.

  1. Copy the contents of rS to the XER.
  2. Copy the contents of the LR to rD.
  3. Copy the contents of rS to the CTR.
  4. Copy the contents of rS to CSRR0.
  5. Copy the contents of IVOR0 to rD.
  6. Copy the contents of rS to the MAS1.

address, move register, and complement register). Table 302. Additional simplified mnemonics for accessing SPRGs

RM0004 Simplified mnemonics for PowerPC instructions B.9.2 Load immediate (li) The addi and addis instructions can be used to load an immediate value into a register. Additional mnemonics are provided to convey the idea that no addition is being performed but that data is being moved from the immediate operand of the instruction to a register. 1. Load a 16-bit signed immediate value into rD. li rD,value equivalent to addi rD,0,value 2. Load a 16-bit signed immediate value, shifted left by 16 bits, into rD. lis rD,value equivalent to addis rD,0,value B.9.3 Load address (la) This mnemonic permits computing the value of a base-displacement operand, using the addi instruction that normally requires a separate register and immediate operands. la rD,d(rA) equivalent to addi rD,rA,d The la mnemonic is useful for obtaining the address of a variable specified by name, allowing the assembler to supply the base register number and compute the displacement. If the variable v is located at offset dv bytes from the address in rv, and the assembler has been told to use rv as a base for references to the data structure containing v, the following line causes the address of v to be loaded into rD: la rD,v equivalent to addi rD,rv,dv B.9.4 Move register (mr) Several instructions can be coded to copy the contents of one register to another. A simplified mnemonic is provided that signifies that no computation is being performed, but merely that data is being moved from one register to another. The following instruction copies the contents of rS into rA. This mnemonic can be coded with a dot (.) suffix to cause the Rc bit to be set in the underlying instruction. mr rA,rS equivalent to or rA,rS,rS B.9.5 Complement register (not) Several instructions can be coded in a way that they complement the contents of one register and place the result into another register. Simplified mnemonics allows this operation to be coded easily. The following instruction complements the contents of rS and places the result into rA. This mnemonic can be coded with a dot (.) suffix to cause the Rc bit to be set in the underlying instruction. not rA,rS equivalent to nor rA,rS,rS B.9.6 Move to conditi on register (mtcr) This mnemonic permits copying GPR contents to the CR, using the same syntax as the mfcr instruction. mtcr rS equivalent to mtcrf 0xFF,rS

units (APUs) defined as part of the Motorola Book E implementation standards. additional simplified mnemonics not listed here. Table 303. Simplified mnemonics

Table 303. Simplified mnemonics (continued)

  1. Simplified mnemonics for branch instructions that do not test a CR bit should not specify one; a

programming error may occur.

  1. The value in the BI operand selects CR n[2], the EQ bit.
  2. Instructions for which B0 is either 12 (branch if condition true) or 4 (branch if condition false) do not depend
  3. The value in the BI operand selects CR n[0], the LT bit.
  4. The value in the BI operand selects CR n[1], the GT bit.
  5. The value in the BI operand selects CR n[3], the SO bit.

Programming examples RM0004 Appendix C Programming examples This appendix gives examples of how memory synchronization instructions can be used to emulate various synchronization primitives and to provide more complex forms of synchronization. It also describes multiple precision shifts. C.1 Synchronization Examples in this appendix have a common form. After possible initialization, a conditional sequence begins with a load and reserve instruction that may be followed by memory accesses and computations that include neither a load and reserve nor a store conditional. The sequence ends with a store conditional with the same target address as the initial load and reserve. In most of the examples, failure of the store conditional causes a branch back to the load and reserve for a repeated attempt. On the assumption that contention is low, the conditional branch in the examples is optimized for the case in which the store conditional succeeds, by setting the branch-prediction bit appropriately. These examples focus on techniques for the correct modification of shared memory locations: see note 4 in Chapter C.1.4: Synchronization notes on page 1147,” for a discussion of how the retry strategy can affect performance. Load and reserve and store conditional instructions depend on the coherence mechanism of the system. Stores to a given location are coherent if they are serialized in some order, and no processor is able to observe a subset of those stores as occurring in a conflicting order. Each load operation, whether ordinary or load and reserve, returns a value that has a well- defined source. The source can be the store or store conditional instruction that wrote the value, an operation by some other mechanism that accesses memory (for example, an I/O device), or the initial state of memory. The function of an atomic read/modify/write operation is to read a location and write its next value, possibly as a function of its current value, all as a single atomic operation. We assume that locations accessed by read/modify/write operations are accessed coherently, so the concept of a value being the next in the sequence of values for a location is well defined. The conditional sequence, as defined above, provides the effect of an atomic read/modify/write operation, but not with a single atomic instruction. Let addr be the location that is the common target of the load and reserve and store conditional instructions. Then the guarantee the architecture makes for the successful execution of the conditional sequence is that no store into addr by another processor or mechanism has intervened between the source of the load and reserve and the store conditional. For each of these examples, it is assumed that a similar sequence of instructions is used by all processes requiring synchronization on the accessed data. Note: Because memory synchronization instructions have implementation dependencies (for example, the granularity at which reservations are managed), they must be used with care. The operating system should provide system library programs that use these instructions to implement the high-level synchronization functions (such as, test and set or compare and swap) needed by application programs. Application programs should use these library programs, rather than use memory synchronization instructions directly.

RM0004 Programming examples C.1.1 Synchronization primitives The following examples show how the lwarx and stwcx. instructions can be used to implement various synchronization primitives. The sequences used to emulate the various primitives consist primarily of a loop using lwarx and stwcx.. No additional synchronization is necessary, because the stwcx. will fail, clearing EQ, if the word loaded by lwarx has changed before the stwcx. is executed: see : Atomic update primitives using lwarx and stwcx. on page 176 for details. Fetch and No-op The fetch and no-op primitive atomically loads the current value in a word in memory. In this example it is assumed that the address of the word to be loaded is in GPR3 and the data loaded are returned in GPR4. loop: lwarx r4,0,r3 #load and reserve stwcx. r4,0,r3 #store old value if still reserved bc 4,2,loop #loop if lost reservation If the stwcx. succeeds, it stores to the target location the same value that was loaded by the preceding lwarx. While the store is redundant with respect to the value in the location, its success ensures that the value loaded by the lwarx was the current value, that is, that the source of the value loaded by the lwarx was the last store to the location that preceded the stwcx. in the coherence order for the location. Fetch and store The fetch and store primitive atomically loads and replaces a word in memory. In this example it is assumed that the address of the word to be loaded and replaced is in GPR3, the new value is in GPR4, and the old value is returned in GPR5. loop: lwarx r5,0,r3 #load and reserve stwcx. r4,0,r3 #store new value if still reserved bc 4,2,loop #loop if lost reservation Fetch and add The fetch and add primitive atomically increments a word in memory. In this example it is assumed that the address of the word to be incremented is in GPR3, the increment is in GPR4, and the old value is returned in GPR5. loop: lwarx r5,0,r3 #load and reserve add r0,r4,r5 #increment word stwcx. r0,0,r3 #store new value if still reserved bc 4,2,loop #loop if lost reservation Fetch and AND The Fetch and AND primitive atomically ANDs a value into a word in memory. In this example it is assumed that the address of the word to be ANDed is in GPR3, the value to AND into it is in GPR4, and the old value is returned in GPR5. loop: lwarx r5,0,r3 #load and reserve and r0,r4,r5 #AND word

Programming examples RM0004 stwcx. r0,0,r3 #store new value if still reserved bc 4,2,loop #loop if lost reservation This sequence can be changed to perform another Boolean operation atomically on a word in memory by changing the and to the desired Boolean instruction (or, xor, etc.). Test and set This version of the test and set primitive atomically loads a word from memory, sets the word in memory to a nonzero value if the value loaded is zero, and sets the EQ bit of CR Field 0 to indicate whether the value loaded is zero. In this example it is assumed that the address of the word to be tested is in GPR3, the new value (nonzero) is in GPR4, and the old value is returned in GPR5. loop: lwarx r5,0,r3 #load and reserve cmpwi r5,0 #done if word bc 4,2,done #not equal to 0 stwcx. r4,0,r3 #try to store non-0 bc 4,2,loop #loop if lost reservation done: Compare and swap The compare and swap primitive atomically compares a value in a register with a word in memory, if they are equal stores the value from a second register into the word in memory, if they are unequal loads the word from memory into the first register, and sets CR0[EQ] to indicate the result of the comparison. In this example it is assumed that the address of the word to be tested is in GPR3, the comparand is in GPR4 and the old value is returned there, and the new value is in GPR5. loop: lwarx r6,0,r3 #load and reserve cmpw r4,r6 #1st 2 operands equal? bc 4,2,exit #skip if not stwcx. r5,0,r3 #store new value if still reserved bc 4,2,loop #loop if lost reservation exit: or r4,r6,r6 #return value from memory Note: 1 The semantics given for compare and swap above are based on those of the IBM System/370 compare and swap instruction. Other architectures may define a compare and swap instruction differently. 2 Compare and swap is shown primarily for pedagogical reasons. It is useful on machines that lack the better synchronization facilities provided by lwarx and stwcx.. A major weakness of a System/370-style compare and swap instruction is that, although the instruction itself is atomic, it checks only that the old and current values of the word being tested are equal, with the result that programs that use such a compare and swap to control a shared resource can err if the word has been modified and the old value subsequently restored. The sequence shown above has the same weakness. 3 In some applications the second bc and/or the or can be omitted. The bc is needed only if the application requires that if CR0[EQ] on exit indicates not equal then GPR4 and GPR6 are not equal. The or is needed only if the application requires that if the comparands are not equal then the word from memory is loaded into the register with which it was compared (rather than into a third register). If any of these instructions is omitted, the resulting compare and swap does not obey System/370 semantics.

RM0004 Programming examples C.1.2 Lock acquisition and release This example gives an algorithm for locking that demonstrates the use of synchronization with an atomic read/modify/write operation. A shared memory location, the address of which is an argument of the lock and unlock procedures, given by GPR3, is used as a lock, to control access to some shared resource such as a shared data structure. The lock is open when its value is 0 and closed (locked) when its value is 1. Before accessing the shared resource the program executes the lock procedure, which sets the lock by changing its value from 0 to 1. To do this, the lock procedure calls test_and_set, which executes the code sequence shown in the test and set example of Chapter C.1.1 on page 1144,” thereby atomically loading the old value of the lock, writing to the lock the new value (1) given in GPR4, returning the old value in GPR5 (not used below), and setting the EQ bit of CR Field 0 according to whether the value loaded is 0. The lock procedure repeats the test_and_set until it succeeds in changing the value of the lock from 0 to 1. Because the shared resource must not be accessed until the lock has been set, the lock procedure contains an isync after the bc that checks for the success of test_and_set. The isync delays all subsequent instructions until all preceding instructions have completed. lock: mfspr r6,LR #save Link Register addi r4,r0,1 #obtain lock: loop: bl test_and_set# test-and-set bc 4,2,loop # retry til old = 0 # Delay subsequent instructions til prior instructions finish isync mtspr LR,r6 #restore Link Register blr #return The unlock procedure stores a 0 to the lock location. Most applications that use locking require, for correctness, that if the access to the shared resource includes stores, the program must execute an msync before releasing the lock. The msync ensures that the program’s modifications are performed with respect to other processors before the store that releases the lock is performed with respect to those processors. In this example, the unlock procedure begins with an msync for this purpose. unlock: msync #order prior stores addi r1,r0,0 #before lock release stw r1,0(r3) #store 0 to lock location blr #return C.1.3 List insertion This example shows how lwarx and stwcx. can be used to implement simple insertion into a singly linked list. (Complicated list insertion, in which multiple values must be changed atomically, or in which the correct order of insertion depends on the contents of the elements, cannot be implemented in the manner shown below and requires a more complicated strategy such as using locks.) The next element pointer from the list element after which the new element is to be inserted, here called the parent element, is stored into the new element, so that the new element points to the next element in the list: this store is performed unconditionally. Then the address of the new element is conditionally stored into the parent element, thereby adding the new element to the list. In this example it is assumed that the address of the parent element is in GPR3, the address of the new element is in GPR4, and the next element pointer is at offset 0 from the start of

Programming examples RM0004 the element. It is also assumed that the next element pointer of each list element is in a reservation granule separate from that of the next element pointer of all other list elements: see: Atomic update primitives using lwarx and stwcx. on page 176 loop: lwarx r2,0,r3 #get next pointer stw r2,0(r4) #store in new element msync #order stw before stwcx.(can omit if not MP) stwcx. r4,0,r3 #add new element to list bc 4,2,loop #loop if stwcx. failed In the preceding example, if two list elements have next element pointers in the same reservation granule then, in a multiprocessor, livelock can occur. (Livelock is a state in which processors interact in a way such that no processor makes progress.) If list elements cannot be allocated such that each element’s next element pointer is in a different reservation granule, livelock can be avoided with this more complicated sequence: lwz r2,0(r3) #get next pointer loop1: or r5,r2,r2 #keep a copy stw r2,0(r4) #store in new element msync #order stw before stwcx. loop2: lwarx r2,0,r3 #get it again cmpw r2,r5 #loop if changed (someone bc 4,2,loop1 # else progressed) stwcx. r4,0,r3 #add new element to list bc 4,2,loop #loop if failed C.1.4 Synchronization notes 1. In general, lwarx and stwcx. should be paired, with the same effective address used for both. The only exception is that an unpaired stwcx. to any (scratch) effective address can be used to clear any reservation held by the processor. 2. It is acceptable to execute a lwarx for which no stwcx. is executed. For example, this occurs in the test and set sequence shown above if the value loaded is not zero. 3. To increase the likelihood that forward progress is made, it is important that looping on lwarx/stwcx. pairs be minimized. For example, in the sequence shown above for test and set, this is achieved by testing the old value before attempting the store: were the order reversed, more stwcx. instructions might be executed, and reservations might more often be lost between the lwarx and the stwcx.. 4. The manner in which lwarx and stwcx. are communicated to other processors and mechanisms, and between levels of the memory subsystem within a given processor is implementation-dependent (see: Atomic update primitives using lwarx and stwcx. on page 176). In some implementations performance may be improved by minimizing looping on a lwarx instruction that fails to return a desired value. For example, in the test and set example shown above, to stay in the loop until the word loaded is zero, bne- $+12 can be changed to bne- loop. However, in some implementations better performance may be obtained by using an ordinary load instruction to do the initial checking of the value, as follows. loop: lwz r5,0(r3) #load the word cmpi cr0,0,r5,0 #loop back if word bc 4,2,loop # not equal to 0 lwarx r5,0,r3 #try again, reserving cmpi cr0,0,r5,0 # (likely to succeed)

RM0004 Programming examples bc 4,2,loop stwcx. r4,0,r3 #try to store non-0 bc 4,2,loop #loop if lost reservation 5. In a multiprocessor, livelock is possible if a loop containing a lwarx/stwcx. pair also contains an ordinary store instruction for which any byte of the affected memory area is in the reservation granule: see: Atomic update primitives using lwarx and stwcx. on page 176. For example, the first code sequence shown in Chapter C.1.3 on page 1146,” can cause livelock if two list elements have next element pointers in the same reservation granule. C.2 Multiple-precision shifts This section gives examples of how multiple-precision shifts can be programmed. A multiple-precision shift is defined to be a shift of an N-word quantity (32-bit implementations), where N>1. The quantity to be shifted is contained in N registers. The shift amount is specified either by an immediate value in the instruction or by a value in a register. The examples shown below distinguish between the cases N=2 and N>2. If N=2, the shift amount may be in the range 0–63, which are the maximum ranges supported by the Shift instructions used. However if N>2, the shift amount must be in the range 0–31 for the examples to yield the desired result. The specific instance shown for N>2 is N=3: extending those code sequences to larger N is straightforward, as is reducing them to the case N=2 when the more stringent restriction on shift amount is met. For shifts with immediate shift amounts only the case N=3 is shown, because the more stringent restriction on shift amount is always met. In the examples it is assumed that GPRs 2 and 3 (and 4) contain the quantity to be shifted, and that the result is to be placed into the same registers. In all cases, for both input and result, the lowest-numbered register contains the highest-order part of the data and highest- numbered register contains the lowest-order part. For non-immediate shifts, the shift amount is assumed to be in GPR6. For immediate shifts, the shift amount is assumed to be greater than 0. GPRs 0 and 31 are used as scratch registers. For N>2, the number of instructions required is 2N–1 (immediate shifts) or 3N–1 (non- immediate shifts).

Table 304. Shifts

perform various conversions. IEEE compatibility is required, or if the values being tested can be NaNs or infinities. Table 304. Shifts (continued)

Programming examples RM0004 C.3.2 Conversion from floating-point number to unsigned integer word In a 32-bit implementation The full convert to unsigned integer word function can be implemented with the sequence shown below, assuming the floating-point value to be converted is in FPR1, the value 0 is in FPR0, the value 232–1 is in FPR3, the value 231 is in FPR4, the result is returned in GPR3, and a double word at displacement ‘disp’ from the address in GPR1 can be used as scratch space. C.4 Floating point selection This section gives examples of how the optional floating select instruction (fsel) can be used to implement floating-point minimum and maximum functions, and certain simple forms of if- then-else constructions, without branching. The examples show program fragments in an imaginary, C-like, high-level programming language, and the corresponding program fragment using fsel and other Book E instructions. In the examples, a, b, x, y, and z are floating-point variables, which are assumed to be in FPRs fa, fb, fx, fy, and fz. FPR fs is assumed to be available for scratch space. Warning: Care must be taken in us ing fsel if IEEE compatibility is required, or if the values being tested can be NaNs or infinities: see Section C.4.1, “Notes.” fsel f2,f1,f1,f0 #use 0 if < 0 fsub f5,f3,f1 #use max if > max fsel f2,f5,f2,f3 fsub f5,f2,f4 #subtract 2 fcmpu cr2,f2,f4 #use diff if ≥ 231 fsel f2,f5,f5,f2 fctiw[z] f2,f2 #convert to integer stfd f2,disp(r1) #store float lwz r3,disp+4(r1) #load word bc 12,8,$+8 #add 2 31 if input xoris r3,r3,0x8000 # was ≥ 231

other use of fsel is contemplated. Table 305. Comparison to zero Table 306. Minimum and maximum Table 307. Simple if-then-else constructions

Programming examples RM0004 In these notes, the optimized program is the Book E program shown, and the unoptimized program (not shown) is the corresponding Book E program that uses fcmpu and branch conditional instructions instead of fsel. 1. The unoptimized program affects FPSCR[VXSNAN] and therefore may cause the system error handler to be invoked if the corresponding exception is enabled; the optimized program does not affect this bit. This property of the optimized program is incompatible with the IEEE standard. 2. The optimized program gives the incorrect result if a is a NaN. 3. The optimized program gives the incorrect result if a and/or b is a NaN (except that it may give the correct result in some cases for the minimum and maximum functions, depending on how those functions are defined to operate on NaNs). 4. The optimized program gives the incorrect result if a and b are infinities of the same sign. (Here it is assumed that invalid operation exceptions are disabled, in which case the result of the subtraction is a NaN. The analysis is more complicated if invalid operation exceptions are enabled, because in that case the target register of the subtraction is unchanged.) 5. The optimized program affects FPSCR[OX, UX, XX,VXISI], and therefore may cause the system error handler to be invoked if the corresponding exceptions are enabled; the unoptimized program does not affect these bits. This property of the optimized program is incompatible with the IEEE standard.

RM0004 Guidelines for 32-bit book E Appendix D Guidelines for 32-bit book E This appendix provides guidelines used by 32-bit Book E implementations; a set of guidelines is also outlined for software developers. Application software written to these guidelines can be labeled 32-bit Book E applications and can be expected to execute properly on all implementations of Book E, both 32-bit and 64-bit implementations. 32-bit Book E implementations execute applications that adhere to the software guidelines for 32-bit Book E software outlined in this appendix and are not expected to properly execute 64-bit Book E applications or any applications not adhering to these guidelines (that is, 64-bit Book E applications). D.1 Registers on 32-bit book E implementations Book E defines 32- and 64-bit registers. All 32-bit registers are supported as defined in Book E. However, except for the 64-bit FPRs, only bits 32–63 of Book E’s 64-bit registers are required to be implemented in hardware in 32-bit Book E implementation. Such 64-bit registers include LR, CTR, 32 GPRs, SRR0, and CSRR0. Book E makes no restrictions regarding implementing a subset of the 64-bit floating-point architecture. Likewise, other than floating-point instructions, all instructions defined to return a 64-bit result return only bits 32–63 of the result on a 32-bit Book E implementation. D.2 Addressing on 32-bit book E implementations Only bits 32–63 of the 64-bit Book E instruction and data memory effective addresses need to be calculated and presented to main memory, so a 32-bit implementation can bypass prepending the 32 zeros when implementing these instructions. For branch to LR and branch to CR instructions, given that LR and CTR are implemented as 32-bit registers, only 2 zeros need to be concatenated to the right of bits 32–61 of these registers to form the 32- bit branch target address. The simplest implementation of next sequential instruction address computation suggests allowing effective address computations to wrap from 0xFFFF_FFFC to 0x0000_0000. This wrapping is required of PowerPC implementations. For 32-bit Book E applications, there appears little if any benefit to allowing this wrapping behavior. Book E specifies that the situation where the computation of the next sequential instruction address after address 0xFFFF_FFFC is undefined. (Note that the next sequential instruction address after address 0xFFFF_FFFC on a 64-bit Book E implementation is 0x0000_0001_0000_0000.) D.3 TLB fields on 32-bit book E implementations 32-bit Book E implementations should support bits 32–53 of the effective page number (EPN) field in the TLB. This size provides support for a 32-bit effective address, which PowerPC ABIs may have come to expect to be available. 32-bit Book E implementations may support greater than 32-bit real addresses by supporting more than bits 32–53 of the real page number (RPN) field in the TLB.

Guidelines for 32-bit book E RM0004 D.4 32-bit book E software guidelines D.4.1 32-bit instruction selection Generally speaking, 32-bit software should avoid instructions that depend on any particular setting of bits 0–31 of any 64-bit application-accessible system register, including GPRs, for producing the correct 32-bit results. Context switching is not required to preserve the upper 32 bits of application-accessible 64-bit system registers and insertion of arbitrary settings of those upper 32 bits at arbitrary times during the execution of the 32-bit application must not affect the final result. D.4.2 32-bit addressing Book E provides a complete set of data memory access instructions that perform a modulo 32 on the computed effective address and then prepend 32 zeros to produce the full 64-bit address. Book E also provides a complete set of branch instructions that perform a modulo 32 on the computed branch target effective address and then prepend 32 zeros to produce the full 64-bit branch target address. On a 32-bit Book E implementation, these instructions are executed as defined, but without prepending the 32 zeros (only the low-order 32 bits of the address are calculated). On a 64-bit implementation, executing these instructions as defined provides the effect of restricting the application to the lowest 32-bit address space. However, there is one exception. Next sequential instruction address computations (not a taken branch) are not defined for 32-bit Book E applications when the current instruction address is 0xFFFF_FFFC. On a 32-bit Book E implementation, the instruction address could simply wrap to 0x0000_0000, providing the same effect that is required in the PowerPC Architecture. However, when the 32-bit Book E application is executed on a 64-bit Book E implementation, the next sequential instruction address calculated will be 0x0000_0001_0000_0000 and not 0x0000_0000_0000_0000. To avoid this problem the 32- bit Book E application must either avoid this situation by not allowing code to span this address boundary, or requiring a branch absolute to address 0 be placed at address 0xFFFF_FFFC to emulate the wrap. Either of these approaches allows the application to execute on 32-bit and 64-bit Book E implementations.

combinations of input operands. Flag settings are performed on appropriate element flags. For all tables in this appendix, the annotation and general rules in Table 308 apply. Table 308. Notation conventions and general rules 0x7F7_FFFFF . The encoding for double-precision is 0x7FEF_FFFF_FFFF_FFFF . 0xFF7F_FFFF . The encoding for double-precision is 0xFFEF_FFFF_FFFF_FFFF . 0x00800000. The encoding for double-precision is 0x0010_0000_0000_0000. 0x8080_0000. The encoding for double-precision is 0x8010_0000_0000_0000. a signed integer result force the result to 0x7FFF_FFFF (positive) or 0x8000_0000 (negative). (positive) or 0x0000_0000 (negative).

Table 309 lists results for add, subtract, multiply, and divide operations. Table 309. Floating-point results summary—add, sub, mul, div

Table 309. Floating-point results summ ary—add, sub, mul, div (continued)

Table 310 lists results for double- to single-precision conversion. Table 310. Floating-point results summary—single convert from double

Table 311 lists results for single- to double-precision conversion. Table 312 lists results for conversion to unsigned operations. Table 311. Floating-point results summary—double convert from single Table 312. Floating-point results summary—convert to unsigned

Table 313 lists results for conversion to signed operations. Table 314 lists results for conversion from unsigned operations. Table 315 lists results for conversion from signed operations. Table 313. Floating-point results summary—convert to signed Table 314. Floating-point results summary—convert from unsigned Table 315. Floating-point results summary—convert from signed

Table 316 lists results for *abs, *nabs, and *neg operations. **Table 316. Floating-point results summary—*abs, *nabs, *neg**

15 Glossary

The glossary contains an alphabetical list of terms, phrases, and abbreviations used in this book. Some of the terms and definitions included in the glossary are reprinted from IEEE Standard 754-1985, IEEE Standard for Binary Floating-Point Arithmetic, copyright ©1985 by the Institute of Electrical and Electronics Engineers, Inc. with the permission of the IEEE. A Architecture. A detailed specification of requirements for a processor or computer system. It does not specify details of how the processor or computer system must be implemented; instead it provides a template for a family of compatible implementations. Asynchronous interrupt. interrupts that are caused by events external to the processor’s execution. In this document, the term asynchronous interrupt is used interchangeably with the word interrupt. Atomic access. A bus access that attempts to be part of a read-write operation to the same address uninterrupted by any other access to that address (the term refers to the fact that the transactions are indivisible). The PowerPC architecture implements atomic accesses through the lwarx/stwcx. instruction pair. B Biased exponent. An exponent whose range of values is shifted by a constant (bias). Typically a bias is provided to allow a range of positive values to express a range that includes both positive and negative values. Big-endian. A byte-ordering method in memory where the address n of a word corresponds to the most- significant byte. In an addressed memory word, the bytes are ordered (left to right) 0, 1, 2, 3, with 0 being the most-significant byte. See Little-endian. Boundedly undefined. A characteristic of certain operation results that are not rigidly prescribed by the PowerPC architecture. Boundedly-undefined results for a given operation may vary among implementations and between execution attempts in the same implementation. Although the architecture does not prescribe the exact behavior for when results are allowed to be boundedly undefined, the results of executing instructions in contexts where results are allowed to be boundedly undefined are constrained to ones that could have been achieved by executing an arbitrary sequence of defined instructions, in valid form, starting in the state the machine was in before attempting to execute the given instruction. Branch prediction. The process of guessing whether a branch will be taken. Such predictions can be correct or incorrect; the term ‘predicted’ as it is used here does not imply that the prediction is correct

(successful). The PowerPC architecture defines a means for static branch prediction as part of the instruction encoding. Branch resolution. The determination of whether a branch is taken or not taken. A branch is said to be resolved when the processor can determine which instruction path to take. If the branch is resolved as predicted, the instructions following the predicted branch that may have been speculatively executed can complete (see Completion). If the branch is not resolved as predicted, instructions on the mispredicted path, and any results of speculative execution, are purged from the pipeline and fetching continues from the nonpredicted path. C Cache. High-speed memory containing recently accessed data or instructions (subset of main memory). Cache block. A small region of contiguous memory that is copied from memory into a cache. The size of a cache block may vary among processors; the maximum block size is one page. In PowerPC processors, cache coherency is maintained on a cache-block basis. Note that the term cache block is often used interchangeably with ‘cache line.’ Cache coherency. An attribute wherein an accurate and common view of memory is provided to all devices that share the same memory system. Caches are coherent if a processor performing a read from its cache is supplied with data corresponding to the most recent value written to memory or to another processor’s cache. Cache flush. An operation that removes from a cache any data from a specified address range. This operation ensures that any modified data within the specified address range is written back to main memory. This operation is generated typically by a Data Cache Block Flush (dcbf) instruction. Caching-inhibited. A memory update policy in which the cache is bypassed and the load or store is performed to or from main memory. Cast out. A cache block that must be written to memory when a cache miss causes a cache block to be replaced. Changed bit. One of two page history bits found in each page table entry (PTE). The processor sets the changed bit if any store is performed into the page. See also Page access history bits and Referenced bit. Clean. An operation that causes a cache block to be written to memory, if modified, and then left in a valid, unmodified state in the cache. Clear. To cause a bit or bit field to register a value of zero. See also Set. Completion. Completion occurs when an instruction has finished executing, written back any results, and is removed from the completion queue (CQ). When an instruction completes, it is guaranteed that this instruction and all previous instructions can cause no interrupts. Context synchronization. An operation that ensures that all instructions in execution complete past the point where they can produce an interrupt, that all instructions in execution complete in the context in which they began execution, and that all subsequent instructions are fetched and executed in the new context. Context synchronization may result from executing specific instructions (such as isync or rfi) or when certain events occur (such as an interrupt).

D Denormalized number. A nonzero floating-point number whose exponent has a reserved value, usually the format's minimum, and whose explicit or implicit leading significand bit is zero. E Effective address (EA). The 32-bit address specified for a load, store, or an instruction fetch. This address is then submitted to the MMU for translation to either a physical memory address or an I/O address. Exception. A condition that, if enabled, generates an interrupt. Execution synchronization. A mechanism by which all instructions in execution are architecturally complete before beginning execution (appearing to begin execution) of the next instruction. Similar to context synchronization but doesn't force the contents of the instruction buffers to be deleted and refetched. Exponent. In the binary representation of a floating-point number, the exponent is the component that normally signifies the integer power to which the value two is raised in determining the value of the represented number. See also Biased exponent. F Fetch. Instruction retrieval from either the cache or main memory and placing them into the instruction queue. Finish. Finishing occurs in the last cycle of execution. In this cycle, the CQ entry is updated to indicate that the instruction has finished executing. Floating-point register (FPR). Any of the 32 registers in the floating-point register file. These registers provide the source operands and destination results for floating-point instructions. Load instructions move data from memory to FPRs and store instructions move data from FPRs to memory. The FPRs are 64 bits wide and store floating-point values in double-precision format. Floating-point unit. The functional unit in a processor responsible for executing all floating- point instructions. Flush. An operation that causes a cache block to be invalidated and the data, if modified, to be written to memory. Fraction. In the binary representation of a floating-point number, the field of the significand that lies to the right of its implied binary point. G General-purpose register (GPR). Any of the 32 registers in the general-purpose register file. These registers provide the source operands and destination results for all integer data manipulation instructions. Integer load instructions move data from memory to GPRs and store instructions move data from GPRs to memory. Guarded. The guarded attribute pertains to out-of-order execution. When a page is designated as guarded, instructions and data cannot be accessed out-of-order.

H Harvard architecture. An architectural model featuring separate caches and other memory management resources for instructions and data. I IEEE 754. A standard written by the Institute of Electrical and Electronics Engineers that defines operations and representations of binary floating-point numbers. Illegal instructions. A class of instructions that are not implemented for a particular PowerPC processor. These include instructions not defined by the PowerPC architecture. In addition, for 32-bit implementations, instructions that are defined only for 64-bit implementations are considered to be illegal instructions. For 64-bit implementations instructions that are defined only for 32-bit implementations are considered to be illegal instructions. Implementation. A particular processor that conforms to the PowerPC architecture, but may differ from other architecture-compliant implementations for example in design, feature set, and implementation of optional features. The PowerPC architecture has many different implementations. Imprecise interrupt. A type of synchronous interrupt that is allowed not to adhere to the precise interrupt model (see Precise interrupt). The PowerPC architecture allows only floating-point exceptions to be handled imprecisely. Integer unit. The functional unit responsible for executing all integer instructions. In order. An aspect of an operation that adheres to a sequential model. An operation is said to be performed in-order if, at the time that it is performed, it is known to be required by the sequential execution model. See Out-of-order. Instruction latency. The total number of clock cycles necessary to execute an instruction and make ready the results of that instruction. Interrupt. A condition encountered by the processor that requires special, supervisor-level processing. Interrupt handler. A software routine that executes when an interrupt is taken. Normally, the interrupt handler corrects the condition that caused the interrupt, or performs some other meaningful task (that may include aborting the program that caused the interrupt).

K Kill. An operation that causes a cache block to be invalidated without writing any modified data to memory. L Latency. The number of clock cycles necessary to execute an instruction and make ready the results of that execution for a subsequent instruction. L2 cache. See Secondary cache. Least-significant bit (lsb). The bit of least value in an address, register, field, data element, or instruction encoding. Least-significant byte (LSB). The byte of least value in an address, register, data element, or instruction encoding. Little-endian. A byte-ordering method in memory where the address n of a word corresponds to the least- significant byte. In an addressed memory word, the bytes are ordered (left to right) 3, 2, 1, 0, with 3 being the most-significant byte. See Big-endian. M Mantissa. The decimal part of logarithm. Memory access ordering. The specific order in which the processor performs load and store memory accesses and the order in which those accesses complete. Memory-mapped accesses. Accesses whose addresses use the page or block address translation mechanisms provided by the MMU and that occur externally with the bus protocol defined for memory. Memory coherency. An aspect of caching in which it is ensured that an accurate view of memory is provided to all devices that share system memory. Memory consistency. Refers to agreement of levels of memory with respect to a single processor and system memory (for example, on-chip cache, secondary cache, and system memory). Memory management unit (MMU). The functional unit that is capable of translating an effective (logical) address to a physical address, providing protection mechanisms, and defining caching methods. Most-significant bit (msb).

The highest-order bit in an address, registers, data element, or instruction encoding. Most-significant byte (MSB). The highest-order byte in an address, registers, data element, or instruction encoding. N NaN. An abbreviation for not a number; a symbolic entity encoded in floating-point format. There are two types of NaNs—signa ling NaNs and quiet NaNs. No-op. No-operation. A single-cycle operation that does not affect registers or generate bus activity. Normalization. A process by which a floating-point value is manipulated such that it can be represented in the format for the appropriate precision (single- or double-precision). For a floating-point value to be representable in the single- or double-precision format, the leading implied bit must be a 1. O OEA (operating environment architecture). The level of the architecture that describes PowerPC memory management model, supervisor-level registers, synchronization requirements, and the interrupt model. It also defines the time-base feature from a supervisor-level perspective. Implementations that conform to the PowerPC OEA also conform to the PowerPC UISA and VEA. Optional. A feature, such as an instruction, a register, or an interrupt, that is defined by the PowerPC architecture but not required to be implemented. Out-of-order. An aspect of an operation that allows it to be performed ahead of one that may have preceded it in the sequential model, for example, speculative operations. An operation is said to be performed out-of-order if, at the time that it is performed, it is not known to be required by the sequential execution model. See In-order. Out-of-order execution. A technique that allows instructions to be issued and completed in an order that differs from their sequence in the instruction stream. Overflow. An condition that occurs during arithmetic operations when the result cannot be stored accurately in the destination register(s). For example, if two 32-bit numbers are multiplied, the result may not be representable in 32 bits. Since 32-bit registers cannot represent this sum, an overflow condition occurs.

P Page. A region in memory. The OEA defines a page as a 4-Kbyte area of memory, aligned on a 4- Kbyte boundary. Page fault. A page fault is a condition that occurs when the processor attempts to access a memory location that does not reside within a page not currently resident in physical memory. On PowerPC processors, a page fault interrupt condition occurs when a matching, valid page table entry (PTE[V] = 1) cannot be located. Physical memory. The actual memory that can be accessed through the system’s memory bus. Pipelining. A technique that breaks operations, such as instruction processing or bus transactions, into smaller distinct stages or tenures (respectively) so that a subsequent operation can begin before the previous one has completed. Precise interrupts. A category of interrupt for which the pipeline can be stopped so instructions that preceded the faulting instruction can complete and subsequent instructions can be flushed and redispatched after interrupt handling has completed. See Imprecise interrupts. Primary opcode. The most-significant 6 bits (bits 0–5) of the instruction encoding that identifies the type of instruction. Program order. The order of instructions in an executing program. More specifically, this term is used to refer to the original order in which program instructions are fetched into the instruction queue from the cache. Protection boundary. A boundary between protection domains. Q Quiet NaN. A type of NaN that can propagate through most arithmetic operations without signaling interrupts. A quiet NaN is used to represent the results of certain invalid operations, such as invalid arithmetic operations on infinities or on NaNs, when invalid. See Signaling NaN. R Record bit. Bit 31 (or the Rc bit) in the instruction encoding. When it is set, updates the condition register (CR) to reflect the result of the operation. Referenced bit.

One of two page history bits found in each page table entry. The processor sets the referenced bit whenever the page is accessed for a read or write. See also Page access history bits. Register indirect addressing. A form of addressing that specifies one GPR that contains the address for the load or store. Register indirect with immediate index addressing. A form of addressing that specifies an immediate value to be added to the contents of a specified GPR to form the target address for the load or store. Register indirect with index addressing. A form of addressing that specifies that the contents of two GPRs be added together to yield the target address for the load or store. Rename register. Temporary buffers used by instructions that have finished execution but have not completed. Reservation. The processor establishes a reservation on a cache block of memory space when it executes an lwarx instruction to read a memory semaphore into a GPR. Reservation station. A buffer between the dispatch and execute stages that allows instructions to be dispatched even though the results of instructions on which the dispatched instruction may depend are not available. RISC (reduced instruction set computing). An architecture characterized by fixed-length instructions with nonoverlapping functionality and by a separate set of load and store instructions that perform memory accesses. S Secondary cache. A cache memory that is typically larger and has a longer access time than the primary cache. A secondary cache may be shared by multiple devices. Also referred to as L2, or level-2, cache. Set (v). To write a nonzero value to a bit or bit field; the opposite of clear. The term ‘set’ may also be used to generally describe the updating of a bit or bit field. Set (n). A subdivision of a cache. Cacheable data can be stored in a given location in one of the sets, typically corresponding to its lower-order address bits. Because several memory locations can map to the same location, cached data is typically placed in the set whose cache block corresponding to that address was used least recently. See Set-associative. Set-associative. Aspect of cache organization in which the cache space is divided into sections, called sets. The cache controller associates a particular main memory address with the contents of a particular set, or region, within the cache.

Signaling NaN. A type of NaN that generates an invalid operation program interrupt when it is specified as arithmetic operands. See Quiet NaN. Significand. The component of a binary floating-point number that consists of an explicit or implicit leading bit to the left of its implied binary point and a fraction field to the right. Simplified mnemonics. Assembler mnemonics that represent a more complex form of a common operation. Snooping. Monitoring addresses driven by a bus master to detect the need for coherency actions. Split-transaction. A transaction with independent request and response tenures. Stall. An occurrence when an instruction cannot proceed to the next stage. Static branch prediction. Mechanism by which software (for example, compilers) can hint to the machine hardware about the direction a branch is likely to take. Superscalar. A superscalar processor is one that can dispatch multiple instructions concurrently from a conventional linear instruction stream. In a superscalar implementation, multiple instructions can be in the same stage at the same time. Supervisor mode. The privileged operation state of a processor. In supervisor mode, software, typically the operating system, can access all control registers and can access the supervisor memory space, among other privileged operations. Synchronization. A process to ensure that operations occur strictly in order. See Context synchronization and Execution synchronization. Synchronous interrupt. An interrupt that is generated by the execution of a particular instruction or instruction sequence. There are two types of synchronous interrupts, precise and imprecise. System memory. The physical memory available to a processor.

T TLB (translation lookaside buffer). A cache that holds recently-used page table entries. Throughput. The measure of the number of instructions that are processed per clock cycle. U UISA (user instruction set architecture). The level of the architecture to which user-level software should conform. The UISA defines the base user-level instruction set, user-level registers, data types, floating-point memory conventions and interrupt model as seen by user programs, and the memory and programming models. Underflow. A condition that occurs during arithmetic operations when the result cannot be represented accurately in the destination register. For example, underflow can happen if two floating- point fractions are multiplied and the result requires a smaller exponent and/or mantissa than the single-precision format can provide. In other words, the result is too small to be represented accurately. User mode. The operating state of a processor used typically by application software. In user mode, software can access only certain control registers and can access only user memory space. No privileged operations can be performed. Also referred to as problem state. V VEA (virtual environment architecture). The level of the architecture that describes the memory model for an environment in which multiple devices can access memory, defines aspects of the cache model, defines cache control instructions, and defines the time-base facility from a user-level perspective. Implementations that conform to the PowerPC VEA also adhere to the UISA, but may not necessarily adhere to the OEA. Virtual address. An intermediate address used in the translation of an effective address to a physical address. Virtual memory. The address space created using the memory management facilities of the processor. Program access to virtual memory is possible only when it coinc

W Way. A location in the cache that holds a cache block, its tags and status bits. Word. A 32-bit data element. Write-back. A cache memory update policy in which processor write cycles are directly written only to the cache. External memory is updated only indirectly, for example, when a modified cache block is cast out to make room for newer data. Write-through. A cache memory update policy in which all processor write cycles are written to both the cache and memory.

Table 317. Document revision history 29-Nov-2007 1 Initial release.