The following counters can be examined in PIX.
| Counter | Value | Description |
|---|---|---|
| SRBM_PERF_SEL_COUNT | 0 | Tie High - Count Number of Clocks. |
| SRBM_PERF_SEL_BIF_BUSY | 1 | The Bus Interface (BIF) block is busy. |
| SRBM_PERF_SEL_DRM_BUSY | 2 | The Digital Rights Management (DRM) block is busy. |
| SRBM_PERF_SEL_SDMA0_BUSY | 3 | The System DMA (SDMA0) block is busy. |
| SRBM_PERF_SEL_IH_BUSY | 4 | The Interrupt Handler (IH) block is busy. |
| SRBM_PERF_SEL_MCB_BUSY | 5 | Any of the Memory Controller (MCB) blocks are busy. |
| SRBM_PERF_SEL_MCB_NON_DISPLAY_BUSY | 6 | Any of the Memory Controller (MCB) blocks are busy with non display traffic. |
| SRBM_PERF_SEL_MCC_BUSY | 7 | Any of the Memory Controller (MCC) blocks are busy. |
| SRBM_PERF_SEL_MCD_BUSY | 8 | Any of the Memory Controller (MCD) blocks are busy. |
| SRBM_PERF_SEL_RESERVED | 9 | Reserved |
| SRBM_PERF_SEL_SEM_BUSY | 10 | The Semaphore (SEM) block is busy. |
| SRBM_PERF_SEL_UVD_BUSY | 11 | The Universal Video Decoder (UVD) block is busy |
| SRBM_PERF_SEL_VMC_BUSY | 12 | The Virtual Memory Controller (VMC) block is busy. |
| SRBM_PERF_SEL_XSP_BUSY | 13 | The Crossbar Switch (XSP) block is busy. |
| SRBM_PERF_SEL_SDMA1_BUSY | 14 | The System DMA (SDMA1) block is busy. |
| SRBM_PERF_SEL_SAMMSP_BUSY | 15 | The Secure Access Management Media Security Processor (SAMMSP) block is busy. |
| SRBM_PERF_SEL_VCE_BUSY | 16 | The Video Compression Engine (VCE) block is busy. |
| SRBM_PERF_SEL_XDMA_BUSY | 17 | The Multi-GPU DMA (XDMA) block is busy. |
| SRBM_PERF_SEL_ACP_BUSY | 18 | The Audio CoProcessor (ACP) block is busy. |
| SRBM_PERF_SEL_SDMA2_BUSY | 19 | The System DMA (SDMA2) block is busy. |
| SRBM_PERF_SEL_SDMA3_BUSY | 20 | The System DMA (SDMA3) block is busy. |
| SRBM_PERF_SEL_SAMSCP_BUSY | 21 | The Secure Access Management Streaming Crypto Processor (SAMSCP) block is busy. |
| SRBM_PERF_SEL_VMC1_BUSY | 22 | The Virtual Memory Controller1 (VMC1) block is busy. |
| SRBM_PERF_SEL_SCORPIO_START | 22 | The start of Xbox One X-specific counters. |
| SRBM_PERF_SEL_VCNX_BUSY | 22 | The Video Codex Next (VCNX) block is busy. |
| Counter | Value | Description |
|---|---|---|
| CPF_PERF_SEL_ALWAYS_COUNT | 0 | Always Count. |
| CPF_PERF_SEL_MIU_STALLED_WAITING_RDREQ_FREE | 1 | CPF Miu stalled waiting on Rdreq free. |
| CPF_PERF_SEL_TCIU_STALLED_WAITING_ON_FREE | 2 | CPF Tciu stalled waiting on Free. |
| CPF_PERF_SEL_TCIU_STALLED_WAITING_ON_TAGS | 3 | CPF Tciu stalled waiting on tags. |
| CPF_PERF_SEL_CSF_BUSY_FOR_FETCHING_RING | 4 | CSF busy for fetching Ring data. |
| CPF_PERF_SEL_CSF_BUSY_FOR_FETCHING_IB1 | 5 | CSF busy for fetching Indirect buffer 1. |
| CPF_PERF_SEL_CSF_BUSY_FOR_FETCHING_IB2 | 6 | CSF busy for fetching Indirect buffer 2. |
| CPF_PERF_SEL_CSF_BUSY_FOR_FECTHINC_STATE | 7 | CSF busy for fetching state data. |
| CPF_PERF_SEL_MIU_BUSY_FOR_OUTSTANDING_TAGS | 8 | Miu is busy for outstanding tags. |
| CPF_PERF_SEL_CSF_RTS_MIU_NOT_RTR | 9 | RTS sent from CSF but no RTR. |
| CPF_PERF_SEL_CSF_STATE_FIFO_NOT_RTR | 10 | CSF state FIFO no RTR. |
| CPF_PERF_SEL_CSF_FETCHING_CMD_BUFFERS | 11 | CSF is fetching command buffers. |
| CPF_PERF_SEL_GRBM_DWORDS_SENT | 12 | CPF to GRBM dwords sent. |
| CPF_PERF_SEL_DYNAMIC_CLOCK_VALID | 13 | CPF Dynamic clock is valid. |
| CPF_PERF_SEL_REGISTER_CLOCK_VALID | 14 | CPF register clock is valid. |
| CPF_PERF_SEL_MIU_WRITE_REQUEST_SEND | 15 | CPF Miu write request send. |
| CPF_PERF_SEL_MIU_READ_REQUEST_SEND | 16 | CPF Miu read request send. |
| Counter | Value | Description |
|---|---|---|
| CPG_PERF_SEL_ALWAYS_COUNT | 0 | Always Count. |
| CPG_PERF_SEL_RBIU_FIFO_FULL | 1 | RBIU Transaction FIFO Full. |
| CPG_PERF_SEL_CSF_RTS_BUT_MIU_NOT_RTR | 2 | CSF is Ready to Send data but the MIU is not Ready to Receive it. |
| CPG_PERF_SEL_CSF_ST_BASE_SIZE_FIFO_FULL | 3 | PFP to CSF State Request FIFO is Full. |
| CPG_PERF_SEL_CP_GRBM_DWORDS_SENT | 4 | Count of DWs actually sent to the GRBM. |
| CPG_PERF_SEL_ME_PARSER_BUSY | 5 | Count of MicroEngine Busy Clocks. |
| CPG_PERF_SEL_COUNT_TYPE0_PACKETS | 6 | Count of Type0 Packets processed in NRT. |
| CPG_PERF_SEL_COUNT_TYPE3_PACKETS | 7 | Count of Type3 Packets processed in NRT. |
| CPG_PERF_SEL_CSF_FETCHING_CMD_BUFFERS | 8 | Count of clocks that the RB, I1, I2, C1 & C2 are fetching data. |
| CPG_PERF_SEL_CP_GRBM_OUT_OF_CREDITS | 9 | CP to GRBM path has data to send but is waiting for credits (free signals). |
| CPG_PERF_SEL_CP_PFP_GRBM_OUT_OF_CREDITS | 10 | CP.PFP to GRBM path has data to send but is waiting for credits (free signals). |
| CPG_PERF_SEL_CP_GDS_GRBM_OUT_OF_CREDITS | 11 | CP to GRBM path to GDS has data to send but is waiting for credits (free signals). |
| CPG_PERF_SEL_RCIU_STALLED_ON_ME_READ | 12 | RCIU is stalled waiting for read data to be returned for the MicroEngine. |
| CPG_PERF_SEL_RCIU_STALLED_ON_DMA_READ | 13 | RCIU is stalled waiting for read data to be returned for the DMA. |
| CPG_PERF_SEL_SSU_STALLED_ON_ACTIVE_CNTX | 14 | Surface Sync Unit is waiting on a matching base in an Active Context. |
| CPG_PERF_SEL_SSU_STALLED_ON_CLEAN_SIGNALS | 15 | Surface Sync Unit is waiting on all Clean signals to return. |
| CPG_PERF_SEL_QU_STALLED_ON_EOP_DONE_PULSE | 16 | Query Unit is stalled waiting for an EOP Done signal from the SC or SPI. |
| CPG_PERF_SEL_QU_STALLED_ON_EOP_DONE_WR_CONFIRM | 17 | Query Unit is stalled waiting for a write confirm on EOP Done data. |
| CPG_PERF_SEL_PFP_STALLED_ON_CSF_READY | 18 | PFP is stalled trying to write to the CSF. |
| CPG_PERF_SEL_PFP_STALLED_ON_MEQ_READY | 19 | PFP is stalled trying to write to the MEQ. |
| CPG_PERF_SEL_PFP_STALLED_ON_RCIU_READY | 20 | PFP is stalled trying to write to the RCIU |
| CPG_PERF_SEL_PFP_STALLED_FOR_DATA_FROM_ROQ | 21 | PFP is stalled waiting for data from the ROQ. |
| CPG_PERF_SEL_ME_STALLED_FOR_DATA_FROM_PFP | 22 | MicroEngine is stalled waiting for data from the PFP. |
| CPG_PERF_SEL_ME_STALLED_FOR_DATA_FROM_STQ | 23 | MicroEngine is stalled waiting for data from the State Queue. |
| CPG_PERF_SEL_ME_STALLED_ON_NO_AVAIL_GFX_CNTX | 24 | ME is stalled waiting for an available GFX context. |
| CPG_PERF_SEL_ME_STALLED_WRITING_TO_RCIU | 25 | ME is stalled trying to write to the RCIU (GRBM). |
| CPG_PERF_SEL_ME_STALLED_WRITING_CONSTANTS | 26 | ME is stalled trying to write constants to the RCIU/MIU (GRBM/MC). |
| CPG_PERF_SEL_ME_STALLED_ON_PARTIAL_FLUSH | 27 | The ME sent out a Partial Flush event and is waiting for a response from the SPI. |
| CPG_PERF_SEL_ME_WAIT_ON_CE_COUNTER | 28 | The ME is waiting on the CE to write the constants to memory (and increment the up/down counter). |
| CPG_PERF_SEL_ME_WAIT_ON_AVAIL_BUFFER | 29 | The ME is waiting on the next buffer to be available - it is still in use. |
| CPG_PERF_SEL_SEMAPHORE_BUSY_POLLING_FOR_PASS | 30 | Semaphore Unit is busy polling for a Pass response. |
| CPG_PERF_SEL_LOAD_STALLED_ON_SET_COHERENCY | 31 | LOAD packet is stalled waiting on Coherency Counter=0 (SET packet writes completed). |
| CPG_PERF_SEL_DYNAMIC_CLK_VALID | 32 | Input from CGTT_LOCAL output: dyn_oclk_vld |
| CPG_PERF_SEL_REGISTER_CLK_VALID | 33 | Input from CGTT_LOCAL output: reg_oclk_vld |
| CPG_PERF_SEL_MIU_WRITE_REQUEST_SENT | 34 | Count of MC write requests. |
| CPG_PERF_SEL_MIU_READ_REQUEST_SENT | 35 | Count of MC read requests. |
| CPG_PERF_SEL_CE_STALL_RAM_DUMP | 36 | CE is stalled while trying to issue a RAM DUMP. |
| CPG_PERF_SEL_CE_STALL_RAM_WRITE | 37 | CE is stalled while trying to issue a RAM WRITE. |
| CPG_PERF_SEL_CE_STALL_ON_INC_FIFO | 38 | CE is stalled while trying to write to the increment FIFO. |
| CPG_PERF_SEL_CE_STALL_ON_WR_RAM_FIFO | 39 | CE is stalled while trying to write to the write ram FIFO. |
| CPG_PERF_SEL_CE_STALL_ON_DATA_FROM_MIU | 40 | CE is stalled waiting on data from the MIU. |
| CPG_PERF_SEL_CE_STALL_ON_DATA_FROM_ROQ | 41 | CE is stalled waiting on data from the ROQ. |
| CPG_PERF_SEL_CE_STALL_ON_CE_BUFFER_FLAG | 42 | CE is stalled waiting for an available buffer. |
| CPG_PERF_SEL_CE_STALL_ON_DE_COUNTER | 43 | CE is stalled waiting for the DE counter. |
| CPG_PERF_SEL_TCIU_STALL_WAIT_ON_FREE | 44 | TCIU is stalled waiting on credits from the TC. |
| CPG_PERF_SEL_TCIU_STALL_WAIT_ON_TAGS | 45 | TCIU has used all available tags and is waiting on the TC to return them. |
| CPG_PERF_SEL_SCORPIO_START | 46 | Start of Xbox One X-specific counters. |
| CPG_PERF_SEL_TCIU_WRITE_REQUEST_SENT | 46 | Count of TC write requests sent. |
| CPG_PERF_SEL_CPG_STAT_BUSY | 47 | CPG is busy. |
| CPG_PERF_SEL_CPG_STAT_IDLE | 48 | CPG is idle. |
| CPG_PERF_SEL_CPG_STAT_STALL | 49 | CPG is stalled. |
| CPG_PERF_SEL_CPG_TCIU_BUSY | 50 | CPG TCIU interface is busy. |
| CPG_PERF_SEL_CPG_TCIU_IDLE | 51 | CPG TCIU interface is busy. |
| CPG_PERF_SEL_CPG_TCIU_STALL | 52 | CPG TCIU interface is stalled waiting on Free, Tags. |
| Counter | Value | Description |
|---|---|---|
| CPC_PERF_SEL_ALWAYS_COUNT | 0 | Always Count. |
| CPC_PERF_SEL_RCIU_STALL_WAIT_ON_FREE | 1 | Rciu Stall Waiting on Free. |
| CPC_PERF_SEL_RCIU_STALL_PRIV_VIOLATION | 2 | Rciu Stall on Priv violation. |
| CPC_PERF_SEL_MIU_STALL_ON_RDREQ_FREE | 3 | Miu stall On read req free. |
| CPC_PERF_SEL_MIU_STALL_ON_WRREQ_FREE | 4 | Miu Stall on wrreq free. |
| CPC_PERF_SEL_TCIU_STALL_WAIT_ON_FREE | 5 | Tciu Stall waiting for free. |
| CPC_PERF_SEL_ME1_STALL_WAIT_ON_RCIU_READY | 6 | Me1 stall waiting on Rciu Ready. |
| CPC_PERF_SEL_ME1_STALL_WAIT_ON_RCIU_READY_PERF | 7 | Me1 stall waiting on Rciu Ready Perf. |
| CPC_PERF_SEL_ME1_STALL_WAIT_ON_RCIU_READ | 8 | Me1 stall waiting on Rciu read. |
| CPC_PERF_SEL_ME1_STALL_WAIT_ON_MIU_READ | 9 | Me1 stall waiting on miu read. |
| CPC_PERF_SEL_ME1_STALL_WAIT_ON_MIU_WRITE | 10 | Me1 stall waiting on miu wrute. |
| CPC_PERF_SEL_ME1_STALL_ON_DATA_FROM_ROQ | 11 | Me1 stall on data from roq. |
| CPC_PERF_SEL_ME1_STALL_ON_DATA_FROM_ROQ_PERF | 12 | Me1 stall on data from roq Perf. |
| CPC_PERF_SEL_ME1_BUSY_FOR_PACKET_DECODE | 13 | Me1 busy for packet decode. |
| CPC_PERF_SEL_ME2_STALL_WAIT_ON_RCIU_READY | 14 | Me2 stall waiting on Rciu Ready. (Valid Only if ME2 is available) |
| CPC_PERF_SEL_ME2_STALL_WAIT_ON_RCIU_READY_PERF | 15 | Me2 stall waiting on Rciu Ready Perf. (Valid Only if ME2 is available) |
| CPC_PERF_SEL_ME2_STALL_WAIT_ON_RCIU_READ | 16 | Me2 stall waiting on Rciu read. (Valid Only if ME2 is available) |
| CPC_PERF_SEL_ME2_STALL_WAIT_ON_MIU_READ | 17 | Me2 stall waiting on miu read. (Valid Only if ME2 is available) |
| CPC_PERF_SEL_ME2_STALL_WAIT_ON_MIU_WRITE | 18 | Me2 stall waiting on miu wrute. (Valid Only if ME2 is available) |
| CPC_PERF_SEL_ME2_STALL_ON_DATA_FROM_ROQ | 19 | Me2 stall on data from roq. (Valid Only if ME2 is available) |
| CPC_PERF_SEL_ME2_STALL_ON_DATA_FROM_ROQ_PERF | 20 | Me2 stall on data from roq Perf. (Valid Only if ME2 is available) |
| CPC_PERF_SEL_ME2_BUSY_FOR_PACKET_DECODE | 21 | Me2 busy for packet decode. (Valid Only if ME2 is available) |
| CPC_PERF_SEL_SCORPIO_START | 22 | Start of Xbox One X-specific counters. |
| CPC_PERF_SEL_CPC_STAT_BUSY | 22 | CPC is busy. |
| CPC_PERF_SEL_CPC_STAT_IDEL | 23 | CPC is idle. |
| CPC_PERF_SEL_CPC_STAT_STALL | 24 | CPC is stalled. |
| CPC_PERF_SEL_CPC_TCIU_BUSY | 25 | CPC TCIU interface is busy. |
| CPC_PERF_SEL_CPC_TCIU_IDLE | 26 | CPC TCIU interface is idle. |
| CPC_PERF_SEL_ME1_DC0_SPI_BUSY | 27 | Me1 processor is busy. |
| CPC_PERF_SEL_ME2_DC1_SPI_BUSY | 28 | Me2 processor is busy. (Valid only if ME2 is available) |
| Counter | Value | Description |
|---|---|---|
| CB_PERF_SEL_NONE | 0 | Count nothing |
| CB_PERF_SEL_BUSY | 1 | Number of busy cycles |
| CB_PERF_SEL_CORE_SCLK_VLD | 2 | Number of cycles that the core clock is enabled. |
| CB_PERF_SEL_REG_SCLK0_VLD | 3 | Number of cycles that the register clock is enabled. (Non-harvestable) |
| CB_PERF_SEL_REG_SCLK1_VLD | 4 | Number of cycles that the register clock is enabled. (harvestable) |
| CB_PERF_SEL_DRAWN_QUAD | 5 | This is the number of drawn quads. Use CB_PERF_SEL_DRAWN_BUSY as denominator to get per clock rates. Filtering using CB_PERFCOUNTER_FILTER fields has an effect in this mode. |
| CB_PERF_SEL_DRAWN_PIXEL | 6 | This is the number of drawn pixels. Use CB_PERF_SEL_DRAWN_BUSY as denominator to get per clock rates. Filtering using CB_PERFCOUNTER_FILTER fields has an effect in this mode. |
| CB_PERF_SEL_DRAWN_QUAD_FRAGMENT | 7 | This is the number of drawn quad fragments. Use CB_PERF_SEL_DRAWN_BUSY as denominator to get per clock rates. Filtering using CB_PERFCOUNTER_FILTER fields has an effect in this mode. |
| CB_PERF_SEL_DRAWN_TILE | 8 | This is the number of drawn tiles. Use CB_PERF_SEL_DRAWN_BUSY as denominator to get per clock rates. Filtering using CB_PERFCOUNTER_FILTER fields has an effect in this mode. This counter is slightly broken in that it actually counts the number of tiles reaching fragop. Some of these tiles may be empty when there is fast clear or compression. |
| CB_PERF_SEL_DB_CB_TILE_VALID_READY | 9 | Number of cycles the DB to CB tile interface is valid and ready. This is measured after the input FIFO. It counts tiles + events. Subtract CB_PERF_SEL_EVENT to get the count of tiles |
| CB_PERF_SEL_DB_CB_TILE_VALID_READYB | 10 | Number of cycles the DB to CB tile interface is valid and not ready. This is measured after the input FIFO. |
| CB_PERF_SEL_DB_CB_TILE_VALIDB_READY | 11 | Number of cycles the DB to CB tile interface is not valid and ready. This is measured after the input FIFO. |
| CB_PERF_SEL_DB_CB_TILE_VALIDB_READYB | 12 | Number of cycles the DB to CB tile interface is not valid and not ready. This is measured after the input FIFO. |
| CB_PERF_SEL_CM_FC_TILE_VALID_READY | 13 | Number of cycles the cmask to fmask tile interface is valid and ready. This interface counts tile times enabled render targets plus events. This will be different from CB_PERF_SEL_DB_CB_TILE_VALID_READY because each tile is replicated by the number of enabled render targets (events are not) |
| CB_PERF_SEL_CM_FC_TILE_VALID_READYB | 14 | Number of cycles the cmask to fmask tile interface is valid and not ready |
| CB_PERF_SEL_CM_FC_TILE_VALIDB_READY | 15 | Number of cycles the cmask to fmask tile interface is not valid and ready |
| CB_PERF_SEL_CM_FC_TILE_VALIDB_READYB | 16 | Number of cycles the cmask to fmask tile interface is not valid and not ready |
| CB_PERF_SEL_MERGE_TILE_ONLY_VALID_READY | 17 | Number of cycles the Merge tile interface has a valid tile-only tile and is ready. Measured at tile input to merge. |
| CB_PERF_SEL_MERGE_TILE_ONLY_VALID_READYB | 18 | Number of cycles the Merge tile interface has a valid tile-only tile and is not ready. Measured at tile input to merge. |
| CB_PERF_SEL_DB_CB_LQUAD_VALID_READY | 19 | Number of cycles the DB to CB lquad interface is valid and ready. This is measured after the input FIFO. This is counting lquad which for fast modes are the same as quads but for slow mode are twice the number of quads. To get the number of quads on this interface use: CB_PERF_SEL_DB_CB_LQUAD_VALID_READY - CB_PERF_SEL_LQUAD_FORMAT_IS_EXPORT_32_ABGR/2 |
| CB_PERF_SEL_DB_CB_LQUAD_VALID_READYB | 20 | Number of cycles the DB to CB lquad interface is valid and not ready. This is measured after the input FIFO |
| CB_PERF_SEL_DB_CB_LQUAD_VALIDB_READY | 21 | Number of cycles the DB to CB lquad interface is not valid and ready. This is measured after the input FIFO |
| CB_PERF_SEL_DB_CB_LQUAD_VALIDB_READYB | 22 | Number of cycles the DB to CB lquad interface is not valid and not ready. This is measured after the input FIFO |
| CB_PERF_SEL_LQUAD_NO_TILE | 23 | Number of cycles that a quad has arrived over the lquad interface but the corresponding tile has not arrived yet. |
| CB_PERF_SEL_LQUAD_FORMAT_IS_EXPORT_32_R | 24 | Number of transactions/quads on the DB_CB_lquad interface that use the EXPORT_32R (FAST MODE) format. Sum of all CB_PERF_SEL_LQUAD_FORMAT_IS_EXPORT* should match CB_PERF_SEL_DB_CB_LQUAD_VALID_READY |
| CB_PERF_SEL_LQUAD_FORMAT_IS_EXPORT_32_AR | 25 | Number of transactions/quads on the DB_CB_lquad interface that use the EXPORT_32AR (FAST MODE) format. Sum of all CB_PERF_SEL_LQUAD_FORMAT_IS_EXPORT* should match CB_PERF_SEL_DB_CB_LQUAD_VALID_READY |
| CB_PERF_SEL_LQUAD_FORMAT_IS_EXPORT_32_GR | 26 | Number of transactions/quads on the DB_CB_lquad interface that use the EXPORT_32GR (FAST MODE) format. Sum of all CB_PERF_SEL_LQUAD_FORMAT_IS_EXPORT* should match CB_PERF_SEL_DB_CB_LQUAD_VALID_READY |
| CB_PERF_SEL_LQUAD_FORMAT_IS_EXPORT_32_ABGR | 27 | Number of transactions on the DB_CB_lquad interface that use the EXPORT_32ABGR (SLOW MODE) format. To get the number of quads divide this number by 2. Sum of all CB_PERF_SEL_LQUAD_FORMAT_IS_EXPORT* should match CB_PERF_SEL_DB_CB_LQUAD_VALID_READY |
| CB_PERF_SEL_LQUAD_FORMAT_IS_EXPORT_FP16_ABGR | 28 | Number of transactions/quads on the DB_CB_lquad interface that use the EXPORT_FP16ABGR (FAST MODE) format. Sum of all CB_PERF_SEL_LQUAD_FORMAT_IS_EXPORT* should match CB_PERF_SEL_DB_CB_LQUAD_VALID_READY |
| CB_PERF_SEL_LQUAD_FORMAT_IS_EXPORT_SIGNED16_ABGR | 29 | Number of transactions/quads on the DB_CB_lquad interface that use the EXPORT_SIGNED16ABGR (FAST MODE) format. Sum of all CB_PERF_SEL_LQUAD_FORMAT_IS_EXPORT* should match CB_PERF_SEL_DB_CB_LQUAD_VALID_READY |
| CB_PERF_SEL_LQUAD_FORMAT_IS_EXPORT_UNSIGNED16_ABGR | 30 | Number of transactions/quads on the DB_CB_lquad interface that use the EXPORT_UNSIGNED16ABGR (FAST MODE) format. Sum of all CB_PERF_SEL_LQUAD_FORMAT_IS_EXPORT* should match CB_PERF_SEL_DB_CB_LQUAD_VALID_READY |
| CB_PERF_SEL_QUAD_KILLED_BY_EXTRA_PIXEL_EXPORT | 31 | Number of quads that were killed because they were an extra pixel export from the shader based on CB_TARGET_MASK, CB_SHADER_MASK settings. |
| CB_PERF_SEL_QUAD_KILLED_BY_COLOR_INVALID | 32 | Number of quads that was killed because their render target had CB_COLOR_INFO.FORMAT set to invalid. |
| CB_PERF_SEL_QUAD_KILLED_BY_NULL_TARGET_SHADER_MASK | 33 | Number of quads that were killed because of null (target&shader) mask. |
| CB_PERF_SEL_QUAD_KILLED_BY_NULL_SAMPLE_MASK | 34 | Number of quads that were killed because of a null sample mask. |
| CB_PERF_SEL_QUAD_KILLED_BY_DISCARD_PIXEL | 35 | Number of quads that were killed because they were blending but enabled optimization to discard all the pixels of the quad. This optimization can be disabled by setting DISABLE_BLEND_OPT_DISCARD_PIXEL. |
| CB_PERF_SEL_FC_CLEAR_QUAD_VALID_READY | 36 | Number of cycles in the fmask logic that the read_latency FIFO to the clear interface is valid and ready. This is counting quads + events. The number of quads should match the number of quads on the lquad interface (i.e. CB_PERF_SEL_DB_CB_LQUAD_VALID_READY - CB_PERF_SEL_LQUAD_FORMAT_IS_EXPORT_32ABGR/2) minus any quads that are dropped and show up in the perf counters CB_PERF_SEL_QUAD_KILLED_BY* |
| CB_PERF_SEL_FC_CLEAR_QUAD_VALID_READYB | 37 | Number of cycles in the fmask logic that the read_latency FIFO to the clear interface is valid and not ready. |
| CB_PERF_SEL_FC_CLEAR_QUAD_VALIDB_READY | 38 | Number of cycles in the fmask logic that the read_latency FIFO to the clear interface is not valid and ready. |
| CB_PERF_SEL_FC_CLEAR_QUAD_VALIDB_READYB | 39 | Number of cycles in the fmask logic that the read_latency FIFO to the clear interface is not valid and not ready. |
| CB_PERF_SEL_FOP_IN_VALID_READY | 40 | Number of cycles the fragop unit input interface is valid and ready. This is counting quads + events. The number of quads here should match CB_PERF_SEL_FC_CLEAR_QUAD_VALID_READY plus any quads inserted by the fast_clear_eliminate (fc_clear) block |
| CB_PERF_SEL_FOP_IN_VALID_READYB | 41 | Number of cycles the fragop unit input interface is valid and not ready |
| CB_PERF_SEL_FOP_IN_VALIDB_READY | 42 | Number of cycles the fragop unit input interface is not valid and ready |
| CB_PERF_SEL_FOP_IN_VALIDB_READYB | 43 | Number of cycles the fragop unit input interface is not valid and not ready |
| CB_PERF_SEL_FC_CC_QUADFRAG_VALID_READY | 44 | Number of cycles the fmask to color cache interface is valid and ready. This is counting quad fragments + events. The number of quads here should match CB_PERF_SEL_FOP_IN_VALID_READY times the number of fragments created for each of these quads by the fragop block |
| CB_PERF_SEL_FC_CC_QUADFRAG_VALID_READYB | 45 | Number of cycles the fmask to color cache interface is valid and not ready |
| CB_PERF_SEL_FC_CC_QUADFRAG_VALIDB_READY | 46 | Number of cycles the fmask to color cache interface is not valid and ready |
| CB_PERF_SEL_FC_CC_QUADFRAG_VALIDB_READYB | 47 | Number of cycles the fmask to color cache interface is not valid and not ready |
| CB_PERF_SEL_CC_IB_SR_FRAG_VALID_READY | 48 | Number of cycles the color cache input block’s serializer interface is valid and ready. This is counting quad fragments + events. The number of quads here should match CB_PERF_SEL_FC_CC_QUADFRAG_VALID_READY times the number of cycles needed to handle these quads due to any HW limitations like 32bit floating point blending, linear or linear_aligned array mode etc |
| CB_PERF_SEL_CC_IB_SR_FRAG_VALID_READYB | 49 | Number of cycles the color cache input block’s serializer two-lookup interface is valid and not ready |
| CB_PERF_SEL_CC_IB_SR_FRAG_VALIDB_READY | 50 | Number of cycles the color cache input block’s serializer two-lookup interface is not valid and ready |
| CB_PERF_SEL_CC_IB_SR_FRAG_VALIDB_READYB | 51 | Number of cycles the color cache input block’s serializer to two-lookup interface is not valid and not ready |
| CB_PERF_SEL_CC_IB_TB_FRAG_VALID_READY | 52 | Number of cycles the color cache input block to tag block interface is valid and ready. This is counting quad fragments + events. The number of quads here should match CB_PERF_SEL_CC_IB_SR_FRAG_VALID_READY + CB_PERF_SEL_TWO_PROBE_QUAD_FRAGMENT. |
| CB_PERF_SEL_CC_IB_TB_FRAG_VALID_READYB | 53 | Number of cycles the color cache input block to tag block interface is valid and not ready |
| CB_PERF_SEL_CC_IB_TB_FRAG_VALIDB_READY | 54 | Number of cycles the color cache input block to tag block interface is not valid and ready |
| CB_PERF_SEL_CC_IB_TB_FRAG_VALIDB_READYB | 55 | Number of cycles the color cache input block to tag block interface is not valid and not ready |
| CB_PERF_SEL_CC_RB_BC_EVENFRAG_VALID_READY | 56 | Number of cycles the reorder_buffer to blend_control interface (evenfrag) is valid and ready. This is counting quad fragments + events. The number of quads in (CB_PERF_SEL_CC_RB_BC_EVENFRAG_VALID_READY + CB_PERF_SEL_CC_RB_BC_ODDFRAG_VALID_READY) equals quads in CB_PERF_SEL_CC_IB_SR_FRAG_VALID_READY as the de-serializer collapses the two cache probe quads back into one. The quads are split into two fifos, evenfifo stores the quads with even 32B address. Events go down both paths |
| CB_PERF_SEL_CC_RB_BC_EVENFRAG_VALID_READYB | 57 | Number of cycles the reorder_buffer to blend_control interface (evenfrag) is valid and not ready |
| CB_PERF_SEL_CC_RB_BC_EVENFRAG_VALIDB_READY | 58 | Number of cycles the reorder_buffer to blend_control interface (evenfrag) is not valid and ready |
| CB_PERF_SEL_CC_RB_BC_EVENFRAG_VALIDB_READYB | 59 | Number of cycles the reorder_buffer to blend_control interface (evenfrag) is not valid and not ready |
| CB_PERF_SEL_CC_RB_BC_ODDFRAG_VALID_READY | 60 | Number of cycles the reorder_buffer to blend_control interface (oddfrag) is valid and ready. This is counting quad fragments + events. The number of quads in (CB_PERF_SEL_CC_RB_BC_EVENFRAG_VALID_READY + CB_PERF_SEL_CC_RB_BC_ODDFRAG_VALID_READY) equals quads in CB_PERF_SEL_CC_IB_SR_FRAG_VALID_READY as the de-serializer collapses the two cache probe quads back into one. The quads are split into two fifos, oddfifo stores the quads with odd 32B address. Events go down both paths |
| CB_PERF_SEL_CC_RB_BC_ODDFRAG_VALID_READYB | 61 | Number of cycles the reorder_buffer to blend_control interface (oddfrag) is valid and not ready |
| CB_PERF_SEL_CC_RB_BC_ODDFRAG_VALIDB_READY | 62 | Number of cycles the reorder_buffer to blend_control interface (oddfrag) is not valid and ready |
| CB_PERF_SEL_CC_RB_BC_ODDFRAG_VALIDB_READYB | 63 | Number of cycles the reorder_buffer to blend_control interface (oddfrag) is not valid and not ready |
| CB_PERF_SEL_CC_BC_CS_FRAG_VALID | 64 | Number of cycles the blend_control to cache_storage interface is valid. This is counting quad fragments + events. The number of quads here should match CB_PERF_SEL_CC_IB_SR_FRAG_VALID_READY |
| CB_PERF_SEL_CM_CACHE_HIT | 65 | Number of tile transactions that hit cmask cache (TAG HIT && SECTOR HIT). It is a hit if you hit the tag and a sector in that tag meaning that sector’s data is valid. Total probe transactions = incoming tile transactions (in compressed AA mode or single sample with fast clear mode) times render targets = CB_PERF_SEL_CM_CACHE_HIT+CB_PERF_SEL_CM_CACHE_SECTOR_MISS. |
| CB_PERF_SEL_CM_CACHE_TAG_MISS | 66 | Number of tile transactions that have cmask cache tag misses (TAG MISS). A tag miss is when there isn’t a tag that matches the input. Because of sectoring, a tag hit does not imply that the required data is in the cache (i.e TAG HIT && SECTOR MISS). |
| CB_PERF_SEL_CM_CACHE_SECTOR_MISS | 67 | Number of tile transactions that have cmask cache sector miss (TAG MISS || SECTOR MISS). A sector miss is when the required sector is not in the cache. Tag misses are included in this count because a sector cannot be in the cache if there is no tag for it. Total transactions that have TAG HIT && SECTOR MISS = CB_PERF_SEL_CM_CACHE_SECTOR_MISS - CB_PERF_SEL_CM_CACHE_TAG_MISS |
| CB_PERF_SEL_CM_CACHE_REEVICTION_STALL | 68 | Number of cycles the cmask cache is stalled because it is trying to evict a line that already has a pending evict. This can happen when a cache line has been written out to memory and waiting for the write ack gets a hit again, eventually becoming dirty and needing to be re-evicted. Since only one eviction per cache line can be outstanding, the cache is stalled |
| CB_PERF_SEL_CM_CACHE_EVICT_NONZERO_INFLIGHT_STALL | 69 | Number of cycles the cmask cache is stalled because it is trying to evict a line that has a nonzero in-flight count. When a cache lines LRU count reaches the scaled eviction point set by CM_CACHE_EVICT_POINT and it is still in-flight then this counter get incremented This is usually a sign that the cache is unable to withstand the memory read latency. CM_CACHE_EVICT_POINT can be set to a LOWER value to handle the increased read latency |
| CB_PERF_SEL_CM_CACHE_REPLACE_PENDING_EVICT_STALL | 70 | Number of cycles the cmask cache is stalled because the line being replaced has a pending evict. This happens when the cache line chosen for allocation on a miss is still waiting for its write-ack to return from a previous eviction. This is usually a sign that the cache is unable to withstand the memory write latency. CM_CACHE_EVICT_POINT can be set to a HIGHER value to handle the increased write latency |
| CB_PERF_SEL_CM_CACHE_INFLIGHT_COUNTER_MAXIMUM_STALL | 71 | Number of cycles the cmask cache is stalled because one of the in-flight counters has reached the maximum value. |
| CB_PERF_SEL_CM_CACHE_READ_OUTPUT_STALL | 72 | Number of cycles the cmask cache is stalled because the read request output path is stalled. This can happen due to increased memory read request ask-go latency |
| CB_PERF_SEL_CM_CACHE_WRITE_OUTPUT_STALL | 73 | Number of cycles the cmask cache is stalled because the write request output path is stalled. This can happen due to increased memory write request ask-go latency |
| CB_PERF_SEL_CM_CACHE_ACK_OUTPUT_STALL | 74 | Number of cycles the cmask cache is stalled because the acknowledge output path is stalled. This can happen if the next stage in the pipeline (cmask cache read latency hiding FIFO) is full and stalling the cmask cache so the cache cannot output any transactions. |
| CB_PERF_SEL_CM_CACHE_STALL | 75 | Number of cycles the cmask cache is stalled. This is the combined cache stall signal. |
| CB_PERF_SEL_CM_CACHE_FLUSH | 76 | This is the number of cmask cache flushes. |
| CB_PERF_SEL_CM_CACHE_TAGS_FLUSHED | 77 | Number of dirty tags that are flushed (the cmask cache has 32 tags). |
| CB_PERF_SEL_CM_CACHE_SECTORS_FLUSHED | 78 | Number of valid sectors (i.e they have valid data) per dirty tags that are flushed (the cmask cache has 1 sector per tag). The valid sectors may be dirty. The total valid non-dirty sectors in flushed out tags = CB_PERF_SEL_CM_CACHE_SECTORS_FLUSHED - CB_PERF_SEL_CM_CACHE_DIRTY_SECTORS_FLUSHED |
| CB_PERF_SEL_CM_CACHE_DIRTY_SECTORS_FLUSHED | 79 | Number of dirty sectors per dirty tags that are flushed (the cmask cache has 1 sector per tag). |
| CB_PERF_SEL_FC_CACHE_HIT | 80 | Number of tile transactions that hit the fmask cache (TAG HIT && SECTOR HIT). It is a hit if you hit the tag and a sector in that tag meaning that sector’s data is valid. Total probe transactions = incoming tile transactions times render targets = CB_PERF_SEL_FC_CACHE_HIT+CB_PERF_SEL_FC_CACHE_SECTOR_MISS. |
| CB_PERF_SEL_FC_CACHE_TAG_MISS | 81 | Number of tile transactions that have fmask cache tag misses (TAG MISS). A tag miss is when there isn’t a tag that matches the input. Because of sectoring, a tag hit does not imply that the required data is in the cache (i.e TAG HIT && SECTOR MISS). |
| CB_PERF_SEL_FC_CACHE_SECTOR_MISS | 82 | Number of tile transactions that have fmask cache sector miss (TAG MISS || SECTOR MISS). A sector miss is when the required sector is not in the cache. Tag misses are included in this count because a sector cannot be in the cache if there is no tag for it. Total transactions that have TAG HIT && SECTOR MISS = CB_PERF_SEL_FC_CACHE_SECTOR_MISS - CB_PERF_SEL_FC_CACHE_TAG_MISS. |
| CB_PERF_SEL_FC_CACHE_REEVICTION_STALL | 83 | Number of cycles the fmask cache is stalled because it is trying to evict a line that already has a pending evict. This can happen when a cache line has been written out to memory and waiting for the write ack gets a hit again, eventually becoming dirty and needing to be re-evicted. Since only one eviction per cache line can be outstanding, the cache is stalled. |
| CB_PERF_SEL_FC_CACHE_EVICT_NONZERO_INFLIGHT_STALL | 84 | Number of cycles the fmask cache is stalled because it is trying to evict a line that has a nonzero in-flight count. When a cache lines LRU count reaches the scaled eviction point set by FC_CACHE_EVICT_POINT and it is still in-flight then this counter get incremented This is usually a sign that the cache is unable to withstand the memory read latency. FC_CACHE_EVICT_POINT can be set to a LOWER value to handle the increased read latency. |
| CB_PERF_SEL_FC_CACHE_REPLACE_PENDING_EVICT_STALL | 85 | Number of cycles the fmask cache is stalled because the line being replaced has a pending evict. This happens when the cache line chosen for allocation on a miss is still waiting for its write-ack to return from a previous eviction. This is usually a sign that the cache is unable to withstand the memory write latency. FC_CACHE_EVICT_POINT can be set to a HIGHER value to handle the increased write latency. |
| CB_PERF_SEL_FC_CACHE_INFLIGHT_COUNTER_MAXIMUM_STALL | 86 | Number of cycles the fmask cache is stalled because one of the in-flight counters has reached the maximum value. |
| CB_PERF_SEL_FC_CACHE_READ_OUTPUT_STALL | 87 | Number of cycles the fmask cache is stalled because the read request output path is stalled. This can happen due to increased memory read request ask-go latency |
| CB_PERF_SEL_FC_CACHE_WRITE_OUTPUT_STALL | 88 | Number of cycles the fmask cache is stalled because the write request output path is stalled. This can happen due to increased memory write request ask-go latency |
| CB_PERF_SEL_FC_CACHE_ACK_OUTPUT_STALL | 89 | Number of cycles the fmask cache is stalled because the acknowledge output path is stalled. This can happen if the next stage in the pipeline (fmask cache read latency hiding FIFO) is full and stalling the fmask cache so the cache cannot output any transactions. |
| CB_PERF_SEL_FC_CACHE_STALL | 90 | Number of cycles the fmask cache is stalled. This is the combined cache stall signal. |
| CB_PERF_SEL_FC_CACHE_FLUSH | 91 | This is the number of fmask cache flushes. |
| CB_PERF_SEL_FC_CACHE_TAGS_FLUSHED | 92 | Number of dirty tags that are flushed (the fmask cache has 64 tags). |
| CB_PERF_SEL_FC_CACHE_SECTORS_FLUSHED | 93 | Number of valid sectors (i.e they have valid data) per dirty tags that are flushed (the fmask cache has 4 sectors per tag). The valid sectors may be dirty. The total valid non-dirty sectors in flushed out tags = CB_PERF_SEL_FC_CACHE_SECTORS_FLUSHED - CB_PERF_SEL_FC_CACHE_DIRTY_SECTORS_FLUSHED. |
| CB_PERF_SEL_FC_CACHE_DIRTY_SECTORS_FLUSHED | 94 | Number of dirty sectors per dirty tags that are flushed (the fmask cache has 4 sectors per tag). |
| CB_PERF_SEL_CC_CACHE_HIT | 95 | Number of quad fragment transactions that hit the color cache (TAG HIT && SECTOR HIT). It is a hit if you hit the tag and a sector in that tag meaning that sector’s data is valid. Total probe transactions = quads in CB_PERF_SEL_CC_IB_TB_FRAG_VALID_READY = CB_PERF_SEL_CC_CACHE_HIT+CB_PERF_SEL_CC_CACHE_SECTOR_MISS |
| CB_PERF_SEL_CC_CACHE_TAG_MISS | 96 | Number of quad fragment transactions that have color cache tag misses (TAG MISS). A tag miss is when there isn’t a tag that matches the input. Because of sectoring, a tag hit does not imply that the required data is in the cache (i.e TAG HIT && SECTOR MISS). |
| CB_PERF_SEL_CC_CACHE_SECTOR_MISS | 97 | Number of quad fragment transactions that have color cache sector miss (TAG MISS || SECTOR MISS). A sector miss is when the required sector is not in the cache. Tag misses are included in this count because a sector cannot be in the cache if there is no tag for it. Total transactions that have TAG HIT && SECTOR MISS = CB_PERF_SEL_CC_CACHE_SECTOR_MISS - CB_PERF_SEL_CC_CACHE_TAG_MISS. |
| CB_PERF_SEL_CC_CACHE_REEVICTION_STALL | 98 | Number of cycles the color cache is stalled because it is trying to evict a line that already has a pending evict. This can happen when a cache line has been written out to memory and waiting for the write ack gets a hit again, eventually becoming dirty and needing to be re-evicted. Since only one eviction per cache line can be outstanding, the cache is stalled. |
| CB_PERF_SEL_CC_CACHE_EVICT_NONZERO_INFLIGHT_STALL | 99 | Number of cycles the color cache is stalled because it is trying to evict a line that has a nonzero in-flight count. When a cache lines LRU count reaches the scaled eviction point set by CC_CACHE_EVICT_POINT and it is still in-flight then this counter get incremented This is usually a sign that the cache is unable to withstand the memory read latency. CC_CACHE_EVICT_POINT can be set to a LOWER value to handle the increased read latency. |
| CB_PERF_SEL_CC_CACHE_REPLACE_PENDING_EVICT_STALL | 100 | Number of cycles the color cache is stalled because the line being replaced has a pending evict. This happens when the cache line chosen for allocation on a miss is still waiting for its write-ack to return from a previous eviction. This is usually a sign that the cache is unable to withstand the memory write latency. CC_CACHE_EVICT_POINT can be set to a HIGHER value to handle the increased write latency. |
| CB_PERF_SEL_CC_CACHE_INFLIGHT_COUNTER_MAXIMUM_STALL | 101 | Number of cycles the color cache is stalled because one of the in-flight counters has reached the maximum value. |
| CB_PERF_SEL_CC_CACHE_READ_OUTPUT_STALL | 102 | Number of cycles the color cache is stalled because the read request output path is stalled. This can happen due to increased memory read request ask-go latency |
| CB_PERF_SEL_CC_CACHE_WRITE_OUTPUT_STALL | 103 | Number of cycles the color cache is stalled because the write request output path is stalled. This can happen due to increased memory write request ask-go latency |
| CB_PERF_SEL_CC_CACHE_ACK_OUTPUT_STALL | 104 | Number of cycles the color cache is stalled because the acknowledge output path is stalled. This can happen if the next stage in the pipeline (color cache read latency hiding FIFO) is full and stalling the color cache so the cache cannot output any transactions. |
| CB_PERF_SEL_CC_CACHE_STALL | 105 | Number of cycles the color cache is stalled. This is the combined cache stall signal. |
| CB_PERF_SEL_CC_CACHE_FLUSH | 106 | This is the number of color cache flushes. This is includes surface sync flushes. |
| CB_PERF_SEL_CC_CACHE_TAGS_FLUSHED | 107 | Number of dirty tags that are flushed (the color cache has 64 tags). |
| CB_PERF_SEL_CC_CACHE_SECTORS_FLUSHED | 108 | Number of valid sectors (i.e they have valid data) per dirty tags that are flushed (the color cache has 4 sectors per tag). The valid sectors may be dirty. The total valid non-dirty sectors in flushed out tags = CB_PERF_SEL_CC_CACHE_SECTORS_FLUSHED - CB_PERF_SEL_CC_CACHE_DIRTY_SECTORS_FLUSHED. |
| CB_PERF_SEL_CC_CACHE_DIRTY_SECTORS_FLUSHED | 109 | Number of dirty sectors per dirty tags that are flushed (the color cache has 4 sectors per tag). |
| CB_PERF_SEL_CC_CACHE_WA_TO_RMW_CONVERSION | 110 | The number of color cache probe transactions that hit a write allocate cache line that convert it into a read-modify-write cache line because they need destination buffer data. |
| CB_PERF_SEL_CB_TAP_WRREQ_VALID_READY | 111 | Number of cycles the CB to TAP write request interface is valid and ready. This is measured at the interface converter rather than directly on the interface. The should be equal to the sum of CB_PERF_SEL_*_MC_WRITE_REQUEST |
| CB_PERF_SEL_CB_TAP_WRREQ_VALID_READYB | 112 | Number of cycles the CB to TAP write request interface is valid and not ready. This is measured at the interface converter rather than directly on the interface. |
| CB_PERF_SEL_CB_TAP_WRREQ_VALIDB_READY | 113 | Number of cycles the CB to TAP write request interface is valid and ready. This is measured at the interface converter rather than directly on the interface. |
| CB_PERF_SEL_CB_TAP_WRREQ_VALIDB_READYB | 114 | Number of cycles the CB to TAP write request interface is not valid and not ready. This is measured at the interface converter rather than directly on the interface. |
| CB_PERF_SEL_CM_MC_WRITE_REQUEST | 115 | Number of 32-byte cmask mc write requests. |
| CB_PERF_SEL_FC_MC_WRITE_REQUEST | 116 | Number of 32-byte fmask mc write requests. |
| CB_PERF_SEL_CC_MC_WRITE_REQUEST | 117 | Number of 32-byte color mc write requests. |
| CB_PERF_SEL_CM_MC_WRITE_REQUESTS_IN_FLIGHT | 118 | Number of 32-byte cmask mc write requests in flight. CB_PERF_SEL_CM_MC_WRITE_REQUESTS_IN_FLIGHT/CB_PERF_SEL_CM_MC_WRITE_REQUEST = average latency. |
| CB_PERF_SEL_FC_MC_WRITE_REQUESTS_IN_FLIGHT | 119 | Number of 32-byte fmask mc write requests in flight. CB_PERF_SEL_FC_MC_WRITE_REQUESTS_IN_FLIGHT/CB_PERF_SEL_FC_MC_WRITE_REQUEST = average latency. |
| CB_PERF_SEL_CC_MC_WRITE_REQUESTS_IN_FLIGHT | 120 | Number of 32-byte color mc write requests in flight. CB_PERF_SEL_CC_MC_WRITE_REQUESTS_IN_FLIGHT/CB_PERF_SEL_CC_MC_WRITE_REQUEST = average latency. |
| CB_PERF_SEL_CB_TAP_RDREQ_VALID_READY | 121 | Number of cycles the CB to TAP read request interface is valid and ready. This is measured at the interface converter rather than directly on the interface. |
| CB_PERF_SEL_CB_TAP_RDREQ_VALID_READYB | 122 | Number of cycles the CB to TAP read request interface is valid and not ready. This is measured at the interface converter rather than directly on the interface. |
| CB_PERF_SEL_CB_TAP_RDREQ_VALIDB_READY | 123 | Number of cycles the CB to TAP read request interface is not valid and ready. This is measured at the interface converter rather than directly on the interface. |
| CB_PERF_SEL_CB_TAP_RDREQ_VALIDB_READYB | 124 | Number of cycles the CB to TAP read request interface is not valid and not ready. This is measured at the interface converter rather than directly on the interface. |
| CB_PERF_SEL_CM_MC_READ_REQUEST | 125 | Number of 32-byte cmask mc read requests. Cmask does not make 32-byte requests, so the counter will report the equivalent number of 32-byte requests. |
| CB_PERF_SEL_FC_MC_READ_REQUEST | 126 | Number of 32-byte fmask mc read requests. |
| CB_PERF_SEL_CC_MC_READ_REQUEST | 127 | Number of 32-byte color mc read requests. Color does not make 32-byte requests, so the counter will report the equivalent number of 32-byte requests. |
| CB_PERF_SEL_CM_MC_READ_REQUESTS_IN_FLIGHT | 128 | Number of 32-byte cmask mc read requests in flight. Cmask does not make 32-byte requests, so the counter will report the equivalent number of 32-byte requests. CB_PERF_SEL_CM_MC_READ_REQUESTS_IN_FLIGHT/CB_PERF_SEL_CM_MC_READ_REQUEST = average latency. |
| CB_PERF_SEL_FC_MC_READ_REQUESTS_IN_FLIGHT | 129 | Number of 32-byte fmask mc read requests in flight. CB_PERF_SEL_FC_MC_READ_REQUESTS_IN_FLIGHT/CB_PERF_SEL_FC_MC_READ_REQUEST = average latency. |
| CB_PERF_SEL_CC_MC_READ_REQUESTS_IN_FLIGHT | 130 | Number of 32-byte color mc read requests in flight. Color does not make 32-byte requests, so the counter will report the equivalent number of 32-byte requests. CB_PERF_SEL_CC_MC_READ_REQUESTS_IN_FLIGHT/CB_PERF_SEL_CC_MC_READ_REQUEST = average latency. |
| CB_PERF_SEL_CM_TQ_FULL | 131 | Number of cycles the cmask tile queue is full. This FIFO covers the cmask memory read latency and stalling indicates that this FIFO does not have enough latency hiding capacity |
| CB_PERF_SEL_CM_TQ_FIFO_TILE_RESIDENCY_STALL | 132 | Number of cycles the cmask read latency hiding fifo’s head entry is waiting for cmask read data to return from memory and become resident in cache so the transaction can be popped. This does not mean that cmask logic is stalled, look at CB_PERF_SEL_CM_TQ_FULL for that. |
| CB_PERF_SEL_FC_QUAD_RDLAT_FIFO_FULL | 133 | Number of cycles the fmask read latency quad FIFO is full. This FIFO covers the fmask memory read latency and stalling indicates that this FIFO does not have enough latency hiding capacity |
| CB_PERF_SEL_FC_TILE_RDLAT_FIFO_FULL | 134 | Number of cycles the fmask read latency tile FIFO is full. This FIFO covers the fmask memory read latency and stalling indicates that this FIFO does not have enough latency hiding capacity |
| CB_PERF_SEL_FC_RDLAT_FIFO_QUAD_RESIDENCY_STALL | 135 | Number of cycles the fmask read latency hiding fifo’s head entry is waiting for fmask read data to return from memory and become resident in cache so the transaction can be popped. This does not mean that fmask logic is stalled, look at CB_PERF_SEL_FC_QUAD_RDLAT_FIFO_FULL || CB_PERF_SEL_FC_TILE_RDLAT_FIFO_FULL for that. |
| CB_PERF_SEL_FOP_FMASK_RAW_STALL | 136 | Number of cycles a quadA-quadA sequence (back-to-back) is stalled because of a read-after-write conflict in fragop block. |
| CB_PERF_SEL_FOP_FMASK_BYPASS_STALL | 137 | Number of cycles a quadA-quadB-quadA sequence is stalled because of a read-after-write conflict in fragop block. |
| CB_PERF_SEL_CC_SF_FULL | 138 | Number of cycles the color cache source FIFO is full. This FIFO covers the fmask memory read plus color cache memory read latency and stalling indicates that this FIFO does not have enough latency hiding capacity |
| CB_PERF_SEL_CC_RB_FULL | 139 | Number of cycles the color cache reorder buffer is full. This FIFO covers the color cache memory read latency and stalling indicates that this FIFO does not have enough latency hiding capacity |
| CB_PERF_SEL_CC_EVENFIFO_QUAD_RESIDENCY_STALL | 140 | Number of cycles the color even side read latency hiding FIFO’s head entry is waiting for even 32B color read data to return from memory and become resident in cache so the transaction can be popped. This does not mean that color logic is stalled, look at CB_PERF_SEL_CC_RB_FULL for that. |
| CB_PERF_SEL_CC_ODDFIFO_QUAD_RESIDENCY_STALL | 141 | Number of cycles the color odd side read latency hiding FIFOs head entry is waiting for odd 32B color read data to return from memory and become resident in cache so the transaction can be popped. This does not mean that color logic is stalled, look at CB_PERF_SEL_CC_RB_FULL for that. |
| CB_PERF_SEL_BLENDER_RAW_HAZARD_STALL | 142 | Number of cycles the blend pipeline is stalled to handle read after write hazards. |
| CB_PERF_SEL_EVENT | 143 | Total number of events reaching the CB. This includes events that the CB does not process. |
| CB_PERF_SEL_EVENT_CACHE_FLUSH_TS | 144 | Number of CACHE_FLUSH_TS events |
| CB_PERF_SEL_EVENT_CONTEXT_DONE | 145 | Number of CONTEXT_DONE events |
| CB_PERF_SEL_EVENT_CACHE_FLUSH | 146 | Number of CACHE_FLUSH events |
| CB_PERF_SEL_EVENT_CACHE_FLUSH_AND_INV_TS_EVENT | 147 | Number of CACHE_FLUSH_AND_INV_TS_EVENT events |
| CB_PERF_SEL_EVENT_CACHE_FLUSH_AND_INV_EVENT | 148 | Number of_CACHE_FLUSH_AND_INV_EVENTevents |
| CB_PERF_SEL_EVENT_FLUSH_AND_INV_CB_DATA_TS | 149 | Number of FLUSH_AND_INV_CB_DATA_TS events |
| CB_PERF_SEL_EVENT_FLUSH_AND_INV_CB_META | 150 | Number of FLUSH_AND_INV_CB_META events |
| CB_PERF_SEL_CC_SURFACE_SYNC | 151 | Number of surface syncs |
| CB_PERF_SEL_CMASK_READ_DATA_0XC | 152 | Number of fmask cache probe misses that read cmask with a value of 0xC. A value of 0xC means that a fmask tile requires 0 bit planes. Sum of CB_PERF_SEL_CMASK_READ_DATA_* should match CB_PERF_SEL_FC_CACHE_SECTOR_MISS. |
| CB_PERF_SEL_CMASK_READ_DATA_0XD | 153 | Number of fmask cache probe misses that read cmask with a value of 0xD. A value of 0xD means that a fmask tile requires 1 bit planes. Sum of CB_PERF_SEL_CMASK_READ_DATA_* should match CB_PERF_SEL_FC_CACHE_SECTOR_MISS. |
| CB_PERF_SEL_CMASK_READ_DATA_0XE | 154 | Number of fmask cache probe misses that read cmask with a value of 0xE. A value of 0xE means that a fmask tile requires 2 bit planes. Sum of CB_PERF_SEL_CMASK_READ_DATA_* should match CB_PERF_SEL_FC_CACHE_SECTOR_MISS. |
| CB_PERF_SEL_CMASK_READ_DATA_0XF | 155 | Number of fmask cache probe misses that read cmask with a value of 0xF. A value of 0xF means that a fmask tile requires 3 bit planes. Sum of CB_PERF_SEL_CMASK_READ_DATA_* should match CB_PERF_SEL_FC_CACHE_SECTOR_MISS. |
| CB_PERF_SEL_CMASK_WRITE_DATA_0XC | 156 | Number of tiles that wrote a value of 0xC to cmask. Cmask writes occur at the end of a tile. The hardware does not perform the write if fmask has not changed, but the counter will update as if regardless of whether the hardware wrote or not. A value of 0xC means that a fmask tile requires 0 bit planes. Sum of all CB_PERF_SEL_CMASK_WRITE_DATA_* equals cmask cache probes (CB_PERF_SEL_CM_CACHE_HIT + CB_PERF_SEL_CM_CACHE_SECTOR_MISS). |
| CB_PERF_SEL_CMASK_WRITE_DATA_0XD | 157 | Number of tile that wrote a value of 0xD was written to cmask. Cmask writes occur at the end of a tile. The hardware does not perform the write if fmask has not changed, but the counter will update as if regardless of whether the hardware wrote or not. A value of 0xD means that a fmask tile requires 1 bit planes. Sum of all CB_PERF_SEL_CMASK_WRITE_DATA_* equals cmask cache probes (CB_PERF_SEL_CM_CACHE_HIT + CB_PERF_SEL_CM_CACHE_SECTOR_MISS). |
| CB_PERF_SEL_CMASK_WRITE_DATA_0XE | 158 | Number of tiles that wrote a value of 0xE was written to cmask. Cmask writes occur at the end of a tile. The hardware does not perform the write if fmask has not changed, but the counter will update as if regardless of whether the hardware wrote or not. A value of 0xE means that a fmask tile requires 2 bit planes. Sum of all CB_PERF_SEL_CMASK_WRITE_DATA_* equals cmask cache probes (CB_PERF_SEL_CM_CACHE_HIT + CB_PERF_SEL_CM_CACHE_SECTOR_MISS). |
| CB_PERF_SEL_CMASK_WRITE_DATA_0XF | 159 | Number of tiles that wrote a value of 0xF was written to cmask. Cmask writes occur at the end of a tile. The hardware does not perform the write if fmask has not changed, but the counter will update as if regardless of whether the hardware wrote or not. A value of 0xF means that a fmask tile requires 3 bit planes. Sum of all CB_PERF_SEL_CMASK_WRITE_DATA_* equals cmask cache probes (CB_PERF_SEL_CM_CACHE_HIT + CB_PERF_SEL_CM_CACHE_SECTOR_MISS). |
| CB_PERF_SEL_TWO_PROBE_QUAD_FRAGMENT | 160 | Number of quad fragments that require two cache probes. AA blending can create these when the read fragment does not match the write fragment. |
| CB_PERF_SEL_EXPORT_32_ABGR_QUAD_FRAGMENT | 161 | Number of EXPORT_32_ABGR quad fragments. It takes two clocks to send the src color data for these. |
| CB_PERF_SEL_DUAL_SOURCE_COLOR_QUAD_FRAGMENT | 162 | Number of dual source blending color quad fragments |
| CB_PERF_SEL_QUAD_HAS_1_FRAGMENT_BEFORE_UPDATE | 163 | Number of quads that started with 1 fragment in AA mode in fragop block. |
| CB_PERF_SEL_QUAD_HAS_2_FRAGMENTS_BEFORE_UPDATE | 164 | Number of quads that started with 2 fragments in AA mode in fragop block. |
| CB_PERF_SEL_QUAD_HAS_3_FRAGMENTS_BEFORE_UPDATE | 165 | Number of quads that started with 3 fragments in AA mode in fragop block. |
| CB_PERF_SEL_QUAD_HAS_4_FRAGMENTS_BEFORE_UPDATE | 166 | Number of quads that started with 4 fragments in AA mode in fragop block. |
| CB_PERF_SEL_QUAD_HAS_5_FRAGMENTS_BEFORE_UPDATE | 167 | Number of quads that started with 5 fragments in AA mode in fragop block. |
| CB_PERF_SEL_QUAD_HAS_6_FRAGMENTS_BEFORE_UPDATE | 168 | Number of quads that started with 6 fragments in AA mode in fragop block. |
| CB_PERF_SEL_QUAD_HAS_7_FRAGMENTS_BEFORE_UPDATE | 169 | Number of quads that started with 7 fragments in AA mode in fragop block. |
| CB_PERF_SEL_QUAD_HAS_8_FRAGMENTS_BEFORE_UPDATE | 170 | Number of quads that started with 8 fragments in AA mode in fragop block. |
| CB_PERF_SEL_QUAD_HAS_1_FRAGMENT_AFTER_UPDATE | 171 | Number of quads that ended up with 1 fragment in AA mode in fragop block. |
| CB_PERF_SEL_QUAD_HAS_2_FRAGMENTS_AFTER_UPDATE | 172 | Number of quads that ended up with 2 fragments in AA mode in fragop block. |
| CB_PERF_SEL_QUAD_HAS_3_FRAGMENTS_AFTER_UPDATE | 173 | Number of quads that ended up with 3 fragments in AA mode in fragop block. |
| CB_PERF_SEL_QUAD_HAS_4_FRAGMENTS_AFTER_UPDATE | 174 | Number of quads that ended up with 4 fragments in AA mode in fragop block. |
| CB_PERF_SEL_QUAD_HAS_5_FRAGMENTS_AFTER_UPDATE | 175 | Number of quads that ended up with 5 fragments in AA mode in fragop block. |
| CB_PERF_SEL_QUAD_HAS_6_FRAGMENTS_AFTER_UPDATE | 176 | Number of quads that ended up with 6 fragments in AA mode in fragop block. |
| CB_PERF_SEL_QUAD_HAS_7_FRAGMENTS_AFTER_UPDATE | 177 | Number of quads that ended up with 7 fragments in AA mode in fragop block. |
| CB_PERF_SEL_QUAD_HAS_8_FRAGMENTS_AFTER_UPDATE | 178 | Number of quads that ended up with 8 fragments in AA mode in fragop block. |
| CB_PERF_SEL_QUAD_ADDED_1_FRAGMENT | 179 | Number of quads that added 1 fragment in AA mode in fragop block. |
| CB_PERF_SEL_QUAD_ADDED_2_FRAGMENTS | 180 | Number of quads that added 2 fragments in AA mode in fragop block. |
| CB_PERF_SEL_QUAD_ADDED_3_FRAGMENTS | 181 | Number of quads that added 3 fragments in AA mode in fragop block. |
| CB_PERF_SEL_QUAD_ADDED_4_FRAGMENTS | 182 | Number of quads that added 4 fragments in AA mode in fragop block. |
| CB_PERF_SEL_QUAD_ADDED_5_FRAGMENTS | 183 | Number of quads that added 5 fragments in AA mode in fragop block. |
| CB_PERF_SEL_QUAD_ADDED_6_FRAGMENTS | 184 | Number of quads that added 6 fragments in AA mode in fragop block. |
| CB_PERF_SEL_QUAD_ADDED_7_FRAGMENTS | 185 | Number of quads that added 7 fragments in AA mode in fragop block. |
| CB_PERF_SEL_QUAD_REMOVED_1_FRAGMENT | 186 | Number of quads that removed 1 fragment in AA mode in fragop block. |
| CB_PERF_SEL_QUAD_REMOVED_2_FRAGMENTS | 187 | Number of quads that removed 2 fragments in AA mode in fragop block. |
| CB_PERF_SEL_QUAD_REMOVED_3_FRAGMENTS | 188 | Number of quads that removed 3 fragments in AA mode in fragop block. |
| CB_PERF_SEL_QUAD_REMOVED_4_FRAGMENTS | 189 | Number of quads that removed 4 fragments in AA mode in fragop block. |
| CB_PERF_SEL_QUAD_REMOVED_5_FRAGMENTS | 190 | Number of quads that removed 5 fragments in AA mode in fragop block. |
| CB_PERF_SEL_QUAD_REMOVED_6_FRAGMENTS | 191 | Number of quads that removed 6 fragments in AA mode in fragop block. |
| CB_PERF_SEL_QUAD_REMOVED_7_FRAGMENTS | 192 | Number of quads that removed 7 fragments in AA mode in fragop block. |
| CB_PERF_SEL_QUAD_READS_FRAGMENT_0 | 193 | Number of quads that reads fragment 0 to the color cache. |
| CB_PERF_SEL_QUAD_READS_FRAGMENT_1 | 194 | Number of quads that reads fragment 1 to the color cache. |
| CB_PERF_SEL_QUAD_READS_FRAGMENT_2 | 195 | Number of quads that reads fragment 2 to the color cache. |
| CB_PERF_SEL_QUAD_READS_FRAGMENT_3 | 196 | Number of quads that reads fragment 3 to the color cache. |
| CB_PERF_SEL_QUAD_READS_FRAGMENT_4 | 197 | Number of quads that reads fragment 4 to the color cache. |
| CB_PERF_SEL_QUAD_READS_FRAGMENT_5 | 198 | Number of quads that reads fragment 5 to the color cache. |
| CB_PERF_SEL_QUAD_READS_FRAGMENT_6 | 199 | Number of quads that reads fragment 6 to the color cache. |
| CB_PERF_SEL_QUAD_READS_FRAGMENT_7 | 200 | Number of quads that reads fragment 7 to the color cache. |
| CB_PERF_SEL_QUAD_WRITES_FRAGMENT_0 | 201 | Number of quads that writes fragment 0 to the color cache. |
| CB_PERF_SEL_QUAD_WRITES_FRAGMENT_1 | 202 | Number of quads that writes fragment 1 to the color cache. |
| CB_PERF_SEL_QUAD_WRITES_FRAGMENT_2 | 203 | Number of quads that writes fragment 2 to the color cache. |
| CB_PERF_SEL_QUAD_WRITES_FRAGMENT_3 | 204 | Number of quads that writes fragment 3 to the color cache. |
| CB_PERF_SEL_QUAD_WRITES_FRAGMENT_4 | 205 | Number of quads that writes fragment 4 to the color cache. |
| CB_PERF_SEL_QUAD_WRITES_FRAGMENT_5 | 206 | Number of quads that writes fragment 5 to the color cache. |
| CB_PERF_SEL_QUAD_WRITES_FRAGMENT_6 | 207 | Number of quads that writes fragment 6 to the color cache. |
| CB_PERF_SEL_QUAD_WRITES_FRAGMENT_7 | 208 | Number of quads that writes fragment 7 to the color cache. |
| CB_PERF_SEL_QUAD_BLEND_OPT_DONT_READ_DST | 209 | Number of quads that were blending but enabled optimization to not need read destination data. This optimization can be disabled by setting DISABLE_BLEND_OPT_DONT_RD_DST. |
| CB_PERF_SEL_QUAD_BLEND_OPT_BLEND_BYPASS | 210 | Number of quads that were blending but enabled optimization to bypassed blend operation. This optimization can be disabled by setting DISABLE_BLEND_OPT_BYPASS. |
| CB_PERF_SEL_QUAD_BLEND_OPT_DISCARD_PIXELS | 211 | Number of quads that were blending but enabled optimization to discard at least 1 pixel. This optimization can be disabled by setting DISABLE_BLEND_OPT_DISCARD_PIXEL. |
| CB_PERF_SEL_QUAD_DST_READ_COULD_HAVE_BEEN_OPTIMIZED | 212 | Number of quads that the blender detects that could have optimized away their destination reads (i.e destblend*dst == 0.0f). NOTE: this counter is not completely accurate and does not work for colors that have 32-bit components. |
| CB_PERF_SEL_QUAD_BLENDING_COULD_HAVE_BEEN_BYPASSED | 213 | Number of quads that the blender detects that could have optimized away blending (destblenddst == 0.0f && srcblendsrc == 1.0f). NOTE: this counter is not completely accurate and does not work for colors that have 32-bit components. |
| CB_PERF_SEL_QUAD_COULD_HAVE_BEEN_DISCARDED | 214 | Number of quads that the blender detects that could have discarded some pixels (destblenddst == 1.0f && srcblendsrc == 0.0f). NOTE: this counter is not completely accurate and does not work for colors that have 32-bit components. |
| CB_PERF_SEL_BLEND_OPT_PIXELS_RESULT_EQ_DEST | 215 | Number of blended pixels that the blender detects can be dropped because the blended result is the same as the original destination value. |
| CB_PERF_SEL_DRAWN_BUSY | 216 | Number of busy cycles around the stage calculating the CB_PERF_SEL_DRAWN_* counters. Used as denominator for all the CB_PERF_SEL_DRAWN_* counters. Filtering using CB_PERFCOUNTER_FILTER fields has an effect in this mode. |
| CB_PERF_SEL_TILE_TO_CMR_REGION_BUSY | 217 | Number of busy cycles b/w PERFCOUNTER_START and PERFCOUNTER_STOP for the pipeline region b/w DB_CB_TILE and CMASK CACHE READ. This can be used as denominator for rate calculation for counters in this region |
| CB_PERF_SEL_CMR_TO_FCR_REGION_BUSY | 218 | Number of busy cycles b/w PERFCOUNTER_START and PERFCOUNTER_STOP for the pipeline region b/w CMASK CACHE READ and FMASK CACHE READ. This can be used as denominator for rate calculation for counters in this region |
| CB_PERF_SEL_FCR_TO_CCR_REGION_BUSY | 219 | Number of busy cycles b/w PERFCOUNTER_START and PERFCOUNTER_STOP for the pipeline region b/w FMASK CACHE READ and COLOR CACHE READ. This can be used as denominator for rate calculation for counters in this region |
| CB_PERF_SEL_CCR_TO_CCW_REGION_BUSY | 220 | Number of busy cycles b/w PERFCOUNTER_START and PERFCOUNTER_STOP for the pipeline region b/w COLOR CACHE READ and COLOR CACHE WRITE. This can be used as denominator for rate calculation for counters in this region |
| CB_PERF_SEL_FC_PF_SLOW_MODE_QUAD_EMPTY_HALF_DROPPED | 221 | Number of lquad transactions (halves of slow mode 2-cycle quad) that were dropped in fc_merge logic due to optimization from going into the SRC FIFO because the pixel pair in that half quad was unlit. This optimization can be disabled by setting: DISABLE_SLOW_MODE_EMPTY_HALF_QUAD_KILL |
| CB_PERF_SEL_FC_SEQUENCER_CLEAR | 222 | Number of quads that are generated to expand out the fast clear color as well as number of drawn quads that have partially lit pixels so that unlit samples need to expands their fast clear color. Use CB_PERF_SEL_DRAWN_BUSY as denominator to get per clock rates. |
| CB_PERF_SEL_FC_SEQUENCER_ELIMINATE_FAST_CLEAR | 223 | Number of quads generated during ELIMINATE_FAST_CLEAR pass to expand their fast clear color Use CB_PERF_SEL_DRAWN_BUSY as denominator to get per clock rates. |
| CB_PERF_SEL_FC_SEQUENCER_FMASK_DECOMPRESS | 224 | Number of quads generated during FMASK_DECOMPRESS pass to expand their fast clear color Use CB_PERF_SEL_DRAWN_BUSY as denominator to get per clock rates. |
| CB_PERF_SEL_FC_SEQUENCER_FMASK_COMPRESSION_DISABLE | 225 | Number of quad fragments seen during normal rendering when the state was set to disable fmask compression. This will include drawn quads as well as generated quads to expand their fast clear Use CB_PERF_SEL_DRAWN_BUSY as denominator to get per clock rates. |
| CB_PERF_SEL_SCORPIO_START | 226 | Start of Xbox One X-specific counters. |
| CB_PERF_SEL_CC_CACHE_READS_SAVED_DUE_TO_DCC | 226 | The number of reads prevented by DCC. |
| CB_PERF_SEL_FC_KEYID_RDLAT_FIFO_FULL | 227 | The number of cycles the fmask read latency keyid fifo is full. This fifo covers the fmask memory read latency and stalling indicates that this fifo does not have enough latency hiding capacity. |
| CB_PERF_SEL_FC_DOC_IS_STALLED | 228 | The number of cycles the overwrite combiner is stalled. |
| CB_PERF_SEL_FC_DOC_MRTS_NOT_COMBINED | 229 | The number of times a quad’s mrt was not able to be combined with another one (so as to share the same quarter-tile cam entry). |
| CB_PERF_SEL_FC_DOC_MRTS_COMBINED | 230 | The number of times a quad’s mrt was combined with another one (so as to share the same quarter-tile cam entry). |
| CB_PERF_SEL_FC_DOC_QTILE_CAM_MISS | 231 | The number of misses in the overwrite combiner’s quarter-tile cam. |
| CB_PERF_SEL_FC_DOC_QTILE_CAM_HIT | 232 | The number of hits in the overwrite combiner’s quarter-tile cam. |
| CB_PERF_SEL_FC_DOC_CLINE_CAM_MISS | 233 | The number of misses in the overwrite combiner’s cacheline cam. |
| CB_PERF_SEL_FC_DOC_CLINE_CAM_HIT | 234 | The number of hits in the overwrite combiner’s cacheline cam. |
| CB_PERF_SEL_FC_DOC_QUAD_PTR_FIFO_IS_FULL | 235 | The number of cycles that the overwrite combiner’s quad pointer fifo is full. |
| CB_PERF_SEL_FC_DOC_OVERWROTE_1_SECTOR | 236 | The number of times that the overwrite combiner overwrote just 1 sector. |
| CB_PERF_SEL_FC_DOC_OVERWROTE_2_SECTORS | 237 | The number of times that the overwrite combiner overwrote just 2 sector. |
| CB_PERF_SEL_FC_DOC_OVERWROTE_3_SECTORS | 238 | The number of times that the overwrite combiner overwrote just 3 sector. |
| CB_PERF_SEL_FC_DOC_OVERWROTE_4_SECTORS | 239 | The number of times that the overwrite combiner overwrote all 4 sectors. |
| CB_PERF_SEL_FC_DOC_TOTAL_OVERWRITTEN_SECTORS | 240 | The number of total sectors overwritten by the Overwrite Combiner. |
| CB_PERF_SEL_FC_DCC_CACHE_HIT | 241 | The number of hits in the DCC cache (TAG HIT && SECTOR HIT). |
| CB_PERF_SEL_FC_DCC_CACHE_TAG_MISS | 242 | The number of tag misses in the DCC cache (TAG MISS). |
| CB_PERF_SEL_FC_DCC_CACHE_SECTOR_MISS | 243 | The number of sector misses in the DCC cache (TAG MISS || SECTOR MISS). |
| CB_PERF_SEL_FC_DCC_CACHE_REEVICTION_STALL | 244 | The number of cycles the DCC cache is stalled because it is trying to evict a line that already has a pending evict. This can happen when a cacheline has been written out to memory and waiting for the write ack gets a hit again, eventually becoming dirty and needing to be re-evicted. Since only one eviction per cacheline can be outstanding, the cache is stalled. |
| CB_PERF_SEL_FC_DCC_CACHE_EVICT_NONZERO_INFLIGHT_STALL | 245 | The number of cycles the DCC cache is stalled because it is trying to evict a line that has a nonzero inflight count. When a cachelines LRU count reaches the scaled eviction point set by DCC_CACHE_EVICT_POINT and it is still inflight then this counter get incremented This is usually a sign that the cache is unable to withstand the memory read latency. DCC_CACHE_EVICT_POINT can be set to a LOWER value to handle the increased read latency. This count is the number of times the Key Cache sends a panic to the Color Cache. |
| CB_PERF_SEL_FC_DCC_CACHE_REPLACE_PENDING_EVICT_STALL | 246 | The number of cycles the DCC cache is stalled because the line being replaced has a pending evict. This happens when the cacheline chosen for allocation on a miss is still waiting for its write-ack to return from a previous eviction. This is usually a sign that the cache is unable to withstand the memory write latency. DCC_CACHE_EVICT_POINT can be set to a HIGHER value to handle the increased write latency. |
| CB_PERF_SEL_FC_DCC_CACHE_INFLIGHT_COUNTER_MAXIMUM_STALL | 247 | The number of cycles the DCC cache is stalled because one of the inflight counters has reached the maximum value. |
| CB_PERF_SEL_FC_DCC_CACHE_READ_OUTPUT_STALL | 248 | The number of cycles the DCC cache is stalled because the read request output path is stalled. This can happen due to increased memory read request ask-go latency. |
| CB_PERF_SEL_FC_DCC_CACHE_WRITE_OUTPUT_STALL | 249 | The number of cycles the DCC cache is stalled because the write request output path is stalled. This can happen due to increased memory write request ask-go latency. |
| CB_PERF_SEL_FC_DCC_CACHE_ACK_OUTPUT_STALL | 250 | The number of cycles the DCC cache is stalled because the acknowledge output path is stalled. This can happen if the next stage in the pipeline (DCC cache read latency hiding fifo) is full and stalling the DCC cache so the cache cannot output any transactions. |
| CB_PERF_SEL_FC_DCC_CACHE_STALL | 251 | The number of cycles the DCC cache is stalled. This is the combined cache stall signal. |
| CB_PERF_SEL_FC_DCC_CACHE_FLUSH | 252 | This is the number of DCC cache flushes. |
| CB_PERF_SEL_FC_DCC_CACHE_TAGS_FLUSHED | 253 | The number of dirty tags that are flushed (the DCC cache has 32 tags). |
| CB_PERF_SEL_FC_DCC_CACHE_SECTORS_FLUSHED | 254 | The number of valid sectors (i.e they have valid data) per dirty tags that are flushed (the DCC cache has 1 sector per tag). The valid sectors may be dirty. The total valid non-dirty sectors in flushed out tags. |
| CB_PERF_SEL_FC_DCC_CACHE_DIRTY_SECTORS_FLUSHED | 255 | The number of dirty sectors per dirty tags that are flushed (the DCC cache has 1 sector per tag). |
| CB_PERF_SEL_CC_DCC_BEYOND_TILE_SPLIT | 256 | The number of fragments that lie outside the tile split boundary. The color cache controller limits DCC compression support to only those fragments that are within the first tile split. |
| CB_PERF_SEL_FC_MC_DCC_WRITE_REQUEST | 257 | The number of 32-byte fmask mc DCC write requests. |
| CB_PERF_SEL_FC_MC_DCC_WRITE_REQUESTS_IN_FLIGHT | 258 | The number of 32-byte fmask mc DCC write requests in flight. CB_PERF_SEL_FC_MC_DCC_WRITE_REQUESTS_IN_FLIGHT. |
| CB_PERF_SEL_FC_MC_DCC_READ_REQUEST | 259 | The number of 32-byte fmask mc DCC read requests. |
| CB_PERF_SEL_FC_MC_DCC_READ_REQUESTS_IN_FLIGHT | 260 | The number of 32-byte fmask mc DCC read requests in flight. CB_PERF_SEL_FC_MC_DCC_READ_REQUESTS_IN_FLIGHT. |
| CB_PERF_SEL_CC_DCC_RDREQ_STALL | 261 | The number of times DCC read requests are stalled because the internal scoreboard predicts that the skid fifo will be full. |
| CB_PERF_SEL_CC_DCC_DECOMPRESS_TIDS_IN | 262 | The number of tids that come into the DCC’s decompress module. |
| CB_PERF_SEL_CC_DCC_DECOMPRESS_TIDS_OUT | 263 | The number of tids that come out of the DCC’s decompress module. |
| CB_PERF_SEL_CC_DCC_COMPRESS_TIDS_IN | 264 | The number of tids that come into the DCC’s compress module. |
| CB_PERF_SEL_CC_DCC_COMPRESS_TIDS_OUT | 265 | The number of tids that come out of the DCC’s compress module. |
| CB_PERF_SEL_FC_DCC_KEY_VALUE__CLEAR | 266 | From DCC cache: Number of times all sectors of a cacheline were initialized with a register-defined clear value. |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__4_BLOCKS__2TO1 | 267 | From DCC Compressor in Color Cache: Number of times all 4 sectors were compressed as 2:1 blocks. |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__3BLOCKS_2TO1__1BLOCK_2TO2 | 268 | From DCC Compressor in Color Cache: Number of times sectors 1 through 3 were compressed as 2:1 blocks and sector 0 was compressed as 2:2 (uncompressed). |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__2BLOCKS_2TO1__1BLOCK_2TO2__1BLOCK_2TO1 | 269 | From DCC Compressor in Color Cache: Number of times sectors 2 and 3 were compressed as 2:1 blocks, sector 1 was compressed as 2:2 (uncompressed) and sector 0 was compressed as 2:1 blocks. |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_2TO1__1BLOCK_2TO2__2BLOCKS_2TO1 | 270 | From DCC Compressor in Color Cache: Number of times sector 3 was compressed as 2:1 blocks, sector 2 was compressed as 2:2 (uncompressed) and sectors 0 and 1 were compressed as 2:1 blocks. |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_2TO2__3BLOCKS_2TO1 | 271 | From DCC Compressor in Color Cache: Number of times sector 3 was compressed as 2:2 (uncompressed) and sectors 0 through 2 were compressed as 2:1 blocks. |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__2BLOCKS_2TO1__2BLOCKS_2TO2 | 272 | From DCC Compressor in Color Cache: Number of times sectors 2 and 3 were compressed as 2:1 blocks and sectors 0 and 1 were compressed as 2:2 (uncompressed). |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_2TO1__2BLOCKS_2TO2__1BLOCK_2TO1 | 273 | From DCC Compressor in Color Cache: Number of times sector 3 was compressed as 2:1 blocks, sectors 1 and 2 were compressed as 2:2 (uncompressed) and sector 0 was compressed as 2:1 blocks. |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_2TO1__1BLOCK_2TO2__1BLOCK_2TO1__1BLOCK_2TO2 | 274 | From DCC Compressor in Color Cache: Number of times sector 3 was compressed as 2:1 blocks, sector 2 was compressed as 2:2 (uncompressed), sector 1 was compressed as 2:1 blocks and sector 0 was compressed as 2:2 (uncompressed). |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_2TO2__1BLOCK_2TO1__1BLOCK_2TO2__1BLOCK_2TO1 | 275 | From DCC Compressor in Color Cache: Number of times sector 3 was compressed as 2:2 (uncompressed), sector 2 was compressed as 2:1 blocks, sector 1 was compressed as 2:2 (uncompressed) and sector 0 was compressed as 2:1 blocks. |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__2BLOCKS_2TO2__2BLOCKS_2TO1 | 276 | From DCC Compressor in Color Cache: Number of times sectors 2 and 3 were compressed as 2:2 (uncompressed) and sectors 0 and 1 were compressed as 2:1 blocks. |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_2TO2__2BLOCKS_2TO1__1BLOCK_2TO2 | 277 | From DCC Compressor in Color Cache: Number of times sector 3 was compressed as 2:2 (uncompressed), sectors 1 and 2 were compressed as 2:1 blocks and sector 0 was compressed as 2:2 (uncompressed). |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_2TO1__3BLOCKS_2TO2 | 278 | From DCC Compressor in Color Cache: Number of times sector 3 was compressed as 2:1 blocks and sectors 0 through 2 were compressed as 2:2 (uncompressed). |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_2TO2__1BLOCK_2TO1__2BLOCKS_2TO2 | 279 | From DCC Compressor in Color Cache: Number of times sector 3 was compressed as 2:2 (uncompressed), sector 2 was compressed as 2:1 blocks and sectors 0 and 1 were compressed as 2:2 (uncompressed). |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__2BLOCKS_2TO2__1BLOCK_2TO1__1BLOCK_2TO2 | 280 | From DCC Compressor in Color Cache: Number of times sectors 2 and 3 were compressed as 2:2 (uncompressed), sector 1 was compressed as 2:1 blocks and sector 0 was compressed as 2:2 (uncompressed). |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__3BLOCKS_2TO2__1BLOCK_2TO1 | 281 | From DCC Compressor in Color Cache: Number of times sectors 1 through 3 were compressed as 2:2 (uncompressed) and sector 0 was compressed as 2:1 blocks. |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__2BLOCKS_4TO1 | 282 | From DCC Compressor in Color Cache: Number of times sectors 0 and 1 were compressed as 4:1 blocks. |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_4TO1__1BLOCK_4TO2 | 283 | From DCC Compressor in Color Cache: Number of times sector 1 was compressed as 4:1 blocks and sector 0 was compressed as 4:2 blocks. |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_4TO1__1BLOCK_4TO3 | 284 | From DCC Compressor in Color Cache: Number of times sector 1 was compressed as 4:1 blocks and sector 0 was compressed as 4:3 blocks. |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_4TO1__1BLOCK_4TO4 | 285 | From DCC Compressor in Color Cache: Number of times sector 1 was compressed as 4:1 blocks and sector 0 was compressed as 4:4 (uncompressed). |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_4TO2__1BLOCK_4TO1 | 286 | From DCC Compressor in Color Cache: Number of times sector 1 was compressed as 4:2 blocks and sector 0 was compressed as 4:1 blocks. |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__2BLOCKS_4TO2 | 287 | From DCC Compressor in Color Cache: Number of times sectors 0 and 1 were compressed as 4:2 blocks. |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_4TO2__1BLOCK_4TO3 | 288 | From DCC Compressor in Color Cache: Number of times sector 1 was compressed as 4:2 blocks and sector 0 was compressed as 4:3 blocks. |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_4TO2__1BLOCK_4TO4 | 289 | From DCC Compressor in Color Cache: Number of times sector 1 was compressed as 4:2 blocks and sector 0 was compressed as 4:4 (uncompressed). |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_4TO3__1BLOCK_4TO1 | 290 | From DCC Compressor in Color Cache: Number of times sector 1 was compressed as 4:3 blocks and sector 0 was compressed as 4:1 blocks. |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_4TO3__1BLOCK_4TO2 | 291 | From DCC Compressor in Color Cache: Number of times sector 1 was compressed as 4:3 blocks and sector 0 was compressed as 4:2 blocks. |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__2BLOCKS_4TO3 | 292 | From DCC Compressor in Color Cache: Number of times sectors 0 & 1 were compressed as 4:3 blocks. |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_4TO3__1BLOCK_4TO4 | 293 | From DCC Compressor in Color Cache: Number of times sector 1 was compressed as 4:3 blocks and sector 0 was compressed as 4:4 (uncompressed). |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_4TO4__1BLOCK_4TO1 | 294 | From DCC Compressor in Color Cache: Number of times sector 1 was compressed as 4:4 (uncompressed) and sector 0 was compressed as 4:1 blocks. |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_4TO4__1BLOCK_4TO2 | 295 | From DCC Compressor in Color Cache: Number of times sector 1 was compressed as 4:4 (uncompressed) and sector 0 was compressed as 4:2 blocks. |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_4TO4__1BLOCK_4TO3 | 296 | From DCC Compressor in Color Cache: Number of times sector 1 was compressed as 4:4 (uncompressed) and sector 0 was compressed as 4:3 blocks. |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__2BLOCKS_2TO1__1BLOCK_4TO1 | 297 | From DCC Compressor in Color Cache: Number of times sectors 1 and 2 were compressed as 2:1 blocks and sector 0 was compressed as 4:1 blocks. |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__2BLOCKS_2TO1__1BLOCK_4TO2 | 298 | From DCC Compressor in Color Cache: Number of times sectors 1 and 2 were compressed as 2:1 blocks and sector 0 was compressed as 4:2 blocks. |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__2BLOCKS_2TO1__1BLOCK_4TO3 | 299 | From DCC Compressor in Color Cache: Number of times sectors 1 and 2 were compressed as 2:1 blocks and sector 0 was compressed as 4:3 blocks. |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__2BLOCKS_2TO1__1BLOCK_4TO4 | 300 | From DCC Compressor in Color Cache: Number of times sectors 1 and 2 were compressed as 2:1 blocks and sector 0 was compressed as 4:4 (uncompressed). |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_2TO1__1BLOCK_2TO2__1BLOCK_4TO1 | 301 | From DCC Compressor in Color Cache: Number of times sector 2 was compressed as 2:1 blocks, sector 1 was compressed as 2:2 (uncompressed) and sector 0 was compressed as 4:1 blocks. |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_2TO1__1BLOCK_2TO2__1BLOCK_4TO2 | 302 | From DCC Compressor in Color Cache: Number of times sector 2 was compressed as 2:1 blocks, sector 1 was compressed as 2:2 (uncompressed) and sector 0 was compressed as 4:2 blocks. |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_2TO1__1BLOCK_2TO2__1BLOCK_4TO3 | 303 | From DCC Compressor in Color Cache: Number of times sector 2 was compressed as 2:1 blocks, sector 1 was compressed as 2:2 (uncompressed) and sector 0 was compressed as 4:3 blocks. |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_2TO1__1BLOCK_2TO2__1BLOCK_4TO4 | 304 | From DCC Compressor in Color Cache: Number of times sector 2 was compressed as 2:1 blocks, sector 1 was compressed as 2:2 (uncompressed) and sector 0 was compressed as 4:4 (uncompressed). |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_2TO2__1BLOCK_2TO1__1BLOCK_4TO1 | 305 | From DCC Compressor in Color Cache: Number of times sector 2 was compressed as 2:2 (uncompressed), sector 1 was compressed as 2:1 blocks and sector 0 was compressed as 4:1 blocks. |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_2TO2__1BLOCK_2TO1__1BLOCK_4TO2 | 306 | From DCC Compressor in Color Cache: Number of times sector 2 was compressed as 2:2 (uncompressed), sector 1 was compressed as 2:1 blocks and sector 0 was compressed as 4:2 blocks. |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_2TO2__1BLOCK_2TO1__1BLOCK_4TO3 | 307 | From DCC Compressor in Color Cache: Number of times sector 2 was compressed as 2:2 (uncompressed), sector 1 was compressed as 2:1 blocks and sector 0 was compressed as 4:3 blocks. |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_2TO2__1BLOCK_2TO1__1BLOCK_4TO4 | 308 | From DCC Compressor in Color Cache: Number of times sector 2 was compressed as 2:2 (uncompressed), sector 1 was compressed as 2:1 blocks and sector 0 was compressed as 4:4 (uncompressed). |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__2BLOCKS_2TO2__1BLOCK_4TO1 | 309 | From DCC Compressor in Color Cache: Number of times sectors 1 and 2 were compressed as 2:2 (uncompressed) and sector 0 was compressed as 4:1 blocks. |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__2BLOCKS_2TO2__1BLOCK_4TO2 | 310 | From DCC Compressor in Color Cache: Number of times sectors 1 and 2 were compressed as 2:2 (uncompressed) and sector 0 was compressed as 4:2 blocks. |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__2BLOCKS_2TO2__1BLOCK_4TO3 | 311 | From DCC Compressor in Color Cache: Number of times sectors 1 and 2 were compressed as 2:2 (uncompressed) and sector 0 was compressed as 4:3 blocks. |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_2TO1__1BLOCK_4TO1__1BLOCK_2TO1 | 312 | From DCC Compressor in Color Cache: Number of times sector 2 was compressed as 2:1 blocks, sector 1 was compressed as 4:1 blocks and sector 0 was compressed as 2:1 blocks. |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_2TO1__1BLOCK_4TO2__1BLOCK_2TO1 | 313 | From DCC Compressor in Color Cache: Number of times sector 2 was compressed as 2:1 blocks, sector 1 was compressed as 4:2 blocks and sector 0 was compressed as 2:1 blocks. |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_2TO1__1BLOCK_4TO3__1BLOCK_2TO1 | 314 | From DCC Compressor in Color Cache: Number of times sector 2 was compressed as 2:1 blocks, sector 1 was compressed as 4:3 blocks and sector 0 was compressed as 2:1 blocks. |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_2TO1__1BLOCK_4TO4__1BLOCK_2TO1 | 315 | From DCC Compressor in Color Cache: Number of times sector 2 was compressed as 2:1 blocks, sector 1 was compressed as 4:4 (uncompressed) and sector 0 was compressed as 2:1 blocks. |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_2TO2__1BLOCK_4TO1__1BLOCK_2TO1 | 316 | From DCC Compressor in Color Cache: Number of times sector 2 was compressed as 2:2 (uncompressed), sector 1 was compressed as 4:1 blocks and sector 0 was compressed as 2:1 blocks. |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_2TO2__1BLOCK_4TO2__1BLOCK_2TO1 | 317 | From DCC Compressor in Color Cache: Number of times sector 2 was compressed as 2:2 (uncompressed), sector 1 was compressed as 4:2 blocks and sector 0 was compressed as 2:1 blocks. |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_2TO2__1BLOCK_4TO3__1BLOCK_2TO1 | 318 | From DCC Compressor in Color Cache: Number of times sector 2 was compressed as 2:2 (uncompressed), sector 1 was compressed as 4:3 blocks and sector 0 was compressed as 2:1 blocks. |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_2TO2__1BLOCK_4TO4__1BLOCK_2TO1 | 319 | From DCC Compressor in Color Cache: Number of times sector 2 was compressed as 2:2 (uncompressed), sector 1 was compressed as 4:4 (uncompressed) and sector 0 was compressed as 2:1 blocks. |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_2TO1__1BLOCK_4TO1__1BLOCK_2TO2 | 320 | From DCC Compressor in Color Cache: Number of times sector 2 was compressed as 2:1 blocks, sector 1 was compressed as 4:1 blocks and sector 0 was compressed as 2:2 (uncompressed). |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_2TO1__1BLOCK_4TO2__1BLOCK_2TO2 | 321 | From DCC Compressor in Color Cache: Number of times sector 2 was compressed as 2:1 blocks, sector 1 was compressed as 4:2 blocks and sector 0 was compressed as 2:2 (uncompressed). |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_2TO1__1BLOCK_4TO3__1BLOCK_2TO2 | 322 | From DCC Compressor in Color Cache: Number of times sector 2 was compressed as 2:1 blocks, sector 1 was compressed as 4:3 blocks and sector 0 was compressed as 2:2 (uncompressed). |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_2TO1__1BLOCK_4TO4__1BLOCK_2TO2 | 323 | From DCC Compressor in Color Cache: Number of times sector 2 was compressed as 2:1 blocks, sector 1 was compressed as 4:4 (uncompressed) and sector 0 was compressed as 2:2 (uncompressed). |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_2TO2__1BLOCK_4TO1__1BLOCK_2TO2 | 324 | From DCC Compressor in Color Cache: Number of times sector 2 was compressed as 2:2 (uncompressed), sector 1 was compressed as 4:1 blocks and sector 0 was compressed as 2:2 (uncompressed). |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_2TO2__1BLOCK_4TO2__1BLOCK_2TO2 | 325 | From DCC Compressor in Color Cache: Number of times sector 2 was compressed as 2:2 (uncompressed), sector 1 was compressed as 4:2 blocks and sector 0 was compressed as 2:2 (uncompressed). |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_2TO2__1BLOCK_4TO3__1BLOCK_2TO2 | 326 | From DCC Compressor in Color Cache: Number of times sector 2 was compressed as 2:2 (uncompressed), sector 1 was compressed as 4:3 blocks and sector 0 was compressed as 2:2 (uncompressed). |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_4TO1__2BLOCKS_2TO1 | 327 | From DCC Compressor in Color Cache: Number of times sector 2 was compressed as 4:1 blocks and sectors 0 and 1 were compressed as 2:1 blocks. |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_4TO2__2BLOCKS_2TO1 | 328 | From DCC Compressor in Color Cache: Number of times sector 2 was compressed as 4:2 blocks and sectors 0 and 1 were compressed as 2:1 blocks. |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_4TO3__2BLOCKS_2TO1 | 329 | From DCC Compressor in Color Cache: Number of times sector 2 was compressed as 4:3 blocks and sectors 0 and 1 were compressed as 2:1 blocks. |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_4TO4__2BLOCKS_2TO1 | 330 | From DCC Compressor in Color Cache: Number of times sector 2 was compressed as 4:4 (uncompressed) and sectors 0 and 1 were compressed as 2:1 blocks. |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_4TO1__2BLOCKS_2TO2 | 331 | From DCC Compressor in Color Cache: Number of times sector 2 was compressed as 4:1 blocks and sectors 0 and 1 were compressed as 2:1 blocks. |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_4TO2__2BLOCKS_2TO2 | 332 | From DCC Compressor in Color Cache: Number of times sector 2 was compressed as 4:1 blocks and sectors 0 and 1 were compressed as 2:1 blocks. |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_4TO3__2BLOCKS_2TO2 | 333 | From DCC Compressor in Color Cache: Number of times sector 2 was compressed as 4:1 blocks and sectors 0 and 1 were compressed as 2:1 blocks. |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_4TO1__1BLOCK_2TO1__1BLOCK_2TO2 | 334 | From DCC Compressor in Color Cache: Number of times sector 2 was compressed as 4:1 blocks, sector 1 was compressed as 2:1 blocks and sector 0 was compressed as 2:2 (uncompressed). |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_4TO2__1BLOCK_2TO1__1BLOCK_2TO2 | 335 | From DCC Compressor in Color Cache: Number of times sector 2 was compressed as 4:2 blocks, sector 1 was compressed as 2:1 blocks and sector 0 was compressed as 2:2 (uncompressed). |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_4TO3__1BLOCK_2TO1__1BLOCK_2TO2 | 336 | From DCC Compressor in Color Cache: Number of times sector 2 was compressed as 4:3 blocks, sector 1 was compressed as 2:1 blocks and sector 0 was compressed as 2:2 (uncompressed). |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_4TO4__1BLOCK_2TO1__1BLOCK_2TO2 | 337 | From DCC Compressor in Color Cache: Number of times sector 2 was compressed as 4:4 (uncompressed), sector 1 was compressed as 2:1 blocks and sector 0 was compressed as 2:2 (uncompressed). |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_4TO1__1BLOCK_2TO2__1BLOCK_2TO1 | 338 | From DCC Compressor in Color Cache: Number of times sector 2 was compressed as 4:1 blocks, sector 1 was compressed as 2:2 (uncompressed) and sector 0 was compressed as 2:1 blocks. |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_4TO2__1BLOCK_2TO2__1BLOCK_2TO1 | 339 | From DCC Compressor in Color Cache: Number of times sector 2 was compressed as 4:2 blocks, sector 1 was compressed as 2:2 (uncompressed) and sector 0 was compressed as 2:1 blocks. |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_4TO3__1BLOCK_2TO2__1BLOCK_2TO1 | 340 | From DCC Compressor in Color Cache: Number of times sector 2 was compressed as 4:3 blocks, sector 1 was compressed as 2:2 (uncompressed) and sector 0 was compressed as 2:1 blocks. |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_4TO4__1BLOCK_2TO2__1BLOCK_2TO1 | 341 | From DCC Compressor in Color Cache: Number of times sector 2 was compressed as 4:4 (uncompressed), sector 1 was compressed as 2:2 (uncompressed) and sector 0 was compressed as 2:1 blocks. |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_2TO1__1BLOCK_6TO1 | 342 | From DCC Compressor in Color Cache: Number of times sector 1 was compressed as 2:1 blocks and sector 0 was compressed as 6:1 blocks. |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_2TO1__1BLOCK_6TO2 | 343 | From DCC Compressor in Color Cache: Number of times sector 1 was compressed as 2:1 blocks and sector 0 was compressed as 6:2 blocks. |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_2TO1__1BLOCK_6TO3 | 344 | From DCC Compressor in Color Cache: Number of times sector 1 was compressed as 2:1 blocks and sector 0 was compressed as 6:3 blocks. |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_2TO1__1BLOCK_6TO4 | 345 | From DCC Compressor in Color Cache: Number of times sector 1 was compressed as 2:1 blocks and sector 0 was compressed as 6:4 blocks. |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_2TO1__1BLOCK_6TO5 | 346 | From DCC Compressor in Color Cache: Number of times sector 1 was compressed as 2:1 blocks and sector 0 was compressed as 6:5 blocks. |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_2TO1__1BLOCK_6TO6 | 347 | From DCC Compressor in Color Cache: Number of times sector 1 was compressed as 2:1 blocks and sector 0 was compressed as 6:6 (uncompressed). |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_2TO1__INV0 | 348 | From DCC Compressor in Color Cache: Number of times sector 1 was compressed as 2:1 blocks and the key for the 6:x compression blocks is invalid (top 3 bits of the key are 3’b110). |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_2TO1__INV1 | 349 | From DCC Compressor in Color Cache: Number of times sector 1 was compressed as 2:1 blocks and the key for the 6:x compression blocks is invalid (top 3 bits of the key are 3’b111). |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_2TO2__1BLOCK_6TO1 | 350 | From DCC Compressor in Color Cache: Number of times sector 1 was compressed as 2:2 (uncompressed) and sector 0 was compressed as 6:1 blocks. |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_2TO2__1BLOCK_6TO2 | 351 | From DCC Compressor in Color Cache: Number of times sector 1 was compressed as 2:2 (uncompressed) and sector 0 was compressed as 6:2 blocks. |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_2TO2__1BLOCK_6TO3 | 352 | From DCC Compressor in Color Cache: Number of times sector 1 was compressed as 2:2 (uncompressed) and sector 0 was compressed as 6:3 blocks. |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_2TO2__1BLOCK_6TO4 | 353 | From DCC Compressor in Color Cache: Number of times sector 1 was compressed as 2:2 (uncompressed) and sector 0 was compressed as 6:4 blocks. |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_2TO2__1BLOCK_6TO5 | 354 | From DCC Compressor in Color Cache: Number of times sector 1 was compressed as 2:2 (uncompressed) and sector 0 was compressed as 6:5 blocks. |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_2TO2__INV0 | 355 | From DCC Compressor in Color Cache: Number of times sector 1 was compressed as 2:2 (uncompressed) and the key for the 6:x compression blocks is invalid (top 3 bits of the key are 3’b110). |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_2TO2__INV1 | 356 | From DCC Compressor in Color Cache: Number of times sector 1 was compressed as 2:2 (uncompressed) and the key for the 6:x compression blocks is invalid (top 3 bits of the key are 3’b111). |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_6TO1__1BLOCK_2TO1 | 357 | From DCC Compressor in Color Cache: Number of times sector 1 was compressed as 6:1 blocks and sector 0 was compressed as 2:1 blocks. |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_6TO2__1BLOCK_2TO1 | 358 | From DCC Compressor in Color Cache: Number of times sector 1 was compressed as 6:2 blocks and sector 0 was compressed as 2:1 blocks. |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_6TO3__1BLOCK_2TO1 | 359 | From DCC Compressor in Color Cache: Number of times sector 1 was compressed as 6:3 blocks and sector 0 was compressed as 2:1 blocks. |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_6TO4__1BLOCK_2TO1 | 360 | From DCC Compressor in Color Cache: Number of times sector 1 was compressed as 6:4 blocks and sector 0 was compressed as 2:1 blocks. |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_6TO5__1BLOCK_2TO1 | 361 | From DCC Compressor in Color Cache: Number of times sector 1 was compressed as 6:5 blocks and sector 0 was compressed as 2:1 blocks. |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_6TO6__1BLOCK_2TO1 | 362 | From DCC Compressor in Color Cache: Number of times sector 1 was compressed as 6:6 (uncompressed) and sector 0 was compressed as 2:1 blocks. |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__INV0__1BLOCK_2TO1 | 363 | From DCC Compressor in Color Cache: Number of times sector 0 was compressed as 2:1 blocks and the key for the 6:x compression blocks is invalid (top 3 bits of the key are 3’b110). |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__INV1__1BLOCK_2TO1 | 364 | From DCC Compressor in Color Cache: Number of times sector 0 was compressed as 2:1 blocks and the key for the 6:x compression blocks is invalid (top 3 bits of the key are 3’b111). |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_6TO1__1BLOCK_2TO2 | 365 | From DCC Compressor in Color Cache: Number of times sector 1 was compressed as 6:1 blocks and sector 0 was compressed as 2:2 (uncompressed). |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_6TO2__1BLOCK_2TO2 | 366 | From DCC Compressor in Color Cache: Number of times sector 1 was compressed as 6:2 blocks and sector 0 was compressed as 2:2 (uncompressed). |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_6TO3__1BLOCK_2TO2 | 367 | From DCC Compressor in Color Cache: Number of times sector 1 was compressed as 6:3 blocks and sector 0 was compressed as 2:2 (uncompressed). |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_6TO4__1BLOCK_2TO2 | 368 | From DCC Compressor in Color Cache: Number of times sector 1 was compressed as 6:4 blocks and sector 0 was compressed as 2:2 (uncompressed). |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_6TO5__1BLOCK_2TO2 | 369 | From DCC Compressor in Color Cache: Number of times sector 1 was compressed as 6:5 blocks and sector 0 was compressed as 2:2 (uncompressed). |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__INV0__1BLOCK_2TO2 | 370 | From DCC Compressor in Color Cache: Number of times sector 0 was compressed as 2:2 (uncompressed) and the key for the 6:x compression blocks is invalid (top 3 bits of the key are 3’b110). |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__INV1__1BLOCK_2TO2 | 371 | From DCC Compressor in Color Cache: Number of times sector 0 was compressed as 2:2 (uncompressed) and the key for the 6:x compression blocks is invalid (top 3 bits of the key are 3’b111). |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_8TO1 | 372 | From DCC Compressor in Color Cache: Number of times sector 0 was compressed as 8:1 blocks. |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_8TO2 | 373 | From DCC Compressor in Color Cache: Number of times sector 0 was compressed as 8:2 blocks. |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_8TO3 | 374 | From DCC Compressor in Color Cache: Number of times sector 0 was compressed as 8:3 blocks. |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_8TO4 | 375 | From DCC Compressor in Color Cache: Number of times sector 0 was compressed as 8:4 blocks. |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_8TO5 | 376 | From DCC Compressor in Color Cache: Number of times sector 0 was compressed as 8:5 blocks. |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_8TO6 | 377 | From DCC Compressor in Color Cache: Number of times sector 0 was compressed as 8:6 blocks. |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__1BLOCK_8TO7 | 378 | From DCC Compressor in Color Cache: Number of times sector 0 was compressed as 8:7 blocks. |
| CB_PERF_SEL_CC_DCC_KEY_VALUE__UNCOMPRESSED | 379 | From DCC Compressor in Color Cache: Number of times all sectors of a cacheline were uncompressed with a 2:2 ratio. |
| CB_PERF_SEL_CC_DCC_COMPRESS_RATIO_2TO1 | 380 | From DCC Compressor in Color Cache: Number of times a sector was compressed as 2:1. |
| CB_PERF_SEL_CC_DCC_COMPRESS_RATIO_4TO1 | 381 | From DCC Compressor in Color Cache: Number of times a sector was compressed as 4:1. |
| CB_PERF_SEL_CC_DCC_COMPRESS_RATIO_4TO2 | 382 | From DCC Compressor in Color Cache: Number of times a sector was compressed as 4:2. |
| CB_PERF_SEL_CC_DCC_COMPRESS_RATIO_4TO3 | 383 | From DCC Compressor in Color Cache: Number of times a sector was compressed as 4:3. |
| CB_PERF_SEL_CC_DCC_COMPRESS_RATIO_6TO1 | 384 | From DCC Compressor in Color Cache: Number of times a sector was compressed as 6:1. |
| CB_PERF_SEL_CC_DCC_COMPRESS_RATIO_6TO2 | 385 | From DCC Compressor in Color Cache: Number of times a sector was compressed as 6:2. |
| CB_PERF_SEL_CC_DCC_COMPRESS_RATIO_6TO3 | 386 | From DCC Compressor in Color Cache: Number of times a sector was compressed as 6:3. |
| CB_PERF_SEL_CC_DCC_COMPRESS_RATIO_6TO4 | 387 | From DCC Compressor in Color Cache: Number of times a sector was compressed as 6:4. |
| CB_PERF_SEL_CC_DCC_COMPRESS_RATIO_6TO5 | 388 | From DCC Compressor in Color Cache: Number of times a sector was compressed as 6:5. |
| CB_PERF_SEL_CC_DCC_COMPRESS_RATIO_8TO1 | 389 | From DCC Compressor in Color Cache: Number of times a sector was compressed as 8:1. |
| CB_PERF_SEL_CC_DCC_COMPRESS_RATIO_8TO2 | 390 | From DCC Compressor in Color Cache: Number of times a sector was compressed as 8:2. |
| CB_PERF_SEL_CC_DCC_COMPRESS_RATIO_8TO3 | 391 | From DCC Compressor in Color Cache: Number of times a sector was compressed as 8:3. |
| CB_PERF_SEL_CC_DCC_COMPRESS_RATIO_8TO4 | 392 | From DCC Compressor in Color Cache: Number of times a sector was compressed as 8:4. |
| CB_PERF_SEL_CC_DCC_COMPRESS_RATIO_8TO5 | 393 | From DCC Compressor in Color Cache: Number of times a sector was compressed as 8:5. |
| CB_PERF_SEL_CC_DCC_COMPRESS_RATIO_8TO6 | 394 | From DCC Compressor in Color Cache: Number of times a sector was compressed as 8:6. |
| CB_PERF_SEL_CC_DCC_COMPRESS_RATIO_8TO7 | 395 | From DCC Compressor in Color Cache: Number of times a sector was compressed as 8:7. |
| CB_PERF_SEL_CC_BB_BLEND_PIXEL_VLD | 396 | The number of transfers into CB’s blender stage. |
| CB_PERF_SEL_DB_CB_CONTEXT_DONE | 397 | The number of times the DB_CB_context_done signal, from DB, was asserted. |
| CB_PERF_SEL_DB_CB_EOP_DONE | 398 | The number of times the DB_CB_eop_done signal, from DB, was asserted. |
| CB_PERF_SEL_CC_MC_WRITE_REQUEST_PARTIAL | 399 | The number of color to MC write requests with partial write mask. (Note: CMask, Fmask and DCC are always full write mask writes). |
| CB_PERF_SEL_EVENT_BOTTOM_OF_PIPE_TS | 400 | The number of BOTTOM_OF_PIPE_TS events. |
| CB_PERF_SEL_EVENT_FLUSH_AND_INV_DB_DATA_TS | 401 | The number of FLUSH_AND_INV_DB_DATA_TS events. |
| CB_PERF_SEL_EVENT_FLUSH_AND_INV_CB_PIXEL_DATA | 402 | The number of FLUSH_AND_INV_CB_PIXEL_DATA events. |
| CB_PERF_SEL_DB_CB_TILE_TILENOTEVENT | 403 | The number of tiles (not events) received from DB. |
| CB_PERF_SEL_MERGE_PIXELS_WITH_BLEND_ENABLED | 404 | The number of pixels, after tiles and quads are merged, that have blending enabled. |
| Counter | Value | Description |
|---|---|---|
| DB_PERF_SEL_SC_DB_TILE_SENDS | 0 | Cycles Interface is sending |
| DB_PERF_SEL_SC_DB_TILE_BUSY | 1 | Cycles Interface is busy |
| DB_PERF_SEL_SC_DB_TILE_STALLS | 2 | Cycles Interface is stalled |
| DB_PERF_SEL_SC_DB_TILE_EVENTS | 3 | Events sent over interface |
| DB_PERF_SEL_SC_DB_TILE_TILES | 4 | Tiles sent over interface |
| DB_PERF_SEL_SC_DB_TILE_COVERED | 5 | Fully covered tiles |
| DB_PERF_SEL_HIZ_TC_READ_STARVED | 6 | HiZ starved waiting for htile data from cache |
| DB_PERF_SEL_HIZ_TC_WRITE_STALL | 7 | HiZ stalled writing to htile cache |
| DB_PERF_SEL_HIZ_QTILES_CULLED | 8 | Quarter tiles culled by hiZ |
| DB_PERF_SEL_HIS_QTILES_CULLED | 9 | Quarter tiles culled by HiS |
| DB_PERF_SEL_DB_SC_TILE_SENDS | 10 | Cycles Interface is sending |
| DB_PERF_SEL_DB_SC_TILE_BUSY | 11 | Cycles Interface is busy |
| DB_PERF_SEL_DB_SC_TILE_STALLS | 12 | Cycles Interface is stalled by SC |
| DB_PERF_SEL_DB_SC_TILE_DF_STALLS | 13 | Cycles Interface is stalled by detail walk tile FIFO |
| DB_PERF_SEL_DB_SC_TILE_TILES | 14 | Tiles sent over interface |
| DB_PERF_SEL_DB_SC_TILE_CULLED | 15 | Tiles not killed, but handled entirely at the Hi-Z stage, not returned to the SC |
| DB_PERF_SEL_DB_SC_TILE_HIER_KILL | 16 | Tiles culled due to a hierarchical fail test |
| DB_PERF_SEL_DB_SC_TILE_FAST_OPS | 17 | Tiles culled because they were accelerated fast tile ops. |
| DB_PERF_SEL_DB_SC_TILE_NO_OPS | 18 | Tiles culled because they would not do anything |
| DB_PERF_SEL_DB_SC_TILE_TILE_RATE | 19 | Tiles run at tile rate as opposed to sample rate tile |
| DB_PERF_SEL_DB_SC_TILE_SSAA_KILL | 20 | Tiles culled because they were supersample tiles that are merged into a fast tile op |
| DB_PERF_SEL_DB_SC_TILE_FAST_Z_OPS | 21 | Tiles that operate on Z in 1 clock. Can be inside a slow, pixel rate, or fast tile op |
| DB_PERF_SEL_DB_SC_TILE_FAST_STENCIL_OPS | 22 | Tiles that operate on stencil in 1 clock. Can be inside a slow, pixel rate, or fast tile op |
| DB_PERF_SEL_SC_DB_QUAD_SENDS | 23 | Cycles Interface is sending |
| DB_PERF_SEL_SC_DB_QUAD_BUSY | 24 | Cycles Interface is busy |
| DB_PERF_SEL_SC_DB_QUAD_SQUADS | 25 | Squads transferred over interface |
| DB_PERF_SEL_SC_DB_QUAD_TILES | 26 | Tiles sent over interface |
| DB_PERF_SEL_SC_DB_QUAD_PIXELS | 27 | Pixels transfered over interface |
| DB_PERF_SEL_SC_DB_QUAD_KILLED_TILES | 28 | Number of detail killed tiles |
| DB_PERF_SEL_DB_SC_QUAD_SENDS | 29 | Cycles Interface is sending |
| DB_PERF_SEL_DB_SC_QUAD_BUSY | 30 | Cycles Interface is busy |
| DB_PERF_SEL_DB_SC_QUAD_STALLS | 31 | Cycles Interface is stalled |
| DB_PERF_SEL_DB_SC_QUAD_TILES | 32 | Tiles sent over interface |
| DB_PERF_SEL_DB_SC_QUAD_LIT_QUAD | 33 | Quads transferred over the DB_SC_quad interface that are lit |
| DB_PERF_SEL_DB_CB_TILE_SENDS | 34 | Cycles sending tiles/events to CB. Tiles not shaded or needed by the CB are not sent. |
| DB_PERF_SEL_DB_CB_TILE_BUSY | 35 | |
| DB_PERF_SEL_DB_CB_TILE_STALLS | 36 | |
| DB_PERF_SEL_SX_DB_QUAD_SENDS | 37 | Cycles Interface is sending |
| DB_PERF_SEL_SX_DB_QUAD_BUSY | 38 | Cycles Interface is busy |
| DB_PERF_SEL_SX_DB_QUAD_STALLS | 39 | Cycles Interface is stalled |
| DB_PERF_SEL_SX_DB_QUAD_QUADS | 40 | Quads sent over interface |
| DB_PERF_SEL_SX_DB_QUAD_PIXELS | 41 | Pixels sent over interface |
| DB_PERF_SEL_SX_DB_QUAD_EXPORTS | 42 | Each MRT of a quad sent over interface |
| DB_PERF_SEL_SH_QUADS_OUTSTANDING_SUM | 43 | Multiply by 128 and divide by DB_SC_quad_quads to get PS latency. |
| DB_PERF_SEL_DB_CB_LQUAD_SENDS | 44 | Cycles Interface is sending |
| DB_PERF_SEL_DB_CB_LQUAD_BUSY | 45 | Cycles Interface is busy |
| DB_PERF_SEL_DB_CB_LQUAD_STALLS | 46 | Cycles Interface is stalled |
| DB_PERF_SEL_DB_CB_LQUAD_QUADS | 47 | Quads sent over interface |
| DB_PERF_SEL_TILE_RD_SENDS | 48 | HTile reads. Each is 256B |
| DB_PERF_SEL_MI_TILE_RD_OUTSTANDING_SUM | 49 | Multiply by 16 and divide by tile_rd_sends*8 to get htile memory latency |
| DB_PERF_SEL_QUAD_RD_SENDS | 50 | Quad read reqs. Each is 32B to 256B |
| DB_PERF_SEL_QUAD_RD_BUSY | 51 | Cycles quad read interface is trying to send requests |
| DB_PERF_SEL_QUAD_RD_MI_STALL | 52 | Cycles quad read interface is stalled by the memory interface |
| DB_PERF_SEL_QUAD_RD_RW_COLLISION | 53 | Cycles a quad read is stalled waiting for a write to finish |
| DB_PERF_SEL_QUAD_RD_TAG_STALL | 54 | Cycles a quad read is stalled because the read latency hiding FIFO is full. |
| DB_PERF_SEL_QUAD_RD_32BYTE_REQS | 55 | Number of 32 Byte quad read requests |
| DB_PERF_SEL_QUAD_RD_PANIC | 56 | Cycles DB is panicking for quad read data |
| DB_PERF_SEL_MI_QUAD_RD_OUTSTANDING_SUM | 57 | Multiply by 16 and divide by quad_rd_32byte_reqs to get depth buffer memory latency |
| DB_PERF_SEL_QUAD_RDRET_SENDS | 58 | Number of 32 byte quad read returns |
| DB_PERF_SEL_QUAD_RDRET_BUSY | 59 | Cycles the quad read data is returning |
| DB_PERF_SEL_TILE_WR_SENDS | 60 | 32 Byte HTile writes |
| DB_PERF_SEL_TILE_WR_ACKS | 61 | 32 Byte Htile write acks |
| DB_PERF_SEL_MI_TILE_WR_OUTSTANDING_SUM | 62 | Multiply by 16 and divide by tile_wr_sends to get tile write memory latency |
| DB_PERF_SEL_QUAD_WR_SENDS | 63 | Cycles quad is sending write requests to the memory interface block of the DB |
| DB_PERF_SEL_QUAD_WR_BUSY | 64 | Cycles quad is trying to write to the memory interface |
| DB_PERF_SEL_QUAD_WR_MI_STALL | 65 | Cycles quad is stalled while writing to the memory interface |
| DB_PERF_SEL_QUAD_WR_COHERENCY_STALL | 66 | Cycles quad write is stalled waiting for a previous write to finish on the same address |
| DB_PERF_SEL_QUAD_WR_ACKS | 67 | Number of 32 Byte quad write acks |
| DB_PERF_SEL_MI_QUAD_WR_OUTSTANDING_SUM | 68 | Multiply by 16 and divide by the quad_wr_sends to get quad memory write latency |
| DB_PERF_SEL_TILE_CACHE_MISSES | 69 | Htile Cache misses |
| DB_PERF_SEL_TILE_CACHE_HITS | 70 | Htile Cache hits |
| DB_PERF_SEL_TILE_CACHE_FLUSHES | 71 | Htile Cache flushes |
| DB_PERF_SEL_TILE_CACHE_SURFACE_STALL | 72 | Tile stalls waiting for an htile surface to flush and free |
| DB_PERF_SEL_TILE_CACHE_STARVES | 73 | Tile stalls waiting for an htile cache line to free |
| DB_PERF_SEL_TILE_CACHE_MEM_RETURN_STARVE | 74 | Tile stalls waiting for an htile memory read to return |
| DB_PERF_SEL_TCP_DISPATCHER_READS | 75 | Number of htile cache lines fetched by the normal tile stream |
| DB_PERF_SEL_TCP_PREFETCHER_READS | 76 | Number of htile cache lines fetched by the prefetcher |
| DB_PERF_SEL_TCP_PRELOADER_READS | 77 | Number of htile cache lines fetched by the preloader |
| DB_PERF_SEL_TCP_DISPATCHER_FLUSHES | 78 | Number of htile flushes caused byt the normal tile stream |
| DB_PERF_SEL_TCP_PREFETCHER_FLUSHES | 79 | Number of htile flushes caused byt the normal prefetcher |
| DB_PERF_SEL_TCP_PRELOADER_FLUSHES | 80 | Number of htile flushes caused byt the normal preloader |
| DB_PERF_SEL_DEPTH_TILE_CACHE_SENDS | 81 | Tiles/Events through the Depth Tile Cache |
| DB_PERF_SEL_DEPTH_TILE_CACHE_BUSY | 82 | Cycles the Depth Tile Cache is busy testing hit/miss on tiles/events |
| DB_PERF_SEL_DEPTH_TILE_CACHE_STARVES | 83 | Cycles starved waiting for a depth surface tile tag to free up |
| DB_PERF_SEL_DEPTH_TILE_CACHE_DTILE_LOCKED | 84 | Cycles stalled waiting for a forced flush or invalidate to finish |
| DB_PERF_SEL_DEPTH_TILE_CACHE_ALLOC_STALL | 85 | Cycles depth tile stalled while allocating data |
| DB_PERF_SEL_DEPTH_TILE_CACHE_MISSES | 86 | Depth/Stencil Cache tile misses |
| DB_PERF_SEL_DEPTH_TILE_CACHE_HITS | 87 | Depth/Stencil Cache tile hits |
| DB_PERF_SEL_DEPTH_TILE_CACHE_FLUSHES | 88 | Depth/Stencil Cache tile flushes |
| DB_PERF_SEL_DEPTH_TILE_CACHE_NOOP_TILE | 89 | No-op tiles through the depth tile cache |
| DB_PERF_SEL_DEPTH_TILE_CACHE_DETAILED_NOOP | 90 | Detail walked no-ops though the depth tile cache |
| DB_PERF_SEL_DEPTH_TILE_CACHE_EVENT | 91 | Events |
| DB_PERF_SEL_DEPTH_TILE_CACHE_TILE_FREES | 92 | Depth tile frees |
| DB_PERF_SEL_DEPTH_TILE_CACHE_DATA_FREES | 93 | Depth tile data frees (may free data without freeing the dtile) |
| DB_PERF_SEL_DEPTH_TILE_CACHE_MEM_RETURN_STARVE | 94 | Cycles depth/stencil cache is waiting for memory to return |
| DB_PERF_SEL_STENCIL_CACHE_MISSES | 95 | 512 bit allocation |
| DB_PERF_SEL_STENCIL_CACHE_HITS | 96 | |
| DB_PERF_SEL_STENCIL_CACHE_FLUSHES | 97 | 256 bit flushes. Two per cache line unless it is compressed, in which case it is 1. |
| DB_PERF_SEL_STENCIL_CACHE_STARVES | 98 | Cache starves when the cache wants to miss but is out of cache lines to allocate |
| DB_PERF_SEL_STENCIL_CACHE_FREES | 99 | |
| DB_PERF_SEL_Z_CACHE_SEPARATE_Z_MISSES | 100 | |
| DB_PERF_SEL_Z_CACHE_SEPARATE_Z_HITS | 101 | |
| DB_PERF_SEL_Z_CACHE_SEPARATE_Z_FLUSHES | 102 | |
| DB_PERF_SEL_Z_CACHE_SEPARATE_Z_STARVES | 103 | Cache starves when the cache wants to miss but is out of cache lines to allocate |
| DB_PERF_SEL_Z_CACHE_PMASK_MISSES | 104 | |
| DB_PERF_SEL_Z_CACHE_PMASK_HITS | 105 | |
| DB_PERF_SEL_Z_CACHE_PMASK_FLUSHES | 106 | |
| DB_PERF_SEL_Z_CACHE_PMASK_STARVES | 107 | Cache starves when the cache wants to miss but is out of cache lines to allocate |
| DB_PERF_SEL_Z_CACHE_FREES | 108 | |
| DB_PERF_SEL_PLANE_CACHE_MISSES | 109 | |
| DB_PERF_SEL_PLANE_CACHE_HITS | 110 | |
| DB_PERF_SEL_PLANE_CACHE_FLUSHES | 111 | |
| DB_PERF_SEL_PLANE_CACHE_STARVES | 112 | Cache starves when the cache wants to miss but is out of cache lines to allocate |
| DB_PERF_SEL_PLANE_CACHE_FREES | 113 | |
| DB_PERF_SEL_FLUSH_EXPANDED_STENCIL | 114 | Tiles flushed with expanded stencil |
| DB_PERF_SEL_FLUSH_COMPRESSED_STENCIL | 115 | Tiles flushed with compressed stencil |
| DB_PERF_SEL_FLUSH_SINGLE_STENCIL | 116 | Tiles flushed with single stencil |
| DB_PERF_SEL_PLANES_FLUSHED | 117 | Total planes flushed among all tiles |
| DB_PERF_SEL_FLUSH_1PLANE | 118 | Tiles flushed with 1 ZPlane |
| DB_PERF_SEL_FLUSH_2PLANE | 119 | Tiles flushed with 2 ZPlanes |
| DB_PERF_SEL_FLUSH_3PLANE | 120 | Tiles flushed with 3 ZPlanes |
| DB_PERF_SEL_FLUSH_4PLANE | 121 | Tiles flushed with 4 ZPlanes |
| DB_PERF_SEL_FLUSH_5PLANE | 122 | Tiles flushed with 5 ZPlanes |
| DB_PERF_SEL_FLUSH_6PLANE | 123 | Tiles flushed with 6 ZPlanes |
| DB_PERF_SEL_FLUSH_7PLANE | 124 | Tiles flushed with 7 ZPlanes |
| DB_PERF_SEL_FLUSH_8PLANE | 125 | Tiles flushed with 8 ZPlanes |
| DB_PERF_SEL_FLUSH_9PLANE | 126 | Tiles flushed with 9 ZPlanes |
| DB_PERF_SEL_FLUSH_10PLANE | 127 | Tiles flushed with 10 ZPlanes |
| DB_PERF_SEL_FLUSH_11PLANE | 128 | Tiles flushed with 11 ZPlanes |
| DB_PERF_SEL_FLUSH_12PLANE | 129 | Tiles flushed with 12 ZPlanes |
| DB_PERF_SEL_FLUSH_13PLANE | 130 | Tiles flushed with 13 ZPlanes |
| DB_PERF_SEL_FLUSH_14PLANE | 131 | Tiles flushed with 14 ZPlanes |
| DB_PERF_SEL_FLUSH_15PLANE | 132 | Tiles flushed with 15 ZPlanes |
| DB_PERF_SEL_FLUSH_16PLANE | 133 | Tiles flushed with 16 ZPlanes |
| DB_PERF_SEL_FLUSH_EXPANDED_Z | 134 | Tiles flushed with expanded Z |
| DB_PERF_SEL_EARLYZ_WAITING_FOR_POSTZ_DONE | 135 | Cycles stalled while transitioning from Late/ReZ to EarlyZ |
| DB_PERF_SEL_REZ_WAITING_FOR_POSTZ_DONE | 136 | Cycles stalled while transitioning to ReZ from an incompatible Z mode/func/etc. |
| DB_PERF_SEL_DK_TILE_SENDS | 137 | Detail kill block tiles/squads/events sent |
| DB_PERF_SEL_DK_TILE_BUSY | 138 | Cycles Detail Kill block busy |
| DB_PERF_SEL_DK_TILE_QUAD_STARVES | 139 | Cycles Detail Kill has tile, but no quads from the SC_DB_quad |
| DB_PERF_SEL_DK_TILE_STALLS | 140 | Cycles Detail Kill is stalled from below |
| DB_PERF_SEL_DK_SQUAD_SENDS | 141 | Detail Kill squads |
| DB_PERF_SEL_DK_SQUAD_BUSY | 142 | Cycles the squad Detail Kill input is busy |
| DB_PERF_SEL_DK_SQUAD_STALLS | 143 | Cycles squads are stalled from below. |
| DB_PERF_SEL_OP_PIPE_BUSY | 144 | Cycles the quad OP pipe of the DB is busy (including memory fetches, but not initial startup) |
| DB_PERF_SEL_OP_PIPE_MC_READ_STALL | 145 | Cycles the Op Pipe is waiting for memory to return (ignoring initial startup) |
| DB_PERF_SEL_QC_BUSY | 146 | Cycles the quad coherency is busy |
| DB_PERF_SEL_QC_XFC | 147 | Squads going through the quad coherency block |
| DB_PERF_SEL_QC_CONFLICTS | 148 | Stalls on a squad input because of a coherency check conflict (squads ahead of it could affect the dest data) |
| DB_PERF_SEL_QC_FULL_STALL | 149 | Quad Coherency is stalled because it is out of slots. |
| DB_PERF_SEL_QC_IN_PREZ_TILE_STALLS_POSTZ | 150 | Cycles postZ squad inputs are stalled because the OP pipe is operating on a preZ tile, and there is no PreZ squad input available |
| DB_PERF_SEL_QC_IN_POSTZ_TILE_STALLS_PREZ | 151 | Cycles preZ squad inputs are stalled because the OP pipe is operating on a postZ tile, and there is no PostZ squad input available |
| DB_PERF_SEL_TSC_INSERT_SUMMARIZE_STALL | 152 | Cycles the op pipe is stalled by the tile summarizer inserting summarize squads |
| DB_PERF_SEL_TL_BUSY | 153 | Cycles Tile Lookup block busy |
| DB_PERF_SEL_TL_DTC_READ_STARVED | 154 | Cycles Tile Lookup block |
| DB_PERF_SEL_TL_Z_FETCH_STALL | 155 | Cycles Tile Lookup block stalled by zfetch block |
| DB_PERF_SEL_TL_STENCIL_STALL | 156 | Cycles Tile Lookup block stalled by stencil test block |
| DB_PERF_SEL_TL_Z_DECOMPRESS_STALL | 157 | Cycles Tile Lookup block stalled by z decompress block |
| DB_PERF_SEL_TL_STENCIL_LOCKED_STALL | 158 | Cycles Tile Lookup block stalled by stencil being locked |
| DB_PERF_SEL_TL_EVENTS | 159 | Cycles Tile Lookup block sends out events |
| DB_PERF_SEL_TL_SUMMARIZE_SQUADS | 160 | Cycles Tile Lookup block sends out summarize squads |
| DB_PERF_SEL_TL_FLUSH_EXPAND_SQUADS | 161 | Cycles Tile Lookup block sends out flush expand squads |
| DB_PERF_SEL_TL_EXPAND_SQUADS | 162 | Cycles Tile Lookup block sends out expand squads |
| DB_PERF_SEL_TL_PREZ_SQUADS | 163 | Cycles Tile Lookup block sends out preZ squads |
| DB_PERF_SEL_TL_POSTZ_SQUADS | 164 | Cycles Tile Lookup block sends out postZ squads |
| DB_PERF_SEL_TL_PREZ_NOOP_SQUADS | 165 | Cycles Tile Lookup block sends out no-ops on the preZ path |
| DB_PERF_SEL_TL_POSTZ_NOOP_SQUADS | 166 | Cycles Tile Lookup block sends out no-ops on the postZ path |
| DB_PERF_SEL_TL_TILE_OPS | 167 | Cycles Tile Lookup block sends out tile op squads |
| DB_PERF_SEL_TL_IN_XFC | 168 | Cycles Tile Lookup block receives any input |
| DB_PERF_SEL_TL_IN_SINGLE_STENCIL_EXPAND_STALL | 169 | Cycles Tile Lookup block is stalling due to expanding single stencil |
| DB_PERF_SEL_TL_IN_FAST_Z_STALL | 170 | Cycles Tile Lookup block is stalling due to fastZ |
| DB_PERF_SEL_TL_OUT_XFC | 171 | Cycles Tile Lookup block is sending anything out |
| DB_PERF_SEL_TL_OUT_SQUADS | 172 | Cycles Tile Lookup block is sending squads out |
| DB_PERF_SEL_ZF_PLANE_MULTICYCLE | 173 | Cycles the z fetch is multi-cycling a squad because it needs to fetch more than 2 planes from the cache for compressed Z |
| DB_PERF_SEL_POSTZ_SAMPLES_PASSING_Z | 174 | Samples passing Z test during a PostZ pass |
| DB_PERF_SEL_POSTZ_SAMPLES_FAILING_Z | 175 | Samples failing Z test during a PostZ pass |
| DB_PERF_SEL_POSTZ_SAMPLES_FAILING_S | 176 | Samples failing Stencil test during a PostZ pass |
| DB_PERF_SEL_PREZ_SAMPLES_PASSING_Z | 177 | Samples passing Z test during a PreZ pass |
| DB_PERF_SEL_PREZ_SAMPLES_FAILING_Z | 178 | Samples failing Z test during a PreZ pass |
| DB_PERF_SEL_PREZ_SAMPLES_FAILING_S | 179 | Samples failing Stencil test during a PreZ pass |
| DB_PERF_SEL_TS_TC_UPDATE_STALL | 180 | Cycles Tile Summarizer to Tile Cache write interface is stalled |
| DB_PERF_SEL_SC_KICK_START | 181 | Times the DB sent a hang panic to the SC |
| DB_PERF_SEL_SC_KICK_END | 182 | Times the DB completed a hang panic |
| DB_PERF_SEL_CLOCK_REG_ACTIVE | 183 | Cycles register part of DB is awake |
| DB_PERF_SEL_CLOCK_MAIN_ACTIVE | 184 | Cycles core part of DB is awake |
| DB_PERF_SEL_CLOCK_MEM_EXPORT_ACTIVE | 185 | Cycles mem export part of DB is awake |
| DB_PERF_SEL_ESR_PS_OUT_BUSY | 186 | Early Squad Router PS Iter Out cycles busy |
| DB_PERF_SEL_ESR_PS_LQF_BUSY | 187 | Early Squad Router to LQuad FIFO cycles busy |
| DB_PERF_SEL_ESR_PS_LQF_STALL | 188 | Early Squad Router to LQuad FIFO cycles stalled |
| DB_PERF_SEL_ETR_OUT_SEND | 189 | Squads being sent through the Early Tile Router |
| DB_PERF_SEL_ETR_OUT_BUSY | 190 | Cycles Early Tile router busy |
| DB_PERF_SEL_ETR_OUT_LTILE_PROBE_FIFO_FULL_STALL | 191 | Cycles Early Tile router is stalled due to ltile probe FIFO being full |
| DB_PERF_SEL_ETR_OUT_CB_TILE_STALL | 192 | Cycles Early Tile router is stalled due to running out of credits on cb tile interface |
| DB_PERF_SEL_ETR_OUT_ESR_STALL | 193 | Cycles Early Tile router is stalled due to Early Squad router back pressure |
| DB_PERF_SEL_ESR_PS_SQQ_BUSY | 194 | Cycles Early Squad Router squad-to-quad block busy |
| DB_PERF_SEL_ESR_PS_SQQ_STALL | 195 | Cycles Early Squad Router squad-to-quad block stalled |
| DB_PERF_SEL_ESR_EOT_FWD_BUSY | 196 | Cycles Early Squad Router end of tile forwarder busy |
| DB_PERF_SEL_ESR_EOT_FWD_HOLDING_SQUAD | 197 | Cycles Early Squad Router end of tile forwarder is holding an squad waiting for end of tile |
| DB_PERF_SEL_ESR_EOT_FWD_FORWARD | 198 | Tiles where an end of tile is being forwarded |
| DB_PERF_SEL_ESR_SQQ_ZI_BUSY | 199 | Cycles Early Squad Router Z Interp busy |
| DB_PERF_SEL_ESR_SQQ_ZI_STALL | 200 | Cycles Early Squad Router Z Interp stalled |
| DB_PERF_SEL_POSTZL_SQ_PT_BUSY | 201 | Cycles PostZ Squad Launcher squad to tile sync interface busy |
| DB_PERF_SEL_POSTZL_SQ_PT_STALL | 202 | Cycles PostZ Squad Launcher squad to tile sync interface stalled |
| DB_PERF_SEL_POSTZL_SE_BUSY | 203 | Cycles PostZ Squad Launcher shader export busy |
| DB_PERF_SEL_POSTZL_SE_STALL | 204 | Cycles PostZ Squad Launcher shader export stalled |
| DB_PERF_SEL_POSTZL_PARTIAL_LAUNCH | 205 | Squads that are partially launched by PostZ Squad Launcher |
| DB_PERF_SEL_POSTZL_FULL_LAUNCH | 206 | Squads that are fully launched by PostZ Squad Launcher |
| DB_PERF_SEL_POSTZL_PARTIAL_WAITING | 207 | Cycles a partial squad is waiting for more quads to fill in the squad |
| DB_PERF_SEL_POSTZL_TILE_MEM_STALL | 208 | Cycles PostZ Squad Launcher stalled waiting for data from memory |
| DB_PERF_SEL_POSTZL_TILE_INIT_STALL | 209 | Cycles PostZ Squad Launcher stalled waiting for cache to be initialized |
| DB_PEFF_SEL_PREZL_TILE_MEM_STALL | 210 | Cycles PreZ Squad Launcher stalled waiting for data from memory |
| DB_PERF_SEL_PREZL_TILE_INIT_STALL | 211 | Cycles PreZ Squad Launcher stalled waiting for cache to be initialized |
| DB_PERF_SEL_DTT_SM_CLASH_STALL | 212 | Cycles Depth Tile Tag Surface Manager stalled due to address clash with existing surface |
| DB_PERF_SEL_DTT_SM_SLOT_STALL | 213 | Cycles Depth Tile Tag Surface Manager stalled due to no slots being available |
| DB_PERF_SEL_DTT_SM_MISS_STALL | 214 | Cycles Depth Tile Tag Surface Manager stalled due to surface miss |
| DB_PERF_SEL_MI_RDREQ_BUSY | 215 | Cycles Memory Interface read request busy |
| DB_PERF_SEL_MI_RDREQ_STALL | 216 | Cycles Memory Interface read request stalled |
| DB_PERF_SEL_MI_WRREQ_BUSY | 217 | Cycles Memory Interface write request busy |
| DB_PERF_SEL_MI_WRREQ_STALL | 218 | Cycles Memory Interface write request stalled |
| DB_PERF_SEL_RECOMP_TILE_TO_1ZPLANE_NO_FASTOP | 219 | Tiles that went to 1 zplane without a fast z op |
| DB_PERF_SEL_DKG_TILE_RATE_TILE | 220 | Tiles running at tile rate through detail kill block |
| DB_PERF_SEL_PREZL_SRC_IN_SENDS | 221 | Transactions going into Prez Squad Launcher Sample Rate converter |
| DB_PERF_SEL_PREZL_SRC_IN_STALL | 222 | Cycles squad is stalled in the Prez Squad Launcher Sample Rate converter |
| DB_PERF_SEL_PREZL_SRC_IN_SQUADS | 223 | Squads going into Prez Squad Launcher Sample Rate converter |
| DB_PERF_SEL_PREZL_SRC_IN_SQUADS_UNROLLED | 0 | Squads being unrolled in Prez Squad Launcher Sample Rate converter |
| DB_PERF_SEL_PREZL_SRC_IN_TILE_RATE | 0 | Tile rate tiles going into the Prez Squad Launcher Sample Rate converter |
| DB_PERF_SEL_PREZL_SRC_IN_TILE_RATE_UNROLLED | 0 | Tile rate tiles being unrolled in Prez Squad Launcher Sample Rate converter |
| DB_PERF_SEL_PREZL_SRC_OUT_STALL | 0 | Cycles transaction is stalled on the output of Prez Squad Launcher Sample Rate converter |
| DB_PERF_SEL_POSTZL_SRC_IN_SENDS | 0 | Transactions going into Postz Squad Launcher Sample Rate converter |
| DB_PERF_SEL_POSTZL_SRC_IN_STALL | 0 | Cycles squad is stalled in the Postz Squad Launcher Sample Rate converter |
| DB_PERF_SEL_POSTZL_SRC_IN_SQUADS | 0 | Squads going into Postz Squad Launcher Sample Rate converter |
| DB_PERF_SEL_POSTZL_SRC_IN_SQUADS_UNROLLED | 0 | Squads being unrolled in Postz Squad Launcher Sample Rate converter |
| DB_PERF_SEL_POSTZL_SRC_IN_TILE_RATE | 0 | Tile rate tiles going into the Postz Squad Launcher Sample Rate converter |
| DB_PERF_SEL_POSTZL_SRC_IN_TILE_RATE_UNROLLED | 0 | Tile rate tiles being unrolled in Postz Squad Launcher Sample Rate converter |
| DB_PERF_SEL_POSTZL_SRC_OUT_STALL | 234 | Cycles transaction is stalled on the output of Postz Squad Launcher Sample Rate converter |
| DB_PERF_SEL_ESR_PS_SRC_IN_SENDS | 235 | Transactions going into Early Squad Router PS Iter Sample Rate converter |
| DB_PERF_SEL_ESR_PS_SRC_IN_STALL | 236 | Cycles squad is stalled in the Early Squad Router PS Iter Sample Rate converter |
| DB_PERF_SEL_ESR_PS_SRC_IN_SQUADS | 237 | Squads going into Early Squad Router PS Iter Sample Rate converter |
| DB_PERF_SEL_ESR_PS_SRC_IN_SQUADS_UNROLLED | 238 | Squads being unrolled in Early Squad Router PS Iter Sample Rate converter |
| DB_PERF_SEL_ESR_PS_SRC_IN_TILE_RATE | 239 | Tile rate tiles going into the Early Squad Router PS Iter Sample Rate converter |
| DB_PERF_SEL_ESR_PS_SRC_IN_TILE_RATE_UNROLLED | 240 | Tile rate tiles being unrolled in Early Squad Router PS Iter Sample Rate converter |
| DB_PERF_SEL_ESR_PS_SRC_IN_TILE_RATE_UNROLLED_TO_PIXEL_RATE | 241 | Tile rate tiles being unrolled to pixel rate in Early Squad Router PS Iter Sample Rate converter |
| DB_PERF_SEL_ESR_PS_SRC_OUT_STALL | 242 | Cycles transaction is stalled on the output of Early Squad Router PS Iter Sample Rate converter |
| DB_PERF_SEL_DEPTH_BOUNDS_QTILES_CULLED | 243 | Quarter tiles culled at hiz time due to depth bounds |
| DB_PERF_SEL_PREZ_SAMPLES_FAILING_DB | 244 | PreZ Samples failing depth bounds test |
| DB_PERF_SEL_POSTZ_SAMPLES_FAILING_DB | 245 | PostZ Samples failing depth bounds test |
| DB_PERF_SEL_FLUSH_COMPRESSED | 246 | Tiles that flushed compressed |
| DB_PERF_SEL_FLUSH_PLANE_LE4 | 247 | Tiles that flushed compressed with less than or equal to 4 planes |
| DB_PERF_SEL_TILES_Z_FULLY_SUMMARIZED | 248 | Tiles that fully summarized z |
| DB_PERF_SEL_TILES_STENCIL_FULLY_SUMMARIZED | 249 | Tiles that fully summarized stencil |
| DB_PERF_SEL_TILES_Z_CLEAR_ON_EXPCLEAR | 250 | Tiles that do a z clear on expanded and clear |
| DB_PERF_SEL_TILES_S_CLEAR_ON_EXPCLEAR | 251 | Tiles that do a stencil clear on expanded and clear |
| DB_PERF_SEL_TILES_DECOMP_ON_EXPCLEAR | 252 | Tiles that do a z decompress on expanded and clear |
| DB_PERF_SEL_TILES_COMPRESSED_TO_DECOMPRESSED | 253 | Tiles that transition from compressed to expanded |
| DB_PERF_SEL_OP_PIPE_PREZ_BUSY | 254 | Cycles the op pipe is busy with prez squads |
| DB_PERF_SEL_OP_PIPE_POSTZ_BUSY | 255 | Cycles the op pipe is busy with postz squads |
| DB_PERF_SEL_DI_DT_STALL | 256 | Cycles the op pipe stalls due to di/dt controller |
| DB_PERF_SEL_SCORPIO_START | 257 | Start of Xbox One X-specific counters. |
| DB_PERF_SEL_DB_SC_Z_TILE_RATE | 257 | Tiles that are z-eligible to run at tile rate. |
| DB_PERF_SEL_DB_SC_S_TILE_RATE | 258 | Tiles that are s-eligible to run at tile rate. |
| DB_PERF_SEL_DB_SC_C_TILE_RATE | 259 | Tiles that are c-eligible to run at tile rate. |
| Counter | Value | Description |
|---|---|---|
| GRBM_PERF_SEL_COUNT | 0 | Tie High - Count Number of Clocks |
| GRBM_PERF_SEL_USER_DEFINED | 1 | User defined performance select. |
| GRBM_PERF_SEL_GUI_ACTIVE | 2 | The GUI is Active |
| GRBM_PERF_SEL_CP_BUSY | 3 | Any of the Command Processor (CPG/CPC/CPF) blocks are busy. |
| GRBM_PERF_SEL_CP_COHER_BUSY | 4 | The Command Processor Graphics (CPG) Surface Coherency Logic is busy. |
| GRBM_PERF_SEL_CP_DMA_BUSY | 5 | The Command Processor Graphics (CPG) DMA Logic is busy. |
| GRBM_PERF_SEL_CB_BUSY | 6 | Any of the Color Blocks (CB) are busy in the shader engine(s). |
| GRBM_PERF_SEL_DB_BUSY | 7 | Any of the Depth Blocks (DB) are busy in the shader engine(s). |
| GRBM_PERF_SEL_PA_BUSY | 8 | Any of the Primitive Assembly Blocks (PA) are busy in the shader engine(s). |
| GRBM_PERF_SEL_SC_BUSY | 9 | Any of the Scan Converter Blocks (SC) are busy in the shader engine(s). |
| GRBM_PERF_SEL_RESERVED_6 | 10 | Reserved to maintain backwards compatibility. |
| GRBM_PERF_SEL_SPI_BUSY | 11 | Any of the Shader Pipe Interpolators (SPI) are busy in the shader engine(s). |
| GRBM_PERF_SEL_SX_BUSY | 12 | Any of the Shader Export Blocks (SX) are busy. |
| GRBM_PERF_SEL_TA_BUSY | 13 | Any of the Texture Pipes (TA) are busy in the shader engine(s). |
| GRBM_PERF_SEL_CB_CLEAN | 14 | Any of the Color Blocks (CB) are not clean in all shader engine(s). |
| GRBM_PERF_SEL_DB_CLEAN | 15 | Any of the Depth Blocks (DB) are not clean in all shader engine(s). |
| GRBM_PERF_SEL_RESERVED_5 | 16 | Reserved to maintain backwards compatibility. |
| GRBM_PERF_SEL_VGT_BUSY | 17 | Any of the Vertex Geometry Tessellator Blocks (VGT) are busy in the shader engine(s). |
| GRBM_PERF_SEL_RESERVED_4 | 18 | Reserved to maintain backwards compatibility. |
| GRBM_PERF_SEL_RESERVED_3 | 19 | Reserved to maintain backwards compatibility. |
| GRBM_PERF_SEL_RESERVED_2 | 20 | Reserved to maintain backwards compatibility. |
| GRBM_PERF_SEL_RESERVED_1 | 21 | Reserved to maintain backwards compatibility. |
| GRBM_PERF_SEL_RESERVED_0 | 22 | Reserved to maintain backwards compatibility. |
| GRBM_PERF_SEL_IA_BUSY | 23 | The Input Assembler (IA) is busy. |
| GRBM_PERF_SEL_IA_NO_DMA_BUSY | 24 | The Input Assembler (IA) is busy; Does not include the Index DMA engine status. |
| GRBM_PERF_SEL_GDS_BUSY | 25 | The Global Data Share (GDS) is busy. |
| GRBM_PERF_SEL_BCI_BUSY | 26 | Any of the Barycentric Interpolators (BCI) are busy in the shader engine(s). |
| GRBM_PERF_SEL_RLC_BUSY | 27 | The Ring List Controller (RLC) is busy. |
| GRBM_PERF_SEL_TC_BUSY | 28 | Any of the Texture Cache Blocks (TCP/TCI/TCA/TCC) are busy. |
| GRBM_PERF_SEL_CPG_BUSY | 29 | The Command Processor Graphics (CPG) is busy. |
| GRBM_PERF_SEL_CPC_BUSY | 30 | The Command Processor Compute (CPC) is busy. |
| GRBM_PERF_SEL_CPF_BUSY | 31 | The Command Processor Fetchers (CPF) is busy. |
| GRBM_PERF_SEL_WD_BUSY | 32 | The Work Distributor (WD) is busy. |
| GRBM_PERF_SEL_WD_NO_DMA_BUSY | 33 | The Work Distributor (WD) is busy; Does not include the Index DMA engine status. |
| Counter | Value | Description |
|---|---|---|
| PAPC_PERF_PASX_REQ | 0 | Number of PA->SX requests; increment rate-one per clock (each req leads to 4 cycles of data from SX); range-1/clk, it does not indicate bad performance; it cannot be used to locate bottlenecks; all instances report the same result; can be combined with PASX_FIRST_VECTOR,PASX_SECOND_VECTOR, PASX_VTX_KILL_DISCARD, PASX_VTX_NAN_DISCARD, PASX_FIRST_DEAD, PASX_FIRST_DEAD |
| PAPC_PERF_PASX_DISABLE_PIPE | 1 | Number of transfers lost due to disabled pipe; increment rate-one per clock ; range-1/clk ; does not indicate bad performance; cannot be used for bottleneck location; all instances report the same result; no combinations |
| PAPC_PERF_PASX_FIRST_VECTOR | 2 | Number of First Vectors from SX to PA; increment rate-one per clock ; range-1/clk ; it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_PASX_SECOND_VECTOR | 3 | Number of Second Vectors from SX to PA; increment rate-one per clock ; range-1/clk ; it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_PASX_FIRST_DEAD | 4 | Number of Unused First Vectors (due to granularity of 4); increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_PASX_SECOND_DEAD | 5 | Number of Unused Second Vectors (due to granularity of 4); increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_PASX_VTX_KILL_DISCARD | 6 | Number of vertices which have VTX KILL Enabled and Set; increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_PASX_VTX_NAN_DISCARD | 7 | Number of vertices which have NaN and corresponding NaN discard; increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_PA_INPUT_PRIM | 8 | Number of Primitives input to PA; increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_PA_INPUT_NULL_PRIM | 9 | Number of Null Primitives input to PA; increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; can be combined with _PA_INPUT_PRIM |
| PAPC_PERF_PA_INPUT_EVENT_FLAG | 10 | Number of Events input to PA; increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_PA_INPUT_FIRST_PRIM_SLOT | 11 | Number of First-Prim-Of-Slots input to PA; increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_PA_INPUT_END_OF_PACKET | 12 | Number of End-Of-Packets input to PA; increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_PA_INPUT_EXTENDED_EVENT | 13 | Number of Extended Events input to PA; increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; combined with _PA_INPUT_EVENT_FLAG |
| PAPC_PERF_CLPR_CULL_PRIM | 14 | Number of Prims Culled by Clipper for VV, UCP, VTX_KILL, VTX_NAN; increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; combined with _CLPR_VVUCP_CULL_PRIM , _CLPR_VV_CULL_PRIM, _VV_CULL_PRIM ,_UCP_CULL_PRIM, _VTX_KILL_CULL_PRIM, _VTX_NAN_CULL_PRIM |
| PAPC_PERF_CLPR_VVUCP_CULL_PRIM | 15 | Number of Prims Culled by Clipper for VV and UCP; increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_CLPR_VV_CULL_PRIM | 16 | Number of Prims Culled by Clipper for VV; increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_CLPR_UCP_CULL_PRIM | 17 | Number of Prims Culled by Clipper for UCP; increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_CLPR_VTX_KILL_CULL_PRIM | 18 | Number of Prims Culled by Clipper for VTX_KILL; increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_CLPR_VTX_NAN_CULL_PRIM | 19 | Number of Prims Culled by Clipper for VTX_NAN; increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_CLPR_CULL_TO_NULL_PRIM | 20 | Number of Clipper Culled Prims Retained for Pipe Info; increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_CLPR_VVUCP_CLIP_PRIM | 21 | Number of Prims Clipped by Clipper for VV and/or UCP; increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_CLPR_VV_CLIP_PRIM | 22 | Number of Prims Clipped by Clipper for VV; increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_CLPR_UCP_CLIP_PRIM | 23 | Number of Prims Clipped by Clipper for UCP; increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_CLPR_POINT_CLIP_CANDIDATE | 24 | Number of Points which require detailed clip checked ; increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_CLPR_CLIP_PLANE_CNT_1 | 25 | Number of Prims with 1 Clip Plane Intersection (includes VV and UCP); increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; can be combined with _CLPR_POINT_CLIP_CANDIDATE |
| PAPC_PERF_CLPR_CLIP_PLANE_CNT_2 | 26 | Number of Prims with 2 Clip Plane Intersections (includes VV and UCP); increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; can be combined with _CLPR_POINT_CLIP_CANDIDATE |
| PAPC_PERF_CLPR_CLIP_PLANE_CNT_3 | 27 | Number of Prims with 3 Clip Plane Intersections (includes VV and UCP); increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; can be combined with _CLPR_POINT_CLIP_CANDIDATE |
| PAPC_PERF_CLPR_CLIP_PLANE_CNT_4 | 28 | Number of Prims with 4 Clip Plane Intersections (includes VV and UCP); increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; can be combined with _CLPR_POINT_CLIP_CANDIDATE |
| PAPC_PERF_CLPR_CLIP_PLANE_CNT_5_8 | 29 | Number of Prims with 5-8 Clip Plane Intersections (includes VV and UCP); increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; can be combined with _CLPR_POINT_CLIP_CANDIDATE |
| PAPC_PERF_CLPR_CLIP_PLANE_CNT_9_12 | 30 | Number of Prims with 9-12 Clip Plane Intersections (includes VV and UCP); increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; can be combined with _CLPR_POINT_CLIP_CANDIDATE |
| PAPC_PERF_CLPR_CLIP_PLANE_NEAR | 31 | Number of Prims which intersect the NEAR VV Plane; increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; can be combined with _CLPR_POINT_CLIP_CANDIDATE |
| PAPC_PERF_CLPR_CLIP_PLANE_FAR | 32 | Number of Prims which intersect the FAR VV Plane; increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; can be combined with _CLPR_POINT_CLIP_CANDIDATE |
| PAPC_PERF_CLPR_CLIP_PLANE_LEFT | 33 | Number of Prims which intersect the LEFT VV Plane; increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; can be combined with _CLPR_POINT_CLIP_CANDIDATE |
| PAPC_PERF_CLPR_CLIP_PLANE_RIGHT | 34 | Number of Prims which intersect the RIGHT VV Plane; increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; can be combined with _CLPR_POINT_CLIP_CANDIDATE |
| PAPC_PERF_CLPR_CLIP_PLANE_TOP | 35 | Number of Prims which intersect the TOP VV Plane ; increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; can be combined with _CLPR_POINT_CLIP_CANDIDATE |
| PAPC_PERF_CLPR_CLIP_PLANE_BOTTOM | 36 | Number of Prims which intersect the BOTTOM VV Plane ; increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; can be combined with _CLPR_POINT_CLIP_CANDIDATE |
| PAPC_PERF_CLPR_GSC_KILL_CULL_PRIM | 37 | Number of Prims Culled by Clipper for Geometry Shader Scenario C Cuts; increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_CLPR_RASTER_KILL_CULL_PRIM | 38 | Number of Prims Culled by Clipper for DX10 Rasterization Kill (null PS); increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_CLSM_NULL_PRIM | 39 | Number of null primitives at Clip State Machine pipe stage; increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combination |
| PAPC_PERF_CLSM_TOTALLY_VISIBLE_PRIM | 40 | Number of totally visible (no-clipping) prims; increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; can be combined with PAPC_PERF_CLPR_GSC_KILL_CULL_PRIM,PAPC_PERF_CLPR_RASTER_KILL_CULL_PRIM |
| PAPC_PERF_CLSM_CULL_TO_NULL_PRIM | 41 | Number of primitives which are culled during clip process; increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_CLSM_OUT_PRIM_CNT_1 | 42 | Number of primitives which were clipped and result in 1 primitive; increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_CLSM_OUT_PRIM_CNT_2 | 43 | Number of primitives which were clipped and result in 2 primitives; increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_CLSM_OUT_PRIM_CNT_3 | 44 | Number of primitives which were clipped and result in 3 primitives; increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_CLSM_OUT_PRIM_CNT_4 | 45 | Number of primitives which were clipped and result in 4 primitives; increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_CLSM_OUT_PRIM_CNT_5_8 | 46 | Number of primitives which were clipped and result in 5-8 primitives; increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_CLSM_OUT_PRIM_CNT_9_13 | 47 | Number of primitives which were clipped and result in 9-13 primitives; increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_CLIPGA_VTE_KILL_PRIM | 48 | Number of primitives which are culled by VTE nan/inf logic; increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_SU_INPUT_PRIM | 49 | Number of primitives input to the Setup block; increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_SU_INPUT_CLIP_PRIM | 50 | Number of clipped primitives input to the Setup block; increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_SU_INPUT_NULL_PRIM | 51 | Number of null primitives input to the Setup block; increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_SU_INPUT_PRIM_DUAL | 52 | Number of dual gradient primitives input to the Setup block; increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_SU_INPUT_CLIP_PRIM_DUAL | 53 | Number of dual gradient clipped primitives input to the Setup block; increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_SU_ZERO_AREA_CULL_PRIM | 54 | Number of primitives culled due to zero area; increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_SU_BACK_FACE_CULL_PRIM | 55 | Number of back-face primitives culled due to facedness; increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_SU_FRONT_FACE_CULL_PRIM | 56 | Number of front-face primitives culled due to facedness; increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_SU_POLYMODE_FACE_CULL | 57 | Number of polymode cull-determination primitives culled; increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_SU_POLYMODE_BACK_CULL | 58 | Number of polymode primitives discarded due to Back-Face Cull; increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_SU_POLYMODE_FRONT_CULL | 59 | Number of polymode primitives discarded due to Front-Face Cull; increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_SU_POLYMODE_INVALID_FILL | 60 | Number of polymode lines and/or points which are culled because they are an internal edge or point; increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_SU_OUTPUT_PRIM | 61 | Number of primitives output from the Setup block; increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_SU_OUTPUT_CLIP_PRIM | 62 | Number of clipped primitives output from the Setup block; increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_SU_OUTPUT_NULL_PRIM | 63 | Number of null primitives output from the Setup block; increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_SU_OUTPUT_EVENT_FLAG | 64 | Number of events output from the Setup block; increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_SU_OUTPUT_FIRST_PRIM_SLOT | 65 | Number of First-Prim-Of-Slots output from the Setup block; increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_SU_OUTPUT_END_OF_PACKET | 66 | Number of End-Of-Packets output from the Setup block; increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_SU_OUTPUT_POLYMODE_FACE | 67 | Number of polymode facing primitives output from the Setup block; increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_SU_OUTPUT_POLYMODE_BACK | 68 | Number of polymode back-face primitives output from the Setup block; increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_SU_OUTPUT_POLYMODE_FRONT | 69 | Number of polymode front-face primitives output from the Setup block; increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_SU_OUT_CLIP_POLYMODE_FACE | 70 | Number of clipped polymode facing primitives output from the Setup block; increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_SU_OUT_CLIP_POLYMODE_BACK | 71 | Number of clipped polymode back-face primitives output from the Setup block; increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_SU_OUT_CLIP_POLYMODE_FRONT | 72 | Number of clipped polymode front-face primitives output from the Setup block; increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_SU_OUTPUT_PRIM_DUAL | 73 | Number of dual gradient primitives output from the Setup block; increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_SU_OUTPUT_CLIP_PRIM_DUAL | 74 | Number of dual gradient clipped primitives output from the Setup block; increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_SU_OUTPUT_POLYMODE_DUAL | 75 | Number of dual gradient polymode primitives output from the Setup block; increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_SU_OUTPUT_CLIP_POLYMODE_DUAL | 76 | Number of dual gradient clip polymode primitives output from the Setup block; increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_PASX_REQ_IDLE | 77 | Number of clocks PASX Requestor is Idle; increment rate-one per clock ; range-1/clk;it can indicate bad performance (depending on the test); can be used for bottleneck detection;all instances report the same result; can be combined with a perf counter which indicates the number of clocks required to complete the test |
| PAPC_PERF_PASX_REQ_BUSY | 78 | Number of clocks PASX Requestor is Busy; increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_PASX_REQ_STALLED | 79 | Number of clocks PASX Requestor is Stalled; increment rate-one per clock ; range-1/clk;it can indicate bad performance (depending on the test); can be used for bottleneck detection;all instances report the same result; can be combined with a perf counter which indicates the number of clocks required to complete the test |
| PAPC_PERF_PASX_REC_IDLE | 80 | Number of clocks PASX Receiver is Idle; increment rate-one per clock ; range-1/clk;it can indicate bad performance (depending on the test); can be used for bottleneck detection;all instances report the same result; can be combined with a perf counter which indicates the number of clocks required to complete the test |
| PAPC_PERF_PASX_REC_BUSY | 81 | Number of clocks PASX Receiver is Busy; increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_PASX_REC_STARVED_SX | 82 | Number of clocks PASX Receiver is Stalled by SX; increment rate-one per clock ; range-1/clk;it can indicate bad performance (depending on the test); can be used for bottleneck detection;all instances report the same result; can be combined with a perf counter which indicates the number of clocks required to complete the test |
| PAPC_PERF_PASX_REC_STALLED | 83 | Number of clocks PASX Receiver is Stalled by Position Memory or Clip Code Generator; increment rate-one per clock ; range-1/clk;it can indicate bad performance (depending on the test); can be used for bottleneck detection;all instances report the same result; can be combined with- |
| PAPC_PERF_PASX_REC_STALLED_POS_MEM | 84 | Number of clocks PASX Receiver is Stalled by Position Memory; increment rate-one per clock ; range-1/clk;it can indicate bad performance (depending on the test); can be used for bottleneck detection;all instances report the same result; can be combined with- |
| PAPC_PERF_PASX_REC_STALLED_CCGSM_IN | 85 | Number of clocks PASX Receiver is Stalled by Clip Code Generator ; increment rate-one per clock ; range-1/clk;it can indicate bad performance (depending on the test); can be used for bottleneck detection;all instances report the same result; can be combined with- the other CLIP counters |
| PAPC_PERF_CCGSM_IDLE | 86 | Number of clocks Clip Code Gen is Idle; increment rate-one per clock ; range-1/clk;it can indicate bad performance (depending on the test); can be used for bottleneck detection;all instances report the same result; can be combined with- |
| PAPC_PERF_CCGSM_BUSY | 87 | Number of clocks Clip Code Gen is Busy; increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_CCGSM_STALLED | 88 | Number of clocks Clip Code Gen is Stalled; increment rate-one per clock ; range-1/clk;it can indicate bad performance (depending on the test); can be used for bottleneck detection;all instances report the same result; can be combined with- |
| PAPC_PERF_CLPRIM_IDLE | 89 | Number of clocks Clip Primitive Machine is Idle; increment rate-one per clock ; range-1/clk;it can indicate bad performance (depending on the test); can be used for bottleneck detection;all instances report the same result; can be combined with- |
| PAPC_PERF_CLPRIM_BUSY | 90 | Number of clocks Clip Primitive Machine is Busy; increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_CLPRIM_STALLED | 91 | Number of clocks Clip Primitive Machine is stalled by Clip State Machines; range-1/clk;it can indicate bad performance (depending on the test); can be used for bottleneck detection;all instances report the same result; can be combined with- |
| PAPC_PERF_CLPRIM_STARVED_CCGSM | 92 | Number of clocks Clip Primitive Machine is starved by Clip Code Generator; range-1/clk;it can indicate bad performance (depending on the test); can be used for bottleneck detection;all instances report the same result; can be combined with- |
| PAPC_PERF_CLIPSM_IDLE | 93 | Number of clocks Clip State Machines are Idle; range-1/clk;it can indicate bad performance (depending on the test); can be used for bottleneck detection;all instances report the same result; can be combined with- |
| PAPC_PERF_CLIPSM_BUSY | 94 | Number of clocks Clip State Machines are Busy; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_CLIPSM_WAIT_CLIP_VERT_ENGH | 95 | Number of clocks Clip State Machines are waiting for Clip Vert storage resources; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_CLIPSM_WAIT_HIGH_PRI_SEQ | 96 | Number of clocks Clip State Machines are waiting for High Priority Sequencer; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_CLIPSM_WAIT_CLIPGA | 97 | Number of clocks Clip State Machines are waiting for ClipGA; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_CLIPSM_WAIT_AVAIL_VTE_CLIP | 98 | Number of clocks Clip State Machines are waiting for VTE cycles; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_CLIPSM_WAIT_CLIP_OUTSM | 99 | Number of clocks Clip State Machines are waiting for Clip Output State Machine; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_CLIPGA_IDLE | 100 | Number of clocks Clip Ga is Idle; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_CLIPGA_BUSY | 101 | Number of clocks Clip Ga is Busy; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_CLIPGA_STARVED_VTE_CLIP | 102 | Number of clocks Clip Ga is Starved by VTE or Clipper; range-1/clk;it can potentially be used to detect bad performance;all instances report the same result |
| PAPC_PERF_CLIPGA_STALLED | 103 | Number of clocks Clip Ga is stalled; range-1/clk;it can potentially be used to detect bad performance;all instances report the same result |
| PAPC_PERF_CLIP_IDLE | 104 | Number of clocks Clip is Idle; range-1/clk;it can potentially be used to detect bad performance;all instances report the same result; can be used to detect bottlenecks in combination with other signals |
| PAPC_PERF_CLIP_BUSY | 105 | Number of clocks Clip is Busy; range-1/clk;it can potentially be used to detect bad performance;all instances report the same result; can be used to detect bottlenecks in combination with other signals |
| PAPC_PERF_SU_IDLE | 106 | Number of clocks Setup is Idle; range-1/clk;it can potentially be used to detect bad performance;all instances report the same result; can be used to detect bottlenecks in combination with other signals |
| PAPC_PERF_SU_BUSY | 107 | Number of clocks Setup is Busy; range-1/clk;it can potentially be used to detect bad performance;all instances report the same result; can be used to detect bottlenecks in combination with other signals |
| PAPC_PERF_SU_STARVED_CLIP | 108 | Number of clocks Setup is starved by Clipper; range-1/clk;it can potentially be used to detect bad performance;all instances report the same result |
| PAPC_PERF_SU_STALLED_SC | 109 | Number of clocks Setup is stalled by SC; range-1/clk;it can potentially be used to detect bad performance;all instances report the same result |
| PAPC_PERF_CL_DYN_SCLK_VLD | 110 | Number of clocks the CL dynamic clock is active;range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_SU_DYN_SCLK_VLD | 111 | Number of clocks the SU dynamic clock is active;range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_PA_REG_SCLK_VLD | 112 | Number of clocks the PA register clock is active;range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_SU_MULTI_GPU_PRIM_FILTER_CULL | 113 | Number of primitives culled by multi-gpu filter;range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_PASX_SE0_REQ | 114 | Number of PA to SX requests to shader engine 0;range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_PASX_SE1_REQ | 115 | Number of PA to SX requests to shader engine 1;range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_PASX_SE0_FIRST_VECTOR | 116 | Number of Shader engine 0 first vectors (position);range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_PASX_SE0_SECOND_VECTOR | 117 | Number of Shader engine 0 second vectors (position);range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_PASX_SE1_FIRST_VECTOR | 118 | Number of Shader engine 1 first vectors (position);range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_PASX_SE1_SECOND_VECTOR | 119 | Number of Shader engine 1 second vectors (position);range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_SU_SE0_PRIM_FILTER_CULL | 120 | Number of primitives culled by shader engine 0 filter;range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_SU_SE1_PRIM_FILTER_CULL | 121 | Number of primitives culled by shader engine 1 filter;range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_SU_SE01_PRIM_FILTER_CULL | 122 | Number of primitives culled by SE0 and SE1 filter;range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_SU_SE0_OUTPUT_PRIM | 123 | Number of primitives output to shader engine 0;range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_SU_SE1_OUTPUT_PRIM | 124 | Number of primitives output to shader engine 1;range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_SU_SE01_OUTPUT_PRIM | 125 | Number of primitives output to SE0 and SE1;range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_SU_SE0_OUTPUT_NULL_PRIM | 126 | Number of null primitives output to shader engine 0;range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_SU_SE1_OUTPUT_NULL_PRIM | 127 | Number of null primitives output to shader engine 1;range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_SU_SE01_OUTPUT_NULL_PRIM | 128 | Number of null primitives output to SE0 and SE1;range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_SU_SE0_OUTPUT_FIRST_PRIM_SLOT | 129 | Number of shader engine 0 first primitive of slot;range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_SU_SE1_OUTPUT_FIRST_PRIM_SLOT | 130 | Number of shader engine 1 first primitive of slot;range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_SU_SE0_STALLED_SC | 131 | Number of clocks stalled by shader engine 0 scan converter; range-1/clk;it can potentially be used to detect bad performance;all instances report the same result |
| PAPC_PERF_SU_SE1_STALLED_SC | 132 | Number of clocks stalled by shader engine 1 scan converter; range-1/clk;it can potentially be used to detect bad performance;all instances report the same result |
| PAPC_PERF_SU_SE01_STALLED_SC | 133 | Number of clocks stalled by SE0 and SE1 scan converter; range-1/clk;it can potentially be used to detect bad performance;all instances report the same result |
| PAPC_PERF_CLSM_CLIPPING_PRIM | 134 | Number of clocks the PA is actively clipping primitives on any of the clipper state machines; range-1/clk; cannot be used to detect bad performance |
| PAPC_PERF_SU_CULLED_PRIM | 135 | Number of prims culled in the SU due to zero area, front or back facing, or polymode front or back facing; range-1/clk; cannot be used to detect bad performance |
| PAPC_PERF_SU_OUTPUT_EOPG | 136 | Number of EOPGs output from the Setup block; increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_SU_SE0_OUTPUT_END_OF_PACKET | 143 | Number of End-Of-Packets output to shader engine 0;increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_SU_SE1_OUTPUT_END_OF_PACKET | 144 | Number of End-Of-Packets output to shader engine 1;increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_SU_SE0_OUTPUT_EOPG | 147 | Number of EOPGs output to shader engine 0;increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| PAPC_PERF_SU_SE1_OUTPUT_EOPG | 148 | Number of EOPGs output to shader engine 1;increment rate-one per clock ; range-1/clk;it does not indicate bad performance; no bottleneck detection;all instances report the same result; no combinations |
| Counter | Value | Description |
|---|---|---|
| SC_SRPS_WINDOW_VALID | 0 | Number of clocks event-window is valid at stage register/primitive setup |
| SC_PSSW_WINDOW_VALID | 1 | Number of clocks event-window is valid at primitive setup/supertile walker |
| SC_TPQZ_WINDOW_VALID | 2 | Number of clocks event-window is valid at tile picker/quad-z |
| SC_QZQP_WINDOW_VALID | 3 | Number of clocks event-window is valid at quad-z/quad processor |
| SC_TRPK_WINDOW_VALID | 4 | Number of clocks event-window is valid at tile reorder/packer |
| SC_SRPS_WINDOW_VALID_BUSY | 5 | Number of clocks event-window is valid at stage register/primitive setup with SC busy |
| SC_PSSW_WINDOW_VALID_BUSY | 6 | Number of clocks event-window is valid at primitive setup/supertile walker with SC busy |
| SC_TPQZ_WINDOW_VALID_BUSY | 7 | Number of clocks event-window is valid at tile picker/quad-z with SC busy |
| SC_QZQP_WINDOW_VALID_BUSY | 8 | Number of clocks event-window is valid at quad-z/quad processor with SC busy |
| SC_TRPK_WINDOW_VALID_BUSY | 9 | Number of clocks event-window is valid at tile reorder/packer with SC busy |
| SC_STARVED_BY_PA | 10 | sc starved by pa |
| SC_STALLED_BY_PRIMFIFO | 11 | sc stalled by primitive FIFO |
| SC_STALLED_BY_DB_TILE | 12 | sc stalled by db tile |
| SC_STARVED_BY_DB_TILE | 13 | sc starved by db tile |
| SC_STALLED_BY_TILEORDERFIFO | 14 | sc stalled by tile order FIFO |
| SC_STALLED_BY_TILEFIFO | 15 | sc stalled by tile FIFO |
| SC_STALLED_BY_DB_QUAD | 16 | sc stalled by db quad |
| SC_STARVED_BY_DB_QUAD | 17 | sc starved by db quad |
| SC_STALLED_BY_QUADFIFO | 18 | sc stalled by quad FIFO |
| SC_STALLED_BY_BCI | 19 | sc stalled by bci |
| SC_STALLED_BY_SPI | 20 | sc stalled by spi |
| SC_SCISSOR_DISCARD | 21 | primitive completely discarded by scissor |
| SC_BB_DISCARD | 22 | primitive discarded by bounding-box check, no pixels hit |
| SC_SUPERTILE_COUNT | 23 | supertile count |
| SC_SUPERTILE_PER_PRIM_H0 | 24 | prims with < 2 supertiles |
| SC_SUPERTILE_PER_PRIM_H1 | 25 | prims with < 4 supertiles |
| SC_SUPERTILE_PER_PRIM_H2 | 26 | prims with < 8 supertiles |
| SC_SUPERTILE_PER_PRIM_H3 | 27 | prims with < 16 supertiles |
| SC_SUPERTILE_PER_PRIM_H4 | 28 | prims with < 32 supertiles |
| SC_SUPERTILE_PER_PRIM_H5 | 29 | prims with < 64 supertiles |
| SC_SUPERTILE_PER_PRIM_H6 | 30 | prims with < 128 supertiles |
| SC_SUPERTILE_PER_PRIM_H7 | 31 | prims with < 256 supertiles |
| SC_SUPERTILE_PER_PRIM_H8 | 32 | prims with < 512 supertiles |
| SC_SUPERTILE_PER_PRIM_H9 | 33 | prims with < 1K supertiles |
| SC_SUPERTILE_PER_PRIM_H10 | 34 | prims with < 2K supertiles |
| SC_SUPERTILE_PER_PRIM_H11 | 35 | prims with < 4K supertiles |
| SC_SUPERTILE_PER_PRIM_H12 | 36 | prims with < 8K supertiles |
| SC_SUPERTILE_PER_PRIM_H13 | 37 | prims with < 16K supertiles |
| SC_SUPERTILE_PER_PRIM_H14 | 38 | prims with < 32K supertiles |
| SC_SUPERTILE_PER_PRIM_H15 | 39 | prims with < 64K supertiles |
| SC_SUPERTILE_PER_PRIM_H16 | 40 | prims with < 1M supertiles |
| SC_TILE_PER_PRIM_H0 | 41 | prims with < 2 tiles |
| SC_TILE_PER_PRIM_H1 | 42 | prims with < 4 tiles |
| SC_TILE_PER_PRIM_H2 | 43 | prims with < 8 tiles |
| SC_TILE_PER_PRIM_H3 | 44 | prims with < 16 tiles |
| SC_TILE_PER_PRIM_H4 | 45 | prims with < 32 tiles |
| SC_TILE_PER_PRIM_H5 | 46 | prims with < 64 tiles |
| SC_TILE_PER_PRIM_H6 | 47 | prims with < 128 tiles |
| SC_TILE_PER_PRIM_H7 | 48 | prims with < 256 tiles |
| SC_TILE_PER_PRIM_H8 | 49 | prims with < 512 tiles |
| SC_TILE_PER_PRIM_H9 | 50 | prims with < 1K tiles |
| SC_TILE_PER_PRIM_H10 | 51 | prims with < 2K tiles |
| SC_TILE_PER_PRIM_H11 | 52 | prims with < 4K tiles |
| SC_TILE_PER_PRIM_H12 | 53 | prims with < 8K tiles |
| SC_TILE_PER_PRIM_H13 | 54 | prims with < 16K tiles |
| SC_TILE_PER_PRIM_H14 | 55 | prims with < 32K tiles |
| SC_TILE_PER_PRIM_H15 | 56 | prims with < 64K tiles |
| SC_TILE_PER_PRIM_H16 | 57 | prims with < 1M tiles |
| SC_TILE_PER_SUPERTILE_H0 | 58 | supertiles walked with 0 tiles hit |
| SC_TILE_PER_SUPERTILE_H1 | 59 | supertiles walked with 1 tiles hit |
| SC_TILE_PER_SUPERTILE_H2 | 60 | supertiles walked with 2 tiles hit |
| SC_TILE_PER_SUPERTILE_H3 | 61 | supertiles walked with 3 tiles hit |
| SC_TILE_PER_SUPERTILE_H4 | 62 | supertiles walked with 4 tiles hit |
| SC_TILE_PER_SUPERTILE_H5 | 63 | supertiles walked with 5 tiles hit |
| SC_TILE_PER_SUPERTILE_H6 | 64 | supertiles walked with 6 tiles hit |
| SC_TILE_PER_SUPERTILE_H7 | 65 | supertiles walked with 7 tiles hit |
| SC_TILE_PER_SUPERTILE_H8 | 66 | supertiles walked with 8 tiles hit |
| SC_TILE_PER_SUPERTILE_H9 | 67 | supertiles walked with 9 tiles hit |
| SC_TILE_PER_SUPERTILE_H10 | 68 | supertiles walked with 10 tiles hit |
| SC_TILE_PER_SUPERTILE_H11 | 69 | supertiles walked with 11 tiles hit |
| SC_TILE_PER_SUPERTILE_H12 | 70 | supertiles walked with 12 tiles hit |
| SC_TILE_PER_SUPERTILE_H13 | 71 | supertiles walked with 13 tiles hit |
| SC_TILE_PER_SUPERTILE_H14 | 72 | supertiles walked with 14 tiles hit |
| SC_TILE_PER_SUPERTILE_H15 | 73 | supertiles walked with 15 tiles hit |
| SC_TILE_PER_SUPERTILE_H16 | 74 | supertiles walked with 16 tiles hit |
| SC_TILE_PICKED_H1 | 75 | number of times 1 tile picked |
| SC_TILE_PICKED_H2 | 76 | number of times 2 tiles picked |
| SC_TILE_PICKED_H3 | 77 | number of times 3 tile picked |
| SC_TILE_PICKED_H4 | 78 | number of times 4 tiles picked |
| SC_QZ0_MULTI_GPU_TILE_DISCARD | 79 | tiles discarded by optimization; quad-z pipe 0 |
| SC_QZ1_MULTI_GPU_TILE_DISCARD | 80 | tiles discarded by optimization; quad-z pipe 1 |
| SC_QZ0_TILE_COUNT | 83 | tile count; quad-z pipe 0 |
| SC_QZ1_TILE_COUNT | 84 | tile count; quad-z pipe 1 |
| SC_QZ0_TILE_COVERED_COUNT | 87 | tile covered count; quad-z pipe 0 |
| SC_QZ1_TILE_COVERED_COUNT | 88 | tile covered count; quad-z pipe 1 |
| SC_QZ0_TILE_NOT_COVERED_COUNT | 91 | tile not covered count; quad-z pipe 0 |
| SC_QZ1_TILE_NOT_COVERED_COUNT | 92 | tile not covered count; quad-z pipe 1 |
| SC_QZ0_QUAD_PER_TILE_H0 | 95 | tiles walked with 0 quads hit; quad-z pipe 0 |
| SC_QZ0_QUAD_PER_TILE_H1 | 96 | tiles walked with 1 quads hit; quad-z pipe 0 |
| SC_QZ0_QUAD_PER_TILE_H2 | 97 | tiles walked with 2 quads hit; quad-z pipe 0 |
| SC_QZ0_QUAD_PER_TILE_H3 | 98 | tiles walked with 3 quads hit; quad-z pipe 0 |
| SC_QZ0_QUAD_PER_TILE_H4 | 99 | tiles walked with 4 quads hit; quad-z pipe 0 |
| SC_QZ0_QUAD_PER_TILE_H5 | 100 | tiles walked with 5 quads hit; quad-z pipe 0 |
| SC_QZ0_QUAD_PER_TILE_H6 | 101 | tiles walked with 6 quads hit; quad-z pipe 0 |
| SC_QZ0_QUAD_PER_TILE_H7 | 102 | tiles walked with 7 quads hit; quad-z pipe 0 |
| SC_QZ0_QUAD_PER_TILE_H8 | 103 | tiles walked with 8 quads hit; quad-z pipe 0 |
| SC_QZ0_QUAD_PER_TILE_H9 | 104 | tiles walked with 9 quads hit; quad-z pipe 0 |
| SC_QZ0_QUAD_PER_TILE_H10 | 105 | tiles walked with 10 quads hit; quad-z pipe 0 |
| SC_QZ0_QUAD_PER_TILE_H11 | 106 | tiles walked with 11 quads hit; quad-z pipe 0 |
| SC_QZ0_QUAD_PER_TILE_H12 | 107 | tiles walked with 12 quads hit; quad-z pipe 0 |
| SC_QZ0_QUAD_PER_TILE_H13 | 108 | tiles walked with 13 quads hit; quad-z pipe 0 |
| SC_QZ0_QUAD_PER_TILE_H14 | 109 | tiles walked with 14 quads hit; quad-z pipe 0 |
| SC_QZ0_QUAD_PER_TILE_H15 | 110 | tiles walked with 15 quads hit; quad-z pipe 0 |
| SC_QZ0_QUAD_PER_TILE_H16 | 111 | tiles walked with 16 quads hit; quad-z pipe 0 |
| SC_QZ1_QUAD_PER_TILE_H0 | 112 | tiles walked with 0 quads hit; quad-z pipe 1 |
| SC_QZ1_QUAD_PER_TILE_H1 | 113 | tiles walked with 1 quads hit; quad-z pipe 1 |
| SC_QZ1_QUAD_PER_TILE_H2 | 114 | tiles walked with 2 quads hit; quad-z pipe 1 |
| SC_QZ1_QUAD_PER_TILE_H3 | 115 | tiles walked with 3 quads hit; quad-z pipe 1 |
| SC_QZ1_QUAD_PER_TILE_H4 | 116 | tiles walked with 4 quads hit; quad-z pipe 1 |
| SC_QZ1_QUAD_PER_TILE_H5 | 117 | tiles walked with 5 quads hit; quad-z pipe 1 |
| SC_QZ1_QUAD_PER_TILE_H6 | 118 | tiles walked with 6 quads hit; quad-z pipe 1 |
| SC_QZ1_QUAD_PER_TILE_H7 | 119 | tiles walked with 7 quads hit; quad-z pipe 1 |
| SC_QZ1_QUAD_PER_TILE_H8 | 120 | tiles walked with 8 quads hit; quad-z pipe 1 |
| SC_QZ1_QUAD_PER_TILE_H9 | 121 | tiles walked with 9 quads hit; quad-z pipe 1 |
| SC_QZ1_QUAD_PER_TILE_H10 | 122 | tiles walked with 10 quads hit; quad-z pipe 1 |
| SC_QZ1_QUAD_PER_TILE_H11 | 123 | tiles walked with 11 quads hit; quad-z pipe 1 |
| SC_QZ1_QUAD_PER_TILE_H12 | 124 | tiles walked with 12 quads hit; quad-z pipe 1 |
| SC_QZ1_QUAD_PER_TILE_H13 | 125 | tiles walked with 13 quads hit; quad-z pipe 1 |
| SC_QZ1_QUAD_PER_TILE_H14 | 126 | tiles walked with 14 quads hit; quad-z pipe 1 |
| SC_QZ1_QUAD_PER_TILE_H15 | 127 | tiles walked with 15 quads hit; quad-z pipe 1 |
| SC_QZ1_QUAD_PER_TILE_H16 | 128 | tiles walked with 16 quads hit; quad-z pipe 1 |
| SC_QZ0_QUAD_COUNT | 163 | quad count; quad-z pipe 0 |
| SC_QZ1_QUAD_COUNT | 164 | quad count; quad-z pipe 1 |
| SC_P0_HIZ_TILE_COUNT | 167 | total tiles surviving hi-z |
| SC_P1_HIZ_TILE_COUNT | 168 | total tiles surviving hi-z |
| SC_P0_HIZ_QUAD_PER_TILE_H0 | 171 | tiles with 0 quads surviving hi-z |
| SC_P0_HIZ_QUAD_PER_TILE_H1 | 172 | tiles with 1 quads surviving hi-z |
| SC_P0_HIZ_QUAD_PER_TILE_H2 | 173 | tiles with 2 quads surviving hi-z |
| SC_P0_HIZ_QUAD_PER_TILE_H3 | 174 | tiles with 3 quads surviving hi-z |
| SC_P0_HIZ_QUAD_PER_TILE_H4 | 175 | tiles with 4 quads surviving hi-z |
| SC_P0_HIZ_QUAD_PER_TILE_H5 | 176 | tiles with 5 quads surviving hi-z |
| SC_P0_HIZ_QUAD_PER_TILE_H6 | 177 | tiles with 6 quads surviving hi-z |
| SC_P0_HIZ_QUAD_PER_TILE_H7 | 178 | tiles with 7 quads surviving hi-z |
| SC_P0_HIZ_QUAD_PER_TILE_H8 | 179 | tiles with 8 quads surviving hi-z |
| SC_P0_HIZ_QUAD_PER_TILE_H9 | 180 | tiles with 9 quads surviving hi-z |
| SC_P0_HIZ_QUAD_PER_TILE_H10 | 181 | tiles with 10 quads surviving hi-z |
| SC_P0_HIZ_QUAD_PER_TILE_H11 | 182 | tiles with 11 quads surviving hi-z |
| SC_P0_HIZ_QUAD_PER_TILE_H12 | 183 | tiles with 12 quads surviving hi-z |
| SC_P0_HIZ_QUAD_PER_TILE_H13 | 184 | tiles with 13 quads surviving hi-z |
| SC_P0_HIZ_QUAD_PER_TILE_H14 | 185 | tiles with 14 quads surviving hi-z |
| SC_P0_HIZ_QUAD_PER_TILE_H15 | 186 | tiles with 15 quads surviving hi-z |
| SC_P0_HIZ_QUAD_PER_TILE_H16 | 187 | tiles with 16 quads surviving hi-z |
| SC_P1_HIZ_QUAD_PER_TILE_H0 | 188 | SC_P0_HIZ_QUAD_PER_TILE_H0 tiles with 0 quads surviving hi-z |
| SC_P1_HIZ_QUAD_PER_TILE_H1 | 189 | SC_P0_HIZ_QUAD_PER_TILE_H1 tiles with 1 quads surviving hi-z |
| SC_P1_HIZ_QUAD_PER_TILE_H2 | 190 | tiles with 2 quads surviving hi-z |
| SC_P1_HIZ_QUAD_PER_TILE_H3 | 191 | tiles with 3 quads surviving hi-z |
| SC_P1_HIZ_QUAD_PER_TILE_H4 | 192 | tiles with 4 quads surviving hi-z |
| SC_P1_HIZ_QUAD_PER_TILE_H5 | 193 | tiles with 5 quads surviving hi-z |
| SC_P1_HIZ_QUAD_PER_TILE_H6 | 194 | tiles with 6 quads surviving hi-z |
| SC_P1_HIZ_QUAD_PER_TILE_H7 | 195 | tiles with 7 quads surviving hi-z |
| SC_P1_HIZ_QUAD_PER_TILE_H8 | 196 | tiles with 8 quads surviving hi-z |
| SC_P1_HIZ_QUAD_PER_TILE_H9 | 197 | tiles with 9 quads surviving hi-z |
| SC_P1_HIZ_QUAD_PER_TILE_H10 | 198 | tiles with 10 quads surviving hi-z |
| SC_P1_HIZ_QUAD_PER_TILE_H11 | 199 | tiles with 11 quads surviving hi-z |
| SC_P1_HIZ_QUAD_PER_TILE_H12 | 200 | tiles with 12 quads surviving hi-z |
| SC_P1_HIZ_QUAD_PER_TILE_H13 | 201 | tiles with 13 quads surviving hi-z |
| SC_P1_HIZ_QUAD_PER_TILE_H14 | 202 | tiles with 14 quads surviving hi-z |
| SC_P1_HIZ_QUAD_PER_TILE_H15 | 203 | tiles with 15 quads surviving hi-z |
| SC_P1_HIZ_QUAD_PER_TILE_H16 | 204 | tiles with 16 quads surviving hi-z |
| SC_P0_HIZ_QUAD_COUNT | 239 | total quads surviving hi-z |
| SC_P1_HIZ_QUAD_COUNT | 240 | total quads surviving hi-z |
| SC_P0_DETAIL_QUAD_COUNT | 243 | total quads surviving detail sampler |
| SC_P1_DETAIL_QUAD_COUNT | 244 | total quads surviving detail sampler |
| SC_P0_DETAIL_QUAD_WITH_1_PIX | 247 | quads with 1 pixel surviving detail |
| SC_P0_DETAIL_QUAD_WITH_2_PIX | 248 | quads with 2 pixels surviving detail |
| SC_P0_DETAIL_QUAD_WITH_3_PIX | 249 | quads with 3 pixels surviving detail |
| SC_P0_DETAIL_QUAD_WITH_4_PIX | 250 | quads with 4 pixels surviving detail |
| SC_P1_DETAIL_QUAD_WITH_1_PIX | 251 | quads with 1 pixel surviving detail |
| SC_P1_DETAIL_QUAD_WITH_2_PIX | 252 | quads with 2 pixels surviving detail |
| SC_P1_DETAIL_QUAD_WITH_3_PIX | 253 | quads with 3 pixels surviving detail |
| SC_P1_DETAIL_QUAD_WITH_4_PIX | 254 | quads with 4 pixels surviving detail |
| SC_EARLYZ_QUAD_COUNT | 263 | total quads surviving early-z |
| SC_EARLYZ_QUAD_WITH_1_PIX | 264 | quads with 1 pixel surviving early-z |
| SC_EARLYZ_QUAD_WITH_2_PIX | 265 | quads with 2 pixels surviving early-z |
| SC_EARLYZ_QUAD_WITH_3_PIX | 266 | quads with 3 pixels surviving early-z |
| SC_EARLYZ_QUAD_WITH_4_PIX | 267 | quads with 4 pixels surviving early-z |
| SC_PKR_QUAD_PER_ROW_H1 | 268 | packer row outputs with 1 valid quad |
| SC_PKR_QUAD_PER_ROW_H2 | 269 | packer row outputs with 2 valid quad |
| SC_PKR_QUAD_PER_ROW_H3 | 270 | packer row outputs with 3 valid quad |
| SC_PKR_QUAD_PER_ROW_H4 | 271 | packer row outputs with 4 valid quad |
| SC_PKR_END_OF_VECTOR | 272 | number of pixel vectors |
| SC_PKR_CONTROL_XFER | 273 | number of control transfers |
| SC_PKR_DBHANG_FORCE_EOV | 274 | number of times partial vector ejected b/c of DB hang condition |
| SC_REG_SCLK_BUSY | 275 | number of cycles register clock is busy |
| SC_GRP0_DYN_SCLK_BUSY | 276 | number of cycles group0 dynamic clock is busy |
| SC_GRP1_DYN_SCLK_BUSY | 277 | number of cycles group1 dynamic clock is busy |
| SC_GRP2_DYN_SCLK_BUSY | 278 | number of cycles group2 dynamic clock is busy |
| SC_GRP3_DYN_SCLK_BUSY | 279 | number of cycles group3 dynamic clock is busy |
| SC_GRP4_DYN_SCLK_BUSY | 280 | number of cycles group4 dynamic clock is busy |
| SC_PA0_SC_DATA_FIFO_RD | 281 | number of cycles of read of pa0 FIFO |
| SC_PA0_SC_DATA_FIFO_WE | 282 | number of cycles of write of pa0 FIFO |
| SC_PA1_SC_DATA_FIFO_RD | 283 | number of cycles of read of pa1 FIFO |
| SC_PA1_SC_DATA_FIFO_WE | 284 | number of cycles of write of pa1 FIFO |
| SC_PS_ARB_XFC_ALL_EVENT_OR_PRIM_CYCLES | 285 | number of cycles of arb xfc’s for all event or prims |
| SC_PS_ARB_XFC_ONLY_PRIM_CYCLES | 286 | number of cycles of arb xfc’s for only prims |
| SC_PS_ARB_XFC_ONLY_ONE_INC_PER_PRIM | 287 | number of cycles of arb xfc’s for one inc per prim only |
| SC_PS_ARB_STALLED_FROM_BELOW | 288 | number of cycles of stalled post arbiter for valid data |
| SC_PS_ARB_STARVED_FROM_ABOVE | 289 | number of cycles of PA starving SC in sop/eop range |
| SC_PS_ARB_SC_BUSY | 290 | number of cycles of SC busy |
| SC_PS_ARB_PA_SC_BUSY | 291 | number of cycles in between start of packet/end of packet |
| SC_PA_SC_DEALLOC_0_0_WE | 296 | number of cycles of pa sc FIFO write of pa0 dealloc binary bit 0 |
| SC_PA_SC_DEALLOC_0_1_WE | 297 | number of cycles of pa sc FIFO write of pa0 dealloc binary bit 1 |
| SC_PA_SC_DEALLOC_1_0_WE | 298 | number of cycles of pa sc FIFO write of pa1 dealloc binary bit 0 |
| SC_PA_SC_DEALLOC_1_1_WE | 299 | number of cycles of pa sc FIFO write of pa1 dealloc binary bit 1 |
| SC_PA_SC_DEALLOC_2_0_WE | 300 | number of cycles of pa sc FIFO write of pa2 dealloc binary bit 0 |
| SC_PA_SC_DEALLOC_2_1_WE | 301 | number of cycles of pa sc FIFO write of pa2 dealloc binary bit 1 |
| SC_PA_SC_DEALLOC_3_0_WE | 302 | number of cycles of pa sc FIFO write of pa3 dealloc binary bit 0 |
| SC_PA_SC_DEALLOC_3_1_WE | 303 | number of cycles of pa sc FIFO write of pa3 dealloc binary bit 1 |
| SC_PA0_SC_EOP_WE | 304 | number of cycles of pa sc FIFO write of pa0 end of packet |
| SC_PA0_SC_EOPG_WE | 305 | number of cycles of pa sc FIFO write of pa0 end of prim group |
| SC_PA0_SC_EVENT_WE | 306 | number of cycles of pa sc FIFO write of pa0 event |
| SC_PA1_SC_EOP_WE | 307 | number of cycles of pa sc FIFO write of pa1 end of packet |
| SC_PA1_SC_EOPG_WE | 308 | number of cycles of pa sc FIFO write of pa1 end of prim group |
| SC_PA1_SC_EVENT_WE | 309 | number of cycles of pa sc FIFO write of pa1 event |
| SC_PS_ARB_OOO_THRESHOLD_SWITCH_TO_DESIRED_FIFO | 316 | number of cycles of pa sc arbiter front end switches based on hitting the out of order FIFO skew threshold |
| SC_PS_ARB_OOO_FIFO_EMPTY_SWITCH | 317 | number of cycles of cycles of pa sc arbiter front end switches based on the currently selected FIFO going empty |
| SC_PS_ARB_NULL_PRIM_BUBBLE_POP | 318 | number of cycles of pa sc arbiter null prims without eop read in the eop broadcast window between eop on selected FIFO and eop sync achieved across all fifos |
| SC_PS_ARB_EOP_POP_SYNC_POP | 319 | number of cycles of pa sc arbiter eop reads at the eop sync point with all FIFOs synced |
| SC_PS_ARB_EVENT_SYNC_POP | 320 | number of cycles of pa sc event reads |
| SC_SC_PS_ENG_MULTICYCLE_BUBBLE | 321 | number of cycles of multicycles prims sent with a empty FIFO slot behind the current phase, this is not expected to occur |
| SC_PA0_SC_FPOV_WE | 322 | number of cycles of pa sc FIFO write of pa0 first prim of vector |
| SC_PA1_SC_FPOV_WE | 323 | number of cycles of pa sc FIFO write of pa1 first prim of vector |
| SC_PA0_SC_LPOV_WE | 326 | number of cycles of pa sc FIFO write of pa0 last prim of vector |
| SC_PA1_SC_LPOV_WE | 327 | number of cycles of pa sc FIFO write of pa1 last prim of vector |
| SC_SC_SPI_DEALLOC_0_0 | 330 | number of spi cycles of pa0 originated dealloc bit 0 set |
| SC_SC_SPI_DEALLOC_0_1 | 331 | number of spi cycles of pa0 originated dealloc bit 1 set |
| SC_SC_SPI_DEALLOC_0_2 | 332 | number of spi cycles of pa0 originated dealloc bit 2 set |
| SC_SC_SPI_DEALLOC_1_0 | 333 | number of spi cycles of pa1 originated dealloc bit 0 set |
| SC_SC_SPI_DEALLOC_1_1 | 334 | number of spi cycles of pa1 originated dealloc bit 1 set |
| SC_SC_SPI_DEALLOC_1_2 | 335 | number of spi cycles of pa1 originated dealloc bit 2 set |
| SC_SC_SPI_DEALLOC_2_0 | 336 | number of spi cycles of pa2 originated dealloc bit 0 set |
| SC_SC_SPI_DEALLOC_2_1 | 337 | number of spi cycles of pa2 originated dealloc bit 1 set |
| SC_SC_SPI_DEALLOC_2_2 | 338 | number of spi cycles of pa2 originated dealloc bit 2 set |
| SC_SC_SPI_DEALLOC_3_0 | 339 | number of spi cycles of pa3 originated dealloc bit 0 set |
| SC_SC_SPI_DEALLOC_3_1 | 340 | number of spi cycles of pa3 originated dealloc bit 1 set |
| SC_SC_SPI_DEALLOC_3_2 | 341 | number of spi cycles of pa3 originated dealloc bit 2 set |
| SC_SC_SPI_FPOV_0 | 342 | number of spi cycles of pa0 originated first prim of vector |
| SC_SC_SPI_FPOV_1 | 343 | number of spi cycles of pa1 originated first prim of vector |
| SC_SC_SPI_FPOV_2 | 344 | number of spi cycles of pa2 originated first prim of vector |
| SC_SC_SPI_FPOV_3 | 345 | number of spi cycles of pa3 originated first prim of vector |
| SC_SC_SPI_EVENT | 346 | number of spi cycles of events |
| SC_PS_TS_EVENT_FIFO_PUSH | 347 | number of cycles of sc timestamp event FIFO writes |
| SC_PS_TS_EVENT_FIFO_POP | 348 | number of cycles of sc timestamp event FIFO reads |
| SC_PS_CTX_DONE_FIFO_PUSH | 349 | number of cycles of sc context done event FIFO writes |
| SC_PS_CTX_DONE_FIFO_POP | 350 | number of cycles of sc context done event FIFO reads |
| SC_MULTICYCLE_BUBBLE_FREEZE | 351 | number of cycles of sc prim setup back-pressure for multicycles prims sent with a empty FIFO slot behind the current phase with ENABLE_MULTICYCLE_BUBBLE_FREEZE set |
| SC_EOP_SYNC_WINDOW | 352 | number of cycles of after end of packet is read from the selected pa sc FIFO until end of packet broadcast synchronization is achieved across all pa sc FIFOs |
| SC_PA0_SC_NULL_WE | 353 | number of cycles of pa sc FIFO write of pa0 null primitives |
| SC_PA0_SC_NULL_DEALLOC_WE | 354 | number of cycles of pa sc FIFO write of pa0 null primitives with dealloc tokens set |
| SC_PA0_SC_DATA_FIFO_EOPG_RD | 355 | number of cycles of pa sc FIFO read of pa0 end of prim group |
| SC_PA0_SC_DATA_FIFO_EOP_RD | 356 | number of cycles of pa sc FIFO read of pa0 end of packet |
| SC_PA0_SC_DEALLOC_0_RD | 357 | number of cycles of pa sc FIFO read of pa0 dealloc bit 0 set |
| SC_PA0_SC_DEALLOC_1_RD | 358 | number of cycles of pa sc FIFO read of pa0 dealloc bit 1 set |
| SC_PA1_SC_DATA_FIFO_EOPG_RD | 359 | number of cycles of pa sc FIFO read of pa1 end of prim group |
| SC_PA1_SC_DATA_FIFO_EOP_RD | 360 | number of cycles of pa sc FIFO read of pa1 end of packet |
| SC_PA1_SC_DEALLOC_0_RD | 361 | number of cycles of pa sc FIFO read of pa1 dealloc bit 0 set |
| SC_PA1_SC_DEALLOC_1_RD | 362 | number of cycles of pa sc FIFO read of pa1 dealloc bit 1 set |
| SC_PA1_SC_NULL_WE | 363 | number of cycles of pa sc FIFO write of pa1 null primitives |
| SC_PA1_SC_NULL_DEALLOC_WE | 364 | number of cycles of pa sc FIFO write of pa1 null primitives with dealloc tokens set |
| SC_PS_PA0_SC_FIFO_EMPTY | 377 | number of cycles of pa0_sc_fifo_empty |
| SC_PS_PA0_SC_FIFO_FULL | 378 | number of cycles of pa0_sc_fifo_full |
| SC_PA0_PS_DATA_SEND | 379 | number of cycles of pa0 sc FIFO write |
| SC_PS_PA1_SC_FIFO_EMPTY | 380 | number of cycles of pa1_sc_fifo_empty |
| SC_PS_PA1_SC_FIFO_FULL | 381 | number of cycles of pa1_sc_fifo_full |
| SC_PA1_PS_DATA_SEND | 382 | number of cycles of pa1 sc FIFO write |
| SC_BUSY_PROCESSING_MULTICYCLE_PRIM | 389 | number of cycles of sc arbiter busy processing multi-cycle prim |
| SC_BUSY_CNT_NOT_ZERO | 390 | number of cycles of sc busy counter non zero |
| SC_BM_BUSY | 391 | number of cycles of sc backend mapper busy |
| SC_BACKEND_BUSY | 392 | number of cycles of sc backend busy |
| SC_SCF_SCB_INTERFACE_BUSY | 393 | number of cycles of scf scb interface busy |
| SC_SCB_BUSY | 394 | number of cycles of scb busy (second packer if present) |
| SC_PERF_SEL_SCORPIO_START | 395 | Start of Xbox One X-specific counters. |
| SC_STARVED_BY_PA_WITH_UNSELECTED_PA_NOT_EMPTY | 395 | SC is starved by PA with an unselected FIFO that is not empty (only use if more than one PA). |
| SC_STARVED_BY_PA_WITH_UNSELECTED_PA_FULL | 396 | SC is starved by PA with an unselected FIFO that is full (only use if more than one PA). |
| Counter | Value | Description |
|---|---|---|
| SX_PERF_SEL_PA_IDLE_CYCLES | 0 | Nr of cycles where PA was idle waiting to accept vectors from SX |
| SX_PERF_SEL_PA_REQ | 1 | Nr of PA requests received |
| SX_PERF_SEL_PA_POS | 2 | Nr of positions sent to the PA |
| SX_PERF_SEL_CLOCK | 3 | Nr of clocks where SX was busy in any way shape or form |
| SX_PERF_SEL_GATE_EN1 | 4 | Nr of clocks for register accesses |
| SX_PERF_SEL_GATE_EN2 | 5 | Nr of clocks for bus destination sort module |
| SX_PERF_SEL_GATE_EN3 | 6 | Nr of clocks for color exports |
| SX_PERF_SEL_GATE_EN4 | 7 | Nr of clocks for position exports |
| SX_PERF_SEL_SH_POS_STARVE | 8 | Nr of clocks SX is starved for Position data |
| SX_PERF_SEL_SH_COLOR_STARVE | 9 | Nr of clocks SX is starved for Color data |
| SX_PERF_SEL_SH_POS_STALL | 10 | Nr of clocks SX is being stalled by PA |
| SX_PERF_SEL_SH_COLOR_STALL | 11 | Nr of clocks SX is being stalled by at least one DB |
| SX_PERF_SEL_DB0_PIXELS | 12 | Number of pixels sent to the DB0 |
| SX_PERF_SEL_DB0_HALF_QUADS | 13 | Number of half quads sent to the DB0 |
| SX_PERF_SEL_DB0_PIXEL_STALL | 14 | Number of cycles where pixel traffic is stalled due to the DB0 |
| SX_PERF_SEL_DB0_PIXEL_IDLE | 15 | Number of cycles where the pixel traffic was idle to DB0 |
| SX_PERF_SEL_DB0_PRED_PIXELS | 16 | Nr of non predicated pixels sent to the DB0 |
| SX_PERF_SEL_DB1_PIXELS | 17 | Number of pixels sent to the DB1 |
| SX_PERF_SEL_DB1_HALF_QUADS | 18 | Number of half quads sent to the DB1 |
| SX_PERF_SEL_DB1_PIXEL_STALL | 19 | Number of cycles where pixel traffic is stalled due to the DB1 |
| SX_PERF_SEL_DB1_PIXEL_IDLE | 20 | Number of cycles where the pixel traffic was idle to DB1 |
| SX_PERF_SEL_DB1_PRED_PIXELS | 21 | Nr of non predicated pixels sent to the DB1 |
| SX_PERF_SEL_SCORPIO_START | 22 | Start of Xbox One X-specific counters. |
| SX_PERF_SEL_DB2_PIXELS | 22 | The number of pixels sent to DB2. |
| SX_PERF_SEL_DB2_HALF_QUADS | 23 | The number of half-quads sent to DB2. |
| SX_PERF_SEL_DB2_PIXEL_STALL | 24 | The number of cycles that pixel traffic is stalled due to DB2. |
| SX_PERF_SEL_DB2_PIXEL_IDLE | 25 | The number of cycles that pixel traffic is idle due to DB2. |
| SX_PERF_SEL_DB2_PRED_PIXELS | 26 | The number of non-predicated pixels sent to DB2. |
| SX_PERF_SEL_DB3_PIXELS | 27 | The number of pixels sent to DB3. |
| SX_PERF_SEL_DB3_HALF_QUADS | 28 | The number of half-quads sent to DB3. |
| SX_PERF_SEL_DB3_PIXEL_STALL | 29 | The number of cycles that pixel traffic is stalled due to DB3. |
| SX_PERF_SEL_DB3_PIXEL_IDLE | 30 | The number of cycles that pixel traffic is idle due to DB3. |
| SX_PERF_SEL_DB3_PRED_PIXELS | 31 | The number of non-predicated pixels sent to DB3. |
| Counter | Value | Description |
|---|---|---|
| SPI_PERF_VS_WINDOW_VALID | 0 | Clock count enabled by perfcounter_start event. |
| SPI_PERF_VS_BUSY | 1 | Number of clocks with outstanding waves (SPI or SH). |
| SPI_PERF_VS_FIRST_WAVE | 2 | Number of VS_IS_DS first waves |
| SPI_PERF_VS_LAST_WAVE | 3 | Number of VS_IS_DS last waves |
| SPI_PERF_VS_LSHS_DEALLOC | 4 | Number of VS_IS_DS lds dealloc waves |
| SPI_PERF_VS_PC_STALL | 5 | Number of clocks stalled due to pc space. |
| SPI_PERF_VS_POS0_STALL | 6 | Number of clocks stalled due to pos buf space in SH0. |
| SPI_PERF_VS_POS1_STALL | 7 | Number of clocks stalled due to pos buf space in SH1. |
| SPI_PERF_VS_CRAWLER_STALL | 8 | Number of clocks event/wave order FIFO is full |
| SPI_PERF_VS_EVENT_WAVE | 9 | Number of events and waves |
| SPI_PERF_VS_WAVE | 10 | Number of waves |
| SPI_PERF_VS_PERS_UPD_FULL0 | 11 | Number of clks VS persistent state update fifo0 is full. |
| SPI_PERF_VS_PERS_UPD_FULL1 | 12 | Number of clks VS persistent state update fifo1 is full. |
| SPI_PERF_VS_LATE_ALLOC_FULL | 13 | Number of clks VS late alloc FIFO is full. |
| SPI_PERF_VS_FIRST_SUBGRP | 14 | Number of first of subgroup waves |
| SPI_PERF_VS_LAST_SUBGRP | 15 | Number of last of subgroup waves |
| SPI_PERF_GS_WINDOW_VALID | 16 | Clock count enabled by perfcounter_start event. |
| SPI_PERF_GS_BUSY | 17 | Number of clocks with outstanding waves (SPI or SH). |
| SPI_PERF_GS_CRAWLER_STALL | 18 | Number of clocks event/wave order FIFO is full |
| SPI_PERF_GS_EVENT_WAVE | 19 | Number of events and waves |
| SPI_PERF_GS_WAVE | 20 | Number of waves |
| SPI_PERF_GS_PERS_UPD_FULL0 | 21 | Number of clks GS persistent state update fifo0 is full. |
| SPI_PERF_GS_PERS_UPD_FULL1 | 22 | Number of clks GS persistent state update fifo1 is full. |
| SPI_PERF_GS_FIRST_SUBGRP | 23 | Number of first of subgroup waves |
| SPI_PERF_GS_LAST_SUBGRP | 24 | Number of last of subgroup waves |
| SPI_PERF_ES_WINDOW_VALID | 25 | Clock count enabled by perfcounter_start event. |
| SPI_PERF_ES_BUSY | 26 | Number of clocks with outstanding waves (SPI or SH). |
| SPI_PERF_ES_CRAWLER_STALL | 27 | Number of clocks event/wave order FIFO is full |
| SPI_PERF_ES_FIRST_WAVE | 28 | Number of ES_IS_DS first waves |
| SPI_PERF_ES_LAST_WAVE | 29 | Number of ES_IS_DS last waves |
| SPI_PERF_ES_LSHS_DEALLOC | 30 | Number of ES_IS_DS lds dealloc waves |
| SPI_PERF_ES_EVENT_WAVE | 31 | Number of events and waves |
| SPI_PERF_ES_WAVE | 32 | Number of waves |
| SPI_PERF_ES_PERS_UPD_FULL0 | 33 | Number of clks ES persistent state update fifo0 is full. |
| SPI_PERF_ES_PERS_UPD_FULL1 | 34 | Number of clks ES persistent state update fifo1 is full. |
| SPI_PERF_ES_FIRST_SUBGRP | 35 | Number of first of subgroup waves |
| SPI_PERF_ES_LAST_SUBGRP | 36 | Number of last of subgroup waves |
| SPI_PERF_HS_WINDOW_VALID | 37 | Clock count enabled by perfcounter_start event. |
| SPI_PERF_HS_BUSY | 38 | Number of clocks with outstanding waves (SPI or SH). |
| SPI_PERF_HS_CRAWLER_STALL | 39 | Number of clocks event/wave order FIFO is full |
| SPI_PERF_HS_FIRST_WAVE | 40 | Number of first waves |
| SPI_PERF_HS_LAST_WAVE | 41 | Number of last waves |
| SPI_PERF_HS_LSHS_DEALLOC | 42 | Number of offchipHS lds dealloc waves |
| SPI_PERF_HS_EVENT_WAVE | 43 | Number of events and waves |
| SPI_PERF_HS_WAVE | 44 | Number of waves |
| SPI_PERF_HS_PERS_UPD_FULL0 | 45 | Number of clks HS persistent state update fifo0 is full. |
| SPI_PERF_HS_PERS_UPD_FULL1 | 46 | Number of clks HS persistent state update fifo1 is full. |
| SPI_PERF_LS_WINDOW_VALID | 47 | Clock count enabled by perfcounter_start event. |
| SPI_PERF_LS_BUSY | 48 | Number of clocks with outstanding waves (SPI or SH). |
| SPI_PERF_LS_CRAWLER_STALL | 49 | Number of clocks event/wave order FIFO is full |
| SPI_PERF_LS_FIRST_WAVE | 50 | Number of first waves |
| SPI_PERF_LS_LAST_WAVE | 51 | Number of last waves |
| SPI_PERF_OFFCHIP_LDS_STALL_LS | 52 | Number of clocks ls is stalled due to off-chip LDS |
| SPI_PERF_LS_EVENT_WAVE | 53 | Number of events and waves |
| SPI_PERF_LS_WAVE | 54 | Number of waves |
| SPI_PERF_LS_PERS_UPD_FULL0 | 55 | Number of clks LS persistent state update fifo0 is full. |
| SPI_PERF_LS_PERS_UPD_FULL1 | 56 | Number of clks LS persistent state update fifo1 is full. |
| SPI_PERF_CSG_WINDOW_VALID | 57 | Clock count enabled by perfcounter_start event. |
| SPI_PERF_CSG_BUSY | 58 | Number of clocks with outstanding waves (SPI or SH). |
| SPI_PERF_CSG_NUM_THREADGROUPS | 59 | Number of thread groups launched |
| SPI_PERF_CSG_CRAWLER_STALL | 60 | Number of clocks event/wave order FIFO is full |
| SPI_PERF_CSG_EVENT_WAVE | 61 | Number of events and waves |
| SPI_PERF_CSG_WAVE | 62 | Number of waves |
| SPI_PERF_CSN_WINDOW_VALID | 63 | Clock count enabled by perfcounter_start event. |
| SPI_PERF_CSN_BUSY | 64 | Number of clocks with outstanding waves (SPI or SH). |
| SPI_PERF_CSN_NUM_THREADGROUPS | 65 | Number of thread groups launched |
| SPI_PERF_CSN_CRAWLER_STALL | 66 | Number of clocks event/wave order FIFO is full |
| SPI_PERF_CSN_EVENT_WAVE | 67 | Number of events and waves |
| SPI_PERF_CSN_WAVE | 68 | Number of waves |
| SPI_PERF_PS_CTL_WINDOW_VALID | 69 | Clock count enabled by perfcounter_start event. |
| SPI_PERF_PS_CTL_BUSY | 70 | Number of clocks with outstanding waves (SPI or SH). |
| SPI_PERF_PS_CTL_ACTIVE | 71 | Number of clks ptr_buff is processing waves. |
| SPI_PERF_PS_CTL_DEALLOC_BIN0 | 72 | Count deallocs for SE matching bin0 range |
| SPI_PERF_PS_CTL_FPOS_BIN1_STALL | 73 | Number of clks stalled waiting for a VS done from SE matching bin1 |
| SPI_PERF_PS_CTL_EVENT_WAVE | 74 | Number of events and waves |
| SPI_PERF_PS_CTL_WAVE | 75 | Number of waves |
| SPI_PERF_PS_CTL_OPT_WAVE | 76 | Number of waves with center+centroid+bc_optimize |
| SPI_PERF_PS_CTL_PASS_BIN0 | 77 | Count waves with bin0 pc reads / attr, range is 1-16. |
| SPI_PERF_PS_CTL_PASS_BIN1 | 78 | Count waves with bin1 pc reads / attr, range is 1-16. |
| SPI_PERF_PS_CTL_FPOS_BIN2 | 79 | Count fpos for SE matching bin2 range |
| SPI_PERF_PS_CTL_PRIM_BIN0 | 80 | Count waves with bin0 prims, range is 1-16. |
| SPI_PERF_PS_CTL_PRIM_BIN1 | 81 | Count waves with bin1 prims, range is 1-16. |
| SPI_PERF_PS_CTL_CNF_BIN2 | 82 | Count waves with bin2 conflicts, range is 0-15 but max is 8. |
| SPI_PERF_PS_CTL_CNF_BIN3 | 83 | Count waves with bin3 conflicts, range is 0-15 but max is 8. |
| SPI_PERF_PS_CTL_CRAWLER_STALL | 84 | Number of clocks event/wave order FIFO is full |
| SPI_PERF_PS_CTL_LDS_RES_FULL | 85 | Number of clks PS per-wavefront storage is full |
| SPI_PERF_PS_PERS_UPD_FULL0 | 86 | Number of clks PS persistent state update fifo0 is full |
| SPI_PERF_PS_PERS_UPD_FULL1 | 87 | Number of clks PS persistent state update fifo1 is full |
| SPI_PERF_PIX_ALLOC_PEND_CNT | 88 | Sum of number of pixel alloc requests pending per clk |
| SPI_PERF_PIX_ALLOC_SCB_STALL | 89 | Number of clks pix alloc stalled due to color scoreboard. |
| SPI_PERF_PIX_ALLOC_DB0_STALL | 90 | Number of clks pix alloc stalled due to db0 color buffer. |
| SPI_PERF_PIX_ALLOC_DB1_STALL | 91 | Number of clks pix alloc stalled due to db1 color buffer. |
| SPI_PERF_PIX_ALLOC_DB2_STALL | 92 | Number of clks pix alloc stalled due to db2 color buffer. |
| SPI_PERF_PIX_ALLOC_DB3_STALL | 93 | Number of clks pix alloc stalled due to db3 color buffer. |
| SPI_PERF_LDS0_PC_VALID | 94 | Number of param cache reads sent to PC from SH0. |
| SPI_PERF_LDS1_PC_VALID | 95 | Number of param cache reads sent to PC from SH1. |
| SPI_PERF_RA_PIPE_REQ_BIN2 | 96 | Arb cycles with bin2 mqcs pipe_arb req =sum(csprobe0, csproben) |
| SPI_PERF_RA_TASK_REQ_BIN3 | 97 | Arb cycles with bin3 requests =sum(mqcs0, mqcsn) + dxcs +sum(gfx) |
| SPI_PERF_RA_WR_CTL_FULL | 98 | Arb cycles where RA is stalled due to wave launch fifo_full. |
| SPI_PERF_RA_REQ_NO_ALLOC | 99 | Arb cycles with requests but no allocation. |
| SPI_PERF_RA_REQ_NO_ALLOC_PS | 100 | Arb cycles with PS req and no PS alloc. |
| SPI_PERF_RA_REQ_NO_ALLOC_VS | 101 | Arb cycles with VS req and no VS alloc. |
| SPI_PERF_RA_REQ_NO_ALLOC_GS | 102 | Arb cycles with GS req and no GS alloc. |
| SPI_PERF_RA_REQ_NO_ALLOC_ES | 103 | Arb cycles with ES req and no ES alloc. |
| SPI_PERF_RA_REQ_NO_ALLOC_HS | 104 | Arb cycles with HS req and no HS alloc. |
| SPI_PERF_RA_REQ_NO_ALLOC_LS | 105 | Arb cycles with LS req and no LS alloc. |
| SPI_PERF_RA_REQ_NO_ALLOC_CSG | 106 | Arb cycles with CSg req and no CSg alloc. |
| SPI_PERF_RA_REQ_NO_ALLOC_CSN | 107 | Arb cycles with CSn req and no CSn alloc. |
| SPI_PERF_RA_RES_STALL_PS | 108 | Arb cycles with PS req and no PS fits. |
| SPI_PERF_RA_RES_STALL_VS | 109 | Arb cycles with VS req and no VS fits. |
| SPI_PERF_RA_RES_STALL_GS | 110 | Arb cycles with GS req and no GS fits. |
| SPI_PERF_RA_RES_STALL_ES | 111 | Arb cycles with ES req and no ES fits. |
| SPI_PERF_RA_RES_STALL_HS | 112 | Arb cycles with HS req and no HS fits. |
| SPI_PERF_RA_RES_STALL_LS | 113 | Arb cycles with LS req and no LS fits. |
| SPI_PERF_RA_RES_STALL_CSG | 114 | Arb cycles with CSg req and no CSg fits. |
| SPI_PERF_RA_RES_STALL_CSN | 115 | Arb cycles with CSn req and no CSn fits. |
| SPI_PERF_RA_TMP_STALL_PS | 116 | Cycles where ps wants to req but does not fit in temp space. |
| SPI_PERF_RA_TMP_STALL_VS | 117 | Cycles where vs wants to req but does not fit in temp space. |
| SPI_PERF_RA_TMP_STALL_GS | 118 | Cycles where gs wants to req but does not fit in temp space. |
| SPI_PERF_RA_TMP_STALL_ES | 119 | Cycles where es wants to req but does not fit in temp space. |
| SPI_PERF_RA_TMP_STALL_HS | 120 | Cycles where hs wants to req but does not fit in temp space. |
| SPI_PERF_RA_TMP_STALL_LS | 121 | Cycles where ls wants to req but does not fit in temp space. |
| SPI_PERF_RA_TMP_STALL_CSG | 122 | Cycles where csg wants to req but does not fit in temp space. |
| SPI_PERF_RA_TMP_STALL_CSN | 123 | Cycles where csn wants to req but does not fit in temp space. |
| SPI_PERF_RA_WAVE_SIMD_FULL_PS | 124 | Sum of SIMD where WAVE resource full when !ps_fits. |
| SPI_PERF_RA_WAVE_SIMD_FULL_VS | 125 | Sum of SIMD where WAVE resource full when !vs_fits. |
| SPI_PERF_RA_WAVE_SIMD_FULL_GS | 126 | Sum of SIMD where WAVE resource full when !gs_fits. |
| SPI_PERF_RA_WAVE_SIMD_FULL_ES | 127 | Sum of SIMD where WAVE resource full when !es_fits. |
| SPI_PERF_RA_WAVE_SIMD_FULL_HS | 128 | Sum of SIMD where WAVE can’t take hs wave when !fits. |
| SPI_PERF_RA_WAVE_SIMD_FULL_LS | 129 | Sum of SIMD where WAVE resource full when !ls_fits. |
| SPI_PERF_RA_WAVE_SIMD_FULL_CSG | 130 | Sum of SIMD where WAVE can’t take csg wave when !fits. |
| SPI_PERF_RA_WAVE_SIMD_FULL_CSN | 131 | Sum of SIMD where WAVE can’t take csn wave when !fits. |
| SPI_PERF_RA_VGPR_SIMD_FULL_PS | 132 | Sum of SIMD where VGPR can’t take ps wave when !fits. |
| SPI_PERF_RA_VGPR_SIMD_FULL_VS | 133 | Sum of SIMD where VGPR can’t take vs wave when !fits. |
| SPI_PERF_RA_VGPR_SIMD_FULL_GS | 134 | Sum of SIMD where VGPR can’t take gs wave when !fits. |
| SPI_PERF_RA_VGPR_SIMD_FULL_ES | 135 | Sum of SIMD where VGPR can’t take es wave when !fits. |
| SPI_PERF_RA_VGPR_SIMD_FULL_HS | 136 | Sum of SIMD where VGPR can’t take hs wave when !fits. |
| SPI_PERF_RA_VGPR_SIMD_FULL_LS | 137 | Sum of SIMD where VGPR can’t take ls wave when !fits. |
| SPI_PERF_RA_VGPR_SIMD_FULL_CSG | 138 | Sum of SIMD where VGPR can’t take csg wave when !fits. |
| SPI_PERF_RA_VGPR_SIMD_FULL_CSN | 139 | Sum of SIMD where VGPR can’t take csn wave when !fits. |
| SPI_PERF_RA_SGPR_SIMD_FULL_PS | 140 | Sum of SIMD where SGPR can’t take ps wave when !fits. |
| SPI_PERF_RA_SGPR_SIMD_FULL_VS | 141 | Sum of SIMD where SGPR can’t take vs wave when !fits. |
| SPI_PERF_RA_SGPR_SIMD_FULL_GS | 142 | Sum of SIMD where SGPR can’t take gs wave when !fits. |
| SPI_PERF_RA_SGPR_SIMD_FULL_ES | 143 | Sum of SIMD where SGPR can’t take es wave when !fits. |
| SPI_PERF_RA_SGPR_SIMD_FULL_HS | 144 | Sum of SIMD where SGPR can’t take hs wave when !fits. |
| SPI_PERF_RA_SGPR_SIMD_FULL_LS | 145 | Sum of SIMD where SGPR can’t take ls wave when !fits. |
| SPI_PERF_RA_SGPR_SIMD_FULL_CSG | 146 | Sum of SIMD where SGPR can’t take csg wave when !fits. |
| SPI_PERF_RA_SGPR_SIMD_FULL_CSN | 147 | Sum of SIMD where SGPR can’t take csn wave when !fits. |
| SPI_PERF_RA_LDS_CU_FULL_PS | 148 | Sum of CU where LDS can’t take ps wave when !fits. |
| SPI_PERF_RA_LDS_CU_FULL_LS | 149 | Sum of CU where LDS can’t take ls wave when !fits. |
| SPI_PERF_RA_LDS_CU_FULL_ES | 150 | Sum of CU where LDS can’t take es wave when !fits. |
| SPI_PERF_RA_LDS_CU_FULL_CSG | 151 | Sum of CU where LDS can’t take csg wave when !fits. |
| SPI_PERF_RA_LDS_CU_FULL_CSN | 152 | Sum of CU where LDS can’t take csn wave when !fits. |
| SPI_PERF_RA_BAR_CU_FULL_HS | 153 | Sum of CU where BARRIER can’t take hs wave when !fits. |
| SPI_PERF_RA_BAR_CU_FULL_CSG | 154 | Sum of CU where BARRIER can’t take csg wave when !fits. |
| SPI_PERF_RA_BAR_CU_FULL_CSN | 155 | Sum of CU where BARRIER can’t take csn wave when !fits. |
| SPI_PERF_RA_BULKY_CU_FULL_CSG | 156 | Sum of CU where BULKY can’t take csg wave when !fits. |
| SPI_PERF_RA_BULKY_CU_FULL_CSN | 157 | Sum of CU where BULKY can’t take csn wave when !fits. |
| SPI_PERF_RA_TGLIM_CU_FULL_CSG | 158 | Cycles where csg wants to req but all CU are at tg_limit |
| SPI_PERF_RA_TGLIM_CU_FULL_CSN | 159 | Cycles where csn wants to req but all CU are at tg_limit |
| SPI_PERF_RA_WVLIM_STALL_PS | 160 | Number of clocks ps is stalled due to WAVE LIMIT. |
| SPI_PERF_RA_WVLIM_STALL_VS | 161 | Number of clocks vs is stalled due to WAVE LIMIT. |
| SPI_PERF_RA_WVLIM_STALL_GS | 162 | Number of clocks gs is stalled due to WAVE LIMIT. |
| SPI_PERF_RA_WVLIM_STALL_ES | 163 | Number of clocks es is stalled due to WAVE LIMIT. |
| SPI_PERF_RA_WVLIM_STALL_HS | 164 | Number of clocks hs is stalled due to WAVE LIMIT. |
| SPI_PERF_RA_WVLIM_STALL_LS | 165 | Number of clocks ls is stalled due to WAVE LIMIT. |
| SPI_PERF_RA_WVLIM_STALL_CSG | 166 | Number of clocks csg is stalled due to WAVE LIMIT. |
| SPI_PERF_RA_WVLIM_STALL_CSN | 167 | Number of clocks csn is stalled due to WAVE LIMIT. |
| SPI_PERF_RA_PS_LOCK | 168 | Arb cycles PS has a CU locked when !fits and active_cnt < threshold. |
| SPI_PERF_RA_VS_LOCK | 169 | Arb cycles VS has a CU locked when !fits and active_cnt < threshold. |
| SPI_PERF_RA_GS_LOCK | 170 | Arb cycles GS has a CU locked when !fits and active_cnt < threshold. |
| SPI_PERF_RA_ES_LOCK | 171 | Arb cycles ES has a CU locked when !fits and active_cnt < threshold. |
| SPI_PERF_RA_HS_LOCK | 172 | Arb cycles HS has a CU locked when !fits and active_cnt < threshold. |
| SPI_PERF_RA_LS_LOCK | 173 | Arb cycles LS has a CU locked when !fits and active_cnt < threshold. |
| SPI_PERF_RA_CSG_LOCK | 174 | Arb cycles CSG has a CU locked when !fits and active_cnt < threshold. |
| SPI_PERF_RA_CSN_LOCK | 175 | Arb cycles CSN has a CU locked when !fits and active_cnt < threshold. |
| SPI_PERF_RA_RSV_UPD | 176 | Cycles spent doing updates for reservation changes |
| SPI_PERF_EXP_ARB_COL_CNT | 177 | Sum of number of color exp requests pending each clk |
| SPI_PERF_EXP_ARB_PAR_CNT | 178 | Sum of number of parameter exp requests pending each clk |
| SPI_PERF_EXP_ARB_POS_CNT | 179 | Sum of number of position exp requests pending each clk |
| SPI_PERF_EXP_ARB_GDS_CNT | 180 | Sum of number of gds exp requests pending each clk |
| SPI_PERF_CLKGATE_BUSY_STALL | 181 | Number of clocks with spi busy and not all_clocks_on |
| SPI_PERF_CLKGATE_ACTIVE_STALL | 182 | Number of clocks with spim_active and not all_clocks_on. |
| SPI_PERF_CLKGATE_ALL_CLOCKS_ON | 183 | Number of clocks with all_clocks_on. |
| SPI_PERF_CLKGATE_CGTT_DYN_ON | 184 | Number of clocks with spi cgtt dyn clk on. |
| SPI_PERF_CLKGATE_CGTT_REG_ON | 185 | Number of clocks with spi cgtt reg clk on. |
| Counter | Value | Description |
|---|---|---|
| SQ_PERF_SEL_NONE | 0 | Don’t count anything. |
| SQ_PERF_SEL_ACCUM_PREV | 1 | For counter N, increment by the value of counter N-1. Only accumulates once every 4 cycles. |
| SQ_PERF_SEL_CYCLES | 2 | Clock cycles. (nondeterministic, per-SIMD, global) |
| SQ_PERF_SEL_BUSY_CYCLES | 3 | Clock cycles while SQ is reporting that it is busy. (nondeterministic, per-SIMD, global) |
| SQ_PERF_SEL_WAVES | 4 | Count number of waves sent to SQs. (per-SIMD, emulated, global) |
| SQ_PERF_SEL_LEVEL_WAVES | 5 | Track the number of waves. Set ACCUM_PREV for the next counter to use this. (level, per-SIMD, global) |
| SQ_PERF_SEL_WAVES_EQ_64 | 6 | Count number of waves with exactly 64 active threads sent to SQs. (per-SIMD, emulated, global) |
| SQ_PERF_SEL_WAVES_LT_64 | 7 | Count number of waves with <64 active threads sent to SQs. (per-SIMD, emulated, global) |
| SQ_PERF_SEL_WAVES_LT_48 | 8 | Count number of waves with <48 active threads sent to SQs. (per-SIMD, emulated, global) |
| SQ_PERF_SEL_WAVES_LT_32 | 9 | Count number of waves sent <32 active threads sent to SQs. (per-SIMD, emulated, global) |
| SQ_PERF_SEL_WAVES_LT_16 | 10 | Count number of waves sent <16 active threads sent to SQs. (per-SIMD, emulated, global) |
| SQ_PERF_SEL_WAVES_CU | 11 | Count number of waves sent to CUs. (per-SIMD, emulated) |
| SQ_PERF_SEL_LEVEL_WAVES_CU | 12 | Track the number of waves. Set ACCUM_PREV for the next counter to use this. (level, per-SIMD) |
| SQ_PERF_SEL_BUSY_CU_CYCLES | 13 | Count quad-cycles each CU is busy. (nondeterministic, per-SIMD) |
| SQ_PERF_SEL_ITEMS | 14 | Number of valid items per wave. (per-SIMD, global) |
| SQ_PERF_SEL_QUADS | 15 | Number of completely or partially covered quads per wave. (per-SIMD, emulated, global) |
| SQ_PERF_SEL_EVENTS | 16 | Number of events. (unwindowed, emulated, global) |
| SQ_PERF_SEL_SURF_SYNCS | 17 | Number of surface syncs. (unwindowed, emulated, global) |
| SQ_PERF_SEL_TTRACE_REQS | 18 | Number of thread trace requests. (unwindowed, global, nondeterministic) |
| SQ_PERF_SEL_TTRACE_INFLIGHT_REQS | 19 | Number of in-flight thread trace requests. (nondeterministic, unwindowed, global) |
| SQ_PERF_SEL_TTRACE_STALL | 20 | Number of cycles thread trace stalls stalled execution. (nondeterministic, unwindowed, global) |
| SQ_PERF_SEL_MSG_CNTR | 21 | Number of counter messages. (nondeterministic, unwindowed, global) |
| SQ_PERF_SEL_MSG_PERF | 22 | Number of perfcounter messages. (nondeterministic, unwindowed), global |
| SQ_PERF_SEL_MSG_GSCNT | 23 | Number of geometry messages. (emulated, unwindowed, global) |
| SQ_PERF_SEL_MSG_INTERRUPT | 24 | Number of interrupt messages. (emulated, unwindowed, global) |
| SQ_PERF_SEL_INSTS | 25 | Number of instructions issued. (per-SIMD, emulated) |
| SQ_PERF_SEL_INSTS_VALU | 26 | Number of VALU instructions issued. (per-SIMD, emulated) |
| SQ_PERF_SEL_INSTS_VMEM_WR | 27 | Number of VMEM write instructions issued (including FLAT). (per-SIMD, emulated) |
| SQ_PERF_SEL_INSTS_VMEM_RD | 28 | Number of VMEM read instructions issued (including FLAT). (per-SIMD, emulated) |
| SQ_PERF_SEL_INSTS_VMEM | 29 | Number of VMEM instructions issued. (per-SIMD, emulated) |
| SQ_PERF_SEL_INSTS_SALU | 30 | Number of SALU instructions issued. (per-SIMD, emulated) |
| SQ_PERF_SEL_INSTS_SMEM | 31 | Number of SMEM read instructions issued. (per-SIMD, emulated) |
| SQ_PERF_SEL_INSTS_FLAT | 32 | Number of FLAT instructions issued. (per-SIMD, emulated) |
| SQ_PERF_SEL_INSTS_FLAT_LDS_ONLY | 33 | Number of FLAT instructions issued that read/wrote only from/to LDS (only works if EARLY_TA_DONE is enabled). (per-SIMD, emulated) |
| SQ_PERF_SEL_INSTS_LDS | 34 | Number of LDS instructions issued (including FLAT). (per-SIMD, emulated) |
| SQ_PERF_SEL_INSTS_GDS | 35 | Number of GDS instructions issued. (per-SIMD, emulated) |
| SQ_PERF_SEL_INSTS_EXP | 36 | Number of EXP instructions issued, excluding skipped export instructions. (per-SIMD, emulated) |
| SQ_PERF_SEL_INSTS_EXP_GDS | 37 | Number of EXP and GDS instructions issued, excluding skipped export instructions. (per-SIMD, emulated) |
| SQ_PERF_SEL_INSTS_BRANCH | 38 | Number of Branch instructions issued. (per-SIMD, emulated) |
| SQ_PERF_SEL_INSTS_SENDMSG | 39 | Number of Sendmsg instructions issued. (per-SIMD, emulated) |
| SQ_PERF_SEL_INSTS_VSKIPPED | 40 | Number of vector instructions skipped. (per-SIMD, emulated) |
| SQ_PERF_SEL_INST_LEVEL_VMEM | 41 | Number of in-flight VMEM instructions. Set next counter to ACCUM_PREV and divide by INSTS_VMEM for average latency. Includes FLAT instructions. (per-SIMD, level, nondeterministic) |
| SQ_PERF_SEL_INST_LEVEL_SMEM | 42 | Number of in-flight SMEM read instructions * 2. Set next counter to ACCUM_PREV and divide by INSTS_SMEM for average latency per smem request. Falls slightly short of total request latency because some fetches are divided into two requests that may finish at different times and this counter collects the average latency of the two. (per-SIMD, level, nondeterministic) |
| SQ_PERF_SEL_INST_LEVEL_LDS | 43 | Number of in-flight LDS instructions. Set next counter to ACCUM_PREV and divide by INSTS_LDS for average latency. Includes FLAT instructions. (per-SIMD, level, nondeterministic) |
| SQ_PERF_SEL_INST_LEVEL_GDS | 44 | Number of in-flight GDS instructions. Set next counter to ACCUM_PREV and divide by INSTS_GDS for average latency. (per-SIMD, level, nondeterministic) |
| SQ_PERF_SEL_INST_LEVEL_EXP | 45 | Number of in-flight EXP instructions. Set next counter to ACCUM_PREV and divide by INSTS_EXP for average latency. (per-SIMD, level, nondeterministic) |
| SQ_PERF_SEL_WAVE_CYCLES | 46 | Number of wave-cycles spent by waves in the CUs (per-SIMD, nondeterministic) |
| SQ_PERF_SEL_WAVE_READY | 47 | Number of wave-cycles waves were ready to execute the next instruction (per-SIMD, nondeterministic) |
| SQ_PERF_SEL_WAIT_CNT_VM | 48 | Number of wave-cycles spent waiting for VM counter. In units of 4 cycles. (per-SIMD, nondeterministic) |
| SQ_PERF_SEL_WAIT_CNT_LGKM | 49 | Number of wave-cycles spent waiting for LGKM counter. In units of 4 cycles. (per-SIMD, nondeterministic) |
| SQ_PERF_SEL_WAIT_CNT_EXP | 50 | Number of wave-cycles spent waiting for EXP counter. In units of 4 cycles. (per-SIMD, nondeterministic) |
| SQ_PERF_SEL_WAIT_CNT_ANY | 51 | Number of wave-cycles spent waiting for any counter. In units of 4 cycles. (per-SIMD, nondeterministic) |
| SQ_PERF_SEL_WAIT_BARRIER | 52 | Number of wave-cycles spent waiting for a barrier. In units of 4 cycles. (per-SIMD, nondeterministic) |
| SQ_PERF_SEL_WAIT_EXP_ALLOC | 53 | Number of wave-cycles spent waiting for export allocation (per-SIMD, nondeterministic) |
| SQ_PERF_SEL_WAIT_SLEEP | 54 | Number of wave-cycles spent waiting for sleep to finish (per-SIMD, nondeterministic) |
| SQ_PERF_SEL_WAIT_OTHER | 55 | Number of wave-cycles spent waiting for dependency stalls (per-SIMD, nondeterministic) |
| SQ_PERF_SEL_WAIT_ANY | 56 | Number of wave-cycles spent waiting for anything (per-SIMD, nondeterministic) |
| SQ_PERF_SEL_WAIT_TTRACE | 57 | Number of wave-cycles spent waiting for thread trace stalls (per-SIMD, nondeterministic) |
| SQ_PERF_SEL_WAIT_IFETCH | 58 | Number of wave-cycles spent waiting for instructions to arrive (per-SIMD, nondeterministic) |
| SQ_PERF_SEL_WAIT_INST_VMEM | 59 | Number of wave-cycles spent waiting for VMEM instruction issue. In units of 4 cycles. (per-SIMD, nondeterministic) |
| SQ_PERF_SEL_WAIT_INST_SCA | 60 | Number of wave-cycles spent waiting for SALU or SMEM instruction issue. In units of 4 cycles. (per-SIMD, nondeterministic) |
| SQ_PERF_SEL_WAIT_INST_LDS | 61 | Number of wave-cycles spent waiting for LDS instruction issue. In units of 4 cycles. (per-SIMD, nondeterministic) |
| SQ_PERF_SEL_WAIT_INST_VALU | 62 | Number of wave-cycles spent waiting for VALU instruction issue. In units of 4 cycles. (per-SIMD, nondeterministic) |
| SQ_PERF_SEL_WAIT_INST_EXP_GDS | 63 | Number of wave-cycles spent waiting for EXPORT or GDS instruction issue. In units of 4 cycles. (per-SIMD, nondeterministic) |
| SQ_PERF_SEL_WAIT_INST_MISC | 64 | Number of wave-cycles spent waiting for BRANCH or SENDMSG instruction issue. In units of 4 cycles. (per-SIMD, nondeterministic) |
| SQ_PERF_SEL_WAIT_INST_FLAT | 65 | Number of wave-cycles spent waiting for flat instruction issue. In units of 4 cycles. (per-SIMD, nondeterministic) |
| SQ_PERF_SEL_ACTIVE_INST_ANY | 66 | Number of cycles each wave is working on an instruction. (per-SIMD, nondeterministic) |
| SQ_PERF_SEL_ACTIVE_INST_VMEM | 67 | Number of cycles the SQ instruction arbiter is working on a VMEM instruction. (per-SIMD, nondeterministic) |
| SQ_PERF_SEL_ACTIVE_INST_LDS | 68 | Number of cycles the SQ instruction arbiter is working on a LDS instruction. (per-SIMD, nondeterministic) |
| SQ_PERF_SEL_ACTIVE_INST_VALU | 69 | Number of cycles the SQ instruction arbiter is working on a VALU instruction. (per-SIMD, nondeterministic) |
| SQ_PERF_SEL_ACTIVE_INST_SCA | 70 | Number of cycles the SQ instruction arbiter is working on a SALU or SMEM instruction. (per-SIMD, nondeterministic) |
| SQ_PERF_SEL_ACTIVE_INST_EXP_GDS | 71 | Number of cycles the SQ instruction arbiter is working on an EXPORT or GDS instruction. (per-SIMD, nondeterministic) |
| SQ_PERF_SEL_ACTIVE_INST_MISC | 72 | Number of cycles the SQ instruction arbiter is working on a BRANCH or SENDMSG instruction. (per-SIMD, nondeterministic) |
| SQ_PERF_SEL_ACTIVE_INST_FLAT | 73 | Number of cycles the SQ instruction arbiter is working on a FLAT instruction. (per-SIMD, nondeterministic) |
| SQ_PERF_SEL_INST_CYCLES_VMEM_WR | 74 | Number of cycles needed to send addr and cmd data for VMEM write instructions. (per-SIMD, emulated) |
| SQ_PERF_SEL_INST_CYCLES_VMEM_RD | 75 | Number of cycles needed to send addr and cmd data for VMEM read instructions. (per-SIMD, emulated) |
| SQ_PERF_SEL_INST_CYCLES_VMEM_ADDR | 76 | Number of cycles needed to send VMEM read/write addresses to TA. (per-SIMD, emulated) |
| SQ_PERF_SEL_INST_CYCLES_VMEM_DATA | 77 | Number of cycles needed to send VMEM write data to TA. (per-SIMD, emulated) |
| SQ_PERF_SEL_INST_CYCLES_VMEM_CMD | 78 | Number of cycles needed to send VMEM command to TA. (per-SIMD, emulated) |
| SQ_PERF_SEL_INST_CYCLES_VMEM | 79 | Number of cycles needed to send addr and cmd data for VMEM instructions. (per-SIMD, emulated) |
| SQ_PERF_SEL_INST_CYCLES_LDS | 80 | Number of cycles needed to send instructions to LDS. (per-SIMD, emulated) |
| SQ_PERF_SEL_INST_CYCLES_VALU | 81 | Number of cycles needed to execute VALU instructions. (per-SIMD, emulated) |
| SQ_PERF_SEL_INST_CYCLES_EXP | 82 | Number of cycles needed to export data for EXP instructions. (per-SIMD, emulated) |
| SQ_PERF_SEL_INST_CYCLES_GDS | 83 | Number of cycles needed to export data for GDS instructions. (per-SIMD, emulated) |
| SQ_PERF_SEL_INST_CYCLES_SCA | 84 | Number of cycles needed to execute scalar instructions. (per-SIMD, emulated) |
| SQ_PERF_SEL_INST_CYCLES_SMEM | 85 | Number of cycles needed to execute scalar memory reads. (per-SIMD, emulated) |
| SQ_PERF_SEL_INST_CYCLES_SALU | 86 | Number of cycles needed to execute non-memory read scalar operations. (per-SIMD, emulated) |
| SQ_PERF_SEL_INST_CYCLES_EXP_GDS | 87 | Number of cycles needed to export data for both EXP or GDS instructions (same as than INST_CYCLES_EXP + INST_CYCLES_GDS). (per-SIMD, emulated) |
| SQ_PERF_SEL_INST_CYCLES_MISC | 88 | Number of cycles needed to execute BRANCH or SENDMSG instructions. (per-SIMD, emulated) |
| SQ_PERF_SEL_THREAD_CYCLES_VALU | 89 | Number of thread-cycles used to execute VALU operations (similar to INST_CYCLES_VALU but multiplied by # of active threads). (per-SIMD) |
| SQ_PERF_SEL_THREAD_CYCLES_VALU_MAX | 90 | Maximum number of thread-cycles VALU operations that could have been executed given the instruction mix (similar to INST_CYCLES_VALU but multiplied by # of active threads). (per-SIMD, emulated) |
| SQ_PERF_SEL_IFETCH | 91 | Number of instruction fetch requests from cache. (per-SIMD, emulated) |
| SQ_PERF_SEL_IFETCH_LEVEL | 92 | Number of instruction fetch requests from cache. (per-SIMD, level) |
| SQ_PERF_SEL_CBRANCH_FORK | 93 | Number of conditional branch instructions. (per-SIMD, emulated) |
| SQ_PERF_SEL_CBRANCH_FORK_SPLIT | 94 | Number of conditional branch instructions that take both branches. (per-SIMD, emulated) |
| SQ_PERF_SEL_VALU_LDS_DIRECT_RD | 95 | Number of LDS direct reads issued (excluding interpolation ops). (per-SIMD, emulated) |
| SQ_PERF_SEL_VALU_LDS_INTERP_OP | 96 | Number of LDS interpolation ops issued. (per-SIMD, emulated) |
| SQ_PERF_SEL_LDS_BANK_CONFLICT | 97 | Number of cycles LDS is stalled by bank conflicts. (emulated) |
| SQ_PERF_SEL_LDS_ADDR_CONFLICT | 98 | Number of cycles LDS is stalled by address conflicts. (emulated,nondeterministic) |
| SQ_PERF_SEL_LDS_UNALIGNED_STALL | 99 | Number of cycles LDS is stalled processing flat unaligned load/store ops. (emulated) |
| SQ_PERF_SEL_LDS_MEM_VIOLATIONS | 100 | Number of threads that have a memory violation in the LDS.(emulated) |
| SQ_PERF_SEL_LDS_ATOMIC_RETURN | 101 | Number of atomic return cycles in LDS. (per-SIMD, emulated) |
| SQ_PERF_SEL_LDS_IDX_ACTIVE | 102 | Number of cycles LDS is used for indexed (non-direct,non-interpolation) operations. (per-SIMD, emulated) |
| SQ_PERF_SEL_VALU_DEP_STALL | 103 | Number of cycles VALU is stalled by previous instructions due to dependencies (wait state count). (nondeterministic) |
| SQ_PERF_SEL_VALU_STARVE | 104 | Number of cycles VALU is starved while waves are present. (per-SIMD, nondeterministic) |
| SQ_PERF_SEL_EXP_REQ_FIFO_FULL | 105 | Number of cycles export request FIFO is full. (nondeterministic, unwindowed) |
| SQ_PERF_SEL_LDS_BACK2BACK_STALL | 106 | Number of cycles LDS command stalled due to back to back requests. (nondeterministic, unwindowed) |
| SQ_PERF_SEL_LDS_DATA_FIFO_FULL | 107 | Number of cycles LDS data FIFO is full. (nondeterministic, unwindowed) |
| SQ_PERF_SEL_LDS_CMD_FIFO_FULL | 108 | Number of cycles LDS command FIFO is full. (nondeterministic, unwindowed) |
| SQ_PERF_SEL_VMEM_BACK2BACK_STALL | 109 | Number of cycles texture requests are stalled due to back to back requests. (nondeterministic, unwindowed) |
| SQ_PERF_SEL_VMEM_TA_ADDR_FIFO_FULL | 110 | Number of cycles texture requests are stalled due to full address FIFO in TA. (nondeterministic, unwindowed) |
| SQ_PERF_SEL_VMEM_TA_CMD_FIFO_FULL | 111 | Number of cycles texture requests are stalled due to full cmd FIFO in TA. (nondeterministic, unwindowed). |
| SQ_PERF_SEL_VMEM_EX_DATA_REG_BUSY | 112 | Number of cycles texture requests are stalled due to full data staging register in EX. (nondeterministic, unwindowed). |
| SQ_PERF_SEL_VMEM_WR_BACK2BACK_STALL | 113 | Number of cycles texture writes are stalled due to back to back writes. (nondeterministic, unwindowed) |
| SQ_PERF_SEL_VMEM_WR_TA_DATA_FIFO_FULL | 114 | Number of cycles texture writes are stalled due to full data FIFO in TA. (nondeterministic, unwindowed) |
| SQ_PERF_SEL_VALU_SRC_C_CONFLICT | 115 | Number of cycles VALU is stalled by arbitration due to src c conflict. (nondeterministic, unwindowed) |
| SQ_PERF_SEL_VMEM_RD_SRC_CD_CONFLICT | 116 | Number of cycles VMEM_RD instructions are stalled due to VGPR port conflicts. (nondeterministic, unwindowed) |
| SQ_PERF_SEL_VMEM_WR_SRC_CD_CONFLICT | 117 | Number of cycles VMEM_WR instructions are stalled due to VGPR port conflicts. (nondeterministic, unwindowed) |
| SQ_PERF_SEL_FLAT_SRC_CD_CONFLICT | 118 | Number of cycles FLAT instructions are stalled due to VGPR port conflicts. (nondeterministic, unwindowed) |
| SQ_PERF_SEL_LDS_SRC_CD_CONFLICT | 119 | Number of cycles LDS instructions are stalled due to VGPR port conflicts. (nondeterministic, unwindowed) |
| SQ_PERF_SEL_SRC_CD_BUSY | 120 | Number of total accesses of port C and D (up to 2 per cycle). (per-SIMD) |
| SQ_PERF_SEL_PT_POWER_STALL | 121 | Number of cycles the power throttle indicated ALU should stall (ignores whether something actually got stalled). (nondeterministic, unwindowed) |
| SQ_PERF_SEL_USER0 | 122 | User counter. Incremented by S_INCPERFLEVEL when SIMM16[3:0] == 0. (emulated) |
| SQ_PERF_SEL_USER1 | 123 | User counter. Incremented by S_INCPERFLEVEL when SIMM16[3:0] == 1. (emulated) |
| SQ_PERF_SEL_USER2 | 124 | User counter. Incremented by S_INCPERFLEVEL when SIMM16[3:0] == 2. (emulated) |
| SQ_PERF_SEL_USER3 | 125 | User counter. Incremented by S_INCPERFLEVEL when SIMM16[3:0] == 3. (emulated) |
| SQ_PERF_SEL_USER4 | 126 | User counter. Incremented by S_INCPERFLEVEL when SIMM16[3:0] == 4. (emulated) |
| SQ_PERF_SEL_USER5 | 127 | User counter. Incremented by S_INCPERFLEVEL when SIMM16[3:0] == 5. (emulated) |
| SQ_PERF_SEL_USER6 | 128 | User counter. Incremented by S_INCPERFLEVEL when SIMM16[3:0] == 6. (emulated) |
| SQ_PERF_SEL_USER7 | 129 | User counter. Incremented by S_INCPERFLEVEL when SIMM16[3:0] == 7. (emulated) |
| SQ_PERF_SEL_USER8 | 130 | User counter. Incremented by S_INCPERFLEVEL when SIMM16[3:0] == 8. (emulated) |
| SQ_PERF_SEL_USER9 | 131 | User counter. Incremented by S_INCPERFLEVEL when SIMM16[3:0] == 9. (emulated) |
| SQ_PERF_SEL_USER10 | 132 | User counter. Incremented by S_INCPERFLEVEL when SIMM16[3:0] == 10. (emulated) |
| SQ_PERF_SEL_USER11 | 133 | User counter. Incremented by S_INCPERFLEVEL when SIMM16[3:0] == 11. (emulated) |
| SQ_PERF_SEL_USER12 | 134 | User counter. Incremented by S_INCPERFLEVEL when SIMM16[3:0] == 12. (emulated) |
| SQ_PERF_SEL_USER13 | 135 | User counter. Incremented by S_INCPERFLEVEL when SIMM16[3:0] == 13. (emulated) |
| SQ_PERF_SEL_USER14 | 136 | User counter. Incremented by S_INCPERFLEVEL when SIMM16[3:0] == 14. (emulated) |
| SQ_PERF_SEL_USER15 | 137 | User counter. Incremented by S_INCPERFLEVEL when SIMM16[3:0] == 15. (emulated) |
| SQ_PERF_SEL_USER_LEVEL0 | 138 | User level counter. Incremented by S_INCPERFLEVEL, decremented by S_DECPERFLEVEL when SIMM16[3:0] == 0. (level, emulated) |
| SQ_PERF_SEL_USER_LEVEL1 | 139 | User level counter. Incremented by S_INCPERFLEVEL, decremented by S_DECPERFLEVEL SIMM16[3:0] == 1. (level, emulated) |
| SQ_PERF_SEL_USER_LEVEL2 | 140 | User level counter. Incremented by S_INCPERFLEVEL, decremented by S_DECPERFLEVEL SIMM16[3:0] == 2. (level, emulated) |
| SQ_PERF_SEL_USER_LEVEL3 | 141 | User level counter. Incremented by S_INCPERFLEVEL, decremented by S_DECPERFLEVEL SIMM16[3:0] == 3. (level, emulated) |
| SQ_PERF_SEL_USER_LEVEL4 | 142 | User level counter. Incremented by S_INCPERFLEVEL, decremented by S_DECPERFLEVEL SIMM16[3:0] == 4. (level, emulated) |
| SQ_PERF_SEL_USER_LEVEL5 | 143 | User level counter. Incremented by S_INCPERFLEVEL, decremented by S_DECPERFLEVEL SIMM16[3:0] == 5. (level, emulated) |
| SQ_PERF_SEL_USER_LEVEL6 | 144 | User level counter. Incremented by S_INCPERFLEVEL, decremented by S_DECPERFLEVEL SIMM16[3:0] == 6. (level, emulated) |
| SQ_PERF_SEL_USER_LEVEL7 | 145 | User level counter. Incremented by S_INCPERFLEVEL, decremented by S_DECPERFLEVEL SIMM16[3:0] == 7. (level, emulated) |
| SQ_PERF_SEL_USER_LEVEL8 | 146 | User level counter. Incremented by S_INCPERFLEVEL, decremented by S_DECPERFLEVEL SIMM16[3:0] == 8. (level, emulated) |
| SQ_PERF_SEL_USER_LEVEL9 | 147 | User level counter. Incremented by S_INCPERFLEVEL, decremented by S_DECPERFLEVEL SIMM16[3:0] == 9. (level, emulated) |
| SQ_PERF_SEL_USER_LEVEL10 | 148 | User level counter. Incremented by S_INCPERFLEVEL, decremented by S_DECPERFLEVEL SIMM16[3:0] == 10. (level, emulated) |
| SQ_PERF_SEL_USER_LEVEL11 | 149 | User level counter. Incremented by S_INCPERFLEVEL, decremented by S_DECPERFLEVEL SIMM16[3:0] == 11. (level, emulated) |
| SQ_PERF_SEL_USER_LEVEL12 | 150 | User level counter. Incremented by S_INCPERFLEVEL, decremented by S_DECPERFLEVEL SIMM16[3:0] == 12. (level, emulated) |
| SQ_PERF_SEL_USER_LEVEL13 | 151 | User level counter. Incremented by S_INCPERFLEVEL, decremented by S_DECPERFLEVEL SIMM16[3:0] == 13. (level, emulated) |
| SQ_PERF_SEL_USER_LEVEL14 | 152 | User level counter. Incremented by S_INCPERFLEVEL, decremented by S_DECPERFLEVEL SIMM16[3:0] == 14. (level, emulated) |
| SQ_PERF_SEL_USER_LEVEL15 | 153 | User level counter. Incremented by S_INCPERFLEVEL, decremented by S_DECPERFLEVEL SIMM16[3:0] == 15. (level, emulated) |
| SQ_PERF_SEL_POWER_VALU | 154 | Number of CAC pulses for ALU instructions (nondeterministic, unwindowed) |
| SQ_PERF_SEL_POWER_VALU0 | 155 | Number of CAC pulses for ALU instructions in group 0 (nondeterministic, unwindowed) |
| SQ_PERF_SEL_POWER_VALU1 | 156 | Number of CAC pulses for ALU instructions in group 1 (nondeterministic, unwindowed) |
| SQ_PERF_SEL_POWER_VALU2 | 157 | Number of CAC pulses for ALU instructions in group 2 (nondeterministic, unwindowed) |
| SQ_PERF_SEL_POWER_GPR_RD | 158 | Number of CAC pulses for GPR reads (nondeterministic, unwindowed) |
| SQ_PERF_SEL_POWER_GPR_WR | 159 | Number of CAC pulses for GPR writes (nondeterministic, unwindowed) |
| SQ_PERF_SEL_POWER_LDS_BUSY | 160 | Number of CAC pulses for cycles LDS clocks are active (nondeterministic, unwindowed) |
| SQ_PERF_SEL_POWER_ALU_BUSY | 161 | Number of CAC pulses for cycles ALU clocks are active (nondeterministic, unwindowed) |
| SQ_PERF_SEL_POWER_TEX_BUSY | 162 | Number of CAC pulses for cycles texture clocks are active (nondeterministic, unwindowed) |
| SQ_PERF_SEL_ACCUM_PREV_HIRES | 163 | For counter N, increment by the value of counter N-1. |
| SQ_PERF_SEL_SCORPIO_RANGE_START | 164 | Start of Xbox One X-specific counter range |
| SQ_PERF_SEL_INSTS_VALU_TRANS | ||
| 164 | ||
| The number of VALU transcendental instructions issued (per-SIMD, emulated). | ||
| SQ_PERF_SEL_DUMMY_LAST | 167 | A placeholder to separate SQ from SQC counters. Not a real performance counter. |
| SQC_PERF_SEL_ICACHE_INPUT_VALID_READY | 168 | Successful Transaction (per-SQ, nondeterministic, unwindowed) |
| SQC_PERF_SEL_ICACHE_INPUT_VALID_READYB | 169 | Input stalled by SQC (per-SQ, nondeterministic, unwindowed) |
| SQC_PERF_SEL_ICACHE_INPUT_VALIDB | 170 | SQC starved (per-SQ, nondeterministic, unwindowed) |
| SQC_PERF_SEL_DCACHE_INPUT_VALID_READY | 171 | Successful Transaction (per-SQ, nondeterministic, unwindowed) |
| SQC_PERF_SEL_DCACHE_INPUT_VALID_READYB | 172 | Input stalled by SQC (per-SQ, nondeterministic, unwindowed) |
| SQC_PERF_SEL_DCACHE_INPUT_VALIDB | 173 | SQC starved (per-SQ, nondeterministic, unwindowed) |
| SQC_PERF_SEL_TC_REQ | 174 | Total number of TC requests that were issued by instruction and constant caches. (No-Masking, nondeterministic) |
| SQC_PERF_SEL_TC_INST_REQ | 175 | Number of insruction requests to the TC (No-Masking, nondeterministic) |
| SQC_PERF_SEL_TC_DATA_REQ | 176 | Number of data requests to the TC (No-Masking, nondeterministic) |
| SQC_PERF_SEL_TC_STALL | 177 | Valid request stalled TC request interface (no-credits). (No-Masking, nondeterministic, unwindowed) |
| SQC_PERF_SEL_TC_STARVE | 178 | No requests sent to TC while credits available. (No-Masking, nondeterministic, unwindowed) |
| SQC_PERF_SEL_ICACHE_BUSY_CYCLES | 179 | Clock cycles while cache is reporting that it is busy. (No-Masking, nondeterministic, unwindowed) |
| SQC_PERF_SEL_ICACHE_REQ | 180 | Number of requests. (per-SQ, per-Bank) |
| SQC_PERF_SEL_ICACHE_HITS | 181 | Number of cache hits. (per-SQ, per-Bank, nondeterministic) |
| SQC_PERF_SEL_ICACHE_MISSES | 182 | Number of cache misses, includes uncached requests. (per-SQ, per-Bank, nondeterministic) |
| SQC_PERF_SEL_ICACHE_MISSES_DUPLICATE | 183 | Number of misses that were duplicates (access to a non-resident, miss pending CL). (per-SQ, per-Bank, nondeterministic) |
| SQC_PERF_SEL_ICACHE_UNCACHED | 184 | Number of uncached cache requests. (per-SQ, per-Bank) |
| SQC_PERF_SEL_ICACHE_VOLATILE | 185 | Number of volatile cache requests. (per-SQ, per-Bank) |
| SQC_PERF_SEL_ICACHE_INVAL_INST | 186 | Number of cache invalidations caused by instructions (No-Masking) |
| SQC_PERF_SEL_ICACHE_INVAL_ASYNC | 187 | Number of asynchronous invalidates (surface sync, register write, etc.) (No-Masking, unwindowed) |
| SQC_PERF_SEL_ICACHE_INVAL_VOLATILE_INST | 188 | Number of volatile cache invalidations caused by instructions (No-Masking) |
| SQC_PERF_SEL_ICACHE_INVAL_VOLATILE_ASYNC | 189 | Number of asynchronous volatile invalidates (surface sync, register write, etc.) (No-Masking, unwindowed) |
| SQC_PERF_SEL_ICACHE_INPUT_STALL_ARB_NO_GRANT | 190 | Number of arbitration stalls due to lost arbitration to other SQs. (per-SQ, nondeterministic, unwindowed) |
| SQC_PERF_SEL_ICACHE_INPUT_STALL_BANK_READYB | 191 | Number of arbitration stalls due to bank stalling. (per-SQ, nondeterministic, unwindowed) |
| SQC_PERF_SEL_ICACHE_CACHE_STALLED | 192 | Number of cache stalls. (per-Bank, nondeterministic, unwindowed) |
| SQC_PERF_SEL_ICACHE_CACHE_STALL_INFLIGHT_NONZERO | 193 | Number of cycles stalled because allocated line has nonzero in-flight. (per-Bank, nondeterministic, unwindowed) |
| SQC_PERF_SEL_ICACHE_CACHE_STALL_INFLIGHT_MAX | 194 | Number of cycles stalled on a hit CL but in-flight counter is maxed out. (per-Bank, nondeterministic, unwindowed) |
| SQC_PERF_SEL_ICACHE_CACHE_STALL_VOLATILE_MISMATCH | 195 | Number of cycles stalled because the volatile bit mismatched. (per-Bank, nondeterministic, unwindowed) |
| SQC_PERF_SEL_ICACHE_CACHE_STALL_UNCACHED_HIT | 196 | Number of cycles stalled because an uncached access had a tag match. (per-Bank, nondeterministic, unwindowed) |
| SQC_PERF_SEL_ICACHE_CACHE_STALL_OUTPUT | 197 | Number of cycles stalled at cache controller output. (per-Bank, nondeterministic, unwindowed) |
| SQC_PERF_SEL_ICACHE_CACHE_STALL_OUTPUT_MISS_FIFO | 198 | Number of cycles stalled because the miss FIFO is full. (per-Bank, nondeterministic, unwindowed) |
| SQC_PERF_SEL_ICACHE_CACHE_STALL_OUTPUT_HIT_FIFO | 199 | Number of cycles stalled because a hit FIFO is full. (per-Bank, nondeterministic, unwindowed) |
| SQC_PERF_SEL_ICACHE_CACHE_STALL_OUTPUT_TC_IF | 200 | Number of cycles stalled because the TC request interface is stalled. (per-Bank, nondeterministic, unwindowed) |
| SQC_PERF_SEL_ICACHE_STALL_OUTXBAR_ARB_NO_GRANT | 201 | Number of arbitration stalls due to lost arbitration for the output port. (per-Bank, nondeterministic, unwindowed) |
| SQC_PERF_SEL_DCACHE_BUSY_CYCLES | 202 | Clock cycles while cache is reporting that it is busy. (No-Masking, nondeterministic, unwindowed) |
| SQC_PERF_SEL_DCACHE_REQ | 203 | Number of requests (post-bank-serialization). (per-SQ, per-Bank) |
| SQC_PERF_SEL_DCACHE_HITS | 204 | Number of cache hits. (per-SQ, per-Bank, nondeterministic) |
| SQC_PERF_SEL_DCACHE_MISSES | 205 | Number of cache misses, includes uncached requests. (per-SQ, per-Bank, nondeterministic) |
| SQC_PERF_SEL_DCACHE_MISSES_DUPLICATE | 206 | Number of misses that were duplicates (access to a non-resident, miss pending CL). (per-SQ, per-Bank, nondeterministic) |
| SQC_PERF_SEL_DCACHE_UNCACHED | 207 | Number of uncached cache requests. (per-SQ, per-Bank) |
| SQC_PERF_SEL_DCACHE_VOLATILE | 208 | Number of volatile cache requests. (per-SQ, per-Bank) |
| SQC_PERF_SEL_DCACHE_INVAL_INST | 209 | Number of cache invalidations caused by instructions (No-Masking) |
| SQC_PERF_SEL_DCACHE_INVAL_ASYNC | 210 | Number of asynchronous invalidates (surface sync, register write, etc.) (No-Masking, unwindowed) |
| SQC_PERF_SEL_DCACHE_INVAL_VOLATILE_INST | 211 | Number of volatile cache invalidations caused by instructions (No-Masking) |
| SQC_PERF_SEL_DCACHE_INVAL_VOLATILE_ASYNC | 212 | Number of asynchronous volatile invalidates (surface sync, register write, etc.) (No-Masking, unwindowed) |
| SQC_PERF_SEL_DCACHE_INPUT_STALL_ARB_NO_GRANT | 213 | Number of arbitration stalls due to lost arbitration to other SQs. (per-SQ, nondeterministic, unwindowed) |
| SQC_PERF_SEL_DCACHE_INPUT_STALL_BANK_READYB | 214 | Number of arbitration stalls due to bank stalling. (per-SQ, nondeterministic, unwindowed) |
| SQC_PERF_SEL_DCACHE_CACHE_STALLED | 215 | Number of cache stalls. (per-Bank, nondeterministic, unwindowed) |
| SQC_PERF_SEL_DCACHE_CACHE_STALL_INFLIGHT_NONZERO | 216 | Number of cycles stalled because allocated line has nonzero in-flight. (per-Bank, nondeterministic, unwindowed) |
| SQC_PERF_SEL_DCACHE_CACHE_STALL_INFLIGHT_MAX | 217 | Number of cycles stalled on a hit CL but in-flight counter is maxed out. (per-Bank, nondeterministic, unwindowed) |
| SQC_PERF_SEL_DCACHE_CACHE_STALL_VOLATILE_MISMATCH | 218 | Number of cycles stalled because the volatile bit mismatched. (per-Bank, nondeterministic, unwindowed) |
| SQC_PERF_SEL_DCACHE_CACHE_STALL_UNCACHED_HIT | 219 | Number of cycles stalled because an uncached access had a tag match. (per-Bank, nondeterministic, unwindowed) |
| SQC_PERF_SEL_DCACHE_CACHE_STALL_OUTPUT | 220 | Number of cycles stalled at cache controller output. (per-Bank, nondeterministic, unwindowed) |
| SQC_PERF_SEL_DCACHE_CACHE_STALL_OUTPUT_MISS_FIFO | 221 | Number of cycles stalled because the miss FIFO is full. (per-Bank, nondeterministic, unwindowed) |
| SQC_PERF_SEL_DCACHE_CACHE_STALL_OUTPUT_HIT_FIFO | 222 | Number of cycles stalled because a hit FIFO is full. (per-Bank, nondeterministic, unwindowed) |
| SQC_PERF_SEL_DCACHE_CACHE_STALL_OUTPUT_TC_IF | 223 | Number of cycles stalled because the TC request interface is stalled. (per-Bank, nondeterministic, unwindowed) |
| SQC_PERF_SEL_DCACHE_STALL_OUTXBAR_ARB_NO_GRANT | 224 | Number of arbitration stalls due to lost arbitration for the output port. (per-Bank, nondeterministic, unwindowed) |
| SQC_PERF_SEL_DCACHE_REQ_1 | 225 | Number of constant cache 1 dw requests. (per-SQ) |
| SQC_PERF_SEL_DCACHE_REQ_2 | 226 | Number of constant cache 2 dw requests. (per-SQ) |
| SQC_PERF_SEL_DCACHE_REQ_4 | 227 | Number of constant cache 4 dw requests. (per-SQ) |
| SQC_PERF_SEL_DCACHE_REQ_8 | 228 | Number of constant cache 8 dw requests. (per-SQ) |
| SQC_PERF_SEL_DCACHE_REQ_16 | 229 | Number of constant cache 16 dw requests. (per-SQ) |
| SQC_PERF_SEL_DCACHE_REQ_TIME | 230 | Number of constant cache timestamp requests. (per-SQ) |
| SQC_PERF_SEL_SQ_DCACHE_REQS | 231 | Number of constant requests from SQ before any serialization (per-SQ) |
| SQC_PERF_SEL_DCACHE_FLAT_REQ | 232 | Number of constant flat requests (per-SQ) |
| SQC_PERF_SEL_DCACHE_NONFLAT_REQ | 233 | Number of constant non-flat requests (per-SQ) |
| SQC_PERF_SEL_ICACHE_INFLIGHT_LEVEL | 234 | Level Counter: # total outstanding transactions in instruction cache (per-SQ, nondeterministic) |
| SQC_PERF_SEL_ICACHE_PRE_CC_LEVEL | 235 | Level Counter: Pre-Cache Controller # outstanding transactions (per-SQ, nondeterministic) |
| SQC_PERF_SEL_ICACHE_POST_CC_LEVEL | 236 | Level Counter: Post-Cache Controller # outstanding transactions (per-SQ, nondeterministic) |
| SQC_PERF_SEL_ICACHE_POST_CC_HIT_LEVEL | 237 | Level Counter: Post-Cache Controller # outstanding transactions (per-SQ, nondeterministic) |
| SQC_PERF_SEL_ICACHE_POST_CC_MISS_LEVEL | 238 | Level Counter: Post-Cache Controller # outstanding transactions (per-SQ, nondeterministic) |
| SQC_PERF_SEL_DCACHE_INFLIGHT_LEVEL | 239 | Level Counter: # total outstanding transactions in data cache (per-SQ, nondeterministic) |
| SQC_PERF_SEL_DCACHE_PRE_CC_LEVEL | 240 | Level Counter: Pre-Cache Controller # outstanding transactions (per-SQ, nondeterministic) |
| SQC_PERF_SEL_DCACHE_POST_CC_LEVEL | 241 | Level Counter: Post-Cache Controller # outstanding transactions (per-SQ, nondeterministic) |
| SQC_PERF_SEL_DCACHE_POST_CC_HIT_LEVEL | 242 | Level Counter: Post-Cache Controller # outstanding transactions (per-SQ, nondeterministic) |
| SQC_PERF_SEL_DCACHE_POST_CC_MISS_LEVEL | 243 | Level Counter: Post-Cache Controller # outstanding transactions (per-SQ, nondeterministic) |
| SQC_PERF_SEL_TC_INFLIGHT_LEVEL | 244 | Level Counter: total outstanding requests to TC (No-Masking, nondeterministic) |
| SQC_PERF_SEL_ICACHE_TC_INFLIGHT_LEVEL | 245 | Level Counter: # of outstanding instruction requests to TC (No-Masking, nondeterministic) |
| SQC_PERF_SEL_DCACHE_TC_INFLIGHT_LEVEL | 246 | Level Counter: # of outstanding data requests to TC (No-Masking, nondeterministic) |
| SQC_PERF_SEL_ERR_DCACHE_REQ_2_GPR_ADDR_UNALIGNED | 247 | Count of error event: GPR address not aligned correctly (must be 2 GPR aligned) in a 2 dw constant fetch (No-Masking, unwindowed) |
| SQC_PERF_SEL_ERR_DCACHE_REQ_4_GPR_ADDR_UNALIGNED | 248 | Count of error event: GPR address not aligned correctly (must be 4 GPR aligned) in a 4 dw constant fetch (No-Masking, unwindowed) |
| SQC_PERF_SEL_ERR_DCACHE_REQ_8_GPR_ADDR_UNALIGNED | 249 | Count of error event: GPR address not aligned correctly (must be 4 GPR aligned) in a 8 dw constant fetch (No-Masking, unwindowed) |
| SQC_PERF_SEL_ERR_DCACHE_REQ_16_GPR_ADDR_UNALIGNED | 250 | Count of error event: GPR address not aligned correctly (must be 4 GPR aligned) in a 16 dw constant fetch (No-Masking, unwindowed) |
| Counter | Value | Description |
|---|---|---|
| TA_PERF_SEL_TA_BUSY | 0 | TA block is busy. Perf_Windowing not supported for this counter. |
| TA_PERF_SEL_SH_FIFO_BUSY | 1 | sh_fifo subblock is busy. Perf_Windowing not supported for this counter. |
| TA_PERF_SEL_SH_FIFO_CMD_BUSY | 2 | sh_fifo subblock, cmd FIFO section is busy. Perf_Windowing not supported for this counter. |
| TA_PERF_SEL_SH_FIFO_ADDR_BUSY | 3 | sh_fifo subblock, addr FIFO section is busy. Perf_Windowing not supported for this counter. |
| TA_PERF_SEL_SH_FIFO_DATA_BUSY | 4 | sh_fifo subblock, data FIFO section is busy. Perf_Windowing not supported for this counter. |
| TA_PERF_SEL_SH_FIFO_DATA_SFIFO_BUSY | 5 | sh_fifo subblock, data sfifo section is busy. Perf_Windowing not supported for this counter. |
| TA_PERF_SEL_SH_FIFO_DATA_TFIFO_BUSY | 6 | sh_fifo subblock, data tfifo section is busy. Perf_Windowing not supported for this counter. |
| TA_PERF_SEL_GRADIENT_BUSY | 7 | Deriv/Dispatch/Input subblocks are busy busy. Perf_Windowing not supported for this counter. |
| TA_PERF_SEL_GRADIENT_FIFO_BUSY | 8 | Gradient FIFO subblock is busy. Perf_Windowing not supported for this counter. |
| TA_PERF_SEL_LOD_BUSY | 9 | Aniso subblock is busy. Perf_Windowing not supported for this counter. |
| TA_PERF_SEL_LOD_FIFO_BUSY | 10 | LOD FIFO subblock is busy. Perf_Windowing not supported for this counter. |
| TA_PERF_SEL_ADDRESSER_BUSY | 11 | Addresser subblock is busy. Perf_Windowing not supported for this counter. |
| TA_PERF_SEL_ADDRESSER_FIFO_BUSY | 12 | Addresser FIFO subblock is busy. Perf_Windowing not supported for this counter. |
| TA_PERF_SEL_ALIGNER_BUSY | 13 | Aligner subblock is busy. Perf_Windowing not supported for this counter. |
| TA_PERF_SEL_WRITE_PATH_BUSY | 14 | Write Path subblock is busy. Perf_Windowing not supported for this counter. |
| TA_PERF_SEL_RESERVED_15 | 15 | RESERVED 15. |
| TA_PERF_SEL_SQ_TA_CMD_CYCLES | 16 | Number of input cycles input on SQ_TA_cmd interface. Perf_Windowing not supported for this counter. |
| TA_PERF_SEL_SP_TA_ADDR_CYCLES | 17 | Number of input cycles input on SP_TA_addr interface. Perf_Windowing not supported for this counter. |
| TA_PERF_SEL_SP_TA_DATA_CYCLES | 18 | Number of input cycles input on SP_TA_data interface. Perf_Windowing not supported for this counter. |
| TA_PERF_SEL_TA_FA_DATA_STATE_CYCLES | 19 | Number of input cycles input on TA_FA wrts interface. Perf_Windowing not supported for this counter. |
| TA_PERF_SEL_SH_FIFO_ADDR_WAITING_ON_CMD_CYCLES | 20 | Number of cycles addr waiting on cmd in sh_fifo subblock. Perf_Windowing not supported for this counter. |
| TA_PERF_SEL_SH_FIFO_CMD_WAITING_ON_ADDR_CYCLES | 21 | Number of cycles cmd waiting on addr in sh_fifo subblock. Perf_Windowing not supported for this counter. |
| TA_PERF_SEL_SH_FIFO_ADDR_STARVED_WHILE_BUSY_CYCLES | 22 | Number of cycles addr starved while busy in sh_fifo subblock. Perf_Windowing not supported for this counter. |
| TA_PERF_SEL_SH_FIFO_CMD_STARVED_WHILE_BUSY_CYCLES | 23 | Number of cycles cmd starved while busy in sh_fifo subblock. Perf_Windowing not supported for this counter. |
| TA_PERF_SEL_SH_FIFO_DATA_WAITING_ON_DATA_STATE_CYCLES | 24 | Number of cycles data waiting on data state in sh_fifo subblock. Perf_Windowing not supported for this counter. |
| TA_PERF_SEL_SH_FIFO_DATA_STATE_WAITING_ON_DATA_CYCLES | 25 | Number of cycles data state waiting on data in sh_fifo subblock. Perf_Windowing not supported for this counter. |
| TA_PERF_SEL_SH_FIFO_DATA_STARVED_WHILE_BUSY_CYCLES | 26 | Number of cycles data starved while busy in sh_fifo subblock. Perf_Windowing not supported for this counter. |
| TA_PERF_SEL_SH_FIFO_DATA_STATE_STARVED_WHILE_BUSY_CYCLES | 27 | Number of cycles data state starved while busy in sh_fifo subblock. Perf_Windowing not supported for this counter. |
| TA_PERF_SEL_RESERVED_28 | 28 | RESERVED 28. |
| TA_PERF_SEL_RESERVED_29 | 29 | RESERVED 29. |
| TA_PERF_SEL_SH_FIFO_ADDR_CYCLES | 30 | Number of address cycles issued by sh_fifo to sh_dispatch subblock. |
| TA_PERF_SEL_SH_FIFO_DATA_CYCLES | 31 | Number of data cycles issued by sh_fifo to sh_dispatch subblock. |
| TA_PERF_SEL_TOTAL_WAVEFRONTS | 32 | Total number of wavefronts processed by TA. |
| TA_PERF_SEL_GRADIENT_CYCLES | 33 | Number of cycles issued by the per-pixel-gradient state machine. |
| TA_PERF_SEL_WALKER_CYCLES | 34 | Number of cycles issued by the sampler state machine. |
| TA_PERF_SEL_ALIGNER_CYCLES | 35 | Number of cycles issued by the aligner state machine. |
| TA_PERF_SEL_IMAGE_WAVEFRONTS | 36 | Number of image wavefronts processed by TA. |
| TA_PERF_SEL_IMAGE_READ_WAVEFRONTS | 37 | Number of image read (Sample, Load, Gather4*) wavefronts processed by TA. |
| TA_PERF_SEL_IMAGE_WRITE_WAVEFRONTS | 38 | Number of image write wavefronts processed by TA. |
| TA_PERF_SEL_IMAGE_ATOMIC_WAVEFRONTS | 39 | Number of image atomic wavefronts processed by TA. |
| TA_PERF_SEL_IMAGE_TOTAL_CYCLES | 40 | Number of image cycles issued to TC. |
| TA_PERF_SEL_RESERVED_41 | 41 | RESERVED 41. |
| TA_PERF_SEL_RESERVED_42 | 42 | RESERVED 42. |
| TA_PERF_SEL_RESERVED_43 | 43 | RESERVED 43. |
| TA_PERF_SEL_BUFFER_WAVEFRONTS | 44 | Number of buffer wavefronts processed by TA. |
| TA_PERF_SEL_BUFFER_READ_WAVEFRONTS | 45 | Number of buffer read wavefronts processed by TA. |
| TA_PERF_SEL_BUFFER_WRITE_WAVEFRONTS | 46 | Number of buffer write wavefronts processed by TA. |
| TA_PERF_SEL_BUFFER_ATOMIC_WAVEFRONTS | 47 | Number of buffer atomic wavefronts processed by TA. |
| TA_PERF_SEL_BUFFER_COALESCABLE_WAVEFRONTS | 48 | Number of buffer coalesceable wavefronts processed by TA. |
| TA_PERF_SEL_BUFFER_TOTAL_CYCLES | 49 | Number of buffer cycles issued to TC. |
| TA_PERF_SEL_BUFFER_COALESCABLE_ADDR_MULTICYCLED_CYCLES | 50 | Number of buffer coalesceable cycles issued to TC that were not coalesced due to addresser. |
| TA_PERF_SEL_BUFFER_COALESCABLE_CLAMP_16KDWORD_MULTICYCLED_CYCLES | 51 | Number of buffer coalesceable cycles issued to TC that were not coalesced due to clamping or 64KB bounds check. |
| TA_PERF_SEL_BUFFER_COALESCED_READ_CYCLES | 52 | Number of buffer coalesced read cycles issued to TC. |
| TA_PERF_SEL_BUFFER_COALESCED_WRITE_CYCLES | 53 | Number of buffer coalesced write cycles issued to TC. |
| TA_PERF_SEL_ADDR_STALLED_BY_TC_CYCLES | 54 | Number of cycles addr path stalled by TC. Perf_Windowing not supported for this counter. |
| TA_PERF_SEL_ADDR_STALLED_BY_TD_CYCLES | 55 | Number of cycles addr path stalled by TD. Perf_Windowing not supported for this counter. |
| TA_PERF_SEL_DATA_STALLED_BY_TC_CYCLES | 56 | Number of cycles data path stalled by TC. Perf_Windowing not supported for this counter. |
| TA_PERF_SEL_ADDRESSER_STALLED_BY_ALIGNER_ONLY_CYCLES | 57 | Number of cycles Addresser stalled by Aligner and not further down the pipe. Perf_Windowing not supported for this counter. |
| TA_PERF_SEL_ADDRESSER_STALLED_CYCLES | 58 | Number of cycles Addresser stalled. Perf_Windowing not supported for this counter. |
| TA_PERF_SEL_ANISO_STALLED_BY_ADDRESSER_ONLY_CYCLES | 59 | Number of cycles Aniso stalled by Addresser and not further down the pipe. Perf_Windowing not supported for this counter. |
| TA_PERF_SEL_ANISO_STALLED_CYCLES | 60 | Number of cycles Aniso stalled. Perf_Windowing not supported for this counter. |
| TA_PERF_SEL_DERIV_STALLED_BY_ANISO_ONLY_CYCLES | 61 | Number of cycles Deriv stalled by Aniso and not further down the pipe. Perf_Windowing not supported for this counter. |
| TA_PERF_SEL_DERIV_STALLED_CYCLES | 62 | Number of cycles Deriv stalled. Perf_Windowing not supported for this counter. |
| TA_PERF_SEL_ANISO_GT1_CYCLE_QUADS | 63 | Number of quads requiring more than 1 aniso sample. |
| TA_PERF_SEL_COLOR_1_CYCLE_PIXELS | 64 | Number of pixels requiring sampler state machine to take 1 cycle due to format. |
| TA_PERF_SEL_COLOR_2_CYCLE_PIXELS | 65 | Number of pixels requiring sampler state machine to take 2 cycle due to format. |
| TA_PERF_SEL_COLOR_3_CYCLE_PIXELS | 66 | Number of pixels requiring sampler state machine to take 3 cycle due to format. |
| TA_PERF_SEL_COLOR_4_CYCLE_PIXELS | 67 | Number of pixels requiring sampler state machine to take 4 cycle due to format. |
| TA_PERF_SEL_MIP_1_CYCLE_PIXELS | 68 | Number of pixels requiring sampler state machine to take 1 cycle due to mip filter. |
| TA_PERF_SEL_MIP_2_CYCLE_PIXELS | 69 | Number of pixels requiring sampler state machine to take 2 cycle due to mip filter. |
| TA_PERF_SEL_VOL_1_CYCLE_PIXELS | 70 | Number of pixels requiring sampler state machine to take 1 cycle due to z filter. |
| TA_PERF_SEL_VOL_2_CYCLE_PIXELS | 71 | Number of pixels requiring sampler state machine to take 2 cycle due to z filter. |
| TA_PERF_SEL_BILIN_POINT_1_CYCLE_PIXELS | 72 | Number of pixels requiring sampler state machine to take 1 cycle due to xy filter. |
| TA_PERF_SEL_MIPMAP_LOD_0_SAMPLES | 73 | Number of samples fetched from mip 0. |
| TA_PERF_SEL_MIPMAP_LOD_1_SAMPLES | 74 | Number of samples fetched from mip 1. |
| TA_PERF_SEL_MIPMAP_LOD_2_SAMPLES | 75 | Number of samples fetched from mip 2. |
| TA_PERF_SEL_MIPMAP_LOD_3_SAMPLES | 76 | Number of samples fetched from mip 3. |
| TA_PERF_SEL_MIPMAP_LOD_4_SAMPLES | 77 | Number of samples fetched from mip 4. |
| TA_PERF_SEL_MIPMAP_LOD_5_SAMPLES | 78 | Number of samples fetched from mip 5. |
| TA_PERF_SEL_MIPMAP_LOD_6_SAMPLES | 79 | Number of samples fetched from mip 6. |
| TA_PERF_SEL_MIPMAP_LOD_7_SAMPLES | 80 | Number of samples fetched from mip 7. |
| TA_PERF_SEL_MIPMAP_LOD_8_SAMPLES | 81 | Number of samples fetched from mip 8. |
| TA_PERF_SEL_MIPMAP_LOD_9_SAMPLES | 82 | Number of samples fetched from mip 9. |
| TA_PERF_SEL_MIPMAP_LOD_10_SAMPLES | 83 | Number of samples fetched from mip 10. |
| TA_PERF_SEL_MIPMAP_LOD_11_SAMPLES | 84 | Number of samples fetched from mip 11. |
| TA_PERF_SEL_MIPMAP_LOD_12_SAMPLES | 85 | Number of samples fetched from mip 12. |
| TA_PERF_SEL_MIPMAP_LOD_13_SAMPLES | 86 | Number of samples fetched from mip 13. |
| TA_PERF_SEL_MIPMAP_LOD_14_SAMPLES | 87 | Number of samples fetched from mip 14. |
| TA_PERF_SEL_MIPMAP_INVALID_SAMPLES | 88 | Number of samples marked invalid using mip 15 method. |
| TA_PERF_SEL_ANISO_1_CYCLE_QUADS | 89 | Number of quads requiring 1 aniso sample. |
| TA_PERF_SEL_ANISO_2_CYCLE_QUADS | 90 | Number of quads requiring 2 aniso sample. |
| TA_PERF_SEL_ANISO_4_CYCLE_QUADS | 91 | Number of quads requiring 4 aniso sample. |
| TA_PERF_SEL_ANISO_6_CYCLE_QUADS | 92 | Number of quads requiring 6 aniso sample. |
| TA_PERF_SEL_ANISO_8_CYCLE_QUADS | 93 | Number of quads requiring 8 aniso sample. |
| TA_PERF_SEL_ANISO_10_CYCLE_QUADS | 94 | Number of quads requiring 10 aniso sample. |
| TA_PERF_SEL_ANISO_12_CYCLE_QUADS | 95 | Number of quads requiring 12 aniso sample. |
| TA_PERF_SEL_ANISO_14_CYCLE_QUADS | 96 | Number of quads requiring 14 aniso sample. |
| TA_PERF_SEL_ANISO_16_CYCLE_QUADS | 97 | Number of quads requiring 16 aniso sample. |
| TA_PERF_SEL_WRITE_PATH_INPUT_CYCLES | 98 | Number of cycles received from write datapath from sh_dispatct. |
| TA_PERF_SEL_WRITE_PATH_OUTPUT_CYCLES | 99 | Number of cycles sent from write datapath to TC. |
| TA_PERF_SEL_FLAT_WAVEFRONTS | 100 | Number of flat opcode wavefronts processed by the TA. |
| TA_PERF_SEL_FLAT_READ_WAVEFRONTS | 101 | Number of flat opcode reads processed by the TA. |
| TA_PERF_SEL_FLAT_WRITE_WAVEFRONTS | 102 | Number of flat opcode writes processed by the TA. |
| TA_PERF_SEL_FLAT_ATOMIC_WAVEFRONTS | 103 | Number of flat opcode atomics processed by the TA. |
| TA_PERF_SEL_FLAT_COALESCEABLE_WAVEFRONTS | 104 | Number of flat opcode coalesceable ops processed by the TA. |
| TA_PERF_SEL_REG_SCLK_VLD | 105 | Number of cycles reg_sclk is active. Perf_Windowing not supported for this counter. |
| TA_PERF_SEL_LOCAL_CG_DYN_SCLK_GRP0_EN | 106 | Number of cycles grp0 sclk is active. Perf_Windowing not supported for this counter. |
| TA_PERF_SEL_LOCAL_CG_DYN_SCLK_GRP1_EN | 107 | Number of cycles grp1 sclk is active. Perf_Windowing not supported for this counter. |
| TA_PERF_SEL_LOCAL_CG_DYN_SCLK_GRP1_MEMS_EN | 108 | Number of cycles grp1_mems sclk is active. Perf_Windowing not supported for this counter. |
| TA_PERF_SEL_LOCAL_CG_DYN_SCLK_GRP4_EN | 109 | Number of cycles grp4 sclk is active. Perf_Windowing not supported for this counter. |
| TA_PERF_SEL_LOCAL_CG_DYN_SCLK_GRP5_EN | 110 | Number of cycles grp5 sclk is active. Perf_Windowing not supported for this counter. |
| TA_PERF_SEL_SCORPIO_START111 | ||
| Start of Xbox One X-specific counters. | ||
| TA_PERF_SEL_MSAA_FETCHES111 | ||
| The number of MSAA fetches detected. | ||
| TA_PERF_SEL_TOTAL_CUBEEDGE_CYCLES112 | ||
| The number of cube-edge cycles detected. | ||
| TA_PERF_SEL_TOTAL_CUBECORNER_CYCLES113 | ||
| The number of cube-corner cycles detected. | ||
| Counter | Value | Description |
|---|---|---|
| TD_PERF_SEL_TD_BUSY | 0 | TD is processing or waiting for data. Perf_Windowing not supported for this counter. |
| TD_PERF_SEL_INPUT_BUSY | 1 | TD input subblock is busy or waiting for data. Perf_Windowing not supported for this counter. |
| TD_PERF_SEL_OUTPUT_BUSY | 2 | TD output subblock is busy for waiting for data. Perf_Windowing not supported for this counter. |
| TD_PERF_SEL_LERP_BUSY | 3 | TD filter block is busy. Perf_Windowing not supported for this counter. |
| TD_PERF_SEL_RESERVED_4 | 4 | RESERVED_4 |
| TD_PERF_SEL_REG_SCLK_VLD | 5 | Clock gate enable for GRBM register reads & writes. Perf_Windowing not supported for this counter. |
| TD_PERF_SEL_LOCAL_CG_DYN_SCLK_GRP0_EN | 6 | Clock gate enable for group0 - non-harvestable always_on domain. Perf_Windowing not supported for this counter. |
| TD_PERF_SEL_LOCAL_CG_DYN_SCLK_GRP1_EN | 7 | Clock gate enable for group1 - harvestable texture logic domain. Perf_Windowing not supported for this counter. |
| TD_PERF_SEL_LOCAL_CG_DYN_SCLK_GRP4_EN | 8 | Clock gate enable for group4 - non-harvestable GDS chain domain. Perf_Windowing not supported for this counter. |
| TD_PERF_SEL_LOCAL_CG_DYN_SCLK_GRP5_EN | 9 | Clock gate enable for group5 - non-harvestable texture boundary domain. Perf_Windowing not supported for this counter. |
| TD_PERF_SEL_TC_TD_FIFO_FULL | 10 | TC_TD input FIFO is full. |
| TD_PERF_SEL_CONSTANT_STATE_FULL | 11 | TA_TD instruction FIFO is full. |
| TD_PERF_SEL_SAMPLE_STATE_FULL | 12 | TA_TD sample FIFO is full. |
| TD_PERF_SEL_OUTPUT_FIFO_FULL | 13 | TD output FIFO is full. |
| TD_PERF_SEL_RESERVED_14 | 14 | RESERVED_14 |
| TD_PERF_SEL_TC_STALL | 15 | TD is stalled waiting for TC data. |
| TD_PERF_SEL_PC_STALL | 16 | TD is stalled by PC waiting to send LDS data. |
| TD_PERF_SEL_GDS_STALL | 17 | TD is stalled by GDS data. |
| TD_PERF_SEL_RESERVED_18 | 18 | RESERVED_18 |
| TD_PERF_SEL_RESERVED_19 | 19 | RESERVED_19 |
| TD_PERF_SEL_GATHER4_WAVEFRONT | 20 | Count the wavefronts with opcode = gather4, includes gather4_c. |
| TD_PERF_SEL_SAMPLE_C_WAVEFRONT | 21 | Count the wavefronts with opcode = sample_c, includes gather4_c. |
| TD_PERF_SEL_LOAD_WAVEFRONT | 22 | Count the wavefronts with opcode = load, include atomics and store. |
| TD_PERF_SEL_ATOMIC_WAVEFRONT | 23 | Count the wavefronts with opcode = atomic. |
| TD_PERF_SEL_STORE_WAVEFRONT | 24 | Count the wavefronts with opcode = store. |
| TD_PERF_SEL_LDFPTR_WAVEFRONT | 25 | Count the wavefronts with LDFPTR formats. |
| TD_PERF_SEL_RESERVED_26 | 26 | RESERVED_26 |
| TD_PERF_SEL_RESERVED_27 | 27 | RESERVED_27 |
| TD_PERF_SEL_RESERVED_28 | 28 | RESERVED_28 |
| TD_PERF_SEL_RESERVED_29 | 29 | RESERVED_29 |
| TD_PERF_SEL_BYPASS_FILTER_WAVEFRONT | 30 | Count the wavefronts that bypass the filter, includes bypass opcode or bypass formats. |
| TD_PERF_SEL_MIN_MAX_FILTER_WAVEFRONT | 31 | Count the wavefronts that use min/max filtering. |
| TD_PERF_SEL_COALESCABLE_WAVEFRONT | 32 | Count wavefronts that TA finds coalescable. |
| TD_PERF_SEL_COALESCED_PHASE | 33 | Count up when each phase is coalesced, if all phases in a wavefront were coalesced the count will be 16. |
| TD_PERF_SEL_FOUR_PHASE_WAVEFRONT | 34 | dmask=1,2,4,8, wavefront was packed into a 4 phase packet for SP. |
| TD_PERF_SEL_EIGHT_PHASE_WAVEFRONT | 35 | dmask=3,5,6,9,a,c, wavefront was packed into an 8 phase packet for SP. |
| TD_PERF_SEL_SIXTEEN_PHASE_WAVEFRONT | 36 | dmask=7,b,d,e,f, wavefront was forwarded to SP in 16 phases. |
| TD_PERF_SEL_FOUR_PHASE_FORWARD_WAVEFRONT | 37 | Count the wavefronts that were forwarded to SP in 4 phases (coalescable). |
| TD_PERF_SEL_WRITE_ACK_WAVEFRONT | 38 | Count write acknowledgments, sent to SQ and not to SP. |
| TD_PERF_SEL_RESERVED_39 | 39 | RESERVED_39. |
| TD_PERF_SEL_USER_DEFINED_BORDER | 40 | Count the wavefronts that user defined border color was used. |
| TD_PERF_SEL_WHITE_BORDER | 41 | Count the wavefronts that white border color was used. |
| TD_PERF_SEL_OPAQUE_BLACK_BORDER | 42 | Count the wavefronts that opaque black border color was used. |
| TD_PERF_SEL_RESERVED_43 | 43 | RESERVED_43 |
| TD_PERF_SEL_RESERVED_44 | 44 | RESERVED_44 |
| TD_PERF_SEL_NACK | 45 | Count the number of times an ack packet was generated and sent to SP/SQ. |
| TD_PERF_SEL_TD_SP_TRAFFIC | 46 | Count the number of times this TD sends data to the SP. |
| TD_PERF_SEL_CONSUME_GDS_TRAFFIC | 47 | Count the number of times GDS data was consumed by this TD, send to SP. |
| TD_PERF_SEL_ADDRESSCMD_POISON | 48 | Count the wavefronts that had poisoned address or command. |
| TD_PERF_SEL_DATA_POISON | 49 | Count the wavefronts that had poisoned data. |
| Counter | Value | Description |
|---|---|---|
| TCP_PERF_SEL_TA_TCP_ADDR_STARVE_CYCLES | 0 | TA starves TCP addr interface. Not Windowed. |
| TCP_PERF_SEL_TA_TCP_DATA_STARVE_CYCLES | 1 | TA starves TCP data interface. Not Windowed. |
| TCP_PERF_SEL_TCP_TA_ADDR_STALL_CYCLES | 2 | TCP stalls TA addr interface. |
| TCP_PERF_SEL_TCP_TA_DATA_STALL_CYCLES | 3 | TCP stalls TA data interface. Not Windowed. |
| TCP_PERF_SEL_TD_TCP_STALL_CYCLES | 4 | TD stalls TCP |
| TCP_PERF_SEL_TCR_TCP_STALL_CYCLES | 5 | TCR stalls TCP |
| TCP_PERF_SEL_LOD_STALL_CYCLES | 6 | Per Pixel LOD stall |
| TCP_PERF_SEL_READ_TAGCONFLICT_STALL_CYCLES | 7 | Tagram conflict stall on a read |
| TCP_PERF_SEL_WRITE_TAGCONFLICT_STALL_CYCLES | 8 | Tagram conflict stall on a write |
| TCP_PERF_SEL_ATOMIC_TAGCONFLICT_STALL_CYCLES | 9 | Tagram conflict stall on an atomic |
| TCP_PERF_SEL_ALLOC_STALL_CYCLES | 10 | Alloc on in-flight cache line stall |
| TCP_PERF_SEL_LFIFO_STALL_CYCLES | 11 | Memory Latency FIFOs full stall |
| TCP_PERF_SEL_RFIFO_STALL_CYCLES | 12 | Memory Request FIFOs full stall |
| TCP_PERF_SEL_TCR_RDRET_STALL | 13 | write into cache stalled by read return from tcr |
| TCP_PERF_SEL_WRITE_CONFLICT_STALL | 14 | write stall due to cache bank conflict |
| TCP_PERF_SEL_HOLE_READ_STALL | 15 | read from cache stalled due to hole fill read |
| TCP_PERF_SEL_READCONFLICT_STALL_CYCLES | 16 | Read conflict stall due to cache bank conflict |
| TCP_PERF_SEL_PENDING_STALL_CYCLES | 17 | Pending stall |
| TCP_PERF_SEL_READFIFO_STALL_CYCLES | 18 | Read FIFO stall |
| TCP_PERF_SEL_TCP_LATENCY | 19 | Total TCP wave latency (from first clock of wave entering to first clock of wave leaving), divide by TA_TCP_STATE_READ to avg wave latency |
| TCP_PERF_SEL_TCC_READ_REQ_LATENCY | 20 | Total TCP->TCC request latency for reads and atomics with return. Not Windowed. |
| TCP_PERF_SEL_TCC_WRITE_REQ_LATENCY | 21 | Total TCP->TCC request latency for writes and atomics without return. Not Windowed. |
| TCP_PERF_SEL_TCC_WRITE_REQ_HOLE_LATENCY | 22 | Total TCP req ->TCC hole latency for writes and atomics. Not Windowed. |
| TCP_PERF_SEL_TCC_READ_REQ | 23 | Total read requests from TCP to all TCCs |
| TCP_PERF_SEL_TCC_WRITE_REQ | 24 | Total write requests from TCP to all TCCs |
| TCP_PERF_SEL_TCC_ATOMIC_WITH_RET_REQ | 25 | Total atomic with return requests from TCP to all TCCs |
| TCP_PERF_SEL_TCC_ATOMIC_WITHOUT_RET_REQ | 26 | Total atomic without return requests from TCP to all TCCs |
| TCP_PERF_SEL_TOTAL_LOCAL_READ | 27 | DEPRECATED. Replaced by TCP_PERF_SEL_TOTAL_HIT_LRU_READ. |
| TCP_PERF_SEL_TOTAL_GLOBAL_READ | 28 | DEPRECATED. Replaced by TCP_PERF_SEL_TOTAL_MISS_EVICT_READ. |
| TCP_PERF_SEL_TOTAL_LOCAL_WRITE | 29 | DEPRECATED. Replaced by TCP_PERF_SEL_TOTAL_MISS_LRU_WRITE. |
| TCP_PERF_SEL_TOTAL_GLOBAL_WRITE | 30 | DEPRECATED. Replaced by TCP_PERF_SEL_TOTAL_MISS_EVICT_WRITE. |
| TCP_PERF_SEL_TOTAL_ATOMIC_WITH_RET | 31 | Total number of atomic with return pixels/buffers from TA |
| TCP_PERF_SEL_TOTAL_ATOMIC_WITHOUT_RET | 32 | Total number of atomic without return pixels/buffers from TA |
| TCP_PERF_SEL_TOTAL_WBINVL1 | 33 | Total number of wbinvl1 transactions from TA (from shader WBINVL1 instructions) |
| TCP_PERF_SEL_IMG_READ_FMT_1 | 34 | Count of image read pixels that use 1-bit formats |
| TCP_PERF_SEL_IMG_READ_FMT_8 | 35 | Count of image read pixels that use 8-bit formats |
| TCP_PERF_SEL_IMG_READ_FMT_16 | 36 | Count of image read pixels that use 16-bit formats |
| TCP_PERF_SEL_IMG_READ_FMT_32 | 37 | Count of image read pixels that use 32-bit formats |
| TCP_PERF_SEL_IMG_READ_FMT_32_AS_8 | 38 | Count of image read pixels that use 32_AS_8 formats |
| TCP_PERF_SEL_IMG_READ_FMT_32_AS_16 | 39 | Count of image read pixels that use 32_AS_8_8 formats |
| TCP_PERF_SEL_IMG_READ_FMT_32_AS_128 | 40 | Count of image read pixels that use 32_AS_8_8 formats |
| TCP_PERF_SEL_IMG_READ_FMT_64_2_CYCLE | 41 | Count of image read pixels that use 64-bit 2-cycle formats |
| TCP_PERF_SEL_IMG_READ_FMT_64_1_CYCLE | 42 | Count of image read pixels that use 64-bit 1-cycle formats |
| TCP_PERF_SEL_IMG_READ_FMT_96 | 43 | Count of image read pixels that use 96-bit formats |
| TCP_PERF_SEL_IMG_READ_FMT_128_4_CYCLE | 44 | Count of image read pixels that use 128-bit 4-cycle formats |
| TCP_PERF_SEL_IMG_READ_FMT_128_1_CYCLE | 45 | Count of image read pixels that use 128-bit 1-cycle formats |
| TCP_PERF_SEL_IMG_READ_FMT_BC1 | 46 | Count of image read pixels that use BC1 format |
| TCP_PERF_SEL_IMG_READ_FMT_BC2 | 47 | Count of image read pixels that use BC2 format |
| TCP_PERF_SEL_IMG_READ_FMT_BC3 | 48 | Count of image read pixels that use BC3 format |
| TCP_PERF_SEL_IMG_READ_FMT_BC4 | 49 | Count of image read pixels that use BC4 format |
| TCP_PERF_SEL_IMG_READ_FMT_BC5 | 50 | Count of image read pixels that use BC5 format |
| TCP_PERF_SEL_IMG_READ_FMT_BC6 | 51 | Count of image read pixels that use BC6 format |
| TCP_PERF_SEL_IMG_READ_FMT_BC7 | 52 | Count of image read pixels that use BC7 format |
| TCP_PERF_SEL_IMG_READ_FMT_I8 | 53 | Count of image read pixels that use 8-bit interlaced formats |
| TCP_PERF_SEL_IMG_READ_FMT_I16 | 54 | Count of image read pixels that use 16-bit interlaced formats |
| TCP_PERF_SEL_IMG_READ_FMT_I32 | 55 | Count of image read pixels that use 32-bit interlaced formats |
| TCP_PERF_SEL_IMG_READ_FMT_I32_AS_8 | 56 | Count of image read pixels that use 32_AS_8 interlaced formats |
| TCP_PERF_SEL_IMG_READ_FMT_I32_AS_16 | 57 | Count of image read pixels that use 32_AS_8_8 interlaced formats |
| TCP_PERF_SEL_IMG_READ_FMT_D8 | 58 | DEPRECATED. Do not use. |
| TCP_PERF_SEL_IMG_READ_FMT_D16 | 59 | DEPRECATED. Do not use. |
| TCP_PERF_SEL_IMG_READ_FMT_D32 | 60 | DEPRECATED. Do not use. |
| TCP_PERF_SEL_IMG_WRITE_FMT_8 | 61 | Count of image write pixels that use 8-bit formats |
| TCP_PERF_SEL_IMG_WRITE_FMT_16 | 62 | Count of image write pixels that use 16-bit formats |
| TCP_PERF_SEL_IMG_WRITE_FMT_32 | 63 | Count of image write pixels that use 32-bit formats |
| TCP_PERF_SEL_IMG_WRITE_FMT_64 | 64 | Count of image write pixels that use 64-bit formats |
| TCP_PERF_SEL_IMG_WRITE_FMT_128 | 65 | Count of image write pixels that use 128-bit formats |
| TCP_PERF_SEL_IMG_WRITE_FMT_D8 | 66 | DEPRECATED. Do not use. |
| TCP_PERF_SEL_IMG_WRITE_FMT_D16 | 67 | DEPRECATED. Do not use. |
| TCP_PERF_SEL_IMG_WRITE_FMT_D32 | 68 | DEPRECATED. Do not use. |
| TCP_PERF_SEL_IMG_ATOMIC_WITH_RET_FMT_32 | 69 | Count of image atomic with return pixels that use 32-bit formats |
| TCP_PERF_SEL_IMG_ATOMIC_WITHOUT_RET_FMT_32 | 70 | Count of image atomic without return pixels that use 32-bit formats |
| TCP_PERF_SEL_IMG_ATOMIC_WITH_RET_FMT_64 | 71 | Count of image atomic with return pixels that use 64-bit formats |
| TCP_PERF_SEL_IMG_ATOMIC_WITHOUT_RET_FMT_64 | 72 | Count of image atomic without return pixels that use 64-bit formats |
| TCP_PERF_SEL_BUF_READ_FMT_8 | 73 | Count of buffer reads that use 8-bit vertex format |
| TCP_PERF_SEL_BUF_READ_FMT_16 | 74 | Count of buffer reads that use 16-bit vertex format |
| TCP_PERF_SEL_BUF_READ_FMT_32 | 75 | Count of buffer reads that use 32-bit vertex format (includes 64b and 128b) |
| TCP_PERF_SEL_BUF_WRITE_FMT_8 | 76 | Count of buffer writes that use 8-bit vertex format |
| TCP_PERF_SEL_BUF_WRITE_FMT_16 | 77 | Count of buffer writes that use 16-bit vertex format |
| TCP_PERF_SEL_BUF_WRITE_FMT_32 | 78 | Count of buffer writes that use 32-bit vertex format (includes 64b and 128b) |
| TCP_PERF_SEL_BUF_ATOMIC_WITH_RET_FMT_32 | 79 | Count of buffer atomics with return that use 32-bit formats |
| TCP_PERF_SEL_BUF_ATOMIC_WITHOUT_RET_FMT_32 | 80 | Count of buffer atomics without return that use 32-bit formats |
| TCP_PERF_SEL_BUF_ATOMIC_WITH_RET_FMT_64 | 81 | Count of buffer atomics with return that use 64-bit formats |
| TCP_PERF_SEL_BUF_ATOMIC_WITHOUT_RET_FMT_64 | 82 | Count of buffer atomics without return that use 64-bit formats |
| TCP_PERF_SEL_ARR_LINEAR_GENERAL | 83 | Count of buffers that use linear general memory tiling |
| TCP_PERF_SEL_ARR_LINEAR_ALIGNED | 84 | Count of pixels that use linear aligned memory tiling |
| TCP_PERF_SEL_ARR_1D_THIN1 | 85 | Count of pixels that use 1d thin1 memory tiling |
| TCP_PERF_SEL_ARR_1D_THICK | 86 | Count of pixels that use 1d thick memory tiling |
| TCP_PERF_SEL_ARR_2D_THIN1 | 87 | Count of pixels that use 2d thin1 memory tiling |
| TCP_PERF_SEL_ARR_2D_THICK | 88 | Count of pixels that use 2d thick memory tiling |
| TCP_PERF_SEL_ARR_2D_XTHICK | 89 | Count of pixels that use 2d xthick memory tiling |
| TCP_PERF_SEL_ARR_3D_THIN1 | 90 | Count of pixels that use 3d thin1 memory tiling |
| TCP_PERF_SEL_ARR_3D_THICK | 91 | Count of pixels that use 3d thick memory tiling |
| TCP_PERF_SEL_ARR_3D_XTHICK | 92 | Count of pixels that use 3d xthick memory tiling |
| TCP_PERF_SEL_DIM_1D | 93 | Count of pixels that belong to 1D surfaces |
| TCP_PERF_SEL_DIM_2D | 94 | Count of pixels that belong to 2D surfaces |
| TCP_PERF_SEL_DIM_3D | 95 | Count of pixels that belong to 3D surfaces |
| TCP_PERF_SEL_DIM_1D_ARRAY | 96 | Count of pixels that belong to 1D Array surfaces |
| TCP_PERF_SEL_DIM_2D_ARRAY | 97 | Count of pixels that belong to 2D Array surfaces |
| TCP_PERF_SEL_DIM_2D_MSAA | 98 | Count of pixels that belong to 2D MSAA surfaces |
| TCP_PERF_SEL_DIM_2D_ARRAY_MSAA | 99 | Count of pixels that belong to 2D MSAA Array surfaces |
| TCP_PERF_SEL_DIM_CUBE_ARRAY | 100 | Count of pixels that belong to Cube Array surfaces |
| TCP_PERF_SEL_CP_TCP_INVALIDATE | 101 | Number of cache invalidates from the CP. Not Windowed. |
| TCP_PERF_SEL_TA_TCP_STATE_READ | 102 | Number of state reads |
| TCP_PERF_SEL_TAGRAM0_REQ | 103 | L1 Requests, (Tagram 0) 64B units |
| TCP_PERF_SEL_TAGRAM1_REQ | 104 | L1 Requests, (Tagram 1) 64B units |
| TCP_PERF_SEL_TAGRAM2_REQ | 105 | L1 Requests, (Tagram 2) 64B units |
| TCP_PERF_SEL_TAGRAM3_REQ | 106 | L1 Requests, (Tagram 3) 64B units |
| TCP_PERF_SEL_GATE_EN1 | 107 | TCP interface clocks are turned on. Not Windowed. |
| TCP_PERF_SEL_GATE_EN2 | 108 | TCP core clocks are turned on. Not Windowed. |
| TCP_PERF_SEL_CORE_REG_SCLK_VLD | 109 | TCP reg clocks are turned on. Not Windowed. |
| TCP_PERF_SEL_TCC_REQ | 110 | Total requests from TCP to all TCCs. Equals TCP_PERF_SEL_TCC_READ_REQ + TCP_PERF_SEL_TCC_NON_READ_REQ |
| TCP_PERF_SEL_TCC_NON_READ_REQ | 111 | Total non-read requests from TCP to all TCCs. Equals TCP_PERF_SEL_TCC_WRITE_REQ+ TCP_PERF_SEL_TCC_ATOMIC_WITH_RET_REQ+ TCP_PERF_SEL_TCC_ATOMIC_WITHOUT_RET_REQ |
| TCP_PERF_SEL_TCC_BYPASS_READ_REQ | 112 | Total read requests from TCP to the TCS |
| TCP_PERF_SEL_TCC_MISS_EVICT_READ_REQ | 113 | Total read requests from TCP to all TCCs that were caused by a TCP request using the MISS_EVICT policy |
| TCP_PERF_SEL_TCC_VOLATILE_READ_REQ | 114 | Total volatile read requests from TCP to all TCCs |
| TCP_PERF_SEL_TCC_VOLATILE_BYPASS_READ_REQ | 115 | Total volatile read requests from TCP to the TCS |
| TCP_PERF_SEL_TCC_VOLATILE_MISS_EVICT_READ_REQ | 116 | Total volatile read requests from TCP to all TCCs that were caused by a TCP request using the MISS_EVICT policy |
| TCP_PERF_SEL_TCC_BYPASS_WRITE_REQ | 117 | Total write requests from TCP to the TCS |
| TCP_PERF_SEL_TCC_MISS_EVICT_WRITE_REQ | 118 | Total write requests from TCP to all TCCs that were caused by a TCP request using the MISS_EVICT policy |
| TCP_PERF_SEL_TCC_VOLATILE_BYPASS_WRITE_REQ | 119 | Total volatile write requests from TCP to the TCS |
| TCP_PERF_SEL_TCC_VOLATILE_WRITE_REQ | 120 | Total volatile write requests from TCP to all TCCs |
| TCP_PERF_SEL_TCC_VOLATILE_MISS_EVICT_WRITE_REQ | 121 | Total volatile write requests from TCP to all TCCs that were caused by a TCP request using the MISS_EVICT policy |
| TCP_PERF_SEL_TCC_BYPASS_ATOMIC_REQ | 122 | Total atomic requests from TCP to the TCS |
| TCP_PERF_SEL_TCC_ATOMIC_REQ | 123 | Total atomic requests from TCP to all TCCs |
| TCP_PERF_SEL_TCC_VOLATILE_ATOMIC_REQ | 124 | Total volatile atomic requests from TCP to all TCCs |
| TCP_PERF_SEL_TCC_DATA_BUS_BUSY | 125 | Total cycles the TCC data bus is busy servicing this client. Equals TCP_PERF_SEL_TCC_READ_REQ + TCP_PERF_SEL_TCC_WRITE_REQ + TCP_PERF_SEL_TCC_ATOMIC_WITHOUT_RET_REQ + 2* TCP_PERF_SEL_TCC_ATOMIC_WITH_RET_REQ. Not Windowed. |
| TCP_PERF_SEL_TOTAL_ACCESSES | 126 | Total number of pixels/buffers from TA. Equals TCP_PERF_SEL_TOTAL_READ+TCP_PERF_SEL_TOTAL_NONREAD |
| TCP_PERF_SEL_TOTAL_READ | 127 | Total number of read pixels/buffers from TA. Equals TCP_PERF_SEL_TOTAL_HIT_LRU_READ ALPHA+ TCP_PERF_SEL_TOTAL_HIT_EVICT_READ ALPHA+ TCP_PERF_SEL_TOTAL_MISS_LRU_READ+ TCP_PERF_SEL_TOTAL_MISS_EVICT_READ |
| TCP_PERF_SEL_TOTAL_HIT_LRU_READ | 128 | Total number of read pixels/buffers from TA using the HIT_LRU policy |
| TCP_PERF_SEL_TOTAL_HIT_EVICT_READ | 129 | Total number of read pixels/buffers from TA using the HIT_EVICT policy |
| TCP_PERF_SEL_TOTAL_MISS_LRU_READ | 130 | Total number of read pixels/buffers from TA using the MISS_LRU policy |
| TCP_PERF_SEL_TOTAL_MISS_EVICT_READ | 131 | Total number of read pixels/buffers from TA using the MISS_EVICT policy |
| TCP_PERF_SEL_TOTAL_NON_READ | 132 | Total number of non-read pixels/buffers from TA. Equals TCP_PERF_SEL_WRITE + TCP_PERF_SEL_TOTAL_ATOMIC_WITH_RET + TCP_PERF_SEL_TOTOAL_ATOMIC_WITHOUT_RET |
| TCP_PERF_SEL_TOTAL_WRITE | 133 | Total number of local write pixels/buffers from TA. Equals TCP_PERF_SEL_TOTAL_MISS_LRU_WRITE+ TCP_PERF_SEL_TOTAL_MISS_EVICT_WRITE |
| TCP_PERF_SEL_TOTAL_MISS_LRU_WRITE | 134 | Total number of local write pixels/buffers from TA using the MISS_LRU policy |
| TCP_PERF_SEL_TOTAL_MISS_EVICT_WRITE | 135 | Total number of global write pixels/buffers from TA using the MISS_EVICT policy |
| TCP_PERF_SEL_TOTAL_WBINVL1_VOL | 136 | Total number of volatile wbinvl1 transactions from TA (from shader WBINVL1_VOL instructions) |
| TCP_PERF_SEL_TOTAL_WRITEBACK_INVALIDATES | 137 | Total number of cache invalidates. Equals TCP_PERF_SEL_TOTAL_WBINVL1+ TCP_PERF_SEL_TOTAL_WBINVL1_VOL+ TCP_PERF_SEL_CP_TCP_INVALIDATE+ TCP_PERF_SEL_SQ_TCP_INVALIDATE_VOL. Not Windowed. |
| TCP_PERF_SEL_DISPLAY_MICROTILING | 138 | Count of image pixels using display microtiling |
| TCP_PERF_SEL_THIN_MICROTILING | 139 | Count of image pixels using thin microtiling |
| TCP_PERF_SEL_DEPTH_MICROTILING | 140 | Count of image pixels using depth microtiling |
| TCP_PERF_SEL_ARR_PRT_THIN1 | 141 | Count of pixels that use prt thin1 memory tiling |
| TCP_PERF_SEL_ARR_PRT_2D_THIN1 | 142 | Count of pixels that use 2d prt thin1 memory tiling |
| TCP_PERF_SEL_ARR_PRT_3D_THIN1 | 143 | Count of pixels that use 3d prt thin1 memory tiling |
| TCP_PERF_SEL_ARR_PRT_THICK | 144 | Count of pixels that use prt thick memory tiling |
| TCP_PERF_SEL_ARR_PRT_2D_THICK | 145 | Count of pixels that use 2d prt thick memory tiling |
| TCP_PERF_SEL_ARR_PRT_3D_THICK | 146 | Count of pixels that use 3d prt thick memory tiling |
| TCP_PERF_SEL_CP_TCP_INVALIDATE_VOL | 147 | Number of volatile cache invalidates from the CP. Not Windowed. |
| TCP_PERF_SEL_SQ_TCP_INVALIDATE_VOL | 148 | Number of volatile cache invalidates from the SQ. Not Windowed. |
| TCP_PERF_SEL_UNALIGNED | 149 | Count of unaligned buffer fetches |
| TCP_PERF_SEL_ROTATED_MICROTILING | 150 | Count of image pixels using rotated microtiling |
| TCP_PERF_SEL_THICK_MICROTILING | 151 | Count of image pixels using thick microtiling |
| TCP_PERF_SEL_ATC | 152 | Count of pixels/buffers that use ATC |
| TCP_PERF_SEL_POWER_STALL | 153 | Count of stalls due to power throttling |
| TCP_PERF_SEL_IMG_READ_FMT_CTX1 | 154 | Count of image read pixels that use CTX1 format |
| TCP_PERF_SEL_IMG_READ_FMT_DXT3A | 155 | Count of image read pixels that use DXT3A or DXT3A_AS_1_1_1_1 formats |
| TCP_PERF_SEL_IMG_READ_FMT_YCBCR | 156 | Count of image read pixels that use YCBCR format |
| TCP_PERF_SEL_ALT_TILING_1D_LINEAR | 157 | Count of pixels that use 1D linear alt tiling |
| TCP_PERF_SEL_ALT_TILING_2D_LINEAR | 158 | Count of pixels that use 2D linear alt tiling |
| TCP_PERF_SEL_ALT_TILING_3D_LINEAR | 159 | Count of pixels that use 3D linear alt tiling |
| TCP_PERF_SEL_ALT_TILING_2D_TILED | 160 | Count of pixels that use 2D tiled alt tiling |
| TCP_PERF_SEL_ALT_TILING_3D_TILED | 161 | Count of pixels that use 3D tiled alt tiling |
| TCP_PERF_SEL_SCORPIO_START162 | ||
| Start of Xbox One X-specific counters. | ||
| TCP_PERF_SEL_TOTAL_BUF_READS162 | ||
| Count of buffer reads (any format). | ||
| TCP_PERF_SEL_TOTAL_BUF_WRITES163 | ||
| Count of buffer writes (any format). | ||
| TCP_PERF_SEL_TOTAL_BUF_READS_OR_WRITES164 | ||
| Count of buffer reads or writes (any format). | ||
| Counter | Value | Description |
|---|---|---|
| TCC_PERF_SEL_NONE | 0 | Don’t count anything. |
| TCC_PERF_SEL_CYCLE | 1 | Number of cycles. Not windowable. |
| TCC_PERF_SEL_BUSY | 2 | Number of cycles we have a request pending. Not windowable. |
| TCC_PERF_SEL_REQ | 3 | Number of requests of all types. |
| TCC_PERF_SEL_STREAMING_REQ | 4 | Number of streaming requests |
| TCC_PERF_SEL_READ | 5 | Number of read requests. |
| TCC_PERF_SEL_WRITE | 6 | Number of write requests. |
| TCC_PERF_SEL_ATOMIC | 7 | Number of atomic requests of all types. |
| TCC_PERF_SEL_WBINVL2 | 8 | Number of wbinvl2 requests. Deprecated. |
| TCC_PERF_SEL_WBINVL2_CYCLE | 9 | Number of cycles spent performing wbinvl2 operations. Deprecated. |
| TCC_PERF_SEL_HIT | 10 | Number of cache hits. |
| TCC_PERF_SEL_MISS | 11 | Number of cache misses. |
| TCC_PERF_SEL_DEWRITE_ALLOCATE_HIT | 12 | Number of times a read was performed on a cache line that had only been written into before. This can only occur once for the life of a cache line. |
| TCC_PERF_SEL_FULLY_WRITTEN_HIT | 13 | Number of times a read from the mc was avoided because the cache line was already fully written with data. This can only occur once for the life of a cache line. |
| TCC_PERF_SEL_WRITEBACK | 14 | Number of lines written back to main memory. |
| TCC_PERF_SEL_LATENCY_FIFO_FULL | 15 | Number of cycles the latency FIFO was full. |
| TCC_PERF_SEL_SRC_FIFO_FULL | 16 | Number of cycles the src FIFO was expected to be full as measured at the IB block. |
| TCC_PERF_SEL_HOLE_FIFO_FULL | 17 | Number of cycles the hole FIFOs in the TCAs were expected to be full as measured at the IB block. This is usually an indication that there was insufficient bandwidth available on the data bus for source data. |
| TCC_PERF_SEL_MC_WRREQ | 18 | Number of 32-byte writes |
| TCC_PERF_SEL_SCORPIO_START | 19 | Start of Xbox One X-specific counters. |
| TCC_PERF_SEL_MC_WRREQ_UNCACHED | 19 | Number of 32-byte transactions going over the TC_MC_wrreq interface due to uncached traffic. Note that CC mtypes can produce uncached requests, and those are included in this. |
| TCC_PERF_SEL_MC_WRREQ_STALL | 20 | Number of cycles a write request was stalled. |
| TCC_PERF_SEL_MC_WRREQ_CREDIT_STALL | 21 | Number of cycles a mc write request was stalled because the interface was out of credits. |
| TCC_PERF_SEL_MC_WRREQ_MC_HALT_STALL | 22 | Number of cycles a mc write request was stalled because the MC halted the interface. |
| TCC_PERF_SEL_TOO_MANY_MC_WRREQS_STALL | 23 | Number of cycles the TCC could not send a mc write request because it already reached its maximum number of pending mc write requests. |
| TCC_PERF_SEL_MC_WRREQ_LEVEL | 24 | The sum of the number of 32-byte mc write requests in flight. This is primarily meant for measure average mc write latency. Average write latency = TCC_PERF_SEL_MC_WRREQ_LEVEL/TCC_PERF_SEL_MC_WRREQ. |
| TCC_PERF_SEL_MC_RDREQ | 25 | Number of 32-byte reads. The hardware actually does 64-byte reads but the number is adjusted to provide uniformity. |
| TCC_PERF_SEL_MC_RDREQ_UNCACHED | 26 | Number of 32-byte reads due to uncached traffic. |
| TCC_PERF_SEL_MC_RDREQ_CREDIT_STALL | 27 | Number of cycles there was a stall because the read request interface was out of credits. Stalls occur regardless of whether a read needed to be performed or not. |
| TCC_PERF_SEL_MC_RDREQ_MC_HALT_STALL | 28 | Number of cycles there was a stall because the read request interface was halted by the MC. Stalls occur regardless of whether a read needed to be performed or not. |
| TCC_PERF_SEL_MC_RDREQ_LEVEL | 29 | The sum of the number of 32-byte mc read requests in flight. This is primarily meant for measure average mc read latency. Average read latency = TCC_PERF_SEL_MC_RDREQ_LEVEL/TCC_PERF_SEL_MC_RDREQ. |
| TCC_PERF_SEL_TAG_STALL | 30 | Number of cycles the tag is stalled for any reason. This also includes TCC_PERF_SEL_MC_RDREQ_CREDIT_STALL and TCC_PERF_SEL_MC_RDREQ_MC_HALT_STALL stall cycles. |
| TCC_PERF_SEL_TAG_WRITEBACK_FIFO_FULL | 31 | Number of cycles that the write-back FIFO within the tag block is full. This does not immediately cause a stall, but it will cause problems in the long run. |
| TCC_PERF_SEL_TAG_MISS_NOTHING_REPLACEABLE_STALL | 32 | Number of cycles there is a cache miss and a line that can be replaced is not available. |
| TCC_PERF_SEL_TAG_UNCACHED_WRITE_ATOMIC_FIFO_FULL_STALL | 33 | Number of cycles the normal request pipeline in the tag was stalled due to the uncached write/atomic fifo being full. |
| TCC_PERF_SEL_TAG_NO_UNCACHED_WRITE_ATOMIC_ENTRIES_STALL | 34 | Number of cycles the normal request pipeline in the tag was stalled due to the lack of unused uncached write/atomic tracking entries. |
| TCC_PERF_SEL_READ_RETURN_TIMEOUT | 35 | Number of bubbles requests sent to the TCA because a mc read return waited so long for a cache ram port that it timed out. |
| TCC_PERF_SEL_WRITEBACK_READ_TIMEOUT | 36 | Number of bubbles requests sent to the TCA because a write-back waited so long for a cache ram port that it timed out. |
| TCC_PERF_SEL_READ_RETURN_FULL_BUBBLE | 37 | Number of bubbles requests sent to the TCA to prevent the mc read return FIFOs from overflowing. Not windowable. |
| TCC_PERF_SEL_BUBBLE | 38 | Total number of bubble requests sent to the TCA |
| TCC_PERF_SEL_RETURN_ACK | 39 | Number of times only an ack was sent on the return bus. |
| TCC_PERF_SEL_RETURN_DATA | 40 | Number of times only data was sent on the return bus. |
| TCC_PERF_SEL_RETURN_HOLE | 41 | Number of times only a hole was sent on the return bus. |
| TCC_PERF_SEL_RETURN_ACK_HOLE | 42 | Number of times an ack and a hole were sent at the same time on the return bus. |
| TCC_PERF_SEL_IB_STALL | 43 | Number of cycles the IB output was stalled. |
| TCC_PERF_SEL_TCA_LEVEL | 44 | The sum of the number of requests sent to the TCA for output arbitration in flight. Average TCA arbitration latency = TCC_PERF_SEL_TCA_LEVEL/TCC_PERF_SEL_REQ. |
| TCC_PERF_SEL_HOLE_LEVEL | 45 | The sum of the number of hole requests in flight. Average hole latency = TCC_PERF_SEL_HOLE_LEVEL/(TCC_PERF_SEL_WRITE+TCC_PERF_SEL_ATOMIC). Not windowable. |
| TCC_PERF_SEL_MC_RDRET_NACK | 46 | The number of 32-byte mc read returns that were nacked. |
| TCC_PERF_SEL_MC_WRRET_NACK | 47 | The number of 32-byte mc write returns that were nacked. |
| TCC_PERF_SEL_EXE_REQ | 48 | The number of exe requests. |
| TCC_PERF_SEL_CLIENT0_REQ | 64 | Number of cycles client0 sent a request to this TCC. |
| TCC_PERF_SEL_CLIENT1_REQ | 65 | |
| TCC_PERF_SEL_CLIENT2_REQ | 66 | |
| TCC_PERF_SEL_CLIENT3_REQ | 67 | |
| TCC_PERF_SEL_CLIENT4_REQ | 68 | |
| TCC_PERF_SEL_CLIENT5_REQ | 69 | |
| TCC_PERF_SEL_CLIENT6_REQ | 70 | |
| TCC_PERF_SEL_CLIENT7_REQ | 71 | |
| TCC_PERF_SEL_CLIENT8_REQ | 72 | |
| TCC_PERF_SEL_CLIENT9_REQ | 73 | |
| TCC_PERF_SEL_CLIENT10_REQ | 74 | |
| TCC_PERF_SEL_CLIENT11_REQ | 75 | |
| TCC_PERF_SEL_CLIENT12_REQ | 76 | |
| TCC_PERF_SEL_CLIENT13_REQ | 77 | |
| TCC_PERF_SEL_CLIENT14_REQ | 78 | |
| TCC_PERF_SEL_CLIENT15_REQ | 79 | |
| TCC_PERF_SEL_CLIENT16_REQ | 80 | |
| TCC_PERF_SEL_CLIENT17_REQ | 81 | |
| TCC_PERF_SEL_CLIENT18_REQ | 82 | |
| TCC_PERF_SEL_CLIENT19_REQ | 83 | |
| TCC_PERF_SEL_CLIENT20_REQ | 84 | |
| TCC_PERF_SEL_CLIENT21_REQ | 85 | |
| TCC_PERF_SEL_CLIENT22_REQ | 86 | |
| TCC_PERF_SEL_CLIENT23_REQ | 87 | |
| TCC_PERF_SEL_CLIENT24_REQ | 88 | |
| TCC_PERF_SEL_CLIENT25_REQ | 89 | |
| TCC_PERF_SEL_CLIENT26_REQ | 90 | |
| TCC_PERF_SEL_CLIENT27_REQ | 91 | |
| TCC_PERF_SEL_CLIENT28_REQ | 92 | |
| TCC_PERF_SEL_CLIENT29_REQ | 93 | |
| TCC_PERF_SEL_CLIENT30_REQ | 94 | |
| TCC_PERF_SEL_CLIENT31_REQ | 95 | |
| TCC_PERF_SEL_CLIENT32_REQ | 96 | |
| TCC_PERF_SEL_CLIENT33_REQ | 97 | |
| TCC_PERF_SEL_CLIENT34_REQ | 98 | |
| TCC_PERF_SEL_CLIENT35_REQ | 99 | |
| TCC_PERF_SEL_CLIENT36_REQ | 100 | |
| TCC_PERF_SEL_CLIENT37_REQ | 101 | |
| TCC_PERF_SEL_CLIENT38_REQ | 102 | |
| TCC_PERF_SEL_CLIENT39_REQ | 103 | |
| TCC_PERF_SEL_CLIENT40_REQ | 104 | |
| TCC_PERF_SEL_CLIENT41_REQ | 105 | |
| TCC_PERF_SEL_CLIENT42_REQ | 106 | |
| TCC_PERF_SEL_CLIENT43_REQ | 107 | |
| TCC_PERF_SEL_CLIENT44_REQ | 108 | |
| TCC_PERF_SEL_CLIENT45_REQ | 109 | |
| TCC_PERF_SEL_CLIENT46_REQ | 110 | |
| TCC_PERF_SEL_CLIENT47_REQ | 111 | |
| TCC_PERF_SEL_CLIENT48_REQ | 112 | |
| TCC_PERF_SEL_CLIENT49_REQ | 113 | |
| TCC_PERF_SEL_CLIENT50_REQ | 114 | |
| TCC_PERF_SEL_CLIENT51_REQ | 115 | |
| TCC_PERF_SEL_CLIENT52_REQ | 116 | |
| TCC_PERF_SEL_CLIENT53_REQ | 117 | |
| TCC_PERF_SEL_CLIENT54_REQ | 118 | |
| TCC_PERF_SEL_CLIENT55_REQ | 119 | |
| TCC_PERF_SEL_CLIENT56_REQ | 120 | |
| TCC_PERF_SEL_CLIENT57_REQ | 121 | |
| TCC_PERF_SEL_CLIENT58_REQ | 122 | |
| TCC_PERF_SEL_CLIENT59_REQ | 123 | |
| TCC_PERF_SEL_CLIENT60_REQ | 124 | |
| TCC_PERF_SEL_CLIENT61_REQ | 125 | |
| TCC_PERF_SEL_CLIENT62_REQ | 126 | |
| TCC_PERF_SEL_CLIENT63_REQ | 127 | |
| TCC_PERF_SEL_CLIENT64_REQ | 128 | |
| TCC_PERF_SEL_CLIENT65_REQ | 129 | |
| TCC_PERF_SEL_NORMAL_WRITEBACK | 130 | Number of write-backs due to requests that are not write-back requests. |
| TCC_PERF_SEL_TC_OP_WBL2_VOL_WRITEBACK | 131 | Number of write-backs due to TC_OP_WBL2_VOL requests. |
| TCC_PERF_SEL_TC_OP_WBINVL2_WRITEBACK | 132 | Number of write-backs due to TC_OP_WBINVL2 requests. |
| TCC_PERF_SEL_ALL_TC_OP_WB_WRITEBACK | 133 | Number of write-backs due to all TC_OP write-back requests. |
| TCC_PERF_SEL_NORMAL_EVICT | 134 | Number of evicts due to requests that are not invalidate requests. |
| TCC_PERF_SEL_TC_OP_INVL2_VOL_EVICT | 135 | Number of evicts due to TC_OP_INVL2_VOL requests. |
| TCC_PERF_SEL_TC_OP_INVL1L2_VOL_EVICT | 136 | Number of evicts due to TC_OP_INVL1L2_VOL requests. |
| TCC_PERF_SEL_TC_OP_WBL2_VOL_EVICT | 137 | Number of evicts due to TC_OP_WBL2_VOL requests. |
| TCC_PERF_SEL_TC_OP_WBINVL2_EVICT | 138 | Number of evicts due to TC_OP_WBINVL2 requests. |
| TCC_PERF_SEL_ALL_TC_OP_INV_EVICT | 139 | Number of evicts due to all TC_OP invalidate requests. |
| TCC_PERF_SEL_ALL_TC_OP_INV_VOL_EVICT | 140 | Number of evicts due to all TC_OP volatile invalidate requests. |
| TCC_PERF_SEL_TC_OP_WBL2_VOL_CYCLE | 141 | Number of cycles spent performing TC_OP_WBL2_VOL requests. |
| TCC_PERF_SEL_TC_OP_INVL2_VOL_CYCLE | 142 | Number of cycles spent performing TC_OP_INVL2_VOL requests. |
| TCC_PERF_SEL_TC_OP_INVL1L2_VOL_CYCLE | 143 | Number of cycles spent performing TC_OP_INVL1L2_VOL requests. |
| TCC_PERF_SEL_TC_OP_WBINVL2_CYCLE | 144 | Number of cycles spent performing TC_OP_WBINVL2 requests. |
| TCC_PERF_SEL_ALL_TC_OP_WB_OR_INV_CYCLE | 145 | Number of cycles spent performing all TC_OP write-back or invalidate requests. |
| TCC_PERF_SEL_ALL_TC_OP_WB_OR_INV_VOL_CYCLE | 146 | Number of cycles spent performing all TC_OP write-back or invalidate volatile requests. |
| TCC_PERF_SEL_TC_OP_WBL2_VOL_START | 147 | Number of TC_OP_WBL2_VOL requests started. |
| TCC_PERF_SEL_TC_OP_INVL2_VOL_START | 148 | Number of TC_OP_INVL2_VOL requests started. |
| TCC_PERF_SEL_TC_OP_INVL1L2_VOL_START | 149 | Number of TC_OP_INVL1L2_VOL requests started. |
| TCC_PERF_SEL_TC_OP_WBINVL2_START | 150 | Number of TC_OP_WBINVL2 requests started. |
| TCC_PERF_SEL_ALL_TC_OP_WB_OR_INV_START | 151 | Number of TC_OP write-back or invalidate requests started. |
| TCC_PERF_SEL_ALL_TC_OP_WB_OR_INV_VOL_START | 152 | Number of TC_OP write-back or invalidate volatile requests started. |
| TCC_PERF_SEL_TC_OP_WBL2_VOL_FINISH | 153 | Number of TC_OP_WBL2_VOL requests finished. |
| TCC_PERF_SEL_TC_OP_INVL2_VOL_FINISH | 154 | Number of TC_OP_INVL2_VOL requests finished. |
| TCC_PERF_SEL_TC_OP_INVL1L2_VOL_FINISH | 155 | Number of TC_OP_INVL1L2_VOL requests finished. |
| TCC_PERF_SEL_TC_OP_WBINVL2_FINISH | 156 | Number of TC_OP_WBINVL2 requests finished. |
| TCC_PERF_SEL_ALL_TC_OP_WB_OR_INV_FINISH | 157 | Number of TC_OP write-back or invalidate requests finished. |
| TCC_PERF_SEL_ALL_TC_OP_WB_OR_INV_VOL_FINISH | 158 | Number of TC_OP write-back or invalidate volatile requests finished. |
| TCC_PERF_SEL_VOL_MC_WRREQ | 159 | Number of 32-byte writes due to volatile requests. |
| TCC_PERF_SEL_VOL_MC_RDREQ | 160 | Number of 32-byte reads due to volatile requests. |
| TCC_PERF_SEL_VOL_REQ | 161 | Number of volatile requests. WB or INV related requests are not included. |
| TCC_PERF_SEL_COMPRESSED_REQ | 162 | The number of compressed requests. This includes requests that read 0 bytes of compressed data. Metadata requests are not included. This is measured at the tag block. |
| TCC_PERF_SEL_COMPRESSED_0_REQ | 163 | The number of compressed requests that read 0 bytes of compressed data. Metadata requests are not included. This is measured at the tag block. |
| TCC_PERF_SEL_METADATA_REQ | 164 | The number of metadata requests. This is measured at the tag block. |
| TCC_PERF_SEL_MC_RDREQ_MDC | 165 | The number of 32-byte reads due to MDC traffic. |
| TCC_PERF_SEL_MC_RDREQ_COMPRESSED | 166 | The number of 32-byte reads due to compressed data reads. This does not include MDC traffic. |
| TCC_PERF_SEL_IB_TAG_STALL | 167 | The number of cycles the IB output was stalled on the tag. |
| TCC_PERF_SEL_IB_MDC_STALL | 168 | The number of cycles the IB output was stalled on the mdc. |
| TCC_PERF_SEL_MDC_REQ | 169 | The number of requests passed through the metadata cache. |
| TCC_PERF_SEL_MDC_LEVEL | 170 | The sum of the number of mdc requests in flight. Average MDC latency=TCC_PERF_SEL_MDC_LEVEL. |
| TCC_PERF_SEL_MDC_TAG_HIT | 171 | The number of MDC tag hits. A tag contains multiple sectors, so many tag hits will be sector misses. |
| TCC_PERF_SEL_MDC_SECTOR_HIT | 172 | The number of MDC sector hits. A sector hit means that there actually is data in the cache for this request. |
| TCC_PERF_SEL_MDC_SECTOR_MISS | 173 | The number of MDC sector misses. |
| TCC_PERF_SEL_MDC_TAG_STALL | 174 | The number of cycles that the mdc tag is stalled. |
| TCC_PERF_SEL_MDC_TAG_REPLACEMENT_LINE_IN_USE_STALL | 175 | The number of cycles that the mdc tag is stalled because the line to be replaced is stalled. |
| TCC_PERF_SEL_MDC_TAG_DESECTORIZATION_FIFO_FULL_STALL | 176 | The number of cycles that the mdc tag is stalled because the desctorization fifo is full. |
| TCC_PERF_SEL_MDC_TAG_WAITING_FOR_INVALIDATE_COMPLETION_STALL | 177 | The number of cycles that the mdc tag is stalled because it is waiting for the previous invalidate to complete. |
| Counter | Value | Description |
|---|---|---|
| TCA_PERF_SEL_NONE | 0 | Don’t count anything. |
| TCA_PERF_SEL_CYCLE | 1 | Number of cycles. Not windowable. |
| TCA_PERF_SEL_BUSY | 2 | Number of cycles we have a request pending. Not windowable. |
| TCA_PERF_SEL_FORCED_HOLE_TCC0 | 3 | Number of TCC0 hole requests that waited so long that it was decided that they should be forced rather than opportunistic. The TCC number is based on the TCCs connected to this TCA and not on the global numbering. |
| TCA_PERF_SEL_FORCED_HOLE_TCC1 | 4 | |
| TCA_PERF_SEL_FORCED_HOLE_TCC2 | 5 | |
| TCA_PERF_SEL_FORCED_HOLE_TCC3 | 6 | |
| TCA_PERF_SEL_FORCED_HOLE_TCC4 | 7 | |
| TCA_PERF_SEL_FORCED_HOLE_TCC5 | 8 | |
| TCA_PERF_SEL_FORCED_HOLE_TCC6 | 9 | |
| TCA_PERF_SEL_FORCED_HOLE_TCC7 | 10 | |
| TCA_PERF_SEL_REQ_TCC0 | 11 | Number of requests sent to TCC0. This should report the same thing as TCC_PERF_SEL_REQ in the corresponding TCC in the long run. It is only offered to create a deterministic counter in the TCA so that the counters may be sanity tested. |
| TCA_PERF_SEL_REQ_TCC1 | 12 | |
| TCA_PERF_SEL_REQ_TCC2 | 13 | |
| TCA_PERF_SEL_REQ_TCC3 | 14 | |
| TCA_PERF_SEL_REQ_TCC4 | 15 | |
| TCA_PERF_SEL_REQ_TCC5 | 16 | |
| TCA_PERF_SEL_REQ_TCC6 | 17 | |
| TCA_PERF_SEL_REQ_TCC7 | 18 | |
| TCA_PERF_SEL_CROSSBAR_DOUBLE_ARB_TCC0 | 19 | Number of cycles two requests were sent in one clock to TCC0. The TCC number is based on the TCCs connected to this TCA and not on the global numbering. |
| TCA_PERF_SEL_CROSSBAR_DOUBLE_ARB_TCC1 | 20 | |
| TCA_PERF_SEL_CROSSBAR_DOUBLE_ARB_TCC2 | 21 | |
| TCA_PERF_SEL_CROSSBAR_DOUBLE_ARB_TCC3 | 22 | |
| TCA_PERF_SEL_CROSSBAR_DOUBLE_ARB_TCC4 | 23 | |
| TCA_PERF_SEL_CROSSBAR_DOUBLE_ARB_TCC5 | 24 | |
| TCA_PERF_SEL_CROSSBAR_DOUBLE_ARB_TCC6 | 25 | |
| TCA_PERF_SEL_CROSSBAR_DOUBLE_ARB_TCC7 | 26 | |
| TCA_PERF_SEL_CROSSBAR_STALL_TCC0 | 27 | Number of cycles no requests could be sent to TCC0 because its FIFOs were expected to be full. The TCC number is based on the TCCs connected to this TCA and not on the global numbering. |
| TCA_PERF_SEL_CROSSBAR_STALL_TCC1 | 28 | |
| TCA_PERF_SEL_CROSSBAR_STALL_TCC2 | 29 | |
| TCA_PERF_SEL_CROSSBAR_STALL_TCC3 | 30 | |
| TCA_PERF_SEL_CROSSBAR_STALL_TCC4 | 31 | |
| TCA_PERF_SEL_CROSSBAR_STALL_TCC5 | 32 | |
| TCA_PERF_SEL_CROSSBAR_STALL_TCC6 | 33 | |
| TCA_PERF_SEL_CROSSBAR_STALL_TCC7 | 34 | |
| TCA_PERF_SEL_FORCED_HOLE_TCS | 35 | |
| TCA_PERF_SEL_REQ_TCS | 36 | |
| TCA_PERF_SEL_CROSSBAR_DOUBLE_ARB_TCS | 37 | |
| TCA_PERF_SEL_CROSSBAR_STALL_TCS | 38 |
| Counter | Value | Description |
|---|---|---|
| GDS_PERF_SEL_DS_ADDR_CONFL | 0 | Number of conflicting addresses in DS memory |
| GDS_PERF_SEL_DS_BANK_CONFL | 1 | Number of conflicting banks in DS memory |
| GDS_PERF_SEL_WBUF_FLUSH | 2 | Number of times the RBIU write buffer is flushed |
| GDS_PERF_SEL_WR_COMP | 3 | Number of times a WRITE_COMPLETE is put in write buffer |
| GDS_PERF_SEL_WBUF_WR | 4 | Number of times the RBIU writes to buffer |
| GDS_PERF_SEL_RBUF_HIT | 5 | Number of times an RBIU read hits in the buffer |
| GDS_PERF_SEL_RBUF_MISS | 6 | Number of times an RBIU read misses the buffer |
| GDS_PERF_SEL_SE0_SH0_NORET | 7 | Number of commands that do not return data |
| GDS_PERF_SEL_SE0_SH0_RET | 8 | Number of commands that do return data |
| GDS_PERF_SEL_SE0_SH0_ORD_CNT | 9 | Number of ordered count operations |
| GDS_PERF_SEL_SE0_SH0_2COMP_REQ | 10 | Number of commands that return two components |
| GDS_PERF_SEL_SE0_SH0_ORD_WAVE_VALID | 11 | Number of valid wave ID’s processed by ordered append |
| GDS_PERF_SEL_SE0_SH0_GDS_DATA_VALID | 12 | Number of valid write enables from the data bus |
| GDS_PERF_SEL_SE0_SH0_GDS_STALL_BY_ORD | 13 | Number of cycles GDS data is stalled due to Ordered count data |
| GDS_PERF_SEL_SE0_SH0_GDS_WR_OP | 14 | Number of GDS Write operations |
| GDS_PERF_SEL_SE0_SH0_GDS_RD_OP | 15 | Number of GDS Read operations |
| GDS_PERF_SEL_SE0_SH0_GDS_ATOM_OP | 16 | Number of GDS Atomic operations |
| GDS_PERF_SEL_SE0_SH0_GDS_REL_OP | 17 | Number of GDS Relative operations |
| GDS_PERF_SEL_SE0_SH0_GDS_CMPXCH_OP | 18 | Number of GDS Compare or Exchange operations |
| GDS_PERF_SEL_SE0_SH0_GDS_BYTE_OP | 19 | Number of GDS byte operations |
| GDS_PERF_SEL_SE0_SH0_GDS_SHORT_OP | 20 | Number of GDS short operations |
| GDS_PERF_SEL_SE1_SH0_NORET | 35 | Number of commands that do not return data |
| GDS_PERF_SEL_SE1_SH0_RET | 36 | Number of commands that do return data |
| GDS_PERF_SEL_SE1_SH0_ORD_CNT | 37 | Number of ordered count operations |
| GDS_PERF_SEL_SE1_SH0_2COMP_REQ | 38 | Number of commands that return two components |
| GDS_PERF_SEL_SE1_SH0_ORD_WAVE_VALID | 39 | Number of valid wave ID’s processed by ordered append |
| GDS_PERF_SEL_SE1_SH0_GDS_DATA_VALID | 40 | Number of valid write enables from the data bus |
| GDS_PERF_SEL_SE1_SH0_GDS_STALL_BY_ORD | 41 | Number of cycles GDS data is stalled due to Ordered count data |
| GDS_PERF_SEL_SE1_SH0_GDS_WR_OP | 42 | Number of GDS Write operations |
| GDS_PERF_SEL_SE1_SH0_GDS_RD_OP | 43 | Number of GDS Read operations |
| GDS_PERF_SEL_SE1_SH0_GDS_ATOM_OP | 44 | Number of GDS Atomic operations |
| GDS_PERF_SEL_SE1_SH0_GDS_REL_OP | 45 | Number of GDS Relative operations |
| GDS_PERF_SEL_SE1_SH0_GDS_CMPXCH_OP | 46 | Number of GDS Compare or Exchange operations |
| GDS_PERF_SEL_SE1_SH0_GDS_BYTE_OP | 47 | Number of GDS byte operations |
| GDS_PERF_SEL_SE1_SH0_GDS_SHORT_OP | 48 | Number of GDS short operations |
| GDS_PERF_SEL_GWS_RELEASED | 119 | Number of waves released by GWS |
| GDS_PERF_SEL_GWS_BYPASS | 120 | Number of waves bypassed by GWS |
| GDS_PERF_SEL_SCORPIO_START | 121 | Start of Xbox One X-specific counters. |
| MCD_GFX_GARLIC_CH0_RD | 121 | The number of events from MC. |
| MCD_GFX_GARLIC_CH0_WR | 122 | The number of events from MC. |
| MCD_GFX_GARLIC_CH1_RD | 123 | The number of events from MC. |
| MCD_GFX_GARLIC_CH1_WR | 124 | The number of events from MC. |
| MCD_GFX_ONION_CH0_RD | 125 | The number of events from MC. |
| MCD_GFX_ONION_CH0_WR | 126 | The number of events from MC. |
| MCD_GFX_ONION_CH1_RD | 127 | The number of events from MC. |
| MCD_GFX_ONION_CH1_WR | 128 | The number of events from MC. |
| MCD_GFX_MUXED_PERFCNT_S0 | 129 | The number of events from MC. |
| MCD_GFX_MUXED_PERFCNT_S1 | 130 | The number of events from MC. |
| MCD_GFX_MUXED_PERFCNT_S2 | 131 | The number of events from MC. |
| MCD_GFX_MUXED_PERFCNT_S3 | 132 | The number of events from MC. |
| MCC_GFX_MUXED_PERFCNT_S0 | 133 | The number of events from MC. |
| MCC_GFX_MUXED_PERFCNT_S1 | 134 | The number of events from MC. |
| MCC_GFX_MUXED_PERFCNT_S2 | 135 | The number of events from MC. |
| MCC_GFX_MUXED_PERFCNT_S3 | 136 | The number of events from MC. |
| Counter | Value | Description |
|---|---|---|
| VGT_PERF_VGT_SPI_ESTHREAD_EVENT_WINDOW_ACTIVE | 0 | ES thread event window is active |
| VGT_PERF_VGT_SPI_ESVERT_VALID | 1 | ES Vert is valid |
| VGT_PERF_VGT_SPI_ESVERT_EOV | 2 | ES vert end of vector is active |
| VGT_PERF_VGT_SPI_ESVERT_STALLED | 3 | ES vert pipe is stalled |
| VGT_PERF_VGT_SPI_ESVERT_STARVED_BUSY | 4 | ES vert pipe is starved busy |
| VGT_PERF_VGT_SPI_ESVERT_STARVED_IDLE | 5 | ES vert pipe is starved idle |
| VGT_PERF_VGT_SPI_ESVERT_STATIC | 6 | ES vert pipe is static |
| VGT_PERF_VGT_SPI_ESTHREAD_IS_EVENT | 7 | ES Thread Event Indicator |
| VGT_PERF_VGT_SPI_ESTHREAD_SEND | 8 | ES Thread Send is active |
| VGT_PERF_VGT_SPI_GSPRIM_VALID | 9 | ES GS Primitive send is active |
| VGT_PERF_VGT_SPI_GSPRIM_EOV | 10 | ES GS Primitive end of vector is active |
| VGT_PERF_VGT_SPI_GSPRIM_CONT | 11 | ES GS Primitive Continued Event |
| VGT_PERF_VGT_SPI_GSPRIM_STALLED | 12 | ES GS Primitive is stalled |
| VGT_PERF_VGT_SPI_GSPRIM_STARVED_BUSY | 13 | ES GS Primitive is starved busy |
| VGT_PERF_VGT_SPI_GSPRIM_STARVED_IDLE | 14 | ES GS Primitive is starved idle |
| VGT_PERF_VGT_SPI_GSPRIM_STATIC | 15 | ES GS Primitive is static |
| VGT_PERF_VGT_SPI_GSTHREAD_EVENT_WINDOW_ACTIVE | 16 | GS Thread event window is active |
| VGT_PERF_VGT_SPI_GSTHREAD_IS_EVENT | 17 | GS Thread event is being processed |
| VGT_PERF_VGT_SPI_GSTHREAD_SEND | 18 | GS Thread is being sent |
| VGT_PERF_VGT_SPI_VSTHREAD_EVENT_WINDOW_ACTIVE | 19 | VS Thread event window is active |
| VGT_PERF_VGT_SPI_VSVERT_SEND | 20 | VS vert send |
| VGT_PERF_VGT_SPI_VSVERT_EOV | 21 | VS vert end of vector |
| VGT_PERF_VGT_SPI_VSVERT_STALLED | 22 | VS vert pipe is stalled |
| VGT_PERF_VGT_SPI_VSVERT_STARVED_BUSY | 23 | VS vert pipe is starved busy |
| VGT_PERF_VGT_SPI_VSVERT_STARVED_IDLE | 24 | VS vert pipe is starved idle |
| VGT_PERF_VGT_SPI_VSVERT_STATIC | 25 | VS vert pipe is static |
| VGT_PERF_VGT_SPI_VSTHREAD_IS_EVENT | 26 | VS Thread Event Indicator |
| VGT_PERF_VGT_SPI_VSTHREAD_SEND | 27 | VS Thread is being sent |
| VGT_PERF_VGT_PA_EVENT_WINDOW_ACTIVE | 28 | VGT to Primitive Assembler Event Window is active |
| VGT_PERF_VGT_PA_CLIPV_SEND | 29 | VGT to Primitive Assembler clipv is being sent |
| VGT_PERF_VGT_PA_CLIPV_FIRSTVERT | 30 | VGT to Primitive Assembler clipv is the first vert |
| VGT_PERF_VGT_PA_CLIPV_STALLED | 31 | VGT to Primitive Assembler pipe is stalled |
| VGT_PERF_VGT_PA_CLIPV_STARVED_BUSY | 32 | VGT to Primitive Assembler pipe is starved_busy |
| VGT_PERF_VGT_PA_CLIPV_STARVED_IDLE | 33 | VGT to Primitive Assembler pipe is starved_idle |
| VGT_PERF_VGT_PA_CLIPV_STATIC | 34 | VGT to Primitive Assembler pipe is static |
| VGT_PERF_VGT_PA_CLIPP_SEND | 35 | VGT to Primitive Assembler is being sent |
| VGT_PERF_VGT_PA_CLIPP_EOP | 36 | VGT to Primitive Assembler end of packet |
| VGT_PERF_VGT_PA_CLIPP_IS_EVENT | 37 | VGT to Primitive Assembler event transition |
| VGT_PERF_VGT_PA_CLIPP_NULL_PRIM | 38 | VGT to Primitive Assembler null primitive is present |
| VGT_PERF_VGT_PA_CLIPP_NEW_VTX_VECT | 39 | VGT to Primitive Assembler new vertex vector is present |
| VGT_PERF_VGT_PA_CLIPP_STALLED | 40 | VGT to Primitive Assembler pipe is stalled |
| VGT_PERF_VGT_PA_CLIPP_STARVED_BUSY | 41 | VGT to Primitive Assembler pipe is starved_busy |
| VGT_PERF_VGT_PA_CLIPP_STARVED_IDLE | 42 | VGT to Primitive Assembler pipe is starved_idle |
| VGT_PERF_VGT_PA_CLIPP_STATIC | 43 | VGT to Primitive Assembler pipe is static |
| VGT_PERF_VGT_PA_CLIPS_SEND | 44 | VGT to Primitive Assembler is being sent |
| VGT_PERF_VGT_PA_CLIPS_STALLED | 45 | VGT to Primitive Assembler pipe is stalled |
| VGT_PERF_VGT_PA_CLIPS_STARVED_BUSY | 46 | VGT to Primitive Assembler pipe is starved_busy |
| VGT_PERF_VGT_PA_CLIPS_STARVED_IDLE | 47 | VGT to Primitive Assembler pipe is starved_idle |
| VGT_PERF_VGT_PA_CLIPS_STATIC | 48 | VGT to Primitive Assembler pipe is static |
| VGT_PERF_VSVERT_DS_SEND | 49 | DS vert send across vsvert interface |
| VGT_PERF_VSVERT_API_SEND | 50 | API VS vert send across vsvert interface |
| VGT_PERF_HS_TIF_STALL | 51 | TE11 Input FIFO stall |
| VGT_PERF_HS_INPUT_STALL | 52 | HS Input FIFO stall |
| VGT_PERF_HS_INTERFACE_STALL | 53 | HS Interface (hsvert and hswave) stall |
| VGT_PERF_HS_TFM_STALL | 54 | HS TF Memory stall |
| VGT_PERF_TE11_STARVED | 55 | TE11 starved (waiting for tess factors) |
| VGT_PERF_GS_EVENT_STALL | 56 | GS event stalled due to full GS Event FIFO |
| VGT_PERF_RESERVED0 | 57 | RESERVED |
| VGT_PERF_RESERVED1 | 58 | RESERVED |
| VGT_PERF_RESERVED2 | 59 | RESERVED |
| VGT_PERF_RESERVED3 | 60 | RESERVED |
| VGT_PERF_RESERVED4 | 61 | RESERVED |
| VGT_PERF_RESERVED5 | 62 | RESERVED |
| VGT_PERF_RESERVED6 | 63 | RESERVED |
| VGT_PERF_VGT_BUSY | 64 | Number of cycles VGT is busy |
| VGT_PERF_VGT_GS_BUSY | 65 | Number of cycles VGT GS block is busy |
| VGT_PERF_ESVERT_STALLED_ES_TBL | 66 | esvert transfers are stalled because of ES table being full |
| VGT_PERF_ESVERT_STALLED_GS_TBL | 67 | esvert transfers are stalled because of GS table being full |
| VGT_PERF_ESVERT_STALLED_GS_EVENT | 68 | esvert transfers are stalled because of events |
| VGT_PERF_ESVERT_STALLED_GSPRIM | 69 | esvert transfers are stalled because of GS prim interface is full |
| VGT_PERF_GSPRIM_STALLED_ES_TBL | 70 | gsprim transfers are stalled because of ES table being full |
| VGT_PERF_GSPRIM_STALLED_GS_TBL | 71 | gsprim transfers are stalled because of GS table being full |
| VGT_PERF_GSPRIM_STALLED_GS_EVENT | 72 | gsprim transfers are stalled because of events |
| VGT_PERF_GSPRIM_STALLED_ESVERT | 73 | gsprim transfers are stalled because of ES vert interface is full |
| VGT_PERF_ESTHREAD_STALLED_ES_RB_FULL | 74 | ES thread sends are stalled because the ES ring buffer is full |
| VGT_PERF_ESTHREAD_STALLED_SPI_BP | 75 | ES thread is stalled due to back pressure from the SPI |
| VGT_PERF_COUNTERS_AVAIL_STALLED | 76 | GS thread send is stalled because no counters are available |
| VGT_PERF_GS_RB_SPACE_AVAIL_STALLED | 77 | GS thread send is stalled because the GS/VS ring buffer is full |
| VGT_PERF_GS_ISSUE_RTR_STALLED | 78 | GS thread send is stalled due to something other than the counters or ring buffer being full |
| VGT_PERF_GSTHREAD_STALLED | 79 | GS thread send is stalled. Inclusive of the 3 counters above. |
| VGT_PERF_STRMOUT_STALLED | 80 | Number of cycles vs waves are stalled by a stream-out sync. |
| VGT_PERF_WAIT_FOR_ES_DONE_STALLED | 81 | GS thread SM is ready to move to the next stage as soon as the ES thread finishes |
| VGT_PERF_CM_STALLED_BY_GOG | 82 | the output FIFO to the GOG is full and the CM wants to send data |
| VGT_PERF_CM_READING_STALLED | 83 | the GOG can accept data and the CM should be sending data, but isn’t |
| VGT_PERF_CM_STALLED_BY_GSFETCH_DONE | 84 | all CM state machines are waiting for gsfetch_done |
| VGT_PERF_GOG_VS_TBL_STALLED | 85 | GOG is stalled because the VS table is full |
| VGT_PERF_GOG_OUT_INDX_STALLED | 86 | GOG is stalled by back pressure from the output block |
| VGT_PERF_GOG_OUT_PRIM_STALLED | 87 | GOG is stalled by back pressure from the output block |
| VGT_PERF_WAVEID_STALLED | 88 | |
| VGT_PERF_GOG_BUSY | 89 | Counts number of cycles that the GOG block is busy |
| VGT_PERF_REUSED_VS_INDICES | 90 | Counts number of reused indices, excluding GS scenario G and tessellation |
| VGT_PERF_SCLK_REG_VLD_EVENT | 91 | Counts number of cycles sclk_reg is valid |
| VGT_PERF_RESERVED9 | 92 | RESERVED |
| VGT_PERF_SCLK_CORE_VLD_EVENT | 93 | Counts number of cycles sclk_core is valid |
| VGT_PERF_RESERVED10 | 94 | RESERVED |
| VGT_PERF_SCLK_GS_VLD_EVENT | 95 | Counts number of cycles sclk_gs is valid |
| VGT_PERF_VGT_SPI_LSVERT_VALID | 96 | LS Vert is valid. Sensitive to PERF_SEID_IGNORE_MASK |
| VGT_PERF_VGT_SPI_LSVERT_EOV | 97 | LS vert end of vector is active. Sensitive to PERF_SEID_IGNORE_MASK |
| VGT_PERF_VGT_SPI_LSVERT_STALLED | 98 | LS vert pipe is stalled. Sensitive to PERF_SEID_IGNORE_MASK |
| VGT_PERF_VGT_SPI_LSVERT_STARVED_BUSY | 99 | LS vert pipe is starved busy. Sensitive to PERF_SEID_IGNORE_MASK |
| VGT_PERF_VGT_SPI_LSVERT_STARVED_IDLE | 100 | LS vert pipe is starved idle. Sensitive to PERF_SEID_IGNORE_MASK |
| VGT_PERF_VGT_SPI_LSVERT_STATIC | 101 | LS vert pipe is static. Sensitive to PERF_SEID_IGNORE_MASK |
| VGT_PERF_VGT_SPI_LSWAVE_EVENT_WINDOW_ACTIVE | 102 | LS WAVE Event window is active. Sensitive to PERF_SEID_IGNORE_MASK |
| VGT_PERF_VGT_SPI_LSWAVE_IS_EVENT | 103 | LS Wave Event Indicator. Sensitive to PERF_SEID_IGNORE_MASK |
| VGT_PERF_VGT_SPI_LSWAVE_SEND | 104 | LS Wave Send is active. Sensitive to PERF_SEID_IGNORE_MASK |
| VGT_PERF_VGT_SPI_HSVERT_VALID | 105 | HS Vert is valid. Sensitive to PERF_SEID_IGNORE_MASK |
| VGT_PERF_VGT_SPI_HSVERT_EOV | 106 | HS vert end of vector is active. Sensitive to PERF_SEID_IGNORE_MASK |
| VGT_PERF_VGT_SPI_HSVERT_STALLED | 107 | HS vert pipe is stalled. Sensitive to PERF_SEID_IGNORE_MASK |
| VGT_PERF_VGT_SPI_HSVERT_STARVED_BUSY | 108 | HS vert pipe is starved busy. Sensitive to PERF_SEID_IGNORE_MASK |
| VGT_PERF_VGT_SPI_HSVERT_STARVED_IDLE | 109 | HS vert pipe is starved idle. Sensitive to PERF_SEID_IGNORE_MASK |
| VGT_PERF_VGT_SPI_HSVERT_STATIC | 110 | HS vert pipe is static. Sensitive to PERF_SEID_IGNORE_MASK |
| VGT_PERF_VGT_SPI_HSWAVE_EVENT_WINDOW_ACTIVE | 111 | HS WAVE Event window is active. Sensitive to PERF_SEID_IGNORE_MASK |
| VGT_PERF_VGT_SPI_HSWAVE_IS_EVENT | 112 | HS Wave Event Indicator. Sensitive to PERF_SEID_IGNORE_MASK |
| VGT_PERF_VGT_SPI_HSWAVE_SEND | 113 | HS Wave Send is active. Sensitive to PERF_SEID_IGNORE_MASK |
| VGT_PERF_RESERVED11 | 114 | RESERVED |
| VGT_PERF_NULL_TESS_PATCHES | 115 | Count of patches determined to be null due to tessellation factors of negative, 0 or NAN |
| VGT_PERF_LS_THREAD_GROUPS | 116 | Count of thread groups issued to the LS. Sensitive to PERF_SEID_IGNORE_MASK |
| VGT_PERF_HS_THREAD_GROUPS | 117 | Count of thread groups issued to the HS. Sensitive to PERF_SEID_IGNORE_MASK |
| VGT_PERF_ES_THREAD_GROUPS | 118 | Count of thread groups issued to the ES (these are Domain Shader Thread groups). Sensitive to PERF_SEID_IGNORE_MASK |
| VGT_PERF_VS_THREAD_GROUPS | 119 | Count of thread groups issued to the VS (these are Domain Shader Thread groups). Sensitive to PERF_SEID_IGNORE_MASK |
| VGT_PERF_LS_DONE_LATENCY | 120 | Total latency LS FLUSHES issued to the SX DONES received. Divide by LS FLUSHES issued for average latency. Sensitive to PERF_SEID_IGNORE_MASK |
| VGT_PERF_HS_DONE_LATENCY | 121 | Total latency HS FLUSHES issued to the SX DONES received. Divide by HS FLUSHES issued for average latency. Sensitive to PERF_SEID_IGNORE_MASK |
| VGT_PERF_ES_DONE_LATENCY | 122 | Total latency ES FLUSHES issued to the SX DONES received. Divide by ES FLUSHES issued for average latency. Sensitive to PERF_SEID_IGNORE_MASK |
| VGT_PERF_GS_DONE_LATENCY | 123 | Total latency GS FLUSHES issued to the SX DONES received. Divide by GS FLUSHES issued for average latency. Sensitive to PERF_SEID_IGNORE_MASK |
| VGT_PERF_VGT_HS_BUSY | 124 | Counts number of cycles the HS block is busy (creates CS/LS and HS work) |
| VGT_PERF_VGT_TE11_BUSY | 125 | Counts number of cycles the TE11 block is busy. (DX11 Tessellation Fixed Function Logic) |
| VGT_PERF_LS_FLUSH | 126 | Counts number of LS FLUSH events issues by the VGT. Sensitive to PERF_SEID_IGNORE_MASK |
| VGT_PERF_HS_FLUSH | 127 | Counts number of HS FLUSH events issues by the VGT. Sensitive to PERF_SEID_IGNORE_MASK |
| VGT_PERF_ES_FLUSH | 128 | Counts number of ES FLUSH events issues by the VGT. Sensitive to PERF_SEID_IGNORE_MASK |
| VGT_PERF_GS_FLUSH | 129 | Counts number of GS FLUSH events issues by the VGT. Sensitive to PERF_SEID_IGNORE_MASK |
| VGT_PERF_LS_DONE | 130 | Counts number of SX LS DONES received by the VGT. Sensitive to PERF_SEID_IGNORE_MASK |
| VGT_PERF_HS_DONE | 131 | Counts number of SX HS DONES received by the VGT. Sensitive to PERF_SEID_IGNORE_MASK |
| VGT_PERF_ES_DONE | 132 | Counts number of SX ES DONES received by the VGT. Sensitive to PERF_SEID_IGNORE_MASK |
| VGT_PERF_GS_DONE | 133 | Counts number of SX GS DONES received by the VGT. Sensitive to PERF_SEID_IGNORE_MASK |
| VGT_PERF_VSFETCH_DONE | 134 | Counts number of SH VSFETCH DONES received by the VGT. Sensitive to PERF_SEID_IGNORE_MASK |
| VGT_PERF_RESERVED12 | 135 | RESERVED |
| VGT_PERF_ES_RING_HIGH_WATER_MARK | 136 | Records maximum number of ES waves outstanding in the ES ring. Counter select must be set to MAX mode |
| VGT_PERF_GS_RING_HIGH_WATER_MARK | 137 | Records maximum number of GS waves outstanding in the GS ring. Counter select must be set to MAX mode |
| VGT_PERF_VS_TABLE_HIGH_WATER_MARK | 138 | Records maximum number of VS waves outstanding in the VS table. Counter select must be set to MAX mode |
| VGT_PERF_HS_TGS_ACTIVE_HIGH_WATER_MARK | 139 | Records maximum number of active thread groups. Counter select must be set to MAX mode |
| Counter | Value | Description |
|---|---|---|
| IA_PERF_GRP_INPUT_EVENT_WINDOW_ACTIVE | 0 | Grouper input generated event window is active |
| IA_PERF_RESERVED0 | 1 | RESERVED |
| IA_PERF_RESERVED1 | 2 | RESERVED |
| IA_PERF_RESERVED2 | 3 | RESERVED |
| IA_PERF_RESERVED3 | 4 | RESERVED |
| IA_PERF_RESERVED4 | 5 | RESERVED |
| IA_PERF_RESERVED5 | 6 | RESERVED |
| IA_PERF_MC_LAT_BIN_0 | 7 | Memory Controller Latency Bin 0 |
| IA_PERF_MC_LAT_BIN_1 | 8 | Memory Controller Latency Bin 1 |
| IA_PERF_MC_LAT_BIN_2 | 9 | Memory Controller Latency Bin 2 |
| IA_PERF_MC_LAT_BIN_3 | 10 | Memory Controller Latency Bin 3 |
| IA_PERF_MC_LAT_BIN_4 | 11 | Memory Controller Latency Bin 4 |
| IA_PERF_MC_LAT_BIN_5 | 12 | Memory Controller Latency Bin 5 |
| IA_PERF_MC_LAT_BIN_6 | 13 | Memory Controller Latency Bin 6 |
| IA_PERF_MC_LAT_BIN_7 | 14 | Memory Controller Latency Bin 7 |
| IA_PERF_IA_BUSY | 15 | Number of cycles IA is busy |
| IA_PERF_IA_SCLK_REG_VLD_EVENT | 16 | Counts number of cycles sclk_reg is valid |
| IA_PERF_IA_SCLK_INPUT_VLD_EVENT | 17 | Counts number of cycles sclk_input is valid |
| IA_PERF_IA_SCLK_CORE_VLD_EVENT | 18 | Counts number of cycles sclk_core is valid |
| IA_PERF_IA_SCLK_INVAL_VLD_EVENT | 19 | Counts number of cycles sclk_inval is valid |
| IA_PERF_IA_DMA_RETURN | 20 | Each return is equal to 32 bytes returned to the IA |
| IA_PERF_IA_STALLED | 21 | Counts number of cycles that IA is stalled by full prim output FIFOs |
| Counter | Value | Description |
|---|---|---|
| WD_PERF_RBIU_FIFOS_EVENT_WINDOW_ACTIVE | 0 | |
| WD_PERF_RBIU_DR_FIFO_STARVED | 1 | |
| WD_PERF_RBIU_DR_FIFO_STALLED | 2 | |
| WD_PERF_RBIU_DI_FIFO_STARVED | 3 | |
| WD_PERF_RBIU_DI_FIFO_STALLED | 4 | |
| WD_PERF_WD_BUSY | 5 | |
| WD_PERF_WD_SCLK_REG_VLD_EVENT | 6 | |
| WD_PERF_WD_SCLK_INPUT_VLD_EVENT | 7 | |
| WD_PERF_WD_SCLK_CORE_VLD_EVENT | 8 | |
| WD_PERF_WD_STALLED | 9 |
| Counter | Value | Description |
|---|---|---|
| MC_HV_MCB_L1TLB_PERF_CYCLES | 0 | count cycles |
| MC_HV_MCB_L1TLB_PERF_TLB0_REQUEST_COUNT | 1 | number of tlb requests event |
| MC_HV_MCB_L1TLB_PERF_TLB0_HIT_COUNT | 2 | number of tlb hits event |
| MC_HV_MCB_L1TLB_PERF_TLB0_MISS_COUNT | 3 | number of tlb misses event |
| MC_HV_MCB_L1TLB_PERF_TLB0_DEFAULT_REQUEST_COUNT | 4 | number of default requests event |
| MC_HV_MCB_L1TLB_PERF_TLB0_PHYS_UVD_BYPASS_COUNT | 5 | number of physical and uvd bypass requests event |
| MC_HV_MCB_L1TLB_PERF_TLB0_PRETRANSLATED_COUNT | 6 | number of pre-translated requests event |
| MC_HV_MCB_L1TLB_PERF_TLB0_LATENCY_START | 7 | event start for the VM latency counter event |
| MC_HV_MCB_L1TLB_PERF_TLB0_LATENCY_END | 8 | event end for the VM latency counter |
| MC_HV_MCB_L1TLB_PERF_TLB1_REQUEST_COUNT | 9 | number of tlb requests event |
| MC_HV_MCB_L1TLB_PERF_TLB1_HIT_COUNT | 10 | number of tlb hits event |
| MC_HV_MCB_L1TLB_PERF_TLB1_MISS_COUNT | 11 | number of tlb misses event |
| MC_HV_MCB_L1TLB_PERF_TLB1_DEFAULT_REQUEST_COUNT | 12 | number of default requests event |
| MC_HV_MCB_L1TLB_PERF_TLB1_PHYS_UVD_BYPASS_COUNT | 13 | number of physical and uvd bypass requests event |
| MC_HV_MCB_L1TLB_PERF_TLB1_PRETRANSLATED_COUNT | 14 | number of pre-translated requests event |
| MC_HV_MCB_L1TLB_PERF_TLB1_LATENCY_START | 15 | event start for the VM latency counter event |
| MC_HV_MCB_L1TLB_PERF_TLB1_LATENCY_END | 16 | event end for the VM latency counter |
| MC_HV_MCB_L1TLB_PERF_TLB2_REQUEST_COUNT | 17 | number of tlb requests event |
| MC_HV_MCB_L1TLB_PERF_TLB2_HIT_COUNT | 18 | number of tlb hits event |
| MC_HV_MCB_L1TLB_PERF_TLB2_MISS_COUNT | 19 | number of tlb misses event |
| MC_HV_MCB_L1TLB_PERF_TLB2_DEFAULT_REQUEST_COUNT | 20 | number of default requests event |
| MC_HV_MCB_L1TLB_PERF_TLB2_PHYS_UVD_BYPASS_COUNT | 21 | number of physical and uvd bypass requests event |
| MC_HV_MCB_L1TLB_PERF_TLB2_PRETRANSLATED_COUNT | 22 | number of pre-translated requests event |
| MC_HV_MCB_L1TLB_PERF_TLB2_LATENCY_START | 23 | event start for the VM latency counter event |
| MC_HV_MCB_L1TLB_PERF_TLB2_LATENCY_END | 24 | event end for the VM latency counter |
| MC_HV_MCB_L1TLB_PERF_TLB3_REQUEST_COUNT | 25 | number of tlb requests event |
| MC_HV_MCB_L1TLB_PERF_TLB3_HIT_COUNT | 26 | number of tlb hits event |
| MC_HV_MCB_L1TLB_PERF_TLB3_MISS_COUNT | 27 | number of tlb misses event |
| MC_HV_MCB_L1TLB_PERF_TLB3_DEFAULT_REQUEST_COUNT | 28 | number of default requests event |
| MC_HV_MCB_L1TLB_PERF_TLB3_PHYS_UVD_BYPASS_COUNT | 29 | number of physical and uvd bypass requests event |
| MC_HV_MCB_L1TLB_PERF_TLB3_PRETRANSLATED_COUNT | 30 | number of pre-translated requests event |
| MC_HV_MCB_L1TLB_PERF_TLB3_LATENCY_START | 31 | event start for the VM latency counter event |
| MC_HV_MCB_L1TLB_PERF_TLB3_LATENCY_END | 32 | event end for the VM latency counter |
| Counter | Value | Description |
|---|---|---|
| MC_HV_MCD_L1TLB_PERF_CYCLES | 0 | count cycles |
| MC_HV_MCD_L1TLB_PERF_TLB0_REQUEST_COUNT | 1 | number of tlb requests event |
| MC_HV_MCD_L1TLB_PERF_TLB0_HIT_COUNT | 2 | number of tlb hits event |
| MC_HV_MCD_L1TLB_PERF_TLB0_MISS_COUNT | 3 | number of tlb misses event |
| MC_HV_MCD_L1TLB_PERF_TLB0_DEFAULT_REQUEST_COUNT | 4 | number of default requests event |
| MC_HV_MCD_L1TLB_PERF_TLB0_PHYS_UVD_BYPASS_COUNT | 5 | number of physical and uvd bypass requests event |
| MC_HV_MCD_L1TLB_PERF_TLB0_PRETRANSLATED_COUNT | 6 | number of pre-translated requests event |
| MC_HV_MCD_L1TLB_PERF_TLB0_LATENCY_START | 7 | event start for the VM latency counter event |
| MC_HV_MCD_L1TLB_PERF_TLB0_LATENCY_END | 8 | event end for the VM latency counter |
| MC_HV_MCD_L1TLB_PERF_TLB1_REQUEST_COUNT | 9 | number of tlb requests event |
| MC_HV_MCD_L1TLB_PERF_TLB1_HIT_COUNT | 10 | number of tlb hits event |
| MC_HV_MCD_L1TLB_PERF_TLB1_MISS_COUNT | 11 | number of tlb misses event |
| MC_HV_MCD_L1TLB_PERF_TLB1_DEFAULT_REQUEST_COUNT | 12 | number of default requests event |
| MC_HV_MCD_L1TLB_PERF_TLB1_PHYS_UVD_BYPASS_COUNT | 13 | number of physical and uvd bypass requests event |
| MC_HV_MCD_L1TLB_PERF_TLB1_PRETRANSLATED_COUNT | 14 | number of pre-translated requests event |
| MC_HV_MCD_L1TLB_PERF_TLB1_LATENCY_START | 15 | event start for the VM latency counter event |
| MC_HV_MCD_L1TLB_PERF_TLB1_LATENCY_END | 16 | event end for the VM latency counter |
| MC_HV_MCD_L1TLB_PERF_TLB2_REQUEST_COUNT | 17 | number of tlb requests event |
| MC_HV_MCD_L1TLB_PERF_TLB2_HIT_COUNT | 18 | number of tlb hits event |
| MC_HV_MCD_L1TLB_PERF_TLB2_MISS_COUNT | 19 | number of tlb misses event |
| MC_HV_MCD_L1TLB_PERF_TLB2_DEFAULT_REQUEST_COUNT | 20 | number of default requests event |
| MC_HV_MCD_L1TLB_PERF_TLB2_PHYS_UVD_BYPASS_COUNT | 21 | number of physical and uvd bypass requests event |
| MC_HV_MCD_L1TLB_PERF_TLB2_PRETRANSLATED_COUNT | 22 | number of pre-translated requests event |
| MC_HV_MCD_L1TLB_PERF_TLB2_LATENCY_START | 23 | event start for the VM latency counter event |
| MC_HV_MCD_L1TLB_PERF_TLB2_LATENCY_END | 24 | event end for the VM latency counter |
| MC_HV_MCD_L1TLB_PERF_TLB3_REQUEST_COUNT | 25 | number of tlb requests event |
| MC_HV_MCD_L1TLB_PERF_TLB3_HIT_COUNT | 26 | number of tlb hits event |
| MC_HV_MCD_L1TLB_PERF_TLB3_MISS_COUNT | 27 | number of tlb misses event |
| MC_HV_MCD_L1TLB_PERF_TLB3_DEFAULT_REQUEST_COUNT | 28 | number of default requests event |
| MC_HV_MCD_L1TLB_PERF_TLB3_PHYS_UVD_BYPASS_COUNT | 29 | number of physical and uvd bypass requests event |
| MC_HV_MCD_L1TLB_PERF_TLB3_PRETRANSLATED_COUNT | 30 | number of pre-translated requests event |
| MC_HV_MCD_L1TLB_PERF_TLB3_LATENCY_START | 31 | event start for the VM latency counter event |
| MC_HV_MCD_L1TLB_PERF_TLB3_LATENCY_END | 32 | event end for the VM latency counter |
| Counter | Value | Description |
|---|---|---|
| MC_MCB_L1TLB_PERF_CYCLES | 0 | count cycles |
| MC_MCB_L1TLB_PERF_TLB0_REQUEST_COUNT | 1 | number of tlb requests event |
| MC_MCB_L1TLB_PERF_TLB0_HIT_COUNT | 2 | number of tlb hits event |
| MC_MCB_L1TLB_PERF_TLB0_MISS_COUNT | 3 | number of tlb misses event |
| MC_MCB_L1TLB_PERF_TLB0_DEFAULT_REQUEST_COUNT | 4 | number of default requests event |
| MC_MCB_L1TLB_PERF_TLB0_PHYS_UVD_BYPASS_COUNT | 5 | number of physical and uvd bypass requests event |
| MC_MCB_L1TLB_PERF_TLB0_PRETRANSLATED_COUNT | 6 | number of pre-translated requests event |
| MC_MCB_L1TLB_PERF_TLB0_LATENCY_START | 7 | event start for the VM latency counter event |
| MC_MCB_L1TLB_PERF_TLB0_LATENCY_END | 8 | event end for the VM latency counter |
| MC_MCB_L1TLB_PERF_TLB1_REQUEST_COUNT | 9 | number of tlb requests event |
| MC_MCB_L1TLB_PERF_TLB1_HIT_COUNT | 10 | number of tlb hits event |
| MC_MCB_L1TLB_PERF_TLB1_MISS_COUNT | 11 | number of tlb misses event |
| MC_MCB_L1TLB_PERF_TLB1_DEFAULT_REQUEST_COUNT | 12 | number of default requests event |
| MC_MCB_L1TLB_PERF_TLB1_PHYS_UVD_BYPASS_COUNT | 13 | number of physical and uvd bypass requests event |
| MC_MCB_L1TLB_PERF_TLB1_PRETRANSLATED_COUNT | 14 | number of pre-translated requests event |
| MC_MCB_L1TLB_PERF_TLB1_LATENCY_START | 15 | event start for the VM latency counter event |
| MC_MCB_L1TLB_PERF_TLB1_LATENCY_END | 16 | event end for the VM latency counter |
| MC_MCB_L1TLB_PERF_TLB2_REQUEST_COUNT | 17 | number of tlb requests event |
| MC_MCB_L1TLB_PERF_TLB2_HIT_COUNT | 18 | number of tlb hits event |
| MC_MCB_L1TLB_PERF_TLB2_MISS_COUNT | 19 | number of tlb misses event |
| MC_MCB_L1TLB_PERF_TLB2_DEFAULT_REQUEST_COUNT | 20 | number of default requests event |
| MC_MCB_L1TLB_PERF_TLB2_PHYS_UVD_BYPASS_COUNT | 21 | number of physical and uvd bypass requests event |
| MC_MCB_L1TLB_PERF_TLB2_PRETRANSLATED_COUNT | 22 | number of pre-translated requests event |
| MC_MCB_L1TLB_PERF_TLB2_LATENCY_START | 23 | event start for the VM latency counter event |
| MC_MCB_L1TLB_PERF_TLB2_LATENCY_END | 24 | event end for the VM latency counter |
| MC_MCB_L1TLB_PERF_TLB3_REQUEST_COUNT | 25 | number of tlb requests event |
| MC_MCB_L1TLB_PERF_TLB3_HIT_COUNT | 26 | number of tlb hits event |
| MC_MCB_L1TLB_PERF_TLB3_MISS_COUNT | 27 | number of tlb misses event |
| MC_MCB_L1TLB_PERF_TLB3_DEFAULT_REQUEST_COUNT | 28 | number of default requests event |
| MC_MCB_L1TLB_PERF_TLB3_PHYS_UVD_BYPASS_COUNT | 29 | number of physical and uvd bypass requests event |
| MC_MCB_L1TLB_PERF_TLB3_PRETRANSLATED_COUNT | 30 | number of pre-translated requests event |
| MC_MCB_L1TLB_PERF_TLB3_LATENCY_START | 31 | event start for the VM latency counter event |
| MC_MCB_L1TLB_PERF_TLB3_LATENCY_END | 32 | event end for the VM latency counter |
| Counter | Value | Description |
|---|---|---|
| MC_MCD_L1TLB_PERF_CYCLES | 0 | count cycles |
| MC_MCD_L1TLB_PERF_TLB0_REQUEST_COUNT | 1 | number of tlb requests event |
| MC_MCD_L1TLB_PERF_TLB0_HIT_COUNT | 2 | number of tlb hits event |
| MC_MCD_L1TLB_PERF_TLB0_MISS_COUNT | 3 | number of tlb misses event |
| MC_MCD_L1TLB_PERF_TLB0_DEFAULT_REQUEST_COUNT | 4 | number of default requests event |
| MC_MCD_L1TLB_PERF_TLB0_PHYS_UVD_BYPASS_COUNT | 5 | number of physical and uvd bypass requests event |
| MC_MCD_L1TLB_PERF_TLB0_PRETRANSLATED_COUNT | 6 | number of pre-translated requests event |
| MC_MCD_L1TLB_PERF_TLB0_LATENCY_START | 7 | event start for the VM latency counter event |
| MC_MCD_L1TLB_PERF_TLB0_LATENCY_END | 8 | event end for the VM latency counter |
| MC_MCD_L1TLB_PERF_TLB1_REQUEST_COUNT | 9 | number of tlb requests event |
| MC_MCD_L1TLB_PERF_TLB1_HIT_COUNT | 10 | number of tlb hits event |
| MC_MCD_L1TLB_PERF_TLB1_MISS_COUNT | 11 | number of tlb misses event |
| MC_MCD_L1TLB_PERF_TLB1_DEFAULT_REQUEST_COUNT | 12 | number of default requests event |
| MC_MCD_L1TLB_PERF_TLB1_PHYS_UVD_BYPASS_COUNT | 13 | number of physical and uvd bypass requests event |
| MC_MCD_L1TLB_PERF_TLB1_PRETRANSLATED_COUNT | 14 | number of pre-translated requests event |
| MC_MCD_L1TLB_PERF_TLB1_LATENCY_START | 15 | event start for the VM latency counter event |
| MC_MCD_L1TLB_PERF_TLB1_LATENCY_END | 16 | event end for the VM latency counter |
| MC_MCD_L1TLB_PERF_TLB2_REQUEST_COUNT | 17 | number of tlb requests event |
| MC_MCD_L1TLB_PERF_TLB2_HIT_COUNT | 18 | number of tlb hits event |
| MC_MCD_L1TLB_PERF_TLB2_MISS_COUNT | 19 | number of tlb misses event |
| MC_MCD_L1TLB_PERF_TLB2_DEFAULT_REQUEST_COUNT | 20 | number of default requests event |
| MC_MCD_L1TLB_PERF_TLB2_PHYS_UVD_BYPASS_COUNT | 21 | number of physical and uvd bypass requests event |
| MC_MCD_L1TLB_PERF_TLB2_PRETRANSLATED_COUNT | 22 | number of pre-translated requests event |
| MC_MCD_L1TLB_PERF_TLB2_LATENCY_START | 23 | event start for the VM latency counter event |
| MC_MCD_L1TLB_PERF_TLB2_LATENCY_END | 24 | event end for the VM latency counter |
| MC_MCD_L1TLB_PERF_TLB3_REQUEST_COUNT | 25 | number of tlb requests event |
| MC_MCD_L1TLB_PERF_TLB3_HIT_COUNT | 26 | number of tlb hits event |
| MC_MCD_L1TLB_PERF_TLB3_MISS_COUNT | 27 | number of tlb misses event |
| MC_MCD_L1TLB_PERF_TLB3_DEFAULT_REQUEST_COUNT | 28 | number of default requests event |
| MC_MCD_L1TLB_PERF_TLB3_PHYS_UVD_BYPASS_COUNT | 29 | number of physical and uvd bypass requests event |
| MC_MCD_L1TLB_PERF_TLB3_PRETRANSLATED_COUNT | 30 | number of pre-translated requests event |
| MC_MCD_L1TLB_PERF_TLB3_LATENCY_START | 31 | event start for the VM latency counter event |
| MC_MCD_L1TLB_PERF_TLB3_LATENCY_END | 32 | event end for the VM latency counter |
| Counter | Value | Description |
|---|---|---|
| MC_HV_L2TLB_PERF_CYCLES | 0 | count cycles |
| MC_HV_L2TLB_PERF_BANK0_PTE_REQUEST_COUNT | 1 | number of bank0 pte cache requests event |
| MC_HV_L2TLB_PERF_BANK0_PTE_HIT_COUNT | 2 | number of bank0 pte cache hits event |
| MC_HV_L2TLB_PERF_BANK0_PTE_MISS_COUNT | 3 | number of bank0 pte cache misses event |
| MC_HV_L2TLB_PERF_BANK1_PTE_REQUEST_COUNT | 4 | number of bank1 pte cache requests event |
| MC_HV_L2TLB_PERF_BANK1_PTE_HIT_COUNT | 5 | number of bank1 pte cache hits event |
| MC_HV_L2TLB_PERF_BANK1_PTE_MISS_COUNT | 6 | number of bank1 pte cache misses event |
| MC_HV_L2TLB_PERF_PDE0_REQUEST_COUNT | 7 | number of pde0 cache requests event |
| MC_HV_L2TLB_PERF_PDE0_HIT_COUNT | 8 | number of pde0 cache hits event |
| MC_HV_L2TLB_PERF_PDE0_MISS_COUNT | 9 | number of pde0 cache misses event |
| MC_HV_L2TLB_PERF_FAULT_COUNT | 10 | number of faults event |
| MC_HV_L2TLB_PERF_FAULT_COUNT_PER_CLIENT | 11 | number of faults by client-id event |
| MC_HV_L2TLB_PERF_BANK0_4K_PTE_HIT_COUNT | 12 | number of bank0 4k pte cache hits event |
| MC_HV_L2TLB_PERF_BANK0_4K_PTE_MISS_COUNT | 13 | number of bank0 4k pte cache misses event |
| MC_HV_L2TLB_PERF_BANK0_BIGK_PTE_HIT_COUNT | 14 | number of bank0 bigk pte cache hits event |
| MC_HV_L2TLB_PERF_BANK0_BIGK_PTE_MISS_COUNT | 15 | number of bank0 bigk pte cache misses event |
| MC_HV_L2TLB_PERF_BANK1_4K_PTE_HIT_COUNT | 16 | number of bank1 4k pte cache hits event |
| MC_HV_L2TLB_PERF_BANK1_4K_PTE_MISS_COUNT | 17 | number of bank1 4k pte cache misses event |
| MC_HV_L2TLB_PERF_BANK1_BIGK_PTE_HIT_COUNT | 18 | number of bank1 bigk pte cache hits event |
| MC_HV_L2TLB_PERF_BANK1_BIGK_PTE_MISS_COUNT | 19 | number of bank1 bigk pte cache misses event |
| MC_HV_L2TLB_PERF_INVALIDATION_COUNT | 20 | number of l2 cache invalidations (if VM_INVALIDATE_REQUEST used one for every bit set) |
| Counter | Value | Description |
|---|---|---|
| MC_L2TLB_PERF_CYCLES | 0 | count cycles |
| MC_L2TLB_PERF_BANK0_PTE_REQUEST_COUNT | 1 | number of bank0 pte cache requests event |
| MC_L2TLB_PERF_BANK0_PTE_HIT_COUNT | 2 | number of bank0 pte cache hits event |
| MC_L2TLB_PERF_BANK0_PTE_MISS_COUNT | 3 | number of bank0 pte cache misses event |
| MC_L2TLB_PERF_BANK1_PTE_REQUEST_COUNT | 4 | number of bank1 pte cache requests event |
| MC_L2TLB_PERF_BANK1_PTE_HIT_COUNT | 5 | number of bank1 pte cache hits event |
| MC_L2TLB_PERF_BANK1_PTE_MISS_COUNT | 6 | number of bank1 pte cache misses event |
| MC_L2TLB_PERF_PDE0_REQUEST_COUNT | 7 | number of pde0 cache requests event |
| MC_L2TLB_PERF_PDE0_HIT_COUNT | 8 | number of pde0 cache hits event |
| MC_L2TLB_PERF_PDE0_MISS_COUNT | 9 | number of pde0 cache misses event |
| MC_L2TLB_PERF_FAULT_COUNT | 10 | number of faults event |
| MC_L2TLB_PERF_FAULT_COUNT_PER_CLIENT | 11 | number of faults by client-id event |
| MC_L2TLB_PERF_BANK0_4K_PTE_HIT_COUNT | 12 | number of bank0 4k pte cache hits event |
| MC_L2TLB_PERF_BANK0_4K_PTE_MISS_COUNT | 13 | number of bank0 4k pte cache misses event |
| MC_L2TLB_PERF_BANK0_BIGK_PTE_HIT_COUNT | 14 | number of bank0 bigk pte cache hits event |
| MC_L2TLB_PERF_BANK0_BIGK_PTE_MISS_COUNT | 15 | number of bank0 bigk pte cache misses event |
| MC_L2TLB_PERF_BANK1_4K_PTE_HIT_COUNT | 16 | number of bank1 4k pte cache hits event |
| MC_L2TLB_PERF_BANK1_4K_PTE_MISS_COUNT | 17 | number of bank1 4k pte cache misses event |
| MC_L2TLB_PERF_BANK1_BIGK_PTE_HIT_COUNT | 18 | number of bank1 bigk pte cache hits event |
| MC_L2TLB_PERF_BANK1_BIGK_PTE_MISS_COUNT | 19 | number of bank1 bigk pte cache misses event |
| MC_L2TLB_PERF_INVALIDATION_COUNT | 20 | number of l2 cache invalidations (if VM_INVALIDATE_REQUEST used one for every bit set) |
| Counter | Value | Description |
|---|---|---|
| MC_ARB_PERF_CYCLES | 0 | clock ticks |
| MC_ARB_PERF_FREE_RUN | 1 | free run |
| MC_ARB_PERF_GN_ALL_GRONION_READ_CH0 | 2 | GN all gronion read ch0 |
| MC_ARB_PERF_GN_ALL_GRONION_WRITE_CH0 | 3 | GN all gronion write ch0 |
| MC_ARB_PERF_GN_READ_SEND_REQ_NONSTALLABLE_CH0 | 4 | GN read send req nonstallable ch0 |
| MC_ARB_PERF_GN_READ_SEND_REQ_STALLABLE_CH0 | 5 | GN read send req stallable ch0 |
| MC_ARB_PERF_GN_READ_SEND_REQ_ONION_CH0 | 6 | GN read send req onion ch0 |
| MC_ARB_PERF_GN_WRITE_SEND_REQ_NONSTALLABLE_CH0 | 7 | GN write send req nonstallable ch0 |
| MC_ARB_PERF_GN_WRITE_SEND_REQ_STALLABLE_CH0 | 8 | GN write send req stallable ch0 |
| MC_ARB_PERF_GN_WRITE_SEND_REQ_ONION_CH0 | 9 | GN write send req onion ch0 |
| MC_ARB_PERF_GN_WRITE_SEND_DATA_NONSTALLABLE_CH0 | 10 | GN write send data nonstallable ch0 |
| MC_ARB_PERF_GN_WRITE_SEND_DATA_STALLABLE_CH0 | 11 | GN write send data stallable ch0 |
| MC_ARB_PERF_GN_WRITE_SEND_DATA_ONION_CH0 | 12 | GN write send data onion ch0 |
| MC_ARB_PERF_GN_32B_NONSTALLABLE_WRITE_CH0 | 13 | GN 32B nonstallable write ch0 |
| MC_ARB_PERF_GN_32B_STALLABLE_WRITE_CH0 | 14 | GN 32B stallable write ch0 |
| MC_ARB_PERF_GN_32B_ONION_WRITE_CH0 | 15 | GN 32B onion write ch0 |
| MC_ARB_PERF_GN_64B_NONSTALLABLE_WRITE_CH0 | 16 | GN 64B nonstallable write ch0 |
| MC_ARB_PERF_GN_64B_STALLABLE_WRITE_CH0 | 17 | GN 64B stallable write ch0 |
| MC_ARB_PERF_GN_64B_ONION_WRITE_CH0 | 18 | GN 64B onion write ch0 |
| MC_ARB_PERF_GN_READ_EOB_CH0 | 19 | GN read eob ch0 |
| MC_ARB_PERF_GN_WRITE_EOB_CH0 | 20 | GN write eob ch0 |
| MC_ARB_PERF_GN_WRITE_COMMIT_CH0 | 21 | GN Write commit ch0 |
| MC_ARB_PERF_GN_ALL_READ_RESPONSE_DATA_BEATS_CH0 | 22 | GN all Read response data beats ch0 |
| MC_ARB_PERF_GN_LAST_READ_RESPONSE_DATA_BEAT_CH0 | 23 | GN last read response data beat ch0 |
| MC_ARB_PERF_PRIORITY_URGENT_CH0 | 24 | priority urgent ch0 |
| MC_ARB_PERF_TOKEN_URGENT_CH0 | 25 | token urgent ch0 |
| MC_ARB_PERF_PRIORITY_PROMOTE_CH0 | 26 | priority promote ch0 |
| MC_ARB_PERF_NONSTALLABLE_READ_PRI_2_CH0 | 32 | nonstallable read pri=2 ch0 |
| MC_ARB_PERF_NONSTALLABLE_READ_PRI_1_CH0 | 33 | nonstallable read pri=1 ch0 |
| MC_ARB_PERF_NONSTALLABLE_READ_PRI_0_CH0 | 34 | nonstallable read pri=0 ch0 |
| MC_ARB_PERF_STALLABLE_READ_PRI_2_CH0 | 35 | stallable read pri=2 ch0 |
| MC_ARB_PERF_STALLABLE_READ_PRI_1_CH0 | 36 | stallable read pri=1 ch0 |
| MC_ARB_PERF_STALLABLE_READ_PRI_0_CH0 | 37 | stallable read pri=0 ch0 |
| MC_ARB_PERF_ONION_READ_PRI_2_CH0 | 38 | onion read pri=2 ch0 |
| MC_ARB_PERF_ONION_READ_PRI_1_CH0 | 39 | onion read pri=1 ch0 |
| MC_ARB_PERF_ONION_READ_PRI_0_CH0 | 40 | onion read pri=0 ch0 |
| MC_ARB_PERF_NONSTALLABLE_WRITE_PRI_2_CH0 | 42 | nonstallable write pri=2 ch0 |
| MC_ARB_PERF_NONSTALLABLE_WRITE_PRI_1_CH0 | 43 | nonstallable write pri=1 ch0 |
| MC_ARB_PERF_NONSTALLABLE_WRITE_PRI_0_CH0 | 44 | nonstallable write pri=0 ch0 |
| MC_ARB_PERF_STALLABLE_WRITE_PRI_2_CH0 | 45 | stallable write pri=2 ch0 |
| MC_ARB_PERF_STALLABLE_WRITE_PRI_1_CH0 | 46 | stallable write pri=1 ch0 |
| MC_ARB_PERF_STALLABLE_WRITE_PRI_0_CH0 | 47 | stallable write pri=0 ch0 |
| MC_ARB_PERF_ONION_WRITE_PRI_2_CH0 | 48 | onion write pri=2 ch0 |
| MC_ARB_PERF_ONION_WRITE_PRI_1_CH0 | 49 | onion write pri=1 ch0 |
| MC_ARB_PERF_ONION_WRITE_PRI_0_CH0 | 50 | onion write pri=0 ch0 |
| MC_ARB_PERF_0_LE_GN_READ_TAG_CH0_LE_15 | 52 | 0 <= GN read tag ch0 <= 15 |
| MC_ARB_PERF_16_LE_GN_READ_TAG_CH0_LE_31 | 53 | 16 <= GN read tag ch0 <= 31 |
| MC_ARB_PERF_32_LE_GN_READ_TAG_CH0_LE_47 | 54 | 32 <= GN read tag ch0 <= 47 |
| MC_ARB_PERF_48_LE_GN_READ_TAG_CH0_LE_63 | 55 | 48 <= GN read tag ch0 <= 63 |
| MC_ARB_PERF_64_LE_GN_READ_TAG_CH0_LE_79 | 56 | 64 <= GN read tag ch0 <= 79 |
| MC_ARB_PERF_80_LE_GN_READ_TAG_CH0_LE_95 | 57 | 80 <= GN read tag ch0 <= 95 |
| MC_ARB_PERF_96_LE_GN_READ_TAG_CH0_LE_111 | 58 | 96 <= GN read tag ch0 <= 111 |
| MC_ARB_PERF_112_LE_GN_READ_TAG_CH0_LE_127 | 59 | 112 <= GN read tag ch0 <= 127 |
| MC_ARB_PERF_0_LE_GN_WRITE_TAG_CH0_LE_15 | 62 | 0 <= GN write tag ch0 <= 15 |
| MC_ARB_PERF_16_LE_GN_WRITE_TAG_CH0_LE_31 | 63 | 16 <= GN write tag ch0 <= 31 |
| MC_ARB_PERF_32_LE_GN_WRITE_TAG_CH0_LE_47 | 64 | 32 <= GN write tag ch0 <= 47 |
| MC_ARB_PERF_48_LE_GN_WRITE_TAG_CH0_LE_63 | 65 | 48 <= GN write tag ch0 <= 63 |
| MC_ARB_PERF_64_LE_GN_WRITE_TAG_CH0_LE_79 | 66 | 64 <= GN write tag ch0 <= 79 |
| MC_ARB_PERF_80_LE_GN_WRITE_TAG_CH0_LE_95 | 67 | 80 <= GN write tag ch0 <= 95 |
| MC_ARB_PERF_96_LE_GN_WRITE_TAG_CH0_LE_111 | 68 | 96 <= GN write tag ch0 <= 111 |
| MC_ARB_PERF_112_LE_GN_WRITE_TAG_CH0_LE_127 | 69 | 112 <= GN write tag ch0 <= 127 |
| MC_ARB_PERF_GN_ALL_GRONION_READ_CH1 | 72 | GN all gronion read ch1 |
| MC_ARB_PERF_GN_ALL_GRONION_WRITE_CH1 | 73 | GN all gronion write ch1 |
| MC_ARB_PERF_GN_READ_SEND_REQ_NONSTALLABLE_CH1 | 74 | GN read send req nonstallable ch1 |
| MC_ARB_PERF_GN_READ_SEND_REQ_STALLABLE_CH1 | 75 | GN read send req stallable ch1 |
| MC_ARB_PERF_GN_READ_SEND_REQ_ONION_CH1 | 76 | GN read send req onion ch1 |
| MC_ARB_PERF_GN_WRITE_SEND_REQ_NONSTALLABLE_CH1 | 77 | GN write send req nonstallable ch1 |
| MC_ARB_PERF_GN_WRITE_SEND_REQ_STALLABLE_CH1 | 78 | GN write send req stallable ch1 |
| MC_ARB_PERF_GN_WRITE_SEND_REQ_ONION_CH1 | 79 | GN write send req onion ch1 |
| MC_ARB_PERF_GN_WRITE_SEND_DATA_NONSTALLABLE_CH1 | 80 | GN write send data nonstallable ch1 |
| MC_ARB_PERF_GN_WRITE_SEND_DATA_STALLABLE_CH1 | 81 | GN write send data stallable ch1 |
| MC_ARB_PERF_GN_WRITE_SEND_DATA_ONION_CH1 | 82 | GN write send data onion ch1 |
| MC_ARB_PERF_GN_32B_NONSTALLABLE_WRITE_CH1 | 83 | GN 32B nonstallable write ch1 |
| MC_ARB_PERF_GN_32B_STALLABLE_WRITE_CH1 | 84 | GN 32B stallable write ch1 |
| MC_ARB_PERF_GN_32B_ONION_WRITE_CH1 | 85 | GN 32B onion write ch1 |
| MC_ARB_PERF_GN_64B_NONSTALLABLE_WRITE_CH1 | 86 | GN 64B nonstallable write ch1 |
| MC_ARB_PERF_GN_64B_STALLABLE_WRITE_CH1 | 87 | GN 64B stallable write ch1 |
| MC_ARB_PERF_GN_64B_ONION_WRITE_CH1 | 88 | GN 64B onion write ch1 |
| MC_ARB_PERF_GN_READ_EOB_CH1 | 89 | GN read eob ch1 |
| MC_ARB_PERF_GN_WRITE_EOB_CH1 | 90 | GN write eob ch1 |
| MC_ARB_PERF_GN_WRITE_COMMIT_CH1 | 91 | GN Write commit ch1 |
| MC_ARB_PERF_GN_ALL_READ_RESPONSE_DATA_BEATS_CH1 | 92 | GN all Read response data beats ch1 |
| MC_ARB_PERF_GN_LAST_READ_RESPONSE_DATA_BEAT_CH1 | 93 | GN last read response data beat ch1 |
| MC_ARB_PERF_PRIORITY_URGENT_CH1 | 94 | priority urgent ch1 |
| MC_ARB_PERF_TOKEN_URGENT_CH1 | 95 | token urgent ch1 |
| MC_ARB_PERF_PRIORITY_PROMOTE_CH1 | 96 | priority promote ch1 |
| MC_ARB_PERF_NONSTALLABLE_READ_PRI_2_CH1 | 102 | nonstallable read pri=2 ch1 |
| MC_ARB_PERF_NONSTALLABLE_READ_PRI_1_CH1 | 103 | nonstallable read pri=1 ch1 |
| MC_ARB_PERF_NONSTALLABLE_READ_PRI_0_CH1 | 104 | nonstallable read pri=0 ch1 |
| MC_ARB_PERF_STALLABLE_READ_PRI_2_CH1 | 105 | stallable read pri=2 ch1 |
| MC_ARB_PERF_STALLABLE_READ_PRI_1_CH1 | 106 | stallable read pri=1 ch1 |
| MC_ARB_PERF_STALLABLE_READ_PRI_0_CH1 | 107 | stallable read pri=0 ch1 |
| MC_ARB_PERF_ONION_READ_PRI_2_CH1 | 108 | onion read pri=2 ch1 |
| MC_ARB_PERF_ONION_READ_PRI_1_CH1 | 109 | onion read pri=1 ch1 |
| MC_ARB_PERF_ONION_READ_PRI_0_CH1 | 110 | onion read pri=0 ch1 |
| MC_ARB_PERF_NONSTALLABLE_WRITE_PRI_2_CH1 | 112 | nonstallable write pri=2 ch1 |
| MC_ARB_PERF_NONSTALLABLE_WRITE_PRI_1_CH1 | 113 | nonstallable write pri=1 ch1 |
| MC_ARB_PERF_NONSTALLABLE_WRITE_PRI_0_CH1 | 114 | nonstallable write pri=0 ch1 |
| MC_ARB_PERF_STALLABLE_WRITE_PRI_2_CH1 | 115 | stallable write pri=2 ch1 |
| MC_ARB_PERF_STALLABLE_WRITE_PRI_1_CH1 | 116 | stallable write pri=1 ch1 |
| MC_ARB_PERF_STALLABLE_WRITE_PRI_0_CH1 | 117 | stallable write pri=0 ch1 |
| MC_ARB_PERF_ONION_WRITE_PRI_2_CH1 | 118 | onion write pri=2 ch1 |
| MC_ARB_PERF_ONION_WRITE_PRI_1_CH1 | 119 | onion write pri=1 ch1 |
| MC_ARB_PERF_ONION_WRITE_PRI_0_CH1 | 120 | onion write pri=0 ch1 |
| MC_ARB_PERF_0_LE_GN_READ_TAG_CH1_LE_15 | 122 | 0 <= GN read tag ch1 <= 15 |
| MC_ARB_PERF_16_LE_GN_READ_TAG_CH1_LE_31 | 123 | 16 <= GN read tag ch1 <= 31 |
| MC_ARB_PERF_32_LE_GN_READ_TAG_CH1_LE_47 | 124 | 32 <= GN read tag ch1 <= 47 |
| MC_ARB_PERF_48_LE_GN_READ_TAG_CH1_LE_63 | 125 | 48 <= GN read tag ch1 <= 63 |
| MC_ARB_PERF_64_LE_GN_READ_TAG_CH1_LE_79 | 126 | 64 <= GN read tag ch1 <= 79 |
| MC_ARB_PERF_80_LE_GN_READ_TAG_CH1_LE_95 | 127 | 80 <= GN read tag ch1 <= 95 |
| MC_ARB_PERF_96_LE_GN_READ_TAG_CH1_LE_111 | 128 | 96 <= GN read tag ch1 <= 111 |
| MC_ARB_PERF_112_LE_GN_READ_TAG_CH1_LE_127 | 129 | 112 <= GN read tag ch1 <= 127 |
| MC_ARB_PERF_0_LE_GN_WRITE_TAG_CH1_LE_15 | 132 | 0 <= GN write tag ch1 <= 15 |
| MC_ARB_PERF_16_LE_GN_WRITE_TAG_CH1_LE_31 | 133 | 16 <= GN write tag ch1 <= 31 |
| MC_ARB_PERF_32_LE_GN_WRITE_TAG_CH1_LE_47 | 134 | 32 <= GN write tag ch1 <= 47 |
| MC_ARB_PERF_48_LE_GN_WRITE_TAG_CH1_LE_63 | 135 | 48 <= GN write tag ch1 <= 63 |
| MC_ARB_PERF_64_LE_GN_WRITE_TAG_CH1_LE_79 | 136 | 64 <= GN write tag ch1 <= 79 |
| MC_ARB_PERF_80_LE_GN_WRITE_TAG_CH1_LE_95 | 137 | 80 <= GN write tag ch1 <= 95 |
| MC_ARB_PERF_96_LE_GN_WRITE_TAG_CH1_LE_111 | 138 | 96 <= GN write tag ch1 <= 111 |
| MC_ARB_PERF_112_LE_GN_WRITE_TAG_CH1_LE_127 | 139 | 112 <= GN write tag ch1 <= 127 |
| MC_ARB_PERF_ESRAM_READ_CH0 | 142 | ESRAM read ch0 |
| MC_ARB_PERF_ESRAM_WRITE_SEND_CH0 | 143 | ESRAM write send ch0 |
| MC_ARB_PERF_ESRAM_WRITE_SEND_DATA_CH0 | 144 | ESRAM write send data ch0 |
| MC_ARB_PERF_ESRAM_32B_READ_CH0 | 145 | ESRAM 32B read ch0 |
| MC_ARB_PERF_ESRAM_64B_READ_CH0 | 146 | ESRAM 64B read ch0 |
| MC_ARB_PERF_ESRAM_128B_READ_CH0 | 147 | ESRAM 128B read ch0 |
| MC_ARB_PERF_ESRAM_256B_READ_CH0 | 148 | ESRAM 256B read ch0 |
| MC_ARB_PERF_8_LE_ESRAM_READ_TAG_CH0_LE_15 | 149 | 8 <= ESRAM read tag ch0 <= 15 |
| MC_ARB_PERF_16_LE_ESRAM_READ_TAG_CH0_LE_23 | 150 | 16 <= ESRAM read tag ch0 <= 23 |
| MC_ARB_PERF_24_LE_ESRAM_READ_TAG_CH0_LE_31 | 151 | 24 <= ESRAM read tag ch0 <= 31 |
| MC_ARB_PERF_32_LE_ESRAM_READ_TAG_CH0_LE_39 | 152 | 32 <= ESRAM read tag ch0 <= 39 |
| MC_ARB_PERF_40_LE_ESRAM_READ_TAG_CH0_LE_63 | 153 | 40 <= ESRAM read tag ch0 <= 63 |
| MC_ARB_PERF_8_LE_ESRAM_WRITE_TAG_CH0_LE_15 | 154 | 8 <= ESRAM write tag ch0 <= 15 |
| MC_ARB_PERF_16_LE_ESRAM_WRITE_TAG_CH0_LE_23 | 155 | 16 <= ESRAM write tag ch0 <= 23 |
| MC_ARB_PERF_24_LE_ESRAM_WRITE_TAG_CH0_LE_31 | 156 | 24 <= ESRAM write tag ch0 <= 31 |
| MC_ARB_PERF_32_LE_ESRAM_WRITE_TAG_CH0_LE_39 | 157 | 32 <= ESRAM write tag ch0 <= 39 |
| MC_ARB_PERF_40_LE_ESRAM_WRITE_TAG_CH0_LE_63 | 158 | 40 <= ESRAM write tag ch0 <= 63 |
| MC_ARB_PERF_ESRAM_WRITE_RETURN_CH0 | 159 | ESRAM write return ch0 |
| MC_ARB_PERF_ESRAM_READ_RESPONSE_DATA_BEATS_CH0 | 160 | ESRAM read response data beats ch0 |
| MC_ARB_PERF_ESRAM_LAST_READ_DATA_BEAT_CH0 | 161 | ESRAM last read data beat ch0 |
| MC_ARB_PERF_ESRAM_READ_CH1 | 162 | ESRAM read ch1 |
| MC_ARB_PERF_ESRAM_WRITE_SEND_CH1 | 163 | ESRAM write send ch1 |
| MC_ARB_PERF_ESRAM_WRITE_SEND_DATA_CH1 | 164 | ESRAM write send data ch1 |
| MC_ARB_PERF_ESRAM_32B_READ_CH1 | 165 | ESRAM 32B read ch1 |
| MC_ARB_PERF_ESRAM_64B_READ_CH1 | 166 | ESRAM 64B read ch1 |
| MC_ARB_PERF_ESRAM_128B_READ_CH1 | 167 | ESRAM 128B read ch1 |
| MC_ARB_PERF_ESRAM_256B_READ_CH1 | 168 | ESRAM 256B read ch1 |
| MC_ARB_PERF_8_LE_ESRAM_READ_TAG_CH1_LE_15 | 169 | 8 <= ESRAM read tag ch1 <= 15 |
| MC_ARB_PERF_16_LE_ESRAM_READ_TAG_CH1_LE_23 | 170 | 16 <= ESRAM read tag ch1 <= 23 |
| MC_ARB_PERF_24_LE_ESRAM_READ_TAG_CH1_LE_31 | 171 | 24 <= ESRAM read tag ch1 <= 31 |
| MC_ARB_PERF_32_LE_ESRAM_READ_TAG_CH1_LE_39 | 172 | 32 <= ESRAM read tag ch1 <= 39 |
| MC_ARB_PERF_40_LE_ESRAM_READ_TAG_CH1_LE_63 | 173 | 40 <= ESRAM read tag ch1 <= 63 |
| MC_ARB_PERF_8_LE_ESRAM_WRITE_TAG_CH1_LE_15 | 174 | 8 <= ESRAM write tag ch1 <= 15 |
| MC_ARB_PERF_16_LE_ESRAM_WRITE_TAG_CH1_LE_23 | 175 | 16 <= ESRAM write tag ch1 <= 23 |
| MC_ARB_PERF_24_LE_ESRAM_WRITE_TAG_CH1_LE_31 | 176 | 24 <= ESRAM write tag ch1 <= 31 |
| MC_ARB_PERF_32_LE_ESRAM_WRITE_TAG_CH1_LE_39 | 177 | 32 <= ESRAM write tag ch1 <= 39 |
| MC_ARB_PERF_40_LE_ESRAM_WRITE_TAG_CH1_LE_63 | 178 | 40 <= ESRAM write tag ch1 <= 63 |
| MC_ARB_PERF_ESRAM_WRITE_RETURN_CH1 | 179 | ESRAM write return ch1 |
| MC_ARB_PERF_ESRAM_READ_RESPONSE_DATA_BEATS_CH1 | 180 | ESRAM read response data beats ch1 |
| MC_ARB_PERF_ESRAM_LAST_READ_DATA_BEAT_CH1 | 181 | ESRAM last read data beat ch1 |
| MC_ARB_PERF_HARSH_PRIORITY_VALID_IN_CH1_WRITE | 182 | harsh priority valid in ch1 write |
| MC_ARB_PERF_HARSH_PRIORITY_VALID_IN_CH1_READ | 183 | harsh priority valid in ch1 read |
| MC_ARB_PERF_HARSH_PRIORITY_VALID_IN_CH0_WRITE | 184 | harsh priority valid in ch0 write |
| MC_ARB_PERF_HARSH_PRIORITY_VALID_IN_CH0_READ | 185 | harsh priority valid in ch0 read |
| MC_ARB_PERF_MAX_LATENCY_START_EVENT_CH0 | 186 | max latency start event ch0 |
| MC_ARB_PERF_MAX_LATENCY_END_EVENT_CH0 | 187 | max latency end event ch0 |
| MC_ARB_PERF_MAX_LATENCY_START_EVENT_CH1 | 188 | max latency start event ch1 |
| MC_ARB_PERF_MAX_LATENCY_END_EVENT_CH1 | 189 | max latency end event ch1 |
| MC_ARB_PERF_DISP1_URGENTWVLD_RISE | 190 | disp1_urgentwvld_rise |
| MC_ARB_PERF_DISP0_URGENTWVLD_RISE | 191 | disp0_urgentwvld_rise |
| MC_ARB_PERF_DISP1_URGENTWVLD_D1 | 192 | disp1_urgentwvld_d1 |
| MC_ARB_PERF_DISP0_URGENTWVLD_D1 | 193 | disp0_urgentwvld_d1 |
| Counter | Value | Description |
|---|---|---|
| MC_CITF_PERF_CYCLES | 0 | clock cycles |
| MC_CITF_PERF_CB0_READ_REQUEST_RDNFO_ASK | 1 | CB0 read request rdnfo_ask |
| MC_CITF_PERF_CB0_READ_REQUEST_RDNFO_GO | 2 | CB0 read request rdnfo_go |
| MC_CITF_PERF_CB1_READ_REQUEST_RDNFO_ASK | 3 | CB1 read request rdnfo_ask |
| MC_CITF_PERF_CB1_READ_REQEUST_RDNFO_GO | 4 | CB1 read request rdnfo_go |
| MC_CITF_PERF_DB0_READ_REQUEST_RDNFO_ASK | 5 | DB0 read request rdnfo_ask |
| MC_CITF_PERF_DB0_READ_REQUEST_RDNFO_GO | 6 | DB0 read request rdnfo_go |
| MC_CITF_PERF_DB1_READ_REQUEST_RDNFO_ASK | 7 | DB1 read request rdnfo_ask |
| MC_CITF_PERF_DB1_READ_REQUEST_RDNFO_GO | 8 | DB1 read request rdnfo_go |
| MC_CITF_PERF_ESRAM_DDR_CH0_READ_RETURN_SEND | 9 | ESRAM ch0 read return send && DDR ch0 read return send |
| MC_CITF_PERF_ESRAM_DDR_CH1_READ_RETURN_SEND | 10 | ESRAM ch1 read return send && DDR ch1 read return send |
| MC_CITF_PERF_CB0_WRITE_REQUEST_WRNFO_ASK | 13 | CB0 write request wrnfo_ask |
| MC_CITF_PERF_CB0_WRITE_REQUEST_WRNFO_GO | 14 | CB0 write request wrnfo_go |
| MC_CITF_PERF_CB1_WRITE_REQUEST_WRNFO_ASK | 15 | CB1 write request wrnfo_ask |
| MC_CITF_PERF_CB1_WRITE_REQUEST_WRNFO_GO | 16 | CB1 write request wrnfo_go |
| MC_CITF_PERF_DB0_WRITE_REQUEST_WRNFO_ASK | 17 | DB0 write request wrnfo_ask |
| MC_CITF_PERF_DB0_WRITE_REQUEST_WRNFO_GO | 18 | DB0 write request wrnfo_go |
| MC_CITF_PERF_DB1_WRITE_REQUEST_WRNFO_ASK | 19 | DB1 write request wrnfo_ask |
| MC_CITF_PERF_DB1_WRITE_REQUEST_WRNFO_GO | 20 | DB1 write request wrnfo_go |
| MC_CITF_PERF_TC0_READ_REQUEST_SEND | 25 | TC0 read request send |
| MC_CITF_PERF_TC0_READ_REQUEST_FREE | 26 | TC0 read request free |
| MC_CITF_PERF_TC1_READ_REQUEST_SEND | 27 | TC1 read request send |
| MC_CITF_PERF_TC1_READ_REQUEST_FREE | 28 | TC1 read request free |
| MC_CITF_PERF_TC0_WRITE_REQUEST_SEND | 31 | TC0 write request send |
| MC_CITF_PERF_TC0_WRITE_REQUEST_FREE | 32 | TC0 write request free |
| MC_CITF_PERF_TC1_WRITE_REQUEST_SEND | 33 | TC1 write request send |
| MC_CITF_PERF_TC1_WRITE_REQUEST_FREE | 34 | TC1 write request free |
| MC_CITF_PERF_CB0_READ_RETURN_EOP | 37 | CB0 read return EOP |
| MC_CITF_PERF_CB1_READ_RETURN_EOP | 38 | CB1 read return EOP |
| MC_CITF_PERF_DB0_READ_RETURN_EOP | 39 | DB0 read return EOP |
| MC_CITF_PERF_DB1_READ_RETURN_EOP | 40 | DB1 read return EOP |
| MC_CITF_PERF_CB0_WRITE_RETURN_VLD | 43 | CB0 write return vld |
| MC_CITF_PERF_CB1_WRITE_RETURN_VLD | 44 | CB1 write return vld |
| MC_CITF_PERF_DB0_WRITE_RETURN_VLD | 45 | DB0 write return vld |
| MC_CITF_PERF_DB1_WRITE_RETURN_VLD | 46 | DB1 write return vld |
| MC_CITF_PERF_TC0_READ_RETURN_EOP | 49 | TC0 read return EOP - RTL bug (1`b0) |
| MC_CITF_PERF_TC1_READ_RETURN_EOP | 50 | TC1 read return EOP - RTL bug (1`b0) |
| MC_CITF_PERF_TC0_WRITE_RETURN_VLD | 51 | TC0 write return vld |
| MC_CITF_PERF_TC1_WRITE_RETURN_VLD | 52 | TC1 write return vld |
| MC_CITF_PERF_MCB_MCC_READ_REQUEST_MATCHING_CID | 55 | mcb_mcc read request matching CID specified in MC_CITF_PERF_CNTL2.CID field |
| MC_CITF_PERF_MCC_MCB_READ_RETURN_EOP_MATCHING_CID | 56 | mcc_mcb read return EOP matching CID specified in MC_CITF_PERF_CNTL2.CID field |
| MC_CITF_PERF_MCB_MCC_WRITE_REQUEST_MATCHING_CID | 57 | mcb_mcc write request matching CID specified in MC_CITF_PERF_CNTL2.CID field |
| MC_CITF_PERF_MCC_MCB_WRITE_RETURN_MATCHING_CID | 58 | mcc_mcb write return matching CID specified in MC_CITF_PERF_CNTL2.CID field |
| Counter | Value | Description |
|---|---|---|
| MC_HUB_PERF_CYCLES | 0 | clock cycles |
| MC_HUB_PERF_CPC_WRITE_RETURN | 1 | cpc write return |
| MC_HUB_PERF_CPF_WRITE_RETURN | 2 | cpf write return |
| MC_HUB_PERF_CPG_WRITE_RETURN | 3 | cpg write return |
| MC_HUB_PERF_HDP_WRITE_RETURN | 4 | hdp write return |
| MC_HUB_PERF_IH_WRITE_RETURN | 5 | ih write return |
| MC_HUB_PERF_MCIF_WRITE_RETURN | 6 | mcif write return |
| MC_HUB_PERF_SAMSCP_WRITE_RETURN | 8 | samscp write return |
| MC_HUB_PERF_SDMA1_WRITE_RETURN | 9 | sdma1 write return |
| MC_HUB_PERF_SDMA3_WRITE_RETURN | 10 | sdma3 write return |
| MC_HUB_PERF_SEM_WRITE_RETURN | 11 | sem write return |
| MC_HUB_PERF_SH0_WRITE_RETURN | 12 | sh0 write return |
| MC_HUB_PERF_SMU_WRITE_RETURN | 13 | smu write return |
| MC_HUB_PERF_VIN0_WRITE_RETURN | 14 | vin0 write return |
| MC_HUB_PERF_RLC_WRITE_RETURN | 15 | rlc write return |
| MC_HUB_PERF_SDMA0_WRITE_RETURN | 16 | sdma0 write return |
| MC_HUB_PERF_SDMA2_WRITE_RETURN | 17 | sdma2 write return |
| MC_HUB_PERF_SH1_WRITE_RETURN | 18 | sh1 write return |
| MC_HUB_PERF_UMC_WRITE_RETURN | 20 | umc write return |
| MC_HUB_PERF_UVD__UVD_EXT0__UVD_EXT1_WRITE_RETURN | 21 | uvd, uvd_ext0, uvd_ext1 write return |
| MC_HUB_PERF_VCE_WRITE_RETURN | 22 | vce write return |
| MC_HUB_PERF_VCEU_WRITE_RETURN | 23 | vceu write return |
| MC_HUB_PERF_XDP_WRITE_RETURN | 24 | xdp write return |
| MC_HUB_PERF_MCC_MCB0_WRITE_RETURN | 25 | mcc_mcb0 write return |
| MC_HUB_PERF_MCB_MCC1_WRITE_RETURN | 26 | mcb_mcc1 write return |
| MC_HUB_PERF_CPC_WRITE_ASK_URG_STALL | 81 | cpc write ask+urg+stall |
| MC_HUB_PERF_CPF_WRITE_ASK_URG_STALL | 82 | cpf write ask+urg+stall |
| MC_HUB_PERF_CPG_WRITE_ASK_URG_STALL | 83 | cpg write ask+urg+stall |
| MC_HUB_PERF_HDP_WRITE_ASK_URG_STALL | 84 | hdp write ask+urg+stall |
| MC_HUB_PERF_IH_WRITE_ASK_URG_STALL | 85 | ih write ask+urg+stall |
| MC_HUB_PERF_MCIF_WRITE_ASK_URG_STALL | 86 | mcif write ask+urg+stall |
| MC_HUB_PERF_SAMSCP_WRITE_ASK_URG_STALL | 88 | samscp write ask+urg+stall |
| MC_HUB_PERF_SDMA1_WRITE_ASK_URG_STALL | 89 | sdma1 write ask+urg+stall |
| MC_HUB_PERF_SDMA3_WRITE_ASK_URG_STALL | 90 | sdma3 write ask+urg+stall |
| MC_HUB_PERF_SEM_WRITE_ASK_URG_STALL | 91 | sem write ask+urg+stall |
| MC_HUB_PERF_SH0_WRITE_ASK_URG_STALL | 92 | sh0 write ask+urg+stall |
| MC_HUB_PERF_SMU_WRITE_ASK_URG_STALL | 93 | smu write ask+urg+stall |
| MC_HUB_PERF_VIN0_WRITE_ASK_URG_STALL | 94 | vin0 write ask+urg+stall |
| MC_HUB_PERF_RLC_WRITE_ASK_URG_STALL | 95 | rlc write ask+urg+stall |
| MC_HUB_PERF_SDMA0_WRITE_ASK_URG_STALL | 96 | sdma0 write ask+urg+stall |
| MC_HUB_PERF_SDMA2_WRITE_ASK_URG_STALL | 97 | sdma2 write ask+urg+stall |
| MC_HUB_PERF_SH1_WRITE_ASK_URG_STALL | 98 | sh1 write ask+urg+stall |
| MC_HUB_PERF_UMC_WRITE_ASK_URG_STALL | 100 | umc write ask+urg+stall |
| MC_HUB_PERF_UVD_WRITE_ASK_URG_STALL | 101 | uvd write ask+urg+stall |
| MC_HUB_PERF_VCE_WRITE_ASK_URG_STALL | 102 | vce write ask+urg+stall |
| MC_HUB_PERF_VCEU_WRITE_ASK_URG_STALL | 103 | vceu write ask+urg+stall |
| MC_HUB_PERF_XDP_WRITE_ASK_URG_STALL | 104 | xdp write ask+urg+stall |
| MC_HUB_PERF_GBL0_WRITE_REQUEST_SEND_CLIENT_ID | 105 | gbl0 write request send+client ID |
| MC_HUB_PERF_GBL0_WRITE_REQUEST_STOR_FULL | 106 | gbl0 write request stor full |
| MC_HUB_PERF_GBL0_WRITE_REQUEST_BYPASS_STOR_FULL | 107 | gbl0 write request bypass stor full |
| MC_HUB_PERF_GBL0_WRITE_REQUEST_STOR_EMPTY | 108 | gbl0 write request stor empty |
| MC_HUB_PERF_GBL0_WRITE_REQUEST_STOR_CREDIT_COUNT | 109 | gbl0 write request stor credit count |
| MC_HUB_PERF_GBL0_WRITE_REQUEST_BYPASS_CREDIT_COUNT | 110 | gbl0 write request bypass credit count |
| MC_HUB_PERF_GBL0_WRITE_REQUEST_VM_CREDIT_COUNT | 111 | gbl0 write request vm credit count |
| MC_HUB_PERF_GBL1_WRITE_REQUEST_SEND_CID | 115 | gbl1 write request send+cid |
| MC_HUB_PERF_GBL1_WRITE_REQUEST_STOR_FULL | 116 | gbl1 write request stor full |
| MC_HUB_PERF_GBL1_WRITE_REQUEST_STOR_EMPTY | 118 | gbl1 write request stor empty |
| MC_HUB_PERF_GBL1_WRITE_REQUEST_STOR_CREDIT_COUNT | 119 | gbl1 write request stor credit count |
| MC_HUB_PERF_GBL1_WRITE_REQUEST_VM_CREDIT_COUNT | 121 | gbl1 write request vm credit count |
| MC_HUB_PERF_MCB_MCC0_WRITE_REQUEST_SEND_IS_WRITE_NACK | 125 | mcb_mcc0 write request send+is_write+nack |
| MC_HUB_PERF_MCB_MCC1_WRITE_REQUEST_SEND_IS_WRITE_NACK | 126 | mcb_mcc1 write request send+is+write+nack |
| MC_HUB_PERF_RPB_READ_RETURN_ASK_PRI_GO_GO_PRI_DEST_ID_STALL_EOP = 127, // rpb read return ask+pri+go+go_pri+dest_id+stall+eop | 128 | rpb read return vld+eop |
| MC_HUB_PERF_CPC_READ_RETURN | 129 | cpc read return |
| MC_HUB_PERF_CPF_READ_RETURN | 130 | cpf read return |
| MC_HUB_PERF_CPG_READ_RETURN | 131 | cpg read return |
| MC_HUB_PERF_HDP_READ_RETURN | 132 | hdp read return |
| MC_HUB_PERF_IA_READ_RETURN | 133 | ia read return |
| MC_HUB_PERF_RLC_READ_RETURN | 134 | rlc read return |
| MC_HUB_PERF_SAMSCP_READ_RETURN | 136 | samscp read return |
| MC_HUB_PERF_SDMA0_READ_RETURN | 137 | sdma0 read return |
| MC_HUB_PERF_SDMA1_READ_RETURN | 138 | sdma1 read return |
| MC_HUB_PERF_SDMA2_READ_RETURN | 139 | sdma2 read return |
| MC_HUB_PERF_SDMA3_READ_RETURN | 140 | sdma3 read return |
| MC_HUB_PERF_SEM_READ_RETURN | 141 | sem read return |
| MC_HUB_PERF_SMU_READ_RETURN | 142 | smu read return |
| MC_HUB_PERF_UMC_READ_RETURN | 144 | umc read return |
| MC_HUB_PERF_UVD_UVD_EXT0_UVD_EXT1_READ_RETURN | 145 | uvd|uvd_ext0|uvd_ext1 read return |
| MC_HUB_PERF_VCE_READ_RETURN | 146 | vce read return |
| MC_HUB_PERF_VCEU_READ_RETURN | 147 | vceu read return |
| MC_HUB_PERF_VMC_READ_RETURN__GUEST_ | 148 | vmc read return (guest) |
| MC_HUB_PERF_VMC1READ_RETURN__HYPERVISOR | 149 | vmc1 read return (hypervisor) |
| MC_HUB_PERF_DMIF_READ_RETURN | 150 | dmif read return |
| MC_HUB_PERF_MCIF_READ_RETURN | 151 | mcif read return |
| MC_HUB_PERF_CPC_READ_REQUEST_ASK_URG_STALL | 161 | cpc read request ask+urg+stall |
| MC_HUB_PERF_CPF_READ_REQUEST_ASK_URG_STALL | 162 | cpf read request ask+urg+stall |
| MC_HUB_PERF_CPG_READ_REQUEST_ASK_URG_STALL | 163 | cpg read request ask+urg+stall |
| MC_HUB_PERF_HDP_READ_REQUEST_ASK_URG_STALL | 164 | hdp read request ask+urg+stall |
| MC_HUB_PERF_IA_READ_REQUEST_ASK_URG_STALL | 165 | ia read request ask+urg+stall |
| MC_HUB_PERF_RLC_READ_REQUEST_ASK_URG_STALL | 166 | rlc read request ask+urg+stall |
| MC_HUB_PERF_SAMSCP_READ_REQUEST_ASK_URG_STALL | 168 | samscp read request ask+urg+stall |
| MC_HUB_PERF_SDMA0_READ_REQUEST_ASK_URG_STALL | 169 | sdma0 read request ask+urg+stall |
| MC_HUB_PERF_SDMA1_READ_REQUEST_ASK_URG_STALL | 170 | sdma1 read request ask+urg+stall |
| MC_HUB_PERF_SDMA2_READ_REQUEST_ASK_URG_STALL | 171 | sdma2 read request ask+urg+stall |
| MC_HUB_PERF_SDMA3_READ_REQUEST_ASK_URG_STALL | 172 | sdma3 read request ask+urg+stall |
| MC_HUB_PERF_SEM_READ_REQUEST_ASK_URG_STALL | 173 | sem read request ask+urg+stall |
| MC_HUB_PERF_SMU_READ_REQUEST_ASK_URG_STALL | 174 | smu read request ask+urg+stall |
| MC_HUB_PERF_UMC_READ_REQUEST_ASK_URG_STALL | 176 | umc read request ask+urg+stall |
| MC_HUB_PERF_UVD_READ_REQUEST_ASK_URG_STALL | 177 | uvd read request ask+urg+stall |
| MC_HUB_PERF_VCE_READ_REQUEST_ASK_URG_STALL | 178 | vce read request ask+urg+stall |
| MC_HUB_PERF_VCEU_READ_REQUEST_ASK_URG_STALL | 179 | vceu read request ask+urg+stall |
| MC_HUB_PERF_VMC_READ_REQUEST_ASK_URG_STALL | 180 | vmc read request ask+urg+stall |
| MC_HUB_PERF_VMC1_READ_REQUEST_ASK_URG_STALL | 181 | vmc1 read request ask+urg+stall |
| MC_HUB_PERF_DMIF_READ_REQUEST_ASK_URG_STALL | 182 | dmif read request ask+urg+stall |
| MC_HUB_PERF_MCIF_READ_REQUEST_ASK_URG_STALL | 183 | mcif read request ask+urg+stall |
| MC_HUB_PERF_GBL0_READ_REQUEST_SEND_CID | 184 | gbl0 read request send+cid |
| MC_HUB_PERF_GBL0_READ_REQUEST_STOR_FULL | 185 | gbl0 read request stor full |
| MC_HUB_PERF_GBL0_READ_REQUEST_BYPASS_STOR_FULL | 186 | gbl0 read request bypass stor full |
| MC_HUB_PERF_GBL0_READ_REQUEST_STOR_EMPTY | 187 | gbl0 read request stor empty |
| MC_HUB_PERF_GBL0_READ_REQUEST_STOR_CREDIT_COUNT | 188 | gbl0 read request stor credit count |
| MC_HUB_PERF_GBL0_READ_REQUEST_BYPASS_STOR_CREDIT_COUNT | 189 | gbl0 read request bypass stor credit count |
| MC_HUB_PERF_GBL1_READ_REQUEST_SEND_CID | 194 | gbl1 read request send+cid |
| MC_HUB_PERF_GBL1_READ_REQUEST_STOR_FULL | 195 | gbl1 read request stor full |
| MC_HUB_PERF_GBL1_READ_REQUEST_BYPASS_STOR_FULL | 196 | gbl1 read request bypass stor full |
| MC_HUB_PERF_GBL1_READ_REQUEST_STOR_EMPTY | 197 | gbl1 read request stor empty |
| MC_HUB_PERF_GBL1_READ_REQUEST_STOR_CREDIT_COUNT | 198 | gbl1 read request stor credit count |
| MC_HUB_PERF_GBL1_READ_REQUEST_BYPASS_STOR_CREDIT_COUNT | 199 | gbl1 read request bypass stor credit count |
| MC_HUB_PERF_DMIF_READ_REQUEST_STALL | 204 | dmif read request stall |
| MC_HUB_PERF_MCIF_READ_REQUEST_STALL | 205 | mcif read request stall |
| Counter | Value | Description |
|---|---|---|
| GRN_DUR_PERF_COUNT_CYCLES | 0 | Count Cycles |
| GRN_DUR_PERF_ONION02_REQ_STALLED_DUE_TO_LACK_OF_HEADER_CREDITS | 4 | Onion02 request stalled due to lack of header credits |
| GRN_DUR_PERF_ONION02_REQ_STALLED_DUE_TO_LACK_OF_DATCREDITS | 5 | Onion02 request stalled due to lack of data credits |
| GRN_DUR_PERF_ONION02_READ_GRANTED_AHEAD_OF_WRITE_DUE_TO_LACK_OF_DATCREDITS | 6 | Onion02 read granted ahead of write due to lack of data credits |
| GRN_DUR_PERF_ONION02_REQ_STALLED_DUE_TO_LACK_OF_HEADER_CREDITS2 | 7 | Onion02 request stalled due to lack of header credits |
| GRN_DUR_PERF_ONION02_REQ_STALLED_DUE_TO_LACK_OF_DATCREDITS2 | 8 | Onion02 request stalled due to lack of data credits |
| GRN_DUR_PERF_ONION02_READ_GRANTED_AHEAD_OF_WRITE_DUE_TO_LACK_OF_DATCREDITS2 | 9 | Onion02 read granted ahead of write due to lack of data credits |
| GRN_DUR_PERF_ONION02_TRANSMIT_REQ_HEADER_CREDITS_AVAIL_INVERTED | 10 | Onion02 transmit request header credits available (inverted); the count is inverted so that the perf counter high water mark can be used as a low water mark. |
| GRN_DUR_PERF_ONION13_TRANSMIT_REQ_HEADER_CREDITS_AVAIL_INVERTED | 11 | Onion13 transmit request header credits available (inverted) |
| GRN_DUR_PERF_ONION02_TRANSMIT_DATCREDITS_AVAIL_INVERTED | 12 | Onion02 transmit data credits available (inverted) |
| GRN_DUR_PERF_ONION13_TRANSMIT_DATCREDITS_AVAIL_INVERTED | 13 | Onion13 transmit data credits available (inverted) |
| GRN_DUR_PERF_REQ_ONION_COMMAND_FIFO_BLOCKED_BY_COLLISION_WITH_GARLIC_REQ_CH0 | 32 | Request at the head of the Gronion receivers Onion Command FIFO is blocked by a collision with a Garlic request to the same address. Per Channel. Add 0x20 * the channel number. |
| GRN_DUR_PERF_VALID_REQ_IS_RECV_FROM_GRONION_CH0 | 33 | A valid request is received from Gronion. Per Channel. Add 0x20 * the channel number. |
| GRN_DUR_PERF_VALID_WRITE_IS_RECV_FROM_GRONION_CH0 | 34 | A valid write is received from Gronion. Per Channel. Add 0x20 * the channel number. |
| GRN_DUR_PERF_VALID_32B_REQ_IS_RECV_FROM_GRONION_CH0 | 35 | A valid 32B request is received from Gronion. Per Channel. Add 0x20 * the channel number. |
| GRN_DUR_PERF_VALID_64B_REQ_IS_RECV_FROM_GRONION_CH0 | 36 | A valid 64B request is received from Gronion. Per Channel. Add 0x20 * the channel number. |
| GRN_DUR_PERF_VALID_REQ_TO_THE_ONION_CHANNEL_IS_RECV_FROM_GRONION_CH0 | 37 | A valid request to the Onion channel is received from Gronion. Per Channel. Add 0x20 * the channel number. |
| GRN_DUR_PERF_VALID_REQ_TO_THE_STALLABLE_GARLIC_CHANNEL_IS_RECV_FROM_GRONION_CH0 | 38 | A valid request to the stallable Garlic channel is received from Gronion. Per Channel. Add 0x20 * the channel number. |
| GRN_DUR_PERF_VALID_REQ_TO_THE_NONSTALLABLE_GARLIC_CHANNEL_IS_RECV_FROM_GRONION_CH0 | 39 | A valid request to the nonstallable Garlic channel is received from Gronion. Per Channel. Add 0x20 * the channel number. |
| GRN_DUR_PERF_THE_ENTRIES_USED_IN_THE_OUTSTANDING_WRITE_CAM_CH0 | 40 | The number of entries used in the outstanding write CAM. Per Channel. Add 0x20 * the channel number. |
| GRN_DUR_PERF_REQ_RECV_THAT_COLLIDES_WITH_AN_OUTSTANDING_REQ_WRITE_CAM_CH0 | 41 | A request has been received that collides with an outstanding request to the other destination in the write CAM. Per Channel. Add 0x20 * the channel number. |
| GRN_DUR_PERF_ENTRIES_USED_IN_THE_ONION_PKTGEN_FIFO0_CH0 | 42 | Number of entries used in the Onion PktGen FIFO0. Per Channel. Add 0x20 * the channel number. |
| GRN_DUR_PERF_ENTRIES_USED_IN_THE_ONION_PKTGEN_FIFO1_CH0 | 43 | Number of entries used in the Onion PktGen FIFO1. Per Channel. Add 0x20 * the channel number. |
| GRN_DUR_PERF_ENTRIES_USED_IN_THE_GARLIC_READ_RESPONSE_FIFO_CH0 | 44 | Number of entries used in the Garlic read response FIFO. Per Channel. Add 0x20 * the channel number. |
| GRN_DUR_PERF_ENTRIES_USED_IN_THE_GARLIC_WR_RESPONSE_FIFO_CH0 | 45 | Number of entries used in the Garlic wr response FIFO. Per Channel. Add 0x20 * the channel number. |
| GRN_DUR_PERF_ENTRIES_USED_IN_THE_GRONION_RECVS_NONSTALLABLE_GARLIC_COMMAND_FIFO_CH0 | 46 | Number of entries used in the Gronion receiver’s Nonstallable Garlic Command FIFO. Per Channel. Add 0x20 * the channel number. |
| GRN_DUR_PERF_ENTRIES_USED_IN_THE_GRONION_RECVS_STALLABLE_GARLIC_COMMAND_FIFO_CH0 | 47 | Number of entries used in the Gronion receiver’s Stallable Garlic Command FIFO. Per Channel. Add 0x20 * the channel number. |
| GRN_DUR_PERF_ENTRIES_USED_IN_THE_GRONION_RECVS_ONION_COMMAND_FIFO_CH0 | 48 | Number of entries used in the Gronion receiver’s Onion Command FIFO.. Per Channel. Add 0x20 * the channel number. |
| GRN_DUR_PERF_FLUSHES_QUEUED_IN_THE_MCT_FLUSH_BLOCK_CH0 | 49 | Number of flushes queued in the MCT Flush block. Per Channel. Add 0x20 * the channel number. |
| GRN_DUR_PERF_THE_HEAD_OF_THE_STALLABLE_GARLIC_FIFO_IS_BLOCKED_DUE_TO_COLLISION_WITH_REQ_TO_ONION_CH0 | 50 | The head of the Stallable Garlic FIFO is blocked due to a collision with a request to Onion. Per Channel. Add 0x20 * the channel number. |
| GRN_DUR_PERF_REQ_TO_GARLIC_STALLED_DUE_TO_LACK_OF_SPACE_IN_THE_UNBS_GARLIC_REQ_FIFO_CH0 | 51 | Request to Garlic stalled due to lack of space in the UNB???s Garlic Request FIFO. Per Channel. Add 0x20 * the channel number. |
| GRN_DUR_PERF_PRIURGENT_IS_ASSERTED_FROM_THE_GMC_CH0 | 52 | PriUrgent is asserted from the GMC, giving the nonstallable Garlic channel priority for both requests and responses. Per Channel. Add 0x20 * the channel number. |
| GRN_DUR_PERF_VALID_WRITE_TO_THE_ONION_CHANNEL_IS_RECV_FROM_GRONION_CH0 | 53 | A valid write to the Onion channel is received from Gronion. Per Channel. Add 0x20 * the channel number. |
| GRN_DUR_PERF_VALID_WRITE_TO_THE_NONSTALLABLE_GARLIC_CHANNEL_IS_RECV_FROM_GRONION_CH0 | 54 | A valid write to the nonstallable Garlic channel is received from Gronion. Per Channel. Add 0x20 * the channel number. |
| GRN_DUR_PERF_VALID_WRITE_TO_THE_STALLABLE_GARLIC_CHANNEL_IS_RECV_FROM_GRONION_CH0 | 55 | A valid write to the stallable Garlic channel is received from Gronion. Per Channel. Add 0x20 * the channel number. |
| GRN_DUR_PERF_READ_TO_THE_ONION_CHANNEL_IS_RECV_FROM_GRONION_CH0 | 56 | A read to the Onion channel is received from Gronion. Per Channel. Add 0x20 * the channel number. |
| GRN_DUR_PERF_READ_TO_THE_NONSTALLABLE_GARLIC_CHANNEL_IS_RECV_FROM_GRONION_CH0 | 57 | A read to the nonstallable Garlic channel is received from Gronion. Per Channel. Add 0x20 * the channel number. |
| GRN_DUR_PERF_64B_WRITE_TO_THE_ONION_CHANNEL_IS_RECV_FROM_GRONION_CH0 | 58 | A 64B write to the Onion channel is received from Gronion. Per Channel. Add 0x20 * the channel number. |
| GRN_DUR_PERF_64B_WRITE_TO_THE_NONSTALLABLE_GARLIC_CHANNEL_IS_RECV_FROM_GRONION_CH0 | 59 | A 64B write to the nonstallable Garlic channel is received from Gronion. Per Channel. Add 0x20 * the channel number. |
| GRN_DUR_PERF_64B_WRITE_TO_THE_STALLABLE_GARLIC_CHANNEL_IS_RECV_FROM_GRONION_CH0 | 60 | A 64B write to the stallable Garlic channel is received from Gronion. Per Channel. Add 0x20 * the channel number. |
| GRN_DUR_PERF_64B_READ_TO_THE_ONION_CHANNEL_IS_RECV_FROM_GRONION_CH0 | 61 | A 64B read to the Onion channel is received from Gronion. Per Channel. Add 0x20 * the channel number. |
| GRN_DUR_PERF_64B_READ_TO_THE_NONSTALLABLE_GARLIC_CHANNEL_IS_RECV_FROM_GRONION_CH0 | 62 | A 64B read to the nonstallable Garlic channel is received from Gronion. Per Channel. Add 0x20 * the channel number. |
| Counter | Value | Description |
|---|---|---|
| GRN_SCO_PERF_COUNT_CYCLES | 0 | Count Cycles |
| GRN_SCO_PERF_ONION_EVEN_TAG_MISMATCH | 1 | Onion Even tag re-enumeration received a tag it did not request. This may indicate corruption on the onion path |
| GRN_SCO_PERF_ONION_ODD_TAX_MISMATCH | 2 | Onion Odd tag re-enumeration received a tag it did not request. This may indicate corruption on the onion path |
| GRN_SCO_PERF_ONION_EVEN_TAG_FLIGHT | 3 | Onion Even Number of onion tags currently in flight |
| GRN_SCO_PERF_ONION_ODD_TAG_FLIGHT | 4 | Onion Odd Number of onion tags currently in flight |
| GRN_SCO_PERF_ONION_EVEN_REQUEST_STALLED_HEADER | 5 | Onion Even Request stalled due to lack of header credits |
| GRN_SCO_PERF_ONION_EVEN_REQUEST_STALED_DATA | 6 | Onion Even Request stalled due to lack of data credits |
| GRN_SCO_PERF_ONION_ODD_REQEST_STALLED_HEADER | 7 | Onion Odd Request stalled due to lack of header credits |
| GRN_SCO_PERF_ONION_ODD_REQUEST_STALLED_DATA | 8 | Onion Odd Request stalled due to lack of data credits |
| GRN_SCO_PERF_ONION_EVEN_REQUEST_HEADER_CREDIT_AVAIL | 9 | OnionEven transmit request header credits available (inverted); the count is inverted so that the perf counter high water mark can be used as a low water mark. |
| GRN_SCO_PERF_ONION_ODD_REQUEST_HEADER_CREDIT_AVAIL | 10 | OnionOdd transmit request header credits available (inverted) |
| GRN_SCO_PERF_ONION_EVEN_REQUEST_DATA_AVAIL | 11 | OnionEven transmit data credits available (inverted) |
| GRN_SCO_PERF_ONION_ODD_REQUEST_DATA_AVAIL | 12 | OnionOdd transmit data credits available (inverted) |
| GRN_SCO_PERF_ONION_EVEN_READ_REQUEST_2_CYCLE | 13 | OnionEven read request 2-per cycle |
| GRN_SCO_PERF_ONION_EVEN_READ_REQUEST_SINGLE_UNPACKED | 14 | OnionEven read request single, unpacked |
| GRN_SCO_PERF_ONION_ODD_READ_REQUEST_2_CYCLE | 15 | OnionOdd read request 2-per cycle |
| GRN_SCO_PERF_ONION_ODD_REQD_REQUEST_SINGLE_UNPACKED | 16 | OnionOdd read request single, unpacked |
| GRN_SCO_PERF_ONION_EVEN_RECEIVE_SYNC_FIFO_USAGE | 17 | Onion Even Receive Sync FIFO usage |
| GRN_SCO_PERF_ONION_ODD_RECEIVE_SYNC_FIFO_USAGE | 18 | Onion Odd Receive Sync FIFO usage |
| GRN_SCO_PERF_ONION_EVEN_READ_RESPONSE_HEADER_FIFO_USAGE | 19 | Onion Even Read Response header FIFO usage |
| GRN_SCO_PERF_ONION_ODD_READ_RESPONSE_HEADER_FIFO_USAGE | 20 | Onion Odd Read Response header FIFO usage |
| GRN_SCO_PERF_ONION_EVEN_READ_RESPONSE_DATA_FIFO_USAGE | 21 | Onion Even Read Response date FIFO usage |
| GRN_SCO_PERF_ONION_ODD_READ_RESPONSE_DATA_FIFO_USAGE | 22 | Onion Odd Read Response date FIFO usage |
| GRN_SCO_PERF_RESERVED | 32 | Reserved |
| GRN_SCO_PERF_VALID_REQ_IS_RECV_FROM_GRONION_CH0 | 33 | A valid request is received from Gronion. Per Channel. Add 0x38 * the channel number. |
| GRN_SCO_PERF_VALID_WRITE_IS_RECV_FROM_GRONION_CH0 | 34 | A valid write is received from Gronion. Per Channel. Add 0x38 * the channel number. |
| GRN_SCO_PERF_VALID_32B_REQ_IS_RECV_FROM_GRONION_CH0 | 35 | A valid 32B request is received from Gronion. Per Channel. Add 0x38 * the channel number. |
| GRN_SCO_PERF_VALID_64B_REQ_IS_RECV_FROM_GRONION_CH0 | 36 | A valid 64B request is received from Gronion. Per Channel. Add 0x38 * the channel number. |
| GRN_SCO_PERF_VALID_REQ_TO_THE_ONION_CHANNEL_IS_RECV_FROM_GRONION_CH0 | 37 | A valid request to the Onion channel is received from Gronion. Per Channel. Add 0x38 * the channel number. |
| GRN_SCO_PERF_VALID_REQ_TO_THE_STALLABLE_GARLIC_CHANNEL_IS_RECV_FROM_GRONION_CH0 | 38 | A valid request to the stallable Garlic channel is received from Gronion. Per Channel. Add 0x38 * the channel number. |
| GRN_SCO_PERF_VALID_REQ_TO_THE_NONSTALLABLE_GARLIC_CHANNEL_IS_RECV_FROM_GRONION_CH0 | 39 | A valid request to the nonstallable Garlic channel is received from Gronion. Per Channel. Add 0x38 * the channel number. |
| GRN_SCO_PERF_THE_NUMBER_OF_ENTRIES_USED_IN_THE_OUTSTANDING_WRITE_CAM_CH0 | 40 | The number of entries used in the outstanding write CAM. Per Channel. Add 0x38 * the channel number. |
| GRN_SCO_PERF_REQ_HAS_BEEN_RECEIVED_THAT_COLLIDES_WITH_AN_OUTSTANDING_RMW_CH0 | 41 | A request has been received that collides with an outstanding RMW . Per Channel. Add 0x38 * the channel number. |
| GRN_SCO_PERF_NUMBER_OF_ENTRIES_USED_IN_THE_ONION_PKTGEN_FIFO0_CH0 | 42 | Number of entries used in the Onion PktGen FIFO0. Per Channel. Add 0x38 * the channel number. |
| GRN_SCO_PERF_NUMBER_OF_ENTRIES_USED_IN_THE_ONION_PKTGEN_FIFO1_CH0 | 43 | Number of entries used in the Onion PktGen FIFO1. Per Channel. Add 0x38 * the channel number. |
| GRN_SCO_PERF_NUMBER_OF_ENTRIES_USED_IN_THE_GARLIC_READ_RESPONSE_FIFO_CH0 | 44 | Number of entries used in the Garlic read response FIFO. Per Channel. Add 0x38 * the channel number. |
| GRN_SCO_PERF_NUMBER_OF_ENTRIES_USED_IN_THE_GARLIC_WR_RESPONSE_FIFO_CH0 | 45 | Number of entries used in the Garlic wr response FIFO. Per Channel. Add 0x38 * the channel number. |
| GRN_SCO_PERF_NUMBER_OF_ENTRIES_USED_IN_THE_GRONION_RECEIVER_NONSTALLABLE_GARLIC_CMD_FIFO_CH0 | 46 | Number of entries used in the Gronion receiver’s Nonstallable Garlic Command FIFO. Per Channel. Add 0x38 * the channel number. |
| GRN_SCO_PERF_NUMBER_OF_ENTRIES_USED_IN_THE_GRONION_RECEIVER_STALLABLE_GARLIC_CMD_FIFO_CH0 | 47 | Number of entries used in the Gronion receiver’s Stallable Garlic Command FIFO. Per Channel. Add 0x38 * the channel number. |
| GRN_SCO_PERF_NUMBER_OF_ENTRIES_USED_IN_THE_GRONION_RECEIVER_ONION_COMMAND_FIFO_CH0 | 48 | Number of entries used in the Gronion receiver’s Onion Command FIFO.. Per Channel. Add 0x38 * the channel number. |
| GRN_SCO_PERF_NUMBER_OF_FLUSHES_QUEUED_IN_THE_MCT_FLUSH_BLOCK_CH0 | 49 | Number of flushes queued in the MCT Flush block. Per Channel. Add 0x38 * the channel number. |
| GRN_SCO_PERF_HEAD_OF_THE_STALLABLE_GARLIC_FIFO_IS_BLOCKED_DUE_TO_A_COLLISION_ONION_CH0 | 50 | The head of the Stallable Garlic FIFO is blocked due to a collision with a request to Onion. Per Channel. Add 0x38 * the channel number. |
| GRN_SCO_PERF_REQ_TO_GARLIC_STALLED_DUE_TO_LACK_SPACE_IN_UNB_GARLIC_REQ_FIFO_CH0 | 51 | Request to Garlic stalled due to lack of space in the UNBs Garlic Request FIFO |
| GRN_SCO_PERF_PRIURGENT_IS_ASSERTED_FROM_THE_GMC_CH0 | 52 | PriUrgent is asserted from the GMC, giving the nonstallable Garlic channel priority for both requests and responses. Per Channel. Add 0x38 * the channel number. |
| GRN_SCO_PERF_VALID_WRITE_TO_THE_ONION_CHANNEL_IS_RECV_FROM_GRONION_CH0 | 53 | A valid write to the Onion channel is received from Gronion |
| GRN_SCO_PERF_VALID_WRITE_TO_THE_NONSTALLABLE_GARLIC_CHANNEL_IS_RECV_FROM_GRONION_CH0 | 54 | A valid write to the nonstallable Garlic channel is received from Gronion. Per Channel. Add 0x38 * the channel number. |
| GRN_SCO_PERF_VALID_WRITE_TO_THE_STALLABLE_GARLIC_CHANNEL_IS_RECV_FROM_GRONION_CH0 | 55 | A valid write to the stallable Garlic channel is received from Gronion. Per Channel. Add 0x38 * the channel number. |
| GRN_SCO_PERF_READ_TO_THE_ONION_CHANNEL_IS_RECV_FROM_GRONION_CH0 | 56 | A read to the Onion channel is received from Gronion. Per Channel. Add 0x38 * the channel number. |
| GRN_SCO_PERF_READ_TO_THE_NONSTALLABLE_GARLIC_CHANNEL_IS_RECV_FROM_GRONION_CH0 | 57 | A read to the nonstallable Garlic channel is received from Gronion.. Per Channel. Add 0x38 * the channel number. |
| GRN_SCO_PERF_64B_WRITE_TO_THE_ONION_CHANNEL_IS_RECV_FROM_GRONION_CH0 | 58 | A 64B write to the Onion channel is received from Gronion. Per Channel. Add 0x38 * the channel number. |
| GRN_SCO_PERF_64B_WRITE_TO_THE_NONSTALLABLE_GARLIC_CHANNEL_IS_RECV_FROM_GRONION_CH0 | 59 | a 64B write to the nonstallable Garlic channel is received from Gronion. Per Channel. Add 0x38 * the channel number. |
| GRN_SCO_PERF_64B_WRITE_TO_THE_STALLABLE_GARLIC_CHANNEL_IS_RECV_FROM_GRONION_CH0 | 60 | a 64B write to the stallable Garlic channel is received from Gronion. Per Channel. Add 0x38 * the channel number. |
| GRN_SCO_PERF_64B_READ_TO_THE_ONION_CHANNEL_IS_RECV_FROM_GRONION_CH0 | 61 | a 64B read to the Onion channel is received from Gronion. Per Channel. Add 0x38 * the channel number. |
| GRN_SCO_PERF_64B_READ_TO_THE_NONSTALLABLE_GARLIC_CHANNEL_IS_RECV_FROM_GRONION_CH0 | 62 | a 64B read to the nonstallable Garlic channel is received from Gronion. Per Channel. Add 0x38 * the channel number. |
| GRN_SCO_PERF_GARLIC_DROPPED_READ_RESPONSE_DUE_TO_EDC_ERROR_CH0 | 63 | Garlic dropped Read Response due to EDC error. Per Channel. Add 0x38 * the channel number. |
| GRN_SCO_PERF_RMW_MAIN_FIFO_USAGE_CH0 | 64 | RMW main FIFO usage. Per Channel. Add 0x38 * the channel number. |
| GRN_SCO_PERF_RMW_RDRSP_FIFO_USAGE_CH0 | 65 | RMW RdRsp FIFO usage. Per Channel. Add 0x38 * the channel number. |
| GRN_SCO_PERF_RMW_WRITE_DETECTED_ON_GRONION_CHANNEL_CH0 | 66 | RMW write detected on gronion channel. Per Channel. Add 0x38 * the channel number. |
| GRN_SCO_PERF_REGULAR_NON_CH0 | 67 | Regular, non. Per Channel. Add 0x38 * the channel number. |
| GRN_SCO_PERF_PARTIAL_WRITE_DETECTED_ON_NON_STALLABLE_CHANNEL_CH0 | 68 | Partial write detected on Non Stallable channel. Per Channel. Add 0x38 * the channel number. |
| GRN_SCO_PERF_RMW_QUEUE_IS_FULL_AND_IS_STALLING_THE_STALLABLE_CHANNEL_CH0 | 69 | RMW queue is full and is stalling the stallable channel. Per Channel. Add 0x38 * the channel number. |
| GRN_SCO_PERF_NUMBER_OF_GARLIC_READ_TAGS_IN_FLIGHT_CH0 | 70 | Number of garlic Read tags in flight. Per Channel. Add 0x38 * the channel number. |
| GRN_SCO_PERF_GARLIC_RD_TAG_RE_CH0 | 71 | Garlic Rd Tag re. Per Channel. Add 0x38 * the channel number. |
| GRN_SCO_PERF_ENTRIES_USED_IN_GRONION_ONION_CMD_FIFO_CH0 | 72 | Entries used in Gronion Onion Cmd FIFO. Per Channel. Add 0x38 * the channel number. |
| GRN_SCO_PERF_ENTRIES_USED_IN_GRONION_ONION_ARB_FIFO_CH0 | 73 | Entries used in Gronion Onion Arb FIFO. Per Channel. Add 0x38 * the channel number. |
| GRN_SCO_PERF_ENTRIES_USED_IN_GRONION_ONION_DATA_FIFO_CH0 | 74 | Entries used in Gronion Onion Data FIFO. Per Channel. Add 0x38 * the channel number. |
| GRN_SCO_PERF_GARLIC_READ_TAG_RE_CH0 | 75 | Garlic Read Tag re. Per Channel. Add 0x38 * the channel number. |
| GRN_SCO_PERF_RMW_LOGIC_ACTIVE_SET_SIZE_CH0 | 76 | RMW Logic active set size. Per Channel. Add 0x38 * the channel number. |
| GRN_SCO_PERF_RMW_LATENCY_CH0 | 77 | RMW latency of random RMW request in SCLK cycles divided by 4 as measured at the gronion interface. Use high watermark (MODE_MAX) to get maximum latency. The counter is live, not flopped, at the perf event. For events 0x4D. Per Channel. Add 0x38 * the channel number. |
| GRN_SCO_PERF_ONION_COHERENT_READ_LATENCY_FROM_RANDOM_ONION_READ_REQ_CH0 | 78 | Onion coherent read latency from random onion read request |
| GRN_SCO_PERF_ONION_NON_CH0 | 79 | Onion non. Per Channel. Add 0x38 * the channel number. |
| GRN_SCO_PERF_STALLABLE_GARLIC_READ_LATENCY_FROM_RANDOM_REQ_CH0 | 80 | Stallable garlic read latency from random request. Per Channel. Add 0x38 * the channel number. |
| GRN_SCO_PERF_NON_CH0 | 81 | Non. Per Channel. Add 0x38 * the channel number. |
| GRN_SCO_PERF_ONION_COHERENT_WRITE_LATENCY_CH0 | 82 | Onion coherent write latency. Per Channel. Add 0x38 * the channel number. |
| GRN_SCO_PERF_ONION_NON2_CH0 | 83 | Onion non. Per Channel. Add 0x38 * the channel number. |
| GRN_SCO_PERF_STALLABLE_GARLIC_WRITE_LATENCY_FROM_RANDOM_REQ_CH0 | 84 | Stallable garlic write latency from random request. Per Channel. Add 0x38 * the channel number. |
| GRN_SCO_PERF_NON2_CH0 | 85 | Non. Per Channel. Add 0x38 * the channel number. |
| GRN_SCO_PERF_RMW_MAIN_DATA_FIFO_USAGE_CH0 | 86 | RMW main data FIFO usage.. Per Channel. Add 0x38 * the channel number. |
| GRN_SCO_PERF_TRAFFIC_SHAPING_ENGAGED_CH0 | 87 | Traffic Shaping Engaged. Counts number of cycles where some part of an RMW burst was kept pending because of traffic shaping. This is the number of cycles normal traffic. Per Channel. Add 0x38 * the channel number. |