| 1 | The gem5 Simulator | 6 | NO | 2011 | CAN |
| 2 | Memory Access Scheduling | 5 | NO | 2000 | ISCA |
| 3 | Fine-grained activation for power reduction in DRAM | 5 | NO | 2010 | IEEE Micro |
| 4 | Architecting phase change memory as a scalable DRAM alternative | 5 | NO | 2009 | ISCA |
| 5 | A Scalable Processing-in-memory Accelerator for Parallel Graph Processing | 5 | YES | 2015 | ISCA |
| 6 | Rethinking DRAM design and organization for energy-constrained multi-cores | 4 | NO | 2010 | ISCA |
| 7 | Adaptive granularity memory systems: A tradeoff between storage efficiency and throughput | 4 | YES | 2011 | ISCA |
| 8 | Mini-rank: Adaptive DRAM architecture for improving memory power efficiency | 4 | NO | 2008 | MICRO |
| 9 | A Case for Exploiting Subarray-Level Parallelism (SALP) in DRAM | 4 | YES | 2012 | ISCA |
| 10 | DRAM Circuit Design: Fundamental and High-Speed Topics | 4 | NO | 2007 | Book |
| 11 | Efficient Virtual Memory for Big Memory Servers | 4 | NO | 2013 | ISCA |
| 12 | Rodinia: A benchmark suite for heterogeneous computing | 4 | NO | 2009 | IISWC |
| 13 | EIE: Efficient Inference Engine on Compressed Deep Neural Network | 4 | YES | 2016 | ISCA |
| 14 | Very Deep Convolutional Networks for Large- Scale Image Recognition, | 4 | YES | 2015 | ICLR |
| 15 | Supporting x86-64 address translation for 100s of GPU lanes | 4 | YES | 2014 | HPCA |
| 16 | DRAMSim2: A cycle accurate memory system simulator | 4 | NO | 2011 | CAL |
| 17 | PIM-enabled Instructions: A Low-overhead, Locality-aware Processing-in-memory Architecture | 4 | NO | 2015 | ISCA |
| 18 | Memory Systems: Cache, DRAM, Disk | 3 | NO | 2007 | Book |
| 19 | Future scaling of processor-memory interfaces | 3 | NO | 2009 | SC |
| 20 | Structural aspects of the system/360 model 85: II the cache | 3 | NO | 1968 | IBM Systems Journal |
| 21 | Software caching and computation migration in Olden | 3 | YES | 1995 | Tech Report |
| 22 | TOP-PIM: Throughput-oriented Programmable Processing in Memory | 3 | NO | 2014 | HPDC |
| 23 | Shared Last-level TLBs for Chip Multiprocessors | 3 | NO | 2011 | HPCA |
| 24 | Optimizing NUCA Organizations and Wiring Alternatives for Large Caches with CACTI 6.0 | 3 | NO | 2007 | MICRO |
| 25 | USENIX Association | 3 | NO | 2004 | USENIX |
| 26 | Minimalist Open-page: A DRAM Page-mode Scheduling Policy for the Many-core Era | 3 | NO | 2011 | MICRO |
| 27 | CACTI-3DD: Architecture-level Modeling for 3D Die-stacked DRAM Main Memory | 3 | NO | 2012 | DATE |
| 28 | Understanding the Energy Consumption of Dynamic Random Access Memories | 3 | NO | 2010 | MICRO |
| 29 | Half-DRAM: A High-bandwidth and Low-power DRAM Architecture from the Rethinking of Fine-grained Activation | 3 | YES | 2014 | ISCA |
| 30 | A permutation-based page interleaving scheme to reduce row-buffer conflicts and exploit data locality | 3 | NO | 2000 | MICRO |
| 31 | SPEC CPU2006 Benchmark Descriptions | 3 | NO | 2006 | SIGARCH Computer Architecture News |
| 32 | Eyeriss: A Spatial Architecture for Energy-efficient Dataflow for Convolutional Neural Networks | 3 | YES | 2016 | ISCA |
| 33 | Architectural Support for Address Translation on GPUs: Designing Memory Management Units for CPU/GPUs with Unified Address Spaces | 3 | YES | 2014 | ASPLOS |
| 34 | DRAM Errors in the Wild: A Large-Scale Field Study | 3 | NO | 2009 | SIGMETRICS |
| 35 | Base-Delta-Immediate Compression: Practical Data Compression for On-Chip Caches | 3 | NO | 2012 | PACT |
| 36 | ISAAC: A Convolutional Neural Network Accelerator with In-Situ Analog Arithmetic in Crossbars | 3 | YES | 2016 | ISCA |
| 37 | Disaggregated Memory for Expansion and Sharing in Blade Servers | 2 | NO | 2009 | ISCA |
| 38 | Towards Energy-Proportional Datacenter Memory with Mobile DRAM | 2 | NO | 2012 | ISCA |
| 39 | BOOM: Enabling Mobile Memory based Low-Power Server DIMMs | 2 | NO | 2012 | ISCA |
| 40 | Memory Power Management via Dynamic Voltage/Frequency Scaling | 2 | NO | 2011 | ICAC |
| 41 | Skinflint DRAM System: Minimizing DRAM Chip Writes for Low Power | 2 | YES | 2013 | HPCA |
| 42 | The Dirty-block Index | 2 | NO | 2014 | ISCA |
| 43 | VLSI Memory Chip Design | 2 | NO | 2001 | Springer |
| 44 | Spectre Attacks: Exploiting Speculative Execution | 2 | NO | 2018 | arXiv |
| 45 | Meltdown | 2 | NO | 2018 | arXiv |
| 46 | On the effectiveness of address-space randomization | 2 | NO | 2004 | CCS |
| 47 | Efficient Address Translation for Architectures with Multiple Page Sizes | 2 | NO | 2017 | ASPLOS |
| 48 | Characterization of silent stores | 2 | NO | 2000 | PACT |
| 49 | Improving the reliability of on-chip L2 cache using redundancy | 2 | NO | 2007 | ICCD |
| 50 | Dynamically exploiting narrow width operands to improve processor power and performance | 2 | NO | 1999 | HPCA |
| 51 | Preventing PCM banks from seizing too much power | 2 | NO | 2011 | MICRO |
| 52 | Experimental evaluation of on-chip microprocessor cache memories | 2 | NO | 1984 | ISCA |
| 53 | DRAM Energy reduction by prefetching-based memory traffic clustering | 2 | NO | 2011 | GLSVLSI |
| 54 | On the value locality of store instructions | 2 | NO | 2000 | ISCA |
| 55 | Silent stores for free | 2 | NO | 2000 | MICRO |
| 56 | Understanding and designing new server architectures for emerging warehouse-computing environments | 2 | NO | 2008 | ISCA |
| 57 | Zesto: A cycle-level simulator for highly detailed microarchitecture exploration | 2 | NO | 2009 | ISPASS |
| 58 | PowerNap: Eliminating server idle power | 2 | NO | 2009 | ASPLOS |
| 59 | Memory-link compression schemes: A value locality perspective | 2 | NO | 2008 | IEEE TC |
| 60 | Energy reduction for STT-RAM using early write termination | 2 | NO | 2009 | ICCAD |
| 61 | CoLT: Coalesced Large-Reach TLBs | 2 | NO | 2012 | MICRO |
| 62 | GPUs and the Future of Parallel Computing | 2 | NO | 2011 | IEEE Micro |
| 63 | LazyPIM: An Efficient Cache Coherence Mechanism for Processing-in-Memory | 2 | NO | 2017 | IEEE CAL |
| 64 | The Architecture of the DIVA Processing-in-memory Chip | 2 | NO | 2002 | ICS |
| 65 | 3D-stacked Memory-side Acceleration: Accelerator and System Design | 2 | NO | 2013 | WoNDP |
| 66 | Transparent Offloading and Mapping (TOM): Enabling Programmer-transparent Near-data Processing in GPU Systems | 2 | NO | 2016 | ISCA |
| 67 | Accelerating Pointer Chasing in 3D-stacked Memory: Challenges, Mechanisms, Evaluation | 2 | NO | 2016 | ICCD |
| 68 | Hybrid Memory Cube: New DRAM Architecture Increases Density and Performance | 2 | YES | 2012 | VLSIT |
| 69 | FlexRAM: Toward an Advanced Intelligent Memory System | 2 | NO | 1999 | ICCD |
| 70 | Ramulator: A Fast and Extensible DRAM Simulator | 2 | NO | 2016 | IEEE CAL |
| 71 | EXECUBE: A New Architecture for Scaleable MPPs | 2 | NO | 1994 | ICPP |
| 72 | Simultaneous Multi-Layer Access: Improving 3D-Stacked Memory Bandwidth at Low Cost | 2 | NO | 2016 | ACM TACO |
| 73 | Active Pages: A Computation Model for Intelligent Memory | 2 | NO | 1998 | ISCA |
| 74 | A Case for Intelligent RAM | 2 | NO | 1997 | IEEE Micro |
| 75 | Scheduling Techniques for GPU Architectures with Processing-in-memory Capabilities | 2 | NO | 2016 | PACT |
| 76 | Fast Bulk Bitwise AND and OR in DRAM | 2 | YES | 2015 | IEEE CAL |
| 77 | A Logic-in-Memory Computer | 2 | NO | 1970 | IEEE Trans. Comput. |
| 78 | Beyond the Wall: Near-Data Processing for Databases | 2 | YES | 2015 | DAMON |
| 79 | Compute Caches | 2 | NO | 2017 | HPCA |
| 80 | NDA: Near-DRAM Acceleration Architecture Leveraging Commodity DRAM Devices and Standard Memory Modules | 2 | NO | 2015 | HPCA |
| 81 | Pinatubo: A Processing-in-Memory Architecture for Bulk Bitwise Operations in Emerging Non-Volatile Memories | 2 | YES | 2016 | DAC |
| 82 | In-Datacenter Performance Analysis of a Tensor Processing Unit | 2 | NO | 2017 | ISCA |
| 83 | Di- annao: A small-footprint high-throughput accelerator for ubiquitous machine-learning, | 2 | NO | 2014 | ASPLOS |
| 84 | Dadiannao: A machine-learning supercom- puter, | 2 | NO | 2014 | MICRO |
| 85 | Shidiannao: Shifting vision processing closer to the sensor, | 2 | NO | 2015 | ISCA |
| 86 | Optimizing fpga-based accelerator design for deep convolutional neural networks, | 2 | YES | 2015 | fpga |
| 87 | Flexflow: A flexible dataflow accelerator architecture for convolutional neural networks, | 2 | YES | 2017 | HPCA |
| 88 | NVIDIA Tesla: A Unified Graphics and Computing Architecture | 2 | NO | 2008 | IEEE Micro |
| 89 | Controller for a Synchronous DRAM that Maximizes Throughput by Allowing Memory Requests and Commands to be Issued Out of Order | 2 | YES | 1997 | US Patent 5630096 |
| 90 | Large-reach Memory Management Unit Caches | 2 | NO | 2013 | MICRO |
| 91 | Inter-core Cooperative TLB for Chip Multiprocessors | 2 | NO | 2010 | ASPLOS |
| 92 | Supporting Address Translation for Accelerator-Centric Architectures | 2 | NO | 2017 | HPCA |
| 93 | Redundant Memory Mappings for Fast Access to Large Memories | 2 | NO | 2015 | ISCA |
| 94 | Adaptive Cache Management for Energy-Efficient GPU Computing | 2 | YES | 2014 | MICRO |
| 95 | iGPU: Exception Support and Speculative Execution on GPUs | 2 | NO | 2012 | ISCA |
| 96 | Observations and opportunities in architecting shared virtual memory for heterogeneous systems | 2 | YES | 2016 | ISPASS |
| 97 | ATLAS: A Scalable and High- Performance Scheduling Algorithm for Multiple Memory Controllers, | 2 | YES | 2010 | HPCA |
| 98 | Thread Cluster Memory Scheduling: Exploiting Differences in Memory Access Behavior, | 2 | NO | 2010 | MICRO |
| 99 | Reducing Memory Interference in Multicore Systems via Application-Aware Memory Channel Partitioning, | 2 | NO | 2011 | MICRO |
| 100 | Exploiting inter-warp heterogeneity to improve gpgpu performance, | 2 | NO | 2015 | PACT |
| 101 | Analyzing CUDA workloads using a detailed GPU simulator, | 2 | NO | 2009 | ISPASS |
| 102 | Orchestrated scheduling and prefetching for GPGPUs, | 2 | NO | 2013 | ISCA |
| 103 | Managing GPU concurrency in heterogeneous architectures, | 2 | NO | 2014 | MICRO |
| 104 | Locality-driven dynamic GPU cache bypassing, | 2 | NO | 2015 | - |
| 105 | Cache-conscious wavefront scheduling, | 2 | NO | 2012 | MICRO |
| 106 | IEEE Computer Society | 2 | NO | 2014 | MICRO |
| 107 | Microbank: Architecting Through-Silicon Interposer-Based Main Memory Systems | 2 | NO | 2014 | SC |
| 108 | A White Paper on the Benefits of Chipkill-Correct ECC for PC Server Main Memory | 2 | YES | 1997 | IBM Microelectronics Division |
| 109 | Architecting an Energy-Efficient DRAM System for GPUs, | 2 | NO | 2017 | HPCA |
| 110 | USIMM: the Utah Simulated Memory Module | 2 | NO | 2012 | - |
| 111 | A 1.2v 38nm 2.4gb/s/pin 2gb ddr4 sdram with bank group and 4 half-page architecture | 2 | NO | 2012 | ISSCC |
| 112 | A non-volatile microcontroller with integrated floating-gate transistors | 2 | NO | 2011 | IEEE |
| 113 | A Simpler, Safer Programming and Execution Model for Intermittent Systems | 2 | NO | 2015 | ACM |
| 114 | Arrakis: The operating system is the control plane | 2 | NO | 2014 | OSDI |
| 115 | A Full GPU Virtualization Solution with Mediated Pass-Through, | 2 | NO | 2014 | USENIX |
| 116 | Imagenet classification with deep convolutional neural networks, | 2 | YES | 2012 | NeurIPS |
| 117 | Going deeper with convolutions, | 2 | NO | 2015 | CVPR |
| 118 | Cambricon-x: An accelerator for sparse neural networks, | 2 | NO | 2016 | MICRO |
| 119 | Long short-term memory, | 2 | NO | 1997 | Neural Computation |
| 120 | A Case for Toggle-Aware Compression for GPU Systems, | 2 | NO | 2016 | HPCA |
| 121 | Nv-heaps: making persistent objects fast and safe with next-generation, non-volatile memories | 2 | NO | 2011 | ACM SIGPLAN Notices |
| 122 | High-performance transactions for persistent memories | 2 | NO | 2016 | ASPLOS |
| 123 | Dudetm: Building durable transactions with decoupling for persistent memory | 2 | YES | 2017 | ASPLOS |
| 124 | An analysis of persistent memory use with whisper | 2 | NO | 2017 | ASPLOS |
| 125 | Mnemosyne: Lightweight persistent memory | 2 | NO | 2011 | ACM SIGARCH Computer Architecture News |
| 126 | Efficient Memory Integrity Verification and Encryption for Secure Processors | 2 | YES | 2003 | MICRO |
| 127 | Incidental Computing on IoT Nonvolatile Processors | 2 | NO | 2017 | IEEE |
| 128 | The Dynamic Granularity Memory System | 2 | NO | 2012 | ISCA |
| 129 | Co-architecting Controllers and DRAM to Enhance DRAM Process Scaling | 2 | NO | 2014 | The Memory Forum |
| 130 | DDR4 SDRAM STANDARD | 2 | NO | 2012 | JEDEC |
| 131 | Redundancy techniques for high-density drams | 2 | NO | 1997 | IEEE International Conference on Innovative Systems in Silicon |
| 132 | Archshield: Architectural framework for assisting dram scaling by tolerating high error rates | 2 | NO | 2013 | ACM SIGARCH Computer Architecture News |
| 133 | 8Gb DDR4 SDRAM | 2 | NO | - | SK hynix |
| 134 | The Netflix Prize | 2 | NO | 2017 | KDD |
| 135 | Graphicionado: A high-performance and energy-efficient accelerator for graph analytics | 2 | NO | 2016 | MICRO |
| 136 | Practical Near-Data Processing for In- Memory Analytics Frameworks, | 2 | YES | 2015 | PACT |
| 137 | GraphPIM: Enabling Instruction-Level PIM Offloading in Graph Computing Frameworks, | 2 | NO | 2017 | HPCA |
| 138 | Nvsim: A circuit-level performance, energy, and area model for emerging nonvolatile memory, | 2 | YES | 2012 | IEEE |
| 139 | Cnvlutin: Ineffectual-neuron-free deep neural network computing, | 2 | NO | 2016 | ISCA |
| 140 | DaDianNao: A Machine-Learning Supercomputer | 2 | NO | 2014 | MICRO |
| 141 | Enhancing lifetime and security of PCM-based main memory with start-gap wear leveling | 2 | YES | 2009 | MICRO |
| 142 | Overcoming the challenges of crossbar resistive memory architectures | 2 | NO | 2015 | HPCA |
| 143 | The Datacenter as a Computer | 1 | NO | 2009 | Book |
| 144 | Tiered Memory: An Iso-Power Memory Architecture to Address the Memory Wall | 1 | NO | 2012 | IEEE TC |
| 145 | Smart Refresh: An Enhanced Memory Controller Design for Reducing Energy in Conventional and 3D Die-Stacked DRAMs | 1 | YES | 2007 | MICRO |
| 146 | A Comprehensive Approach to DRAM Power Management | 1 | NO | 2008 | HPCA |
| 147 | Flikker: Saving DRAM Refresh-power through Critical Data Partitioning | 1 | NO | 2011 | ASPLOS |
| 148 | Rethinking DRAM Power Modes for Energy Proportionality | 1 | NO | 2012 | MICRO |
| 149 | NVMain: An Architectural-Level Main Memory Simulator for Emerging Non-volatile Memories | 1 | YES | 2012 | ISVLSI |
| 150 | Power-Supply-Network Design in 3D Integrated Systems | 1 | NO | 2011 | ISQED |
| 151 | Measurement, Analysis and Improvement of Supply Noise in 3D ICs | 1 | NO | 2011 | VLSI Symposium |
| 152 | The Memory System: You Can't Avoid It, You Can't Ignore It, You Can't Fake It | 1 | NO | 2009 | Book |
| 153 | An Optimized 3D-Stacked Memory Architecture by Exploiting Excessive, High-Density TSV Bandwidth | 1 | NO | 2010 | HPCA |
| 154 | Pragmatic Integration of an SRAM Row Cache in Heterogeneous 3-D DRAM Architecture Using TSV | 1 | NO | 2011 | IEEE TVLSI |
| 155 | Thermal Management of High Power Memory Module for Server Platforms | 1 | YES | 2008 | ITHERM |
| 156 | The Datacenter As a Computer: An Introduction to the Design of Warehouse-Scale Machines | 1 | YES | 2009 | Morgan and Claypool Publishers |
| 157 | Decoupled DIMM: Building High-bandwidth Memory System Using Low-speed DRAM Devices | 1 | NO | 2009 | ISCA |
| 158 | More is Less: Improving the Energy Efficiency of Data Movement via Opportunistic Use of Sparse Codes | 1 | NO | 2015 | MICRO |
| 159 | Improving Power and Data Efficiency with Threaded Memory Modules | 1 | NO | 2006 | ICCD |
| 160 | Multicore DIMM: An Energy Efficient Memory Module with Independently Controlled DRAMs | 1 | NO | 2009 | CAL |
| 161 | Multiple Sub-row Buffers in DRAM: Unlocking Performance and Energy Improvement Opportunities | 1 | NO | 2012 | ICS |
| 162 | Energy Efficient Data Encoding in DRAM Channels Exploiting Data Value Similarity | 1 | YES | 2016 | ISCA |
| 163 | Power Protocol: Reducing Power Dissipation on Off-Chip Data Buses | 1 | NO | 2002 | MICRO |
| 164 | Bus-Invert Coding for Low-Power I/O | 1 | NO | 1995 | TVLSI |
| 165 | DRAM-Aware Last-Level Cache Writeback: Reducing Write-Caused Interference in Memory Systems | 1 | NO | 2010 | Tech. Rep. |
| 166 | The Virtual Write Queue: Coordinating DRAM and Last-Level Cache Policies | 1 | NO | 2010 | ISCA |
| 167 | Conditional-Capture Flip-Flop for Statistical Power Reduction | 1 | NO | 2001 | JSSC |
| 168 | Error Control Coding: Fundamentals and Applications | 1 | NO | 2004 | Pearson-Prentice Hall |
| 169 | Decoupled Sectored Caches: Conciliating Low Tag Implementation Cost | 1 | NO | 1994 | ISCA |
| 170 | A Data Cache with Multiple Caching Strategies Tuned to Different Types of Locality | 1 | NO | 1995 | ICS |
| 171 | The Pool of Subsectors Cache Design | 1 | NO | 1999 | ICS |
| 172 | Exploiting Spatial Locality in Data Caches Using Spatial Footprints | 1 | NO | 1998 | ISCA |
| 173 | Accurate and Complexity-Effective Spatial Pattern Prediction | 1 | NO | 2004 | HPCA |
| 174 | SimPoint 3.0: Faster and More Flexible Program Analysis | 1 | NO | 2005 | MoBS |
| 175 | Using Storage Cells to Perform Computation | 1 | NO | 2014 | US Patent 8908465 |
| 176 | In-memory Computational Device | 1 | NO | 2015 | US Patent 9653166 |
| 177 | GateKeeper: A New Hardware Architecture for Accelerating Pre-Alignment in DNA Short Read Mapping | 1 | NO | 2017 | Bioinformatics |
| 178 | A Bit-Parallel, General Integer-Scoring Sequence Alignment Algorithm | 1 | NO | 2013 | CPM |
| 179 | Space/time Trade-offs in Hash Coding with Allowable Errors | 1 | NO | 1970 | ACM Communications |
| 180 | LazyPIM: Efficient Support for Cache Coherence in Processing-in-Memory Architectures | 1 | YES | 2017 | arXiv |
| 181 | Bitmap Index Design and Evaluation | 1 | NO | 1998 | SIGMOD |
| 182 | Improving DRAM Performance by Parallelizing Refreshes with Accesses | 1 | NO | 2014 | HPCA |
| 183 | Understanding Latency Variation in Modern DRAM Chips: Experimental Characterization, Analysis, and Optimization | 1 | YES | 2016 | SIGMETRICS |
| 184 | Low-cost Inter-linked Subarrays (LISA): Enabling Fast Inter-subarray Data Movement in DRAM | 1 | NO | 2016 | HPCA |
| 185 | Understanding Reduced-voltage Operation in Modern DRAM Devices: Experimental Characterization, Analysis, and Mechanisms | 1 | YES | 2017 | SIGMETRICS |
| 186 | Linux Device Drivers | 1 | NO | 2005 | O'Reilly Media |
| 187 | An Efficient and Scalable Semiconductor Architecture for Parallel Automata Processing | 1 | YES | 2014 | IEEE TPDS |
| 188 | Computational RAM: Implementing Processors in Memory | 1 | NO | 1999 | IEEE DT |
| 189 | Suppressing Power Supply Noise Using Data Scrambling in Double Data Rate Memory Systems | 1 | YES | 2009 | US Patent 8503678 |
| 190 | Programming the FlexRAM Parallel Intelligent Memory System | 1 | NO | 2003 | PPoPP |
| 191 | Processing in Memory: The Terasys Massively Parallel PIM Array | 1 | YES | 1995 | Computer |
| 192 | BitFunnel: Revisiting Signatures for Search | 1 | NO | 2017 | SIGIR |
| 193 | A Dichromatic Framework for Balanced Trees | 1 | NO | 1978 | SFCS |
| 194 | Error Detecting and Error Correcting Codes | 1 | NO | 1950 | BSTJ |
| 195 | Optical Image Encryption Based on XOR Operations | 1 | NO | 1999 | SPIE OE |
| 196 | ChargeCache: Reducing DRAM Latency by Exploiting Row Access Locality | 1 | YES | 2016 | HPCA |
| 197 | SoftMC: A Flexible and Practical Open-source Infrastructure for Enabling Experimental DRAM Studies | 1 | YES | 2017 | HPCA |
| 198 | One-Transistor Type DRAM | 1 | NO | 2009 | US Patent 7701751 |
| 199 | An Energy-efficient VLSI Architecture for Pattern Recognition via Deep Embedding of Computation in SRAM | 1 | YES | 2014 | ICASSP |
| 200 | GRIM-filter: Fast Seed Filtering in Read Mapping Using Emerging Memory Technologies | 1 | NO | 2017 | arXiv |