Per Stenström | Chief Scientist Officer & Co-Founder
ZeroPoint Technologies

Per Stenström, Chief Scientist Officer & Co-Founder, ZeroPoint Technologies

Per Stenström is Chief Science Officer of ZeroPoint Technologies and Professor at Chalmers University of Technology in Computer Architecture. His scientific production comprises four textbooks, more than 200 publications and about thirty patents. He has held several leadership roles in the Computer Architecture community such as being program and general chair for the flagship conference IEEE/ACM ISCA. He has worked for Sun Microsystems and had leadership roles in several deep-tech startups over the years. He is a Fellow of the ACM and the IEEE with the citation “for contributions to the design of high-performance memory systems”, of Academia Europaea, of the Royal Swedish Academy of Engineering Sciences, of the Royal Society of Arts and Sciences in Gothenburg and of the Asia-Pacific Artificial Intelligence Association (AAIA).

Appearances:



Future of Memory and Storage - Day 3 @ 15:25

Closing Keynote - AI Memory Compression and Quantization

AI workloads, especially KV cache for Transformer based model serving and embedding vectors and their indices for context engineering, are placing growing pressure on memory resources such as DRAM and HBM for holding KV cache data and NAND flash SSDs for holding embeddings and HNSW indices in vector databases. Memory capacity needed to look up embeddings inline with Generative AI conversations can exceed 100 terabytes of DRAM and NAND Flash, and active KV cache can exceed the parameter memory of a 3B Llama model at a 128K token context window. As we move to extract more embeddings to hold business and end user context in our vector databases, and as our multimodal models switch to thousands of dimensions and deal with more than a million tokens for coding tasks, the pressure on memory is driving up costs of AI infrastructure and ultimately subscription charges for AI-enabled services. In this session, we look at two techniques being used to address this issue: an architectural technique from ZeroPoint Technologies that can inline and transparently compress memory data by 2x even at cacheline grain using hardware engines and an algorithmic technique from Google DeepMind team that exploits properties of AI embedding vector data type to compress memory capacity by more than 4x.

last published: 23/Jul/26 12:15 GMT

back to speakers

 

TO EXHIBIT OR SPONSOR

 

TO SPEAK

 

FMS website sponsored by XCENA

 

Marketing & Press