Nidish Kamath is the Director of Product Management for Memory Interface IP at Rambus. He previously held marketing and product management roles at AMD, Kioxia (formerly Toshiba Memory), Avalanche Technologies, Brocade and Qualcomm, where he worked on computational storage, SmartNICs and GPU cluster networking solutions. He has served in various standards and industry associations such as SNIA, Center for Open Source Software (CROSS), CXL Consortium, UEC and JEDEC.
AI inference spans diverse workloads, from low‑latency chat to long‑context reasoning and large‑scale recommendations—making single, monolithic accelerator and memory designs increasingly inefficient. This talk explains how inference naturally splits into prefill and decode stages with fundamentally different bottlenecks: prefill is compute‑bound, while decode is dominated by memory bandwidth and latency. By matching memory technologies to each stage, using cost‑efficient GDDR or LPDDR for prefill and reserving premium HBM for decode, with pooled memory for KV offload, operators can significantly reduce cost per token without sacrificing latency. The session outlines emerging disaggregated architectures for AI inference workloads.