Agentic AI at scale will require high-capacity, low-latency context memory to store and retrieve key value, or KV, cache data used in the inference workflow. New storage tiers and architectures have been proposed to solve this problem by balancing the reprocessing of inference queries against the retrieval of previously computed tokens. This new tier, positioned between the local GPU system SSDs and network storage, has been called “context memory” and allows long context tokens to persist in a large-scale storage array. This session with VAST Data, Solidigm and Supermicro describes the implementation, uses and trade-offs of the context memory tier.