Alphabet’s Google unveiled a new set of AI compression algorithms designed to dramatically reduce the memory footprint of large language models, triggering a selloff in memory and storage chip stocks on Wednesday.
The new algorithms TurboQuant, PolarQuant, and Quantized Johnson-Lindenstrauss (QJL) target the “key-value cache,” a system that stores frequently accessed information during AI inference. Google said TurboQuant compresses this cache to just 3 bits without additional training or fine-tuning, achieving up to an eightfold speed increase in attention computation on Nvidia H100 GPUs compared with uncompressed versions.
TurboQuant operates in two stages. PolarQuant converts standard data vectors into polar coordinates, simplifying the way data dimensions are represented. QJL then applies a 1-bit error-correction layer using the Johnson-Lindenstrauss Transform to minimize information loss while maintaining model accuracy. Together, the methods significantly cut memory requirements and computational costs, which Google described as “optimally addressing the challenge of memory overhead in vector quantization.”
Shares of major storage manufacturers declined following the announcement. Micron, Western Digital, Seagate, and SanDisk all slipped in trading, while equipment suppliers Lam Research and Applied Materials also moved lower. Analysts at Morgan Stanley called the technology a “breakthrough reshaping the cost curve of AI deployment,” noting it benefits cloud providers by reducing inference costs. The bank said the long-term impact on hardware demand could remain neutral or slightly positive as efficiency gains may enable broader AI adoption.
Google plans to present TurboQuant at the International Conference on Learning Representations in Rio de Janeiro in April and PolarQuant at AISTATS 2026. The research was led by Google’s Amir Zandieh and Vice President Vahab Mirrokni, in collaboration with KAIST and New York University.
Read Article: Meta and YouTube Found Negligent in Social Media Addiction Lawsuit

