AI Computational Inefficiency: Data Bottlenecks in von Neumann Architecture
Large language models (LLMs) must frequently access memory every time they generate a word. Professor Chung Dae-sung explains that the physical distance in the von Neumann architecture, where computation and storage are separate, causes data bottlenecks and power waste.
AI large language models face a structural limitation in that they must frequently access memory every time a word is generated.
The physical distance between the CPU and memory is very far in the electronic world, leading to data bottlenecks and power waste.


