Agent AI is changing the underlying logic of traditional hardware cycles and could potentially turn infrastructure supply constraints from short-term fluctuations into structural constraints lasting several years.
Dell Technologies COO Jeff Clarke recently stated during a meeting with Morgan Stanley that agent AI is distinctly different from past hardware cycles. If related AI workloads continue to grow, the shortage of key components may persist for over five years, with memory and HDD being his biggest concerns at present. In contrast to previous hardware supply squeezes that typically lasted several quarters, this round of demand growth could lead to a much longer-term supply-demand gap.
After the meeting, Morgan Stanley reassessed the current AI infrastructure cycle. Analysts pointed out that the eight booms and busts in hardware over the past 40 years were largely driven by device refreshes and capital expenditure fluctuations, whereas the critical variable this time is whether AI workloads can keep expanding. In other words, what Dell's executives emphasize is not just another hardware replenishment cycle, but a potential long-term infrastructure expansion fueled by the rise in agent AI demand.
The report states that Clarke believes this AI cycle is fundamentally different from the hardware booms and busts of previous decades. The past eight hardware cycles—including servers, storage, PCs, etc.—were primarily driven by device penetration and refresh rates, with demand rising and falling according to upgrade rhythms, but the overall hardware market itself did not see substantive expansion in scale.
For example, during the pandemic, the human-to-PC ratio temporarily increased, but as society reopened, PC demand subsequently declined, leaving the overall market size largely unchanged from the pre-pandemic period. Agent AI, however, operates on a different logic. Clarke believes agent AI is decoupling cognitive output from human input, allowing companies to accomplish more without increasing headcount, resulting in productivity gains of 10x or even 100x.
In this process, AI is no longer just assisting humans in completing tasks—it is itself becoming a “productivity tool” that continuously consumes compute resources. Productivity gains will drive more companies, regardless of size, to begin AI transformation, thereby widening the TAM (Total Addressable Market) of infrastructure.
Dell anticipates that by 2030, datacenter compute power will add 200 gigawatts, ZettaFLOPS will rise 5x to 830, and the growth in inference tokens will be exponential. The number of tokens required by agent AI to complete a unit of work is far higher than that for basic chatbots, and Clarke expects inference token generation to rise 87x by 2030.
A core question in this AI infrastructure cycle is whether rapidly growing token demand will mostly flow to cloud platforms rather than on-prem infrastructures.
Clarke's view is that agent AI workloads will take on a hybrid deployment form. Enterprises will continuously adjust the location of inference workloads between public cloud and on-prem environments, depending on security and cost-effectiveness.
Using Dell as an example: the company will deploy content-related workloads in public cloud but will not move proprietary source code or telemetry data out of its own facilities. This means that AI workloads do not have a single deployment path, and on-prem infrastructure will continue to bear part of the growing compute demand.
In this framework, exponential token consumption growth does not necessarily imply an either/or trade-off between cloud and local infrastructure. Clarke believes the proliferation of open source weighted models will further drive local infrastructure investment, since enterprises can optimize model output for cost within their own environments.
Currently, traditional server shipments are still declining year-over-year, which seems inconsistent with the explosive growth in inference tokens. Clarke explains that sharply increased server performance density is suppressing device shipment numbers.
Dell's 17th/18th-generation servers can replace up to 13 traditional 14th-generation servers, so in the short term, even if server shipments fall, the value and capacity of single devices are still rising.
As datacenters gradually complete architecture restructures centered on accelerated computing, and the spread of agent AI further drives CPU server demand, server shipments could eventually return to growth. Clarke even claims that if inference tokens really grow 87x in five years, traditional server shipments could see exponential changes as well.
Storage, too, will be continuously driven by agent AI. Every operation performed by agent AI—such as memory retention and artifact generation—creates ongoing data storage needs; with the large-scale adoption of KV Cache, storage capacity needs will increase further.
Compared with macroeconomics, geopolitical conflicts, datacenter overbuilding, and power shortages, Clarke lists supply issues as the greatest current concern, describing the supply climate as “We’re in neverland.”
Morgan Stanley points out that, historically, bulk commodity supply cycles usually last 2 to 4 quarters, mainly affected by booms and busts in hardware refresh cycles and supply-side decisions. But Clarke believes that if the growth logic of gigawatts and inference tokens stands, then key components like memory and HDD are no longer just in an ordinary short-term cycle.
Therefore, Dell is currently operating on the assumption that key components may see supply shortages that last for years, especially focusing on memory and HDD. In other words, the core risk of this round of supply constraints is not a short-term insufficiency, but the possibility that, as AI demand continues to expand, the supply side will need several years to catch up.
Clarke also states that Dell has an advantage over peers in supply management, an edge that is translating into share gains across servers, storage, and PC market segments.
Ongoing shortages of key components are also changing the pricing environment for the infrastructure industry.
Morgan Stanley previously attributed part of the gross margin expansion in servers and storage to “profit stacking”: as costs for components like memory rise, vendors add further profit layers. Clarke, however, offers a different explanation.
He acknowledges that profit margins for similar server and storage products are indeed expanding, but sees two structural factors at play: First, when component supply is limited, companies prioritize allocating scarce resources to the highest-margin products; second, as the infrastructure market continuously expands, the pressure to compete for incremental customers drops, reducing the need to cut prices to win new business.
Therefore, margin gains are not just short-term opportunistic profits from supply-demand mismatches, but also reflect a changed pricing benchmark after sustained infrastructure demand expansion.
If agent AI workload growth continues, the core variable in the AI infrastructure cycle will shift from “device refresh” to “incremental demand.” This means demand for compute, servers, and storage may continue to expand even longer, and supply constraints for key components such as memory and HDD may turn from short-term cyclical fluctuations into multi-year structural issues.