The Real Cost of Huang Renxun and AI
The Energy Costs Behind Computing Engines
(Image caption) Jensen Huang, CEO and co-founder of NVIDIA, is a key figure driving the expansion of AI infrastructure and computing power, grappling with global energy and grid challenges.
When the world talks about AI , why does it always circle back to NVIDIA?
In the narrative of the AI era, Jensen Huang is almost an unavoidable name. While the world discusses the future of large models, generative AI, intelligent agents, and general AI, people are superficially talking about parameters, architecture, and inference speed, but are actually discussing a more fundamental and less romanticized question: can computing power be supplied continuously, stably, and at scale? And on the supply side of this question, in the industry's rhythm from 2024 to 2026, the direction often points to the same thing—NVIDIA. When the Blackwell generation was pushed to the forefront, the energy efficiency narrative of GPUs was certainly dazzling, but the more jarring signal was power consumption and power density: a single B200 is generally understood in the industry to reach about 1kW or even higher, meaning that "computing power upgrades" are no longer just about stronger chips, but a systems engineering project that pushes power, cooling, and grid connection to their limits. When you put it on a macro scale, Jensen Huang's role is no longer just that of a chip company CEO, but someone who pushes the computing power curve to the point where the entire society has to supply electricity for it: engineers will follow his roadmap, investors will follow his shipping pace, and energy companies, utilities, and grid operators will be forced to follow the load curve he creates. This also explains why his most insightful statements in recent years are often not "faster," but "more power-consuming": not to create anxiety, but to remind the entire industry that the real boundaries of AI are shifting from algorithms and chips to the slower, more rigid, and more difficult track to be solved by short-term capital—energy and grids.
⸻
From "Components" to "Factory": How Huang Renxun Renamed the Data Center
In Huang's early vision, GPUs were more like high-performance computing tools, serving graphics rendering, scientific simulations, and specific acceleration scenarios. At that time, NVIDIA was more like a supplier of "accelerated computing," with GPUs being a component of servers—important but still subordinate to a larger IT architecture. However, after 2023, generative AI did not ignite a single product line, but a new form of production: model scale increased from tens of billions to trillions, training cycles lengthened from weeks to months, inference went from occasional calls to continuous 24/7 throughput, and computing power demand was no longer a "phased investment" but a "continuous consumption." At this juncture, Huang repeatedly proposed the concept of "AI Factory"—this is not rhetoric, but a structural renaming: data centers are no longer server rooms, but factories; the logic of a factory is continuous operation, stable supply, raw material input, and production capacity output; in an AI factory, the output products can be tokens, inference results, model capabilities, and corporate revenue, but the raw materials input are increasingly clearly pointing to energy. When an industry starts describing data centers as "factories," electricity is no longer just a line of cost in OPEX, but a prerequisite for "whether operations can begin." Simultaneously, cooling and water resources are no longer just operational details, but the lifeline for the long-term operation of the "factory." You'll find that what Huang Renxun truly rewrote wasn't just a speech, but how the entire industry understands data centers: from IT infrastructure to energy-intensive industrial facilities; from an internal corporate issue to a public infrastructure issue. This step is the key move he made to bring "energy" from behind the scenes to the forefront.
⸻
An unavoidable anchor point: IEA incorporates "electricity" into its AI world narrative.
Many people are accustomed to explaining AI's energy anxiety by citing "efficiency improvements": each generation of chips is more energy-efficient, each token costs less, and liquid cooling and interconnects are more advanced, so total energy consumption won't get out of control. But this is a typical "single-point efficiency" myth: when the number of computing power deployed and the frequency of use increase exponentially, the total effect will quickly devour the efficiency dividend. The International Energy Agency (IEA), in its "Energy and AI" report, gave the world a very hard anchor: global data center electricity consumption will be about 415 TWh in 2024, and will approach 945 TWh in its baseline scenario by 2030, clearly pointing out that AI is one of the most important factors driving the growth of data center electricity demand; this scale is equivalent to pushing data center electricity consumption to a "national" scale, rather than the scope that a few companies can absorb on their own through procurement strategies. More importantly, the IEA's narrative isn't just about "more electricity consumption," but about "the pace of electricity growth": data center electricity consumption is growing significantly faster than overall electricity consumption, meaning it will create supply and demand tensions more quickly in certain regions, bringing issues like power transmission, substation, interconnection queuing, and equipment supply chains to the forefront. When the IEA uses such public language to incorporate "electricity" into its AI narrative, it effectively provides macro-level support for Jensen Huang's statement that "energy is the bottleneck": you may not like the conclusion, but it's hard to deny the trend. From this moment on, the story of AI is no longer just about models and chips, but about energy systems and institutional timelines; and the task of the profile is to concretize this enormous systemic pressure into a real cost that readers can understand through the choice and language of a representative figure.
(Image caption) Rack-level AI systems are pushing computing power from "server components" to "industrial capacity": the higher the density, the closer the requirements for power supply, heat dissipation, data center design and grid connection time are to the level of power infrastructure.
⸻
Northern Virginia's Horror: When 60 Data Centers "Simultaneously Got Out"
If the IEA's figures depict "long-term trends," then the "Data Center Alley" incident in Northern Virginia reveals the true face of "systemic risk." During a voltage disturbance in the summer of 2024, approximately 60 data centers switched to backup power under protection mechanisms, causing a massive loss of load from the grid. This forced grid operators to urgently reduce generation to prevent cascading imbalances. Reuters reported that such incidents have heightened grid operators' concerns about data centers' "non-ride-through" behavior under voltage fluctuations, as it amplifies a localized disturbance into a systemic challenge. NERC also released an event review of the large load loss incident, explaining that data center loads are sensitive to voltage disturbances, and their protection and control logic often prioritizes preventing equipment damage, leading to a rapid load shift to backup systems. The key here is not just "power shortage," but "grid stability": data centers, like highly sophisticated and concentrated industrial loads, not only consume a lot of electricity but can also exhibit collective behavior under specific disturbances, placing enormous regulatory pressure on the grid in a short period. This risk becomes acute precisely after AI is industrialized: you can distribute data centers more widely, or you can demand stricter ride-through standards, but either choice involves equipment costs, reliability, local policies, and public-private partnerships. Therefore, you can understand that Huang's "gigawatt-scale AI factory" is not a boast of scale, but a reminder: when AI enters the industrial scale, it will naturally be written into the operating manual of the public power grid, becoming an object that must be regulated, governed, and included in safety assessments.
(Image caption) Northern Virginia (Data Center Alley) is one of the world’s most densely packed data center clusters: as computing power is rapidly stacked in such hubs, power access, transmission capacity and system resilience become hard constraints on the speed of expansion.
⸻
" 25 times energy efficiency" and "total energy output explosion": the gap between advertising and reality
NVIDIA's platforms claim to achieve orders-of-magnitude performance and energy efficiency improvements in inference, a narrative not uncommon in the industry: new precision formats, new interconnects, new system-level designs, coupled with liquid cooling, can indeed push energy efficiency per unit of computing power to higher levels. The problem is that as computing power moves from "laboratory-grade" to "industrial-grade," from kilo-calorie clusters to gigawatt-scale AI factories, the improvement in energy efficiency at a single point will be swallowed up by the explosive growth in deployment scale. The power consumption level of the B200, the power density of rack-level NVL systems, and the rising cost of liquid cooling equipment all point to the same thing: the next stage of competition in AI is not only a performance race, but also an engineering race to "handle heat and electricity well." Recent market reports and media coverage have even emerged regarding the cost of a single rack-level liquid cooling system, showing that the cooling system itself has become a cost item that can be priced, compared, and bottlenecked separately; and in large-scale cloud deployment scenarios, cooling solutions will also involve water resource strategies and sustainability commitments, resulting in trade-offs between "energy saving" and "water saving." While the outside world only sees "how many times faster inference," GFM is more concerned with: after large-scale deployment, how much electricity, heat density, cooling capital, and power supply queuing time are used to achieve this speed? What makes Jensen Huang special is that he did not permanently hide these costs behind the scenes; instead, he repeatedly used terms like "factory" and "gigawatt" to push the promotional language back to engineering language—because he knows that if costs and limitations are not clearly defined in advance, the industry will ultimately hit a wall at the most expensive point: not on the chips, but on the power grid, transformers, cooling, and system time.
⸻
Why the true costs will emerge in 2025–2026 : From CAPEX to "Electricity and Time"
The early costs of AI were indeed primarily capital-intensive investments: chip procurement, server clusters, R&D costs, and supply chain competition. These costs share a common characteristic—they can be rapidly rewritten within a 12-18 month technology iteration cycle: process advancements, architecture upgrades, and software stack optimizations all allow for higher computing power with the same budget. This leads to a habitual belief in the industry: when a bottleneck appears, just wait for the next generation. However, as deployment enters the large-scale phase of 2025-2026, the cost structure subtly shifts: electricity, cooling, infrastructure construction, grid connection, and interconnection queues jump from back-end operating costs to front-end rigid costs. More critically, these costs do not decrease rapidly with chip iterations; they are constrained by physical and institutional factors: power plants require project approval and environmental impact assessments, power transmission requires planning and land acquisition, substations require equipment and construction periods, and grid connection requires standards and coordination. Therefore, "buying more GPUs" is no longer the complete answer, because you might encounter situations like "being able to buy GPUs but not electricity," "having electricity prices but lacking grid connection capabilities," or "having land but having to wait years to connect to the grid." This is the essence of the emerging "real cost": it's not just a bigger bill, but also longer wait times and greater uncertainty. When grid operators and regulators begin discussing whether data centers can be reliably serviced and whether new reliability rules are needed, it means that the expansion of AI is no longer solely determined by the market, but will be reshaped by public regulations. Jensen Huang's role thus becomes more complex: he is both the engine of the computing revolution and a magnifying glass for energy constraints—he has brought the industry to a stage where it must face reality: the ceiling is no longer algorithms or silicon wafers, but whether you can find fuel for the factory and deliver that fuel to the door within regulatory timeframes.
(Image caption) Rack-level AI systems are pushing computing power from "server components" to "industrial capacity": the higher the density, the closer the requirements for power supply, heat dissipation, data center design and grid connection time are to the level of power infrastructure.
⸻
Institutional time lag: Chips are updated annually, while power grids undergo a ten-year cycle.
The chip industry is immersed in a high-speed iteration rhythm: a new architecture every 12-18 months, with performance doubling and energy efficiency leaping. This speed has shaped the psychological expectation of the entire AI industry—"Just wait for the next generation." But the energy system follows a completely different pattern: power generation and transmission are projects on a ten-year timescale. Environmental impact assessments, land acquisition, financing, equipment supply chains, and local political coordination—each of these can extend the construction period. This is the core contradiction you grasp within the "institutional time" framework: the computing power curve is an exponential curve approximating Moore's Law, but the power grid curve does not follow Moore's Law; it is locked by materials, construction, standards, and governance mechanisms. The Northern Virginia incident is symbolic precisely because it embodies the "time difference": data centers can expand rapidly in the short term, but maintaining grid stability amidst disturbances requires long-term institutional and engineering investment. When NERC presents the sensitivity of data center loads to voltage disturbances in a retrospective manner, it is essentially telling the market: no matter how much computing power is expanded, if grid resilience and operating standards are not upgraded in tandem, risks will accumulate in a non-linear way. Jensen Huang has repeatedly stated on multiple occasions that "the power grid can't be as fast as a chip," which isn't a complaint, but rather a shift in industry focus from a "purely technical narrative" to a "systems engineering narrative": Once AI enters the industrial scale, the real competition won't just be about whose GPUs are faster, but about who can deploy computing power on a more reliable, sustainable, and manageable energy foundation. For GFM, this means: we don't write about legends, but about the systems and realities they reveal—because that's the force that will determine the next decade.
⸻
From Corporate Issues to Public Issues: Why AI Is Inevitably Impacting National Security and People's Livelihoods
In the early days of AI computing power, energy issues were largely confined to the internal control of enterprises: signing power purchase agreements, negotiating priority power supply, building backup power generation, and implementing behind-the-meter solutions seemed to solve the problem. However, as computing power leaps from megawatts to gigawatts, and as the peak load of data center clusters becomes equivalent to that of a medium-sized city, energy issues inevitably transcend enterprise boundaries and become public concerns: fluctuations in residential electricity prices, the pressure on industrial electricity consumption, the reliability of local power grids, and power supply security under extreme weather conditions are all amplified. Reuters, in its report on the Northern Virginia incident, pointed out that the sensitivity of data centers to voltage disturbances and their large-scale switching behavior are forcing regulators and grid operators to rethink reliability standards; once these standards are rewritten, they will affect the design, cost, and deployment pace of data centers, thereby affecting the supply curve of AI. At a deeper level, there is national security: when AI enters military analysis, command and decision-making, and critical infrastructure management, the power grid becomes a "risk amplifier"—not because it is fragile, but because it carries too many systems that rely on real-time computing power. Once the power supply is interrupted or the power grid is attacked, the computing power advantage can evaporate within hours. Huang's position in this context is highly symbolic: on the one hand, he promotes denser and more powerful computing platforms, and on the other hand, he repeatedly reminds everyone that "AI starts with energy"—he has transformed a problem that could originally be digested internally by enterprises into a structural challenge that policy circles, congressional hearings, and national security assessments must confront. At this moment, biographical writing truly gains weight: because it is not just about recording a person, but about recording how a person pushes the entire world towards a reality that must be faced.
(Image caption) As AI computing power leaps from megawatts to gigawatts, the power demand of data centers is no longer just an internal engineering problem for enterprises, but begins to be equivalent to the public load of a city.
⸻
With the computing engine running at full speed, the fuel problem has become a question for civilization.
In mainstream media reports, Jensen Huang is often portrayed as a symbol of success: innovation, vision, products, stock price, and market capitalization. However, in the GFM narrative, Huang is neither an energy expert nor a policy official, but rather the person who pushed computing power to the point of forcing out three costs: he wasn't driving a single product, but an industrial machine pushing the foundation of societal energy to its limits; he wasn't releasing a single prophecy, but a pressure field forcing the public and private sectors to re-coordinate, reinvest, and redefine regulations. This also explains why this article isn't "anti-technology": on the contrary, it represents the most mature view of technology—acknowledging boundaries and facing costs are essential to achieving sustainable technological dominance.
As the AI narrative moves from its initial breakthrough phase to its more advanced stage by 2026, the core question is shifting: it's no longer just about "can we do it?" but "can we continue to do it?" Jensen Huang will continue to push for stronger architectures, denser racks, and higher interconnect throughput, and the market will continue to cheer for faster token production. However, as we progress, the deciding factor will no longer be which chip is faster, but who can obtain more reliable power supply, more controllable cooling, more manageable grid connection, and a deployment pace less delayed by institutional timelines. The events in Northern Virginia remind us that AI factories are not just increasing electricity consumption; they are rewriting the standards for grid operation and reliability. IEA predictions suggest that national-scale power consumption for data centers is a likely path. Various market signals also remind us that cooling, power supply, and grid connection are now being priced, compared, and become bottlenecks individually. All of these point to the same conclusion: AI is moving from "technological romance" to "systems engineering," and the ceiling of systems engineering is inevitably the energy foundation. GFM included Jensen Huang in this trio (Maiyue Cheng is the guardian of boundaries and institutional time, and Sam Altman is the accelerator of demand curves) precisely because Huang is the most crucial engine on the supply side: he has made it clear to the world that when computing power becomes a civilization-level infrastructure, the fuel problem is no longer background, but a reality that determines the course of the journey. Thus, the final question left for this era becomes simple yet serious: computing power is in place; are we ready to power it?