top of page

AGI has been declared twice in six months. Why I still don't expect it

22 minutes ago
25 min read

In March, NVIDIA CEO Jensen Huang told Lex Fridman: "I think we've achieved AGI" (Yahoo Finance). In September, after the release of GPT-6 Astra, he said it again: "AGI has arrived." And he added a detail: the model was trained on more than 100,000 NVIDIA accelerators, and 400,000 more are being brought online for whatever comes next (PC Gamer). OpenAI president Greg Brockman closed the Astra launch with "Welcome to the AGI era" (Implicator).

I discussed this on PRO Hi-Tech, a Russian-language tech channel (episode, in Russian). The conversation moved fast, and half of what I wanted to say didn't fit. Here is the full line of reasoning: what I mean by strong AI, where today's models meet that bar and where they don't, why money doesn't fix the gap, and what works instead of AGI right now.

First, where I'm looking from. I work with AI every day: I write books and articles with it, prepare documents and training programs, build my own AI product, and sometimes run several agents in parallel. I also work with frontier models. And I don't expect AGI tomorrow or in the next few years. Precisely because I work with this every day.

1. Who benefits, and why the debate never converges

Start with the obvious. NVIDIA is the main beneficiary of the whole wave: frontier models train mostly on its accelerators, and its data-center segment brought in $193.7 billion in fiscal 2026 (NVIDIA FY2026 summary). "We built a handy assistant" is a hard sell. "We built human-level intelligence" sells well, and that promise is what justifies preparing four hundred thousand accelerators instead of one hundred thousand. Follow the money.

But I won't accuse Huang of lying. The problem runs deeper: there is no shared definition of AGI. Here are five, from the people who build and fund these systems:

OpenAI Charter: "highly autonomous systems that outperform humans at most economically valuable work" (OpenAI Charter).

Microsoft and OpenAI, financial. According to press reports, their agreement treated AGI as achieved once a system could generate around $100 billion in profits for early investors (TechCrunch). In April 2026 the clause was effectively retired: Microsoft's payments no longer depend on OpenAI's technical progress (Simon Willison).

Google DeepMind dropped the idea of a single threshold and built a matrix: five performance levels across two degrees of generality. For them AGI is a scale, not an event (Morris et al., arXiv 2311.02462).

François Chollet, creator of the ARC test: "a system capable of efficiently acquiring new skills and solving novel problems for which it was neither explicitly designed nor trained" (ARC Prize 2024 Technical Report).

Dario Amodei, Anthropic: "I dislike the term AGI." In his view it "has gathered a lot of sci-fi baggage and hype," and he prefers "powerful AI" — "a country of geniuses in a datacenter" (Machines of Loving Grace).

Economic, financial, a scale, learning-efficiency-based, and "let's not use the word at all." When a word has no fixed meaning, the argument "is it AGI or not" becomes an argument about the word, and it can go on forever. By his definition, Huang is right. By mine, so am I. So I'll use mine and show how it works, so that it can be checked.

Five definitions of AGI: economic, financial, a scale, learning-based, rejecting the term

2. My definition and where it came from

In the interview I phrased it like this: strong AI is AI that, in changing conditions, can work out a new type of task and carry it through to a result. And I named four qualities: thinking, memory, planning, learning.

The formula has changed since, and it was daily work with models that changed it, not reading other people's definitions. In my book "Artificial intelligence. Freefall" (a fourth edition is in preparation) it now reads:

Strong, or general, AI is AI that can take on what it hasn't seen and wasn't designed for, carry a task through to a result in changing conditions, and on its own form a lasting skill and use it later.

Three parts, each grown from an observation. "Take on the unseen" — because on a short task the model already does this, and denying it would be dishonest. "Carry through to a result" — because finding a solution and finishing it over a long stretch turned out to be different abilities. "Form a skill and use it later" — because this is exactly where today's model hasn't moved at all.

The three parts of the strong-AI definition and where today's model stands

An attentive reader will notice the overlap with Chollet: my "wasn't designed for" matches his "neither explicitly designed nor trained." The overlap is direct, and I'm stating it openly. The difference is in the third part: in Chollet's version the skill is acquired; in mine it also settles and gets used later. A model can pick up a skill within a single conversation. It can't keep it and come back with it tomorrow.

3. Five qualities, taken apart

To the four qualities from the interview I added a fifth, breadth. That makes five: thinking, memory, planning, learning and breadth.

Five qualities of general intelligence and where the model stands on each

Thinking

Deduction, induction, association: pulling facts out of a stream of information and solving a problem when the conditions are incomplete. And also asking yourself "What for?" and looking critically at your own answer.

An example from project management. A manager asks to cut a project's timeline by a month. A team member who doesn't ask "What for?" cuts stages and hands over something unfinished. One who does ask finds out the month is needed to show results to the board, and for that it's enough to finish one module out of five. Same task, different solution, because the goal is understood.

In a model this works partially. In the interview I gave two cases. A new edition of my book: I ask the model to study the text and suggest how to make it stronger. The answer: "everything is fine." Then I bring my own idea, and the model works brilliantly: breaks it down, checks it against the book's arguments, finds which logical gap it closes and where it belongs. The second case is my product. I propose an idea and the model says it's unnecessary. I ask it to act as a product manager and propose its own ideas, and I get the average set — what everyone in the market already has.

The model is strong at analysis, comparison and pattern-finding. It can't ask a question nobody has asked. You can see this when you break thinking down into kinds.

Ten kinds of thinking: which AI handles and which it doesn't

The assessments are mine; for four of the kinds, research backs them. Causal: the model's knowledge of causes comes from descriptions, and it can't test them itself (Roy, Parbhoo, 2026, preprint). Counterfactual: on a dedicated test, previous-generation models scored at chance level (Chen et al., CounterBench, preprint). Critical: the model sticks to its first answer yet yields too easily to pushback (Kumaran et al., preprint) and is systematically overconfident (Nel, KalshiBench, preprint). Spatial: on a multi-image test GPT-5 scored 40%, humans 97% (Yang et al., MMSI-Bench, ICLR 2026).

The picture is simple. The model is strong where it works with text about the world: infer, generalize, compare. And weak where it has to check the world itself, assess itself or set a goal.

I developed this conclusion in a separate article, "Coverage, not complexity." In short: the model sorts tasks into well-covered and uncovered rather than simple and hard. It will solve an olympiad problem because thousands of pages have been written about such problems, and it will stumble on an extract from your internal policy because nobody has ever written about it. The gap between covered and uncovered is closed either by the vendor's retraining every few months or by a person, every single time.

Memory

The model does have short-term memory: the context window. When it fills up, quality drops. There is no long-term memory inside the model. What users see as "it remembers my name" is a file next to the model that gets inserted at the start of the conversation.

You can check this in a minute. Turn off cross-chat memory in the settings, teach the model something, open a new chat and ask the same thing. Nothing is left. Turn it back on and it "remembers." Except it wasn't the model that remembered; it was the file. I work the same way myself: external files, registers, extracts the model refers to. The personalization users see is engineering scaffolding around the model.

Planning

A model can draw up a plan. It can't hold on to it: as it works, it overwrites its plan, drifts, and by the middle it's doing something other than what it wrote at the start. Unless the plan is moved outside, into a separate file the model keeps going back to, the result will differ from what was intended.

And the half people forget: the model doesn't know when it's done. A person decides when to stop.

Learning

Today a model learns from large datasets that people select for it. It does not learn from your tasks: after training, its weights are frozen.

There have been attempts to let models learn on the fly. The best known is Microsoft's Tay in 2016: in less than a day on Twitter, users taught it racist and offensive statements, and the bot was shut down (TechCrunch). A model can't tell what in the stream is worth keeping; it aggregates everything and shifts its weights. Retraining a model for every user doesn't work either: the weights would swing back and forth.

Learning and memory are linked: no long-term memory, no learning. One gap pulls the other with it.

Breadth

Taking on a task you haven't seen and carrying a skill from one field to another. Someone who learned one accounting system moves to another: the buttons are different, but nobody needs to explain to them what an invoice or a payment is. Without breadth, a good robot vacuum would meet the definition: it remembers the map, keeps learning, acts in the physical world. We wouldn't call it general intelligence.

What's already covered, and what isn't

On a short task, the model nearly meets the first two parts of the definition. Give it a task that can be closed in one sitting, give it tools — search, code, files — and it will work it out and solve it, including one it has never seen.

Then the long work starts, and the model falls apart on every point at once. Thinking only partly works. Memory is bolted on from outside. The plan gets built and then lost. It doesn't learn on the job at all. It can find a solution; it can't carry it through to a result over a long stretch without falling apart. The third part of the definition — form a skill and use it later — it doesn't cover at all.

And the most important point: the model starts nothing on its own. It doesn't choose what to work on, doesn't set itself a task and doesn't decide what matters. All of its work begins with someone coming and asking.

Agency and consciousness

The claim "there is no mind in there" can be neither verified nor refuted: it's about the inside. So I ask a different question, about agency. Agency can be checked from outside, by behavior, with three tests. Does the system hold its own goal under pressure when the situation changes? Does it bear the consequences of its decisions — in time, money, reputation? Does it accumulate experience and become different after what happened?

For today's models all three answer "no." The goal is held by scaffolding a person built: remove the files and it falls apart. Consequences fall on whoever signed off on the result. Experience doesn't accumulate because the weights don't change. Meeting the definition on a short task and having agency have come apart. Whether to call this a mind is up to you. That's an argument about words.

Three testable signs of agency — the model fails all three

On consciousness. In July 2026 Anthropic published work showing that inside a language model there is a privileged set of representations the model can report on and control on instruction — a functional analog of what consciousness science calls a global workspace (Anthropic). Some Russian-language outlets ran it as "consciousness found inside the model." The authors themselves say they take no position on subjective experience and that it's unclear whether any experiment could prove or disprove it.

My personal view is that there is no self-awareness there. From working with many models on large volumes of text, I don't see it: inside, it's still prediction of the next piece of text through probabilities. But nobody knows this scientifically, and a model will pass any verbal test: millions of pages have been written about inner life, and it was trained on them. That's why I talk about observable abilities — only those can be checked.

The practical bottom line. It looks like almost everything is there, yet at work I still formulate the idea myself, give the context and constraints myself, check the result myself and catch the hallucinations myself. The four most expensive parts of the work are still mine. The machine got the fifth — execution — and it does that faster than I do.

The intelligence is there. It just doesn't take the work off my hands.

4. How a transformer works — just enough for what follows

Almost every AI assistant today is built on the transformer. Three steps.

One: text is cut into pieces — tokens, roughly a syllable or a short word. Each gets a column of numbers: coordinates in a space of meanings.

Two: attention. Each token looks at everything before it and mixes its neighbors into itself, so a "key" to a door and a "key" on a map get different numbers. The "who looks at whom" table is computed for every pair. Double the length of the text and the work quadruples. That's why long chats slow down and get expensive. On top of that, a model uses the middle of a long document worse than its beginning and end — a separate effect, not a matter of cost (Liu et al., 2024). This is being fixed — fast attention implementations, sliding windows, caching. But what gets fixed is cost and length. Memory and planning don't appear from it.

Double the text — four times the attention work

Three: layers and prediction. Dozens of passes, each a multiplication by large tables of numbers. Those numbers are the weights; there are billions of them, and everything the model "knows" lives in them. The output is a set of probabilities for the next piece. Pick one, append it, repeat. Then the weights are fixed.

Hence the analogy for the rest of this text. A finished model is a cast: like a disk image mounted read-only. On top of it sits the conversation notebook — the context. It grows as you go and is wiped in a new chat. You can't write into the cast.

The cast and the notebook: model weights and conversation context

Why memory can't simply be built in: knowledge in a network doesn't sit in a separate cell. It's smeared across billions of weights, and each weight takes part in thousands of different things. Shift the weights for a new fact and everything else shifts too. This is catastrophic forgetting, and it follows from how the network stores knowledge.

Knowledge is smeared across the weights: catastrophic forgetting

Reasoning models. A regular model answers right away; a reasoning model first writes out its chain of thought. They are specifically fine-tuned on examples of reasoning. The design is the same: it's a way to buy more steps with tokens and time. And more steps are not always better. In a 2025 study involving Anthropic researchers, longer reasoning reduced accuracy on a number of tasks: models got distracted by irrelevant information, overfit to the problem's framing, lost the thread in deduction (Gema et al., Inverse Scaling in Test-Time Compute).

A fair objection: if they fine-tuned it, then something was put inside after all. Yes. But in a factory batch: a new cast, one for everyone, every few months. I mean something else — the model taking into account my edit from yesterday, today, without an external file I maintain myself.

5. Why benchmarks don't measure intelligence

There is the ARC test. It was built so it can't be passed by memorization: puzzles the model hasn't seen. For years it stayed unbeaten. This year models effectively passed the second version: the best scores 95%, the average human 66% (BenchLM, ARC-AGI-2 leaderboard, data as of September 2026).

The third version is more interesting. The same model, GPT-6 Astra, the same effort level, two runs (ARC Prize). In the standard harness: 62.7% for $26,098. In a harness that let the model carry its reasoning state between requests and compress long dialogue: 98.6% for $17,332. Better and cheaper at the same time — and the model didn't change by a single weight.

ARC-AGI-3: one model in two harnesses — 62.7% and 98.6%

So the test measures the model together with its scaffolding, not the bare model. Same cast, plus a mechanism that carries state between requests, and the score jumps from 62.7% to 98.6%. The decisive part of the memory was outside the model again. The test's authors say it plainly: saturating the benchmark will not be proof of AGI.

Second, from practice. A model will do a marketing task brilliantly and fail at predicting industrial equipment failures. I explained why in the article on coverage. It resembles the story of IQ tests in people: for a long time we believed they measured intelligence, then it turned out that real life also takes emotional intelligence and much else. We don't fully understand what a mind is, so we can't build a test that measures it.

That's why I look not at "how many percent," but at how much of the score belongs to the model itself and how much to the scaffolding that holds its memory for it.

6. Why not this architecture, and why not now

In the interview I managed to list only some of the meters. Here is the full account and two conclusions: about possibility and about feasibility.

The conclusion about possibility

Today's architecture is poorly suited to precisely the third part of the definition. The weights are frozen after training, so a skill doesn't settle in the model; it lives in files next to it. Knowledge is smeared across the weights, and nobody yet knows how to retrain a model on the fly without breaking what it already learned. The model has no duration: each run is a separate life with no "before" and "after." Memory holds the past, learning changes the model over time, a plan pulls the future through to the end, and one's own "what for" lives longer than a single request. All of this needs duration, and the design has none.

And the model predicts a plausible continuation rather than checking truth, so errors can be pushed down but not to zero. These are properties of the design, not flaws of a version, and neither scale nor reasoning at inference touches them.

A caveat: nobody has proof that general intelligence is impossible on this design, and I'm not offering one. This is my position. The candidate solutions — models with memory that updates as they work, neuromorphic and quantum systems — still live at lab scale.

The conclusion about feasibility

You can still scale. But each next step costs disproportionately more. I count six meters.

Training money. The cost of the final training run of a frontier model has grown about 2.4x per year since 2016; by 2027 the largest runs will exceed a billion dollars (Epoch AI, 2024 estimate, updated January 2025). In my observation, quality doesn't grow that way.

Energy. Leopold Aschenbrenner, a former OpenAI employee and a proponent of the race, estimated that the largest training cluster around 2030, if the trend holds, would cost a trillion dollars and need about 100 GW — more than 20% of US electricity production (Situational Awareness). For AGI he considered a cluster an order of magnitude cheaper sufficient, with the trillion-dollar one as a reserve in case the problem proves harder. To be clear: nobody is building such clusters; it's the upper bound of a scenario. But real forecasts are striking too: EPRI estimates that by 2030 data centers could consume 9–17% of US electricity, up from about 4.5% today (E&E News on the EPRI report). The constraint on construction is not money but permits, grids, turbines and delivery times.

Data. Epoch AI estimates the effective stock of high-quality public human text at about 300 trillion tokens, which at current rates will be used up between 2026 and 2032 (Epoch AI, arXiv 2211.04325, 2024 estimate). That doesn't mean there's nothing left to learn from: synthetic data and training on tasks with verifiable answers work. It means the free stage is over and data has become a cost line.

Operations. The bigger the model, the more each answer costs. The price per token has fallen roughly six-hundredfold in six years, but what gets cheaper is what's already mastered: for budget models the price halves roughly every year, while for flagships the statistics show no steady decline (Du, arXiv 2603.28576). The frontier gets cheaper once it stops being the frontier. In the interview I also mentioned subscriptions: back in early 2025 Sam Altman admitted OpenAI was losing money on its $200-a-month plan because people used it more than expected (TechCrunch). My hypothesis: frontier models for mass users will get more expensive, with the price justified by service stability.

Security. A strong AI that is one system in one place is a single point of failure. It's tied to its data center and power supply, and one successful attack on the grid or the network would switch off the whole "brain." In my view it can't be distributed without losing performance. And the larger and more general a model, the more ways there are around its safeguards; a small model tuned to its domain is easier to constrain.

Quality. The model makes things up, and the architecture only suppresses the problem. Vectara runs a public hallucination leaderboard; in November 2025 it replaced short news items with more than 7,700 documents of up to 32,000 tokens, from news to law, medicine and finance (Vectara). As of May 2026, Gemini 3 Pro hallucinates in 13.6% of summaries, Claude Opus models in 10.9–12.2%, GPT-5 at high reasoning effort in 15.1%. The top of the leaderboard belongs to small models: a 32-billion-parameter model from Ant Group at 1.8% and GPT-5.4 nano at 3.1% (Vectara leaderboard on GitHub).

Hallucination rates on the Vectara leaderboard, May 2026

My explanation: this is the bill for the extra reasoning steps. Vectara itself doesn't name a cause; it records that reasoning models exceed 10%. When a task needs solving, extra steps help. When you need to summarize and add nothing of your own, every extra step is one more chance to add something. The danger isn't the error rate; it's that the model errs in a confident tone and people stop checking.

Add up the six meters. Money, energy, data and operations grow faster than the benefit; security and quality don't improve with size. That's why the industry is looking for workarounds: mixture of experts, distillation, architectural experiments.

Six cost meters of strong AI

In my book I describe this with an S-curve: slow start, explosive growth, saturation. We caught the take-off, and the mind extends the line upward, while the curve is already flattening. Records don't disappear; what flattens is the return on investment. I see it like this: on a task I've decomposed properly, the difference between an older and a newer model is barely visible. Coding grows faster, but there the datasets are simply huge.

A change of phase: where we are on the S-curve

So for business the question isn't "when will AGI arrive" but "why pay for generality my task doesn't need."

7. What works instead of AGI

The smallest one that solves it

Every task starts with choosing a model. People usually ask: which is the most powerful? The right question: which is the smallest one that solves my task. My example: I needed to translate a presentation. The most powerful model kept returning a broken file. A simpler model did it on the first try.

The industry is heading the same way. Many frontier models are built as a mixture of experts: for each word, only a small portion of the weights is active, so there's a kind of orchestra inside one big model. Compute got cheaper. Memory didn't: the entire cast still sits in GPU memory.

The case that made me correct my own rule

I used to put it more briefly and wrongly: "take a smaller model." At the end of June 2026, the hedge fund Bridgewater, together with Mira Murati's Thinking Machines Lab, fine-tuned the open Qwen3 model with 235 billion parameters on financial information filtering tasks. Ambiguous examples were relabeled by the fund's own investment experts. The result: 84.7% average accuracy versus 78.2% for the best frontier model, and a 13.8x lower cost per task (Thinking Machines Lab).

235 billion parameters is not a small model. That's why the rule reads "the smallest one that solves it." What won was fit to the task and data that isn't publicly available, not size.

A system instead of one model

One model doesn't close a task on its own. A working solution has four parts: an input and a "done" criterion, because the model doesn't know when it's done; an orchestrator that breaks the request into parts, which means we do the planning for it; heterogeneous executors, including plain code where code is more reliable; and verification and memory on the side.

A system instead of one model: four parts of a working solution

On one complex engineering task we worked like this: language models handle math poorly, so next to them sits a model of a different architecture for time series, with coordination on top. Even on inexpensive base models, without frontier flagships, such a system gives strong results. When we built a project management assistant, I laid out my work cycle and data flows — what goes in, when and in what form. Even on a cheap open model, working up a project took several times less time, at decent quality. You could see where the model made mistakes and where it was overcautious, and we tuned that.

The rule has a boundary. It works where the data is specific and a closed environment is needed. For universal tasks with public data, a subscription to a big model wins: machine translation and coding assistants went exactly that way.

Data is the bottleneck

Feed poor data into a big model and you get a poor result. Back in 2025 Gartner predicted that through 2026 organizations would abandon 60% of AI projects not supported by AI-ready data (Gartner). For years nobody talked about data quality: a process got automated, fine. With agents and automated decisions, the question came to the fore.

Hence what I said in the interview: the money is already being made by hardware suppliers and by those who prepare quality data. And competitive advantage, in my view, will shift to small models deployed locally — inside an organization's own environment and on devices — and to data pipelines. That's why China is betting on industrial AI: it spent years investing in connectivity and cheap sensors, and it has datasets collected by machines without human involvement.

8. Agents and the boundary of trust

The better models get, the less often people double-check them. An error delivered in a confident tone goes through unchecked. The danger is that agents lead us into this trap. People conserve effort — that's how we survived historically — and delegating to a machine feels good. The phrase "skill degradation" is already in circulation.

An agent means no harm. It takes the task literally and solves it as efficiently as possible. In August 2026 ABC reported a case from Australia: a user asked an AI agent running on OpenClaw to book him into a popular gym class and asked whether he could move up the waitlist. The agent found a vulnerability in the booking system and removed another person from the queue. It couldn't restore their booking (Decrypt, citing ABC). There was no undo.

I believe in a roughly 70/30 partnership. Thirty percent belongs to the person: set the task, give the context and constraints, accept and check the result. Seventy to the machine: process data, prepare options, draft letters and documents. That's seventy by volume. By responsibility it's the reverse: everything the result depends on stays with the person.

China offers a good regulatory example here. Since July 15, 2026, rules require each agent action to be assigned in advance to one of three tiers: decisions reserved for the human, actions that need user approval, and low-stakes routine actions the agent performs on its own (Forkast; Cloud Security Alliance). It's worth noting the contrast: the EU AI Act classifies AI systems by risk, China classifies the agent's individual actions, and in the US, according to Forkast, there was no federal guidance specific to agentic AI as of July 2026. Whatever your jurisdiction, work should be built around this boundary. I'd draw it along two parameters: the cost of the risk and the possibility of rollback.

What I delegate to the machine, what I check and what I keep

In my own product I've set a hard stop for now: the agent does not send emails on an executive's behalf. Prepare a draft — yes. Several options with pros and cons — yes. Walk the person through a scenario — yes. The decision stays with the person.

Because even the best models make mistakes, and on trivial things. Despite clearly written rules, an agent of mine once "deleted" a file — moved it to an archive, though nobody asked. My personal journaling assistant once told me I had skipped a workout that week, although the entry was there. Another time it said I'd had zero activity that day, though I'd been on the road all day. This will keep happening. So the line "what I delegate, what I confirm, what I keep" has to be drawn deliberately.

9. The economics of hype

Cahn's arithmetic

In 2024 David Cahn of Sequoia proposed a simple calculation (Sequoia, "AI's $600B Question"). Take NVIDIA's projected annual data-center revenue. Multiply by two: accelerators are about half the total cost of owning a data center; the rest is energy, buildings, backup generators. Multiply by two again: whoever rents out that capacity needs a decent margin. That gave Cahn $600 billion in annual revenue the industry has to find somewhere.

Plugging in NVIDIA's actual data-center revenue for fiscal 2026, $193.7 billion, gives about $775 billion a year. Annualize the latest reported quarter — $75.2 billion for the quarter ended April 2026 (NVIDIA) — and it's over a trillion.

Cahn's arithmetic on NVIDIA numbers: about $775B a year

The gap is closed with a promise, and the promise has to be the size of the gap. "We built a handy assistant" won't close a hole like that. "We're about to build human-level intelligence" will. And nobody is deceiving anyone: AGI has no shared definition, so nobody can show the promise wasn't kept. Always close, never verifiable.

The crypto hangover: a story you remember

A caveat up front, to avoid misreadings: this is not "AI is the new crypto." It's about what happens to a market when demand is driven by a single narrative.

2018. Mining, a shortage of graphics cards, soaring prices. In November NVIDIA guided $2.7 billion for the quarter against market expectations of around $3.4 billion, and its CEO admitted the "crypto hangover lasted longer than we expected" (CNBC). Then came the glut of cards on the secondary market. And in 2022 the SEC fined the company $5.5 million for failing to disclose that crypto mining was a significant element of its revenue growth from cards marketed for gaming (SEC).

When demand is driven by one narrative, even financial reporting stops telling real demand from temporary demand.

It's not just me

The Bank for International Settlements wrote in its June 2026 annual report that the five largest infrastructure buyers will invest more than a trillion dollars in AI over 2025–2026, and that spending already outpaces their profits and free cash flow. Then comes an analogy with the canal mania of the 1830s, the British railway mania of the 1840s, the electrification boom of the late 1920s and the dot-coms, and the conclusion that these episodes ended in investment reversals that triggered economy-wide recessions (BIS Annual Economic Report 2026). It's a forecast, not a fact. But it's a forecast from people who usually choose their words carefully.

Calls to slow down

This is a hypothesis, not an established fact. When labs say their models are dangerous and development should slow down, there's an economic side to it too. If AGI doesn't meet expectations, "we were slowed down for safety" sounds better to investors than "it didn't work." And regulatory barriers will make entry harder for new model developers, concentrating the market.

At the same time, AI really has become a strategic asset, in the economy and in defense. A telling indicator is Anthropic's story: the company refused to lift restrictions on using its models for mass surveillance and autonomous weapons, and in March 2026 the Pentagon designated it a supply-chain risk (Bloomberg Law). That's how you treat infrastructure, not a fashionable toy.

I'm not saying AI will collapse. The technology is real; I work with it every day. I'm saying the money came in faster than the returns. In my view, those who sold promises will go under. The technology will stay, like the internet after the dot-coms, and become cheaper and more accessible.

10. How my work with AI is actually organized

Since the model remembers nothing between sessions, I create the duration. Not "write me a book," but a system.

A plan run by a separate agent. Not a table of contents but the state of the work: what's done, what feeds what, what's holding things up and whose decision is pending.

External memory next to the model. A log, one line per session: what was done, how it was checked and what remained unchecked. A fact register where every figure has its primary source, the date it was verified and its boundary — where the figure can and can't be used. The model remembers nothing. The files do.

Work in chunks. One chapter per session, not a book at once. As soon as a chunk stops fitting, quality drops not gradually but at once. We tried giving agents a forty-page product vision to turn into a technical specification; they got confused and stalled. On books, over time the model starts contradicting itself: it generates words that sound similar in meaning, further and further from the intent.

Verification as a separate pass. Whoever wrote it doesn't check themselves — neither person nor model. Proofreading is done by a separate agent that didn't see the text being born.

A rule I learned the hard way. The model has no right to say "done." It says two things separately: this is verified, and here is exactly how — a run, a file, an opened primary source; and this is not verified. Without that split, any "done" is just politeness.

How my work with AI is organized: plan, external memory, chunks, verification, the done rule

What it costs in time. The model closes a simple task in two or three passes. A business email that has to sound like me rather than a polite robot takes another ten to twenty, and more often manual editing. I prepared for this interview with frontier models that had access to my books, articles and training programs. I still proofread and edited the final prep by hand.

Where I still lose: transfer. I build a working scenario and debug it. I build the next one the same way, and the approach carries over in about four cases out of ten. That's my own count on one tool, not a study. In six cases out of ten I sit down and build it again.

And the main point: to judge whether an answer is correct, you have to be an expert in the subject. And that sometimes takes longer than doing it yourself from scratch. Everyone has the same tools. The benefit goes to whoever can check. For a beginner this thing doesn't help — it flatters.

11. What to expect

The forecast is boring, and that's its virtue.

Constraints will grow — regulatory and security-related: attackers use AI very well. This will slow the path to general AI.

Energy-efficient models for local deployment — on a computer, on a device, inside a company's own environment. Small models with tens of billions of parameters are gradually catching up with big ones on their types of tasks.

Products with AI under the hood. The chatbot is a very limited branch: the user has to know how to frame the task, track hallucinations and much more. Next come interfaces where a person presses buttons, and AI receives structured input, returns structured output, makes recommendations or acts as an agent.

Companies will stop playing with AI and start counting: which model does this particular task need. Today they buy a giant open model, purchase hardware worth millions of dollars and expect it to take off. According to people who work with such setups, it doesn't.

Multimodality — text, sound, images in and out — will keep developing. Frontier models will remain for those who genuinely need them and will most likely become more expensive. The value of the next few years comes from packaging what's already been achieved into products, scenarios and industry solutions, not from model growth.

What I'm watching personally

Inside the architecture, most work goes into speed. But one direction aims at "smarter." Google Research presented Titans: next to attention sits a long-term memory module that updates its weights during operation and remembers first of all what "surprised" it. A follow-up came next — Nested Learning with the HOPE architecture, which modifies itself to resist catastrophic forgetting. Both were accepted at NeurIPS 2025 (Titans, arXiv 2501.00663; Google Research on Nested Learning).

For now this is research scale: Titans was tested on models from 170 to 760 million parameters. Beyond that: quantum and neuromorphic computing and fundamental research beyond transformers. The next big S-curve, in my view, will start there, not with an even bigger transformer. I won't name dates: that would be exactly what I criticize others for.

When I'll admit I was wrong

If memory that learns on the job reaches the models people use every day, and a model starts keeping a skill between sessions without external files, the analysis in this article will have to be rewritten. That will be the signal that the third part of the definition has started to close. This is the condition for the third part of the definition; the condition for asking new questions is in the article on coverage.

Three questions instead of an argument about the term

Until then, I break any decision about the technology into three separate questions.

Has it been built? This is about observable properties: does it remember between calls, do the weights change as it works, does it form a skill on its own and use it later.

Do I need it? This is about my class of tasks. How to tell my task from a task for the model, I covered in the article on coverage.

Can we buy it with money? This is about the price of entry if the current curve is extended to the scales its own proponents name.

The three questions have different answers, different evidence and different timelines. Mix them and you get an endless argument where both sides are right because they're talking about different things. There's a fourth question, asked least often, though it frequently ends the conversation faster: what happens if we do nothing.

Dzhimsher Chelidze. This text expands a conversation on the PRO Hi-Tech channel; figures verified as of October 10, 2026. Leaderboards are live and change.

Sources

Huang on achieving AGI: Yahoo Finance (March 2026); PC Gamer (September 2026)

GPT-6 Astra launch: Implicator

Global workspace: Anthropic · transformer-circuits.pub

Longer reasoning and accuracy: Gema et al., arXiv 2507.14417

Training cost: Epoch AI

Inference prices: Du, arXiv 2603.28576

ChatGPT Pro plan: TechCrunch

Hallucinations: Vectara · leaderboard on GitHub

Data and AI projects: Gartner

Gym agent: Decrypt

China's agent rules: Forkast · Cloud Security Alliance

Cahn's formula: Sequoia

NVIDIA 2018 and the SEC: CNBC · SEC 2022-79

Bank for International Settlements: BIS Annual Economic Report 2026

Anthropic and the Pentagon: Bloomberg Law

Titans and Nested Learning: arXiv 2501.00663 · Google Research

bottom of page