top of page

Beijing's Ten Measures on AI Agents: What a Result Costs and Who Pays When There Isn't One

While the market argues about whether AI will replace employees, Beijing has published a document that answers a different question: what does one result produced by an AI agent cost, and who pays when there is no result?

On July 21, 2026, four agencies of the city administration signed — and on July 23 published — the Several Measures to Accelerate Leading Development Driven by Intelligent Agents. The document is short — ten points. But for the first time in a regulatory act it describes not the technology but the economics of agents: how to measure their effectiveness, how to sell them, and who answers for the outcome.

The vocabulary deserves a separate mention. Terms that a year ago lived only in research papers and engineering forums made it into an official text: Agentic AI, harness engineering, AI operating system, forward deployed engineer, one-person company, token economics, TaaS, AaaS, RaaS. This is the rare case where a government document reads like a specification for a new industry.

What follows is a read of the substance, the link to the two earlier Chinese documents I have covered, an honest list of strengths and weaknesses, and the conclusions we drew for ourselves as a team building an AI product.

The third layer of the Chinese construction

This document did not appear out of nowhere — it is the third in a series, and its meaning is only visible alongside the two before it.

The first layer is the State Council's "AI+" plan, which I covered separately: targets for 2027, 2030 and 2035, six application priorities, eight systemic pillars. That is the answer to "what and why" on a ten-year horizon.

The second layer is the "1397" strategy: one goal, three pillars, nine directions, seven implementation mechanisms. That is the answer to "with what and how" — compute, data, models, talent, the mechanics of adoption.

The Beijing measures are the third layer, and the most practical one: what all of this earns money on and who carries the risk. Neither earlier document contained any economics at all. There was a technology frame, there were application sectors, there were pillars — but no answer to how the cash flow between provider and customer works, and what happens when the agent fails. Here that answer appears. The city acts as the testing ground: run it in on one territory first, replicate later.

There is also a fourth, intermediate document that the measures reference directly: the Implementation Opinions on the Standardized Application and Innovative Development of Intelligent Agents, issued by three national agencies in May 2026. That is where the boundaries of an agent's autonomy are set — which decisions it makes itself, which require human authorization, and which it may execute autonomously.

Points 1 and 2. Foundations: task completion and the harness

The first point is about models, and its key phrase is "steadily raise the ability of large models to carry real tasks through to completion." Not benchmark scores — completion.

The difference is fundamental. An agent receives a vague goal, fills in the missing detail itself, breaks the work into steps, launches the right programs, reads the result and corrects course — dozens of times in a row. Scoring a single step says almost nothing about whether the whole chain succeeds.

The arithmetic is simple. Suppose an agent takes a hundred steps and is right on 99 out of 100 of them. The probability of getting through the whole chain without a single error is about 37%: two attempts out of three break down somewhere in the middle. Raise per-step accuracy to 99.9% and nine tasks out of ten reach the end. Drop it to 95% and fewer than one in a hundred does.

Chart: probability an AI agent completes a chain of steps without a single error at 95%, 99% and 99.9% per-step accuracy
Probability of getting through the chain without a single error, at different per-step accuracy

Which explains why programming was the first area where agents genuinely took off: everything is digital, any change can be rolled back, and the result is verified automatically by tests. Three conditions that most business processes simply do not have.

The second point is about harness engineering. The harness is everything surrounding the model that turns its capabilities into a stable result: what data goes in, what the agent is allowed to use, what it remembers between steps, what it does after a failure, and who checks the output. The model is the engine; the harness is the rest of the car.

The placement of this point is telling: the harness sits above applications and industry scenarios, immediately after models. The text calls for work on how an agent holds a task over a long distance, how several agents divide work between them, and how the system withstands rising load. And a direct statement of the goal: lengthen the chains of complex tasks and increase the stability of their execution.

Diagram: the harness around a language model - context management, tools, memory, access rights, retries, result verification
The agent harness: tools, memory, access rights, retries after failure, result verification

A note on orders of magnitude. In the study Agentic Harness Engineering (April 2026) the base model was not changed at all — only the harness was. Ten rounds of refinement raised the share of successfully solved tasks from 69.7% to 77.0%, and on other model families the gain ranged from 5 to 10 percentage points. That is the kind of gain people normally try to buy by switching to a more expensive model.

The second half of the point is about compatibility: common exchange protocols between agents, opening up key components, a shared skills market and stores of ready-made agent solutions. The inter-agent interaction protocol is explicitly named a candidate for a national standard. For a product owner there is an uncomfortable question hidden here: once agents have their own skill store, competition for the customer shifts. What matters is no longer your position in the app store, but whether somebody else's agent finds you and connects you as a skill by default.

Points 3 and 4. Software gets rewritten, the agent moves into hardware

The third point introduces the notion of "demand intelligence." The idea: the user names the outcome, and the system decides for itself which steps to take, which programs to call and with whom to coordinate — and answers for the final result. Hence the idea of "super software": its strength is not the number of features but the ability to work across the boundaries between existing programs. Reference scenarios are prescribed in six areas: science, medicine, education, public administration, manufacturing, culture.

But the most interesting thing for a practitioner is a separate role the document formalises: the forward deployed engineer. This person goes out to the customer, works through the processes, connects the systems, cleans the data and brings real problems back to the developers. That is state policy acknowledging that in the agent economy value is created on the last mile of deployment, not in the lab. Anyone who has deployed AI inside a living company understands why that line is there.

The fourth point moves the conversation from software to devices. Agent capabilities are to be embedded in next-generation smartphones, smart glasses, headphones and wearables, robots and cars. The key phrase is "the five-in-one unity of chip, model, cloud, terminal and application": hardware makers and model developers should define the product together rather than assembling it from ready-made pieces.

The point ends on an unexpectedly down-to-earth note: qualifying new products are included in appliance trade-in and digital-device upgrade programs. In other words, demand for agentic gadgets is stimulated directly — by a subsidy to the end buyer. That is probably the most pragmatic move in the whole document: don't persuade the market, pay part of the first purchase.

Points 5 and 6. The one-person company and payment for outcomes

The fifth point is the most interesting one for an entrepreneur. Beijing officially backs the OPC — the one-person company — as a new model of entrepreneurship. And this is not a slogan: it promises flexible access to compute, dedicated incubation, mentoring, preferential financing, help with intellectual property, communities of such companies and full-cycle service centres. Plus something very concrete — one-click registration with automatic generation of a standard charter.

The economic logic is transparent. A firm exists because coordinating work internally is cheaper than going to the market every time to find contractors, negotiate and supervise. Agents reduce both the cost of execution and the cost of coordination at once. Add a back office supplied as public infrastructure and the minimum viable size of a business falls. This is a direct state bet that the structure of the economy will change from the bottom up.

The sixth point is about money, and it contains the heaviest sentence in the document: a shift from pricing by volume of tokens consumed to pricing by value is encouraged.

First the document deals with cost: specialised chips for agent workloads, reuse of previously computed results, automatic routing of simple requests to a cheaper model. All in service of a single indicator — how much useful work you get per unit of compute spent. Then it turns to how to sell it. Three models are named.

TaaS — tokens as a service. What is sold is raw material: compute and access to a model. Competition on price, speed and stability.

AaaS — agent as a service. What is sold is a ready-made worker for a class of tasks: model, tools and harness in one package.

RaaS — result as a service. What is sold is the outcome: was the contract signed, was the bug fixed, was the lead converted.

Diagram: three monetization models for AI agents - TaaS, AaaS and RaaS
The higher the tier, the closer to customer value — and the more risk sits with the provider

The higher the tier, the closer to the customer's value — and the heavier it is for the provider: they take on the risk of task failure, of retries, of swings in unit cost and of compensation. The point closes with a measurement system: industry associations and leading companies are tasked with building quality and efficiency indicators instead of simply counting what was spent.

Points 7–10. Safety, resources, export

The seventh point records a change in the nature of the risk. A chatbot produces text; an agent gets the right to act — edit code, initiate payments, control devices, write in the user's name. The risk shifts from "said the wrong thing" to "did the wrong thing." The regulator's answer is category-based supervision: different requirements for different classes of agents depending on what they put at risk. Plus verification infrastructure: test ranges, isolated environments, guard models that watch the working ones, and shared vulnerability-scanning services.

Categorisation works not only as a constraint but as an accelerator: low-risk tasks can be tested and released quickly, high-risk ones go through strict review. Here safety is simultaneously a barrier and a ticket to the market.

The eighth point is about resources. It announces the "Galactic Computing Corridor" project: dedicated infrastructure for agent workloads. The logic matters more than the details. Training a large model is a one-off construction project. Running agents is a permanently operating production network answering a mass of tasks around the clock, with sharply fluctuating load. That is different economics and different hardware. Hence also compute and token vouchers for small companies, "token factories," and instructions to banks and insurers to develop financial products for agent adoption.

The press conference of July 23 also produced numbers: priority projects are selected competitively, with support of up to 100 million yuan per project; in the first half of 2026, funding for Beijing's AI sector exceeded 95 billion yuan.

The ninth point covers the external contour: a cooperation centre with SCO countries, taking models and agents abroad on a "one country, one strategy" basis, and open-sourcing protocols and frameworks. The calculation is clear: in a world of agents the network effect is exceptionally strong, and whoever's protocol is adopted controls the point of entry. The tenth point is organisational: who coordinates, where the money comes from, how projects are selected. Only one thing about it is interesting — there is a real budget underneath everything listed, not a declaration.

What it looks like from the inside: our experience with BAEOS

This document is more interesting for us to read than for an outside observer, for one reason. We are building BAEOS — a personal AI operating system for executives (directors and their deputies) and owners of small and midsize businesses. And on almost every engineering point we made the same decisions, only on our own and through our own mistakes, with nobody instructing us from above. The coincidence is telling: this is not Chinese specificity but a property of the technology itself.

One agent cannot carry it — tasks are split across many. We do not have one universal assistant but a set of specialised agents, each with its own zone, its own set of accessible data and its own result check. The reason is exactly the one described in point one: the longer the chain, the lower the chance of reaching the end. Splitting a task into short verifiable segments is not architectural aesthetics but a way not to lose two attempts out of three.

The model is a replaceable component. For us this is written down as an architectural prohibition: no business logic tied to a specific frontier model, and a mandatory isolation layer from model providers. What the document calls harness engineering in point two is our principle number one — and for the same reason: quality and cost are determined not by the choice of model but by what is built around it. June's episode with access to frontier models being restricted for export reasons showed that this is also a question of product survival.

Agent rights are set by a matrix. No agent has direct access to the database — only through a software layer. Read is the default. Write is possible only with explicit permission and through a draft that a human confirms. Deletion is prohibited entirely. Every call is written to an immutable log. This is precisely the categorisation of decisions that point seven introduces at industry level — only in our case it is implemented at product level, because otherwise the product cannot be sold to an executive who answers for a company.

Token cost is margin. We keep separate accounting of consumption: input and output tokens multiplied by model price, broken down by user and by specific feature. Without it you cannot answer the main question — what one useful result costs, as opposed to one request. And there is a rule that turned out to be the most useful thing in all our experience: before you widen an agent's autonomy, put a meter on it and measure. Autonomy without measurement is not a product, it is a lottery played with your money.

Value lives in memory, not in features. Point three, with its "demand intelligence," describes what we arrived at by our own route: the user should not have to pick a feature — they name a situation. And what distinguishes such a system from a nice chatbot is the accumulated memory of a particular executive's decisions, delegation and recurring patterns. Features get copied within a quarter; accumulated context does not.

And an honest difference. We have no Galactic Corridor, no token vouchers and no full-cycle service centre. Everything the document assigns to the state, a product company has to build for itself or do without. Which, incidentally, is the best illustration of how transferable somebody else's model really is: the engineering part matches, the conditions do not.

Where the document is strong

1. The metric changed to the right one. Completion of a real task instead of benchmark scores. A quiet but fundamental correction: it turns the industry away from a race between models and towards engineering for results.

2. The harness is recognised as a discipline of its own. It stands as point two, ahead of all applications. A rare case of a regulator hitting the real bottleneck rather than the most visible one.

3. Business models are named, not just technologies. Three tiers of selling and an explicit preference for payment by result. None of the earlier documents had that layer.

4. The last mile is recognised as work. The role of the engineer who lives at the customer's site and fixes processes is formalised. That is an acknowledgement that deployment is not an overhead cost but the place where value is created.

5. Safety by category, not by prohibition. Low risk means fast release, high risk means strict review. That approach does not kill experiments and still provides a frame.

6. Demand is stimulated with money, not appeals. Vouchers for small business, a subsidy to device buyers, financial products from banks. The document does not persuade the market — it pays for the first step.

Weak spots and open questions

1. There is no definition of a successful task. This is the main hole. "Payment for value" is proclaimed, but what exactly counts as a completed task and who records it is not stated. Without a shared definition, RaaS turns into an acceptance dispute: the customer says there is no result, the provider says there is.

2. Liability is unresolved. If an agent acts and causes damage, who answers: the agent's provider, the model owner, the customer, or the developer of a connected tool? Categorising decisions sets the frame of the permissible but does not allocate losses.

3. The economics of running agents is unproven. The very fact that demand has to be subsidised with vouchers says a lot: if the numbers added up on their own, there would be nothing to top up. For now this is a bet, not a confirmed model.

4. Horizons are mixed together. Domestic chips and global protocols are a three-to-five-year story. Vouchers and one-click registration are today. In one document they look equivalent, although they are verifiable to very different degrees.

5. The bet on technological sovereignty narrows the choice. Requiring reliance on domestic protocols, frameworks and chips is politically understandable, but in engineering terms it is a constraint: some things will have to be built worse and slower than they could be bought ready-made.

6. People appear only as workforce. Training specialists is there; the question of what to do with those whose tasks go to agents first is not there at all. For a document that explicitly talks about lowering the minimum size of a company, that is a noticeable omission.

7. There are no success criteria for the document itself. No indicators are set by which, two years from now, anyone could say whether it worked. That is a common disease of programmatic documents, but here it contrasts especially sharply with the content: the document teaches you to measure results and has no intention of measuring itself.

What to take from this as an executive

1. Count the length of the chain, not the power of the model. A hundred steps at 99% is 37% success. Do not hand the agent the whole route at once: break it into short segments with checks at the joins.

2. Budget for the harness, not the model. Tools, memory, access rights, logging and result verification deliver more gain than moving to the "smartest" model. The gap between "it worked in the pilot" and "it does not hold up in real work" is almost always a gap in the harness.

3. Start where the result is verifiable and the error is reversible. If you cannot tell automatically that the task is done, you can neither manage quality nor count money.

4. Measure the cost of a successfully completed task. Divide all costs for the period by the number of tasks brought to an acceptable result. "Cost per request" is misleading if some share of requests goes back for rework.

5. Define success in writing — before you sign the contract. If someone is selling you an outcome, fix three things: what counts as a successfully completed task, how that is verified automatically, and what happens in a disputed case. This is exactly what the document itself lacks.

6. Restrict the agent's rights by default. Read freely, write with human confirmation, never delete. Plus an action log. This is cheap at the start and almost impossible to retrofit.

7. Revisit the minimum size of a team. If coordination costs fall, part of a function can be assembled by one person with a set of agents. That applies both to your own structure and to your future competitors — of whom there may be several times more.

And the transfer rule, without which the other seven points are dangerous. Mechanics do not transfer by outcome. Before copying somebody else's model, name the condition under which it worked — and check whether you have that condition. The Beijing measures rest on developer density, access to domestic chips, budget and administrative capacity. The engineering part transfers anywhere; the conditions do not.

Conclusion

The value of this document is not that it is Chinese but that it is the first attempt to describe the economics of AI agents systematically: from requirements for models and harnesses to the sales model, the boundaries of liability and the stimulation of demand. Together with the "AI+" plan and the "1397" strategy it completes the construction — from ten-year goals to the question of who pays whom for what.

The main signal for an executive is simple. The age of agents is not about an even smarter model arriving. It is about work being broken down into dozens of steps, the repeatable and verifiable ones going to agents first, and human value shifting towards setting the goal, exercising judgement and answering for the final result.

Agents do not deliver success automatically. They amplify what a company already has: goals, expertise and a system of execution. If there is no system, there is nothing to amplify.

And now a question for you. Look at the task you are about to hand to AI. How many steps does it contain, and how will you know it has been completed successfully? If there is no answer to the second question, there is nothing yet on which to compute the economics of that project. I would be glad to see your examples and objections in the comments.

Sources: the full text of the measures and the official explainer on the Beijing municipal government portal; the Xinhua report on the press conference; and the study Agentic Harness Engineering.

This material was prepared with the use of artificial intelligence technologies.

Related Posts

See All
bottom of page