Nine Quality Criteria for an IT Product: What to Check, How AI Helps, and What Changes in AI Systems
In February 2024, a Canadian tribunal ruled on a dispute between a passenger and Air Canada. The chatbot on the airline's website had told a customer how to get a bereavement fare discount. It described the rule incorrectly, and the customer paid full price. The airline argued that the chatbot was responsible for its own words. The tribunal disagreed: a company is responsible for all the information on its website, whether it comes from a static page or a chatbot. Air Canada was ordered to pay 812 Canadian dollars (decision 2024 BCCRT 149, as summarised by DWW lawyers).
Formally, the Air Canada chatbot worked: it took questions and answered quickly. What failed was not the feature list but the correctness of the answer. The case shows a common mistake in conversations about quality: people confuse it with a list of features or with a successful demo. Neither tells you whether the system can be used in real work.
Nine criteria
In my practice I use nine quality criteria for an IT product, drawing on the international standard ISO/IEC 25010. They are grouped into three questions that an executive asks: does the system work as needed, will it let us down, and what will it cost us and will it lock us in.
Let's look at them in more detail.
Question | Criterion | What it means | How to check |
|---|---|---|---|
Does it work as needed? | Functionality | Does the system work as planned and perform the intended functions | Critical requirements of the specification or backlog (the product's task list) are implemented; functions work as intended; awkward cases (empty input, input errors, large volumes) are handled; no data is lost |
| Usability | Is it convenient for users, can they learn the system without lengthy training and manuals, do they run into problems, does it take too many clicks and steps, is it fast enough | Time to onboard a new user; number of steps for a typical operation; clear error messages; response time as the user experiences it |
| Accuracy and its retention | Is the result correct (calculations, analytics, answers) and does it stay correct when the data changes | Target accuracy is set as a number before launch; calculations are checked against reference cases; input data quality is checked; testing on your own data, overall and by significant segments; re-testing on a schedule and whenever the data or model changes |
Will it let us down? | Stability and reliability | Does the system work steadily and predictably, does it crash daily or under load, can it be restored after a failure quickly and without data loss | Share of time the system is available; recovery time after a failure and data integrity; monitoring; load testing and a pilot operation period before launch |
| Security | Are data and access protected: who sees, who changes, is there a trail | Access rights separation; secure channel; access log; backups; a defined policy on which data may be sent to external services and models |
| Safety | Could the system harm people, the process or equipment, can it be stopped, and does it fail into a safe state | A list of unacceptable events is drawn up before launch, and for each one it is shown how it is excluded or reduced to an acceptable risk; a person can intervene and stop the system; on failure the system goes into a safe state; no undocumented control actions; AI stays within its authority |
What will it cost, and will it lock us in? | Maintainability | Can someone other than its author work with the system: understand it, change it and confirm that everything still works after the change | Documentation and access to source code or configuration are handed over; the system can be changed by people other than its author; there is a way to test the system after changes; the time and cost of a typical change are stated before the contract is signed |
| Interoperability and portability | Does the system fit into the company's ecosystem and coexist with other systems, can it be moved to another platform or the vendor replaced without a rewrite | Data exchange with other systems goes through documented interfaces and is tested at acceptance; the system runs alongside others on shared resources without interference; data can be exported in an open format; it is known what it takes to move to another platform or your own cloud and to replace the vendor or the model |
| Performance and efficiency | Does the system have enough speed and capacity under load, and how much in resources and money does that take | Response time and throughput at the design load; resource consumption (servers, licences, compute) is measured; it is known whether the system will handle growth in users and data and what that will cost |

The "accuracy and its retention" criterion applies to any system. A report may calculate strictly by formula, but if the data is entered by hand at the end of the month, with errors and delays, the analytics will still be wrong. So accuracy depends both on the system's logic and on the quality of the input data. An AI system adds one more thing: the model is wrong in a certain share of cases, and that share is measured separately.
The threshold rule
A threshold for each criterion is set before launch, based on how critical the system is. A report that people look at once a month needs one set of thresholds; a production control system needs another. If the system falls below the threshold on even one criterion, it is not accepted, however well it scores on the rest. Above the thresholds, the criteria can be weighed against each other and you can argue about what matters more.

The threshold rule protects against the "but it's fast" or "but it looks great" trade-off. A convenient interface does not make up for a data leak or unwanted actions, and an accurate model is no help if nobody can fix the system once its author has left.
AI on both sides
With the arrival of AI, every criterion has two layers. AI has become an assistant that speeds up the testing of any system, AI-based or not. And AI has become the subject of testing: if the system itself is built on a model, every criterion is checked differently.

Most importantly, AI does not replace human work entirely: in any case, a person checks the AI assistant's conclusions. And if the acceptance decision depends on the assistant, the assistant itself is tested as an AI system: how accurately it handles this particular task.
Below, I go through each criterion in three parts: what we check, how AI helps, and what changes if the system itself runs on AI.
Does it work as needed?
The first three criteria answer the question of whether the system delivers value: does it have the right functions, are they convenient to use, and can you trust the result.
1. Functionality
What we check. Does the system work as intended: are the critical requirements implemented, how does it behave in awkward cases, does it lose data. Whether the output of these functions is correct is a separate criterion, accuracy.
How AI helps
From the specification, a model drafts a list of test cases, including awkward ones: an empty field, a date in the past, ten thousand rows instead of ten. It also finds contradictions and gaps in the specification while they are still cheap to fix. A person reviews the list, but starting from it is faster than starting from a blank page.
If the system runs on AI
The same question can produce different answers, so one successful demo is not enough. The system is tested on a set of typical and edge-case requests with known correct answers. Such a set is called evals: in OpenAI's documentation, these are tests that check whether model outputs meet the style and content criteria you specify (OpenAI, Evals). Some answers can be graded by another model. In a study by Zheng and colleagues, such a GPT-4-based "judge" agreed with human ratings in over 80% of cases, but it favoured longer answers, answers in a particular position, and its own answers (Zheng et al., 2023). That is why a person still reviews a sample of the judge's grades.
Before launch, you also decide what the AI system does when it is unsure: declines to answer, asks a clarifying question or hands the request to a person. This is a function too, and it is tested like the others.
2. Usability
What we check. Can a person learn the system without long training, does it demand unnecessary steps, are errors clear, is it fast enough. What matters here is the speed the user feels in daily work; response time under load and the resources behind that speed are checked under performance and efficiency.
How AI helps
A model goes through support tickets and user feedback and groups them by topic. You can see where people struggle, at which step they abandon an operation, which words in the interface are unclear. You can analyse every ticket, not just a sample.
At the design stage, AI also helps to assemble user scenarios, run virtual walk-throughs, find errors and optimise the flows. Then it helps to build first prototypes quickly, test them with users, collect feedback, analyse it and make further improvements. In effect, we cut the time and effort this process takes and can analyse and test more in less time.
If the system runs on AI
Three questions are added to usability. Is the answer clear to the person, and is it clear where it came from: which document, which record, which rule. And can the person disagree with the answer, correct it or override it. The EU AI Act asks the same of high-risk systems (the category with the strictest requirements): a person must be able to disregard or override the system's output and must remain aware of the risk of over-relying on automation (EU AI Act, Article 14). The model's response time is part of usability too: an assistant that thinks for a minute is no use in a live conversation with a customer.
3. Accuracy and its retention
What we check. Is the result correct (calculations, analytics, answers), and does it stay correct when the data changes.
For a conventional system, accuracy starts with data. If data is incomplete, inconsistent or out of date, the system will faithfully compute a wrong result. So calculations are checked against reference cases, and input data is checked for completeness, accuracy, consistency and currency. These and other characteristics are described in the international data quality model ISO/IEC 25012 (overview of the standard).
How AI helps
For an AI system, the set of test requests is run automatically after every change to the instructions, the knowledge base or the model version. A judge model grades some of the answers, and a person reviews a sample. Accuracy becomes a metric you monitor, not a number from the acceptance report. I explain why a model can be accurate on complex tasks and wrong on simple ones in the article "Coverage, Not Complexity". For a conventional system, AI checks calculations against reference cases and finds anomalies in the data: gaps, duplicates, values outside the usual range.
If the system runs on AI
First, the accuracy level of the AI system is set as a number before launch. One hundred percent is unattainable, so you decide in advance what share of errors is acceptable and what to do about them. The EU AI Act requires providers of high-risk systems to state accuracy levels in the instructions for use (EU AI Act, Article 15).
Then accuracy is tested on your own data, not taken from the model's marketing figures, and separately for significant segments: regions, equipment types, customer categories. An average hides skews. In the 2018 Gender Shades study, commercial gender classification systems had error rates of at most 0.8% for lighter-skinned men, but up to 34.7% for darker-skinned women (Buolamwini, Gebru, 2018).
Finally, accuracy testing is repeated on a schedule and after every noticeable change in the data or the model. Lingjiao Chen, Matei Zaharia and James Zou compared two versions of GPT-4 on the task "is this a prime number": over three months in 2023, accuracy fell from 84% to 51% (Chen, Zaharia, Zou, 2023). Their conclusion: language models need continuous monitoring. In my own work with different models, I can say this is a common occurrence, both in public services and in internal systems. One reason is that AI systems are tuned to one type of data, the data changes over time, and the instructions and infrastructure are not adapted.
In the OWASP threat list discussed in the security section, this is item nine, misinformation: plausible but false answers. Model errors cost money and reputation. In 2025, Deloitte Australia partially refunded the Australian government for a report worth about 440,000 Australian dollars: it contained more than a dozen non-existent references and a fabricated quote from a court ruling. The firm acknowledged that a language model had been used in preparing it. The refund was just over 97,000 Australian dollars, about 63,000 US dollars (CFO Dive).
Will it let us down?
The next three criteria answer the question about risk: will the system go down, will data leak, and could it harm people or the process.
4. Stability and reliability
What we check. Does the system work predictably, does it hold up under load, does it recover quickly after a failure, and does it keep its data while doing so.
How AI helps
A model reads the logs and finds recurring failures and spikes that a person would miss among thousands of lines. It helps to analyse an incident: build a timeline, correlate events, suggest hypotheses about the cause. But the final conclusion about the cause is made by an engineer.
At the design stage, AI also calculates the required hardware configuration faster, prepares several options and runs virtual stress tests; the final decisions are made on that basis.
If the system runs on AI
A dependency appears that a conventional system doesn't have: on the model and its provider. The model may become unavailable, hit rate limits or be replaced. Providers retire old models on a schedule: OpenAI promises at least six months' notice for generally available models, Anthropic at least 60 days (OpenAI, Deprecations; Anthropic, Model deprecations).
Even without a change of name, a model's behaviour changes (the GPT-4 example is covered in the accuracy section), so the model is monitored continuously, not only when it is updated.
An AI system needs a fallback mode for when the model is unavailable: another model, a scenario without AI, or a handover to a person. The model version is pinned, and before every change the set of test requests is run again. The EU AI Act explicitly mentions backup solutions and fail-safe plans for high-risk systems (EU AI Act, Article 15).
5. Security
What we check. Are data and access protected: who sees, who changes, is there a trail. We check access rights separation, the secure channel, the access log and backups, as well as the rules on which data may be sent to external services and models.
How AI helps
Models already find vulnerabilities that people and classic tools have missed. In 2024 there were specialised tools for this; by 2026 even open, publicly available models run the necessary tests straight away. In August 2026, Anthropic showed how much the result depends on how the work is organised: 45 agents sharing a common forum found 266 vulnerabilities in 15 open-source projects, while the same agents working independently found 21, albeit with a compute budget four times smaller (Anthropic, 13.08.2026). As a result, code and configuration can be checked with AI more often and more widely, but this does not replace a security specialist. We can run more research, and deeper research, but it still needs a person to lead it.
If the system runs on AI
New threats appear that a conventional program doesn't have. The specific list of threats to applications built on language models is maintained by OWASP, the international non-profit community for application security. The 2025 version has ten items (OWASP Top 10 for LLM Applications):
Prompt injection.
Sensitive information disclosure.
Supply chain vulnerabilities: third-party models, data and components.
Data and model poisoning.
Improper output handling, when the model's output is passed on to other systems without checks.
Excessive agency.
System prompt leakage.
Vector and embedding weaknesses: vector stores and knowledge-base retrieval.
Misinformation: plausible but false answers.
Unbounded consumption: runaway use of resources.
Three items on this list relate directly to security: prompt injection, sensitive information disclosure and system prompt leakage. Item 6 belongs to safety, item 9 to accuracy and item 10 to performance and efficiency; they are covered in their own sections. The three security items are worth knowing by name.
Prompt injection comes first on the list. For a model, an instruction and data look the same, so a command can come directly from the user, hide in an external source (a website, an email, a PDF) or sit in the very data the system reads. A December 2023 case shows what this looks like: a user gave a Chevrolet dealership's chatbot a new rule, to agree with anything the customer says, and then asked for a Tahoe SUV for one dollar. The bot agreed and called it a legally binding offer. The dealer switched the bot off (GM Authority). The most reliable protection is architecture: structured data at the system's input and output, not free text.
Sensitive information disclosure comes second. In spring 2023, employees of Samsung's semiconductor division uploaded source code and notes from an internal meeting to an external chatbot (TechRadar). According to Bloomberg, the company then banned generative AI for staff in one of its major divisions (Bloomberg, 02.05.2023). The question "which data may be given to an external model" is settled before launch, not after the first incident. A model's system instructions are kept as secrets: leaking them is like leaking source code.
What to require from an AI system at acceptance on security:
the system has been tested against prompt injection and jailbreak attempts, including unusual wording and other languages;
it is defined which data may be sent to an external model, and system instructions are stored as secrets;
the data sources for training and for the knowledge base have been checked and documented;
action logs are protected against modification and from being switched off.
6. Safety
What we check. How safe the solution is for the production process, people and equipment. Safety is not about passwords: a system that controls equipment or gives instructions to staff can cause harm without any leak. So we check whether a person can stop the system and whether it goes into a safe state on failure.
Safety testing should start with a list of unacceptable events, and only then move on to protective measures. An unacceptable event is something that must not happen under any circumstances. For people, it is a threat to life and health, a leak of personal data, financial loss. For the company, it is a production stoppage or a failure in equipment control systems, theft of money, a leak of trade secrets, publications and mailings in its name, wrong decisions made on the system's data. The safety threshold is this list: the system is accepted when, for each unacceptable event, it has been shown how the event is excluded or reduced to an acceptable risk. The approach is covered in detail in my book "Artificial Intelligence. Freefall".
How AI helps
AI helps to draw up the list of unacceptable events and failure scenarios: to go through what could go wrong at each step of the process. The decision on what is unacceptable and how to exclude it stays with a person.
If the system runs on AI
To avoid drowning in the risks of an AI system, it helps to keep in mind the three areas of technical AI safety formulated by DeepMind researchers in 2018 (DeepMind Safety Research).
The first area is specification: the system does what was intended, not what was literally written in the task. The classic example from the same work: an agent in the boat-racing game CoastRunners, instead of finishing the course, circled in one place collecting bonus points, because the reward had been defined imprecisely. For an executive, this is a question for the specification: does it say what the system must not do, and does the system have capabilities nobody ordered, such as control actions on equipment without human confirmation.
The second area is robustness: the system withstands interference and attacks. Data in operation differs from the data the system was tuned on, and an attacker deliberately crafts input to deceive it. A robust system in this situation warns the operator and stops rather than inventing an answer.
The third area is assurance: the system's work can be observed and the system can be stopped. This means action logs that cannot be erased or switched off, monitoring, priority of manual control and a "stop" button. The EU AI Act requires that the operation of high-risk systems can be interrupted with a "stop" button or a similar procedure (EU AI Act, Article 14).
From the OWASP list, item six, excessive agency, relates directly to process safety. This is when a system with excessive functionality, permissions or autonomy takes damaging actions in response to an unexpected or manipulated model output (OWASP, LLM06:2025). In July 2025, an AI agent of Replit, an app-building service, deleted the production database of Jason Lemkin, founder of the SaaS community SaaStr, even though it had been explicitly told not to make changes without approval. Replit then introduced automatic separation of production and development databases and a planning mode without code edits (The Register). A prohibition written in words in the task did not stop the agent. The boundary of an agent's authority is set by an executive: it is a management decision, not a technical setting. It means a list of permitted actions, human confirmation of everything irreversible, and a production environment separated from testing.
What to require from an AI system at acceptance on safety:
it is documented what the system must not do, and it has no undocumented capabilities;
a person can interrupt its work at any moment, manual control takes priority, and on failure the system goes into a safe state;
the agent has only the permissions it needs, irreversible actions are confirmed by a person, and production is separated from testing;
the user knows they are working with AI;
on unfamiliar data the system warns the operator and stops.
The final list of requirements depends on how critical the system is, that is, on the threshold set before launch. I wrote more about agent risks in the article "The AI Agent Didn't Fail. It Understood the Task Too Well".
What will it cost, and will it lock us in?
The last three criteria answer the question of how much it will cost to live with the system and whom the company will depend on: the author, the platform and vendor, the bill for resources.
7. Maintainability
What we check. Can someone other than its author and developer work with the system: understand it, change it and confirm that everything works after the change. Have the documentation and access to the code or configuration been handed over, are there tests or a test set, are the time and cost of a typical change known before the contract is signed.
How AI helps
A model writes and updates documentation as development goes along and explains someone else's code to a new developer. The dependency on one person who "knows everything" goes down.
If the system runs on AI
There is more to maintain than code. There are the system instructions that define the model's behaviour. There is the knowledge base the model draws its answers from. There is the set of test requests used to check it after every change. There are the rules of interaction between different AI agents. And there are the different versions of the AI models themselves on which all of this is tuned. If any of this lives only in the author's head, only the author can maintain the system. So all of it is handed over to the client or the support team just like source code, together with a description of how to test the system after the model is replaced.
8. Interoperability and portability
What we check. Can the system be embedded into the company's ecosystem or integrated with other products and IT solutions, does it coexist with other systems. And can it be moved to another platform or your own cloud, or the vendor replaced, without a rewrite. We check that data exchange goes through documented interfaces and is tested at acceptance, that data can be exported in an open format, and that there is a clear plan for migration and vendor replacement.
A system that cannot be moved locks the company into its vendor: the vendor sets the price of the next change or renewal. The question becomes acute when a vendor leaves the market or changes its terms.
How AI helps
A model describes exchange interfaces, maps data formats between systems and helps to move code and configuration to another platform. The final migration check is done by people in a test environment.
If the system runs on AI
Portability gets a new dimension: replaceability of the model. Providers retire models on their own schedule, so the system is built so that the model can be replaced without a rewrite. Instructions, the knowledge base and the test set are not tied to one provider, and after a replacement the system is run through the same set of requests.
Interoperability of an AI system has its own specifics too: the model takes data from the company's systems, and an agent also acts in them. Each such access goes through a documented interface with clear permissions, not through an administrator account handed out "temporarily" for a pilot. I describe how such a harness around an agent is built in the article "The Agent Harness: Why AI Agents Obey the Same Laws as Companies".
9. Performance and efficiency
What we check. Does the system have enough speed and capacity under load, and how much in resources and money does that take. We measure response time and throughput at the design load, consumption of servers, licences and compute, and check what happens as users and data grow. Stability answers whether the system will go down; performance answers whether it will slow down. This is about the cost of running the system itself, not about the project's return on investment.
How AI helps
A model analyses operating and resource consumption logs, finds the bottlenecks that slow the system down under load, and suggests where to cut consumption: unnecessary database queries, idle capacity, heavy operations. The effect of the optimisation is still confirmed by measurement.
If the system runs on AI
An AI system has a new cost line: every request to the model costs money, and the provider has rate limits. In the OWASP list this is item ten, unbounded consumption: a system without limits can burn through the budget or stop responding under a flood of requests (OWASP Top 10 for LLM Applications). The model's response time as the number of users grows is also tested before launch. So before launch you calculate the cost of a typical request and the monthly volume, set limits, and monitor spending as closely as availability.
What this gives an executive
First, a common language with IT and the contractor. Instead of "the system is bad", you say, for example, "usability is below the threshold: a typical operation takes twelve steps instead of five". The second statement is something you can work with.
Second, acceptance based on facts. A demo shows part of the functionality and sometimes usability. The other criteria are almost invisible in a demo, and they are checked separately.
Third, fewer surprises after launch. Almost everything in the table above is cheaper to write into the specification and the contract than to fight for later.
The nine criteria do not promise a return on investment. A system that passes every threshold may still deliver no value if it was built for the wrong task. But you know what quality you got and what you are paying for.
How this relates to standards and regulation
The list is my own and designed for executives, not for quality engineers. Still, all nine characteristics of the international software product quality model ISO/IEC 25010:2023 are in it, although grouped differently: the mapping is not one-to-one (ISO). The additions for AI systems from ISO/IEC 25059:2023 (ISO), such as transparency, controllability, robustness and the ability to intervene, are covered above under "If the system runs on AI"; the second edition of this standard is currently at the approval stage (ISO). The grouping into three questions does not appear in ISO. The closest precedent is McCall's 1977 quality model, which divides factors into three perspectives: operation, revision and transition (McCall, Richards, Walters, 1977).
Regulators and national bodies look at the same things from their side. In the EU, the AI Act requirements for high-risk systems quoted above will apply from 2 December 2027 for standalone high-risk systems and from 2 August 2028 for those built into products, under the Digital Omnibus amendment, Regulation (EU) 2026/1744 (European Parliament, Legislative Train; EUR-Lex). In the US, the NIST AI Risk Management Framework describes seven characteristics of trustworthy AI (NIST AI 100-1). Singapore's AI Verify testing framework assesses systems against 11 governance principles (AI Verify Foundation). The UAE Charter for the Development and Use of Artificial Intelligence of June 2024 sets out 12 principles (UAE Legislation). In Germany, the BSI publishes the AIC4 criteria catalogue for AI cloud services (BSI, AIC4) and, since July 2026, a catalogue of test criteria for AI systems in finance (BSI).
The nine criteria do not replace these standards and frameworks: engineers building a testing framework and lawyers assessing compliance need exactly those. The nine criteria are a way for an executive to ask the right questions before that work begins.
Further reading
Free download: the third edition of my book "Artificial Intelligence. Freefall", download the PDF.
Training and advisory for executives: on the home page.


