Process Architecture

Eight mistakes to avoid when integrating AI into business processes

A faster task can save effort while leaving the wider process unchanged. Eight mistakes to avoid in measurement, controls, ownership, roles and implementation.

In this note

The short answer

A useful first experiment is to improve one task. A common mistake is to treat that improvement as evidence that the whole process is ready to change.

The draft may take minutes instead of hours while the approval still takes a week. The effort saving can be real, but the client may see no faster response. A new output can also create new checking, correction or coordination work elsewhere. These are different outcomes and should be measured separately.

Start from the result the organisation needs. Test whether AI can perform the task, then examine the handoffs, controls, responsibilities and exceptions around it. Redesign only what the evidence and the intended scope require.

Set the scope of the change

Individual assistance and a recurring shared process require different levels of design. An employee drafting a document may remain both author and reviewer. A system that handles every incoming request needs agreed rules for checking, exceptions and handovers.

The boundary is the output's use. Once another person relies on it, the quality expected at that handover must be clear. A quick draft that creates correction work for its recipient has shifted effort rather than removed it.

The eight mistakes below concern that wider process. They do not imply that every use of AI needs a transformation programme.

1. Measuring only the AI step

A process exists to deliver an outcome: a quotation sent, a supplier paid, a repair resolved, an account opened.

That outcome should define the baseline.

A representative sample can establish the baseline without a process-mining system. Follow cases from request to completion, recording active work, waiting, corrections and handoffs. Combine those observations with volume to estimate effort or cost per case. Include difficult cases: a process that looks efficient on ordinary requests may depend on expensive workarounds for its exceptions.

The distinction between elapsed time and active working time is particularly important. Reducing a ten-minute task to two minutes saves eight minutes of effort even if the case still waits two days for approval. It does not, by itself, establish a meaningful improvement in customer turnaround.

This is also where many pilots become misleading. A statement such as "drafting time fell by 80%" may be accurate and still say very little about the economic result. If turnaround time, cost, quality, capacity and rework are not measured, the pilot has established that the model is fast, not that the process is better.

2. Automating work that should be removed or simplified

Once the process is visible, examine why each step exists.

A useful classification is:

The second category deserves particular attention. Processes accumulate workarounds for the limitations of people, systems and organisational structures. If AI or another technology changes one of those limitations, the workaround may no longer need to exist in its current form.

The third category should not be removed merely because it slows the process. A control can be inefficient and still protect against a real risk. The correct question is whether the risk remains and whether the control should be redesigned.

Automating a step that should have been removed first creates an additional problem. The step now has software, configuration, maintenance and an owner, which can make it harder to challenge later.

3. Assuming existing controls will catch AI errors

A control designed for one type of failure does not necessarily catch another.

Routine manual work often produces visible slips: a transposed digit, a missed line, the wrong name copied into a field. Traditional checks, reconciliations and four-eyes reviews are often designed around those errors.

Generative AI can produce a different error profile. An answer can be fluent, internally coherent and wrong. A citation can appear plausible but not exist. A figure can come from the wrong source while the surrounding explanation remains convincing.

Controls need to match the failure mode. Verify consequential facts against their sources and route cases that exceed the tested scope to a person. Keep known test cases to rerun after changes to the model, instructions or data. Systematic sampling and a record of recurring errors can show whether the process is deteriorating; informal impressions cannot provide the same assurance.

Human review needs the same level of design.

A large body of human-factors research predating generative AI has documented automation bias and automation complacency. A review by Parasuraman and Manzey found that people can under-monitor automated systems or over-rely on automated recommendations, particularly when attention is divided (Parasuraman and Manzey, 2010). A separate systematic review of automation bias reached a similar conclusion across multiple fields, while also showing that the effect depends on the user, task and system design (Goddard, Roudsari and Wyatt, 2012).

This matters because a nominal human checkpoint is not automatically an effective control.

A review step should specify what must be verified and give the reviewer the information and time to do it. That can mean comparing named fields with the original, checking contractual terms above a risk threshold or making an independent judgement before seeing the model's recommendation. Record corrections, rejections and escalations, rather than approval alone.

A very low rejection rate can mean the system is performing well. It can also mean that the review has become ceremonial. The organisation needs enough measurement to distinguish between the two.

4. Leaving accountability unclear

AI can make internal accountability less obvious because several people contribute to an output in different ways.

The person releasing an output, the manager owning the process, the system administrator and the vendor have different responsibilities. Calling all of them "human oversight" leaves the actual decisions unresolved.

For most business processes, it is useful to distinguish at least two roles.

The process or output owner is accountable for the business decision, result and escalation path.

The system owner is responsible for configuration, testing, monitoring, changes and known failure modes.

In a small firm, one person may hold both roles. The important point is that they are named.

When an output reaches another team or a customer, they need a clear route for corrections and questions. A vendor's technical responsibility does not answer who can amend the business decision, suspend the process or resolve a complaint. Agree those responsibilities before the first output is released.

5. Treating a model test as a process pilot

A pilot answers the question it was designed to answer.

If the question is whether a model can extract specific fields from an invoice, a technical test may be enough.

If the decision is whether the organisation should change the invoice process, it is not.

A useful pilot therefore has two stages.

Stage A: establish capability

Use representative cases to test whether the system can perform the task at an acceptable level. Include ordinary cases, edge cases and known difficult examples.

Stage B: establish operability

Run cases through the proposed workflow, including handoffs, queues, controls and connected systems. Observe how exceptions reach a person, whether reviewers can keep up and how unresolved decisions are escalated. Measure time and cost from beginning to end, including work required by rejected outputs.

There is also a reasonable limit to redesign at this stage. Rebuilding a process around a technology that has not yet demonstrated basic capability can waste time. Small, reversible experiments should prove the task first. The error is treating that technical proof as sufficient evidence for scale.

Follow a quotation through the proposed process

Consider a distributor testing AI to draft quotations from customer enquiries. The capability test compares product references, quantities and terms against the source. The operating test starts when an enquiry arrives and ends when an approved quotation reaches the customer.

Some enquiries lack a product reference; others require a price exception. The pilot must show who obtains missing information, who approves the exception and whether the draft waits in the same queue as before. Record the effort of those people as well as the model's drafting time.

A faster draft may justify using the assistant even when approval remains unchanged. If the objective is faster customer response, however, that result alone does not meet the acceptance criterion. The pilot decision should state which improvement has been established and which still requires work.

6. Ignoring how people's work changes

When AI removes routine cases, the human role rarely remains unchanged.

Less time may go into producing standard outputs, while more goes into checking, resolving ambiguous cases, maintaining the system and dealing with customers whose problem remains unresolved. Those tasks need skills and decision authority, not simply spare time.

The remaining work can be more demanding because the easier cases have already been removed.

This should be designed explicitly. For each affected role, management should know what work disappears, what work remains, what new tasks appear, what authority changes and what training is required.

7. Having no tested fallback

Removing manual steps reduces cost and duplication. It can also reduce resilience.

Define how the process operates when the service is unavailable. Decide which cases can wait and which must continue, who can handle them, the reduced capacity available and how long it is acceptable. A material change in model behaviour also needs a route to suspend use, investigate and retest before normal operation resumes.

The answer should be more specific than "we can always do it manually". If the manual process has not been used for a year, the people, access rights, instructions or skills required to operate it may no longer be available in practice.

For business processes, the practical case for a fallback is straightforward. Providers fail, integrations break and models change. Continuity is sufficient reason to decide in advance how the process operates without the system. Test a representative case through that route before relying on it.

8. Assuming time saved automatically creates value

Time saved establishes released capacity. Realised business value depends on what the organisation does with it.

If ten employees each save two hours a week, the organisation has not automatically created twenty hours of economic value. The time may become shorter response times, higher quality, additional volume, a cleared backlog, new work or simply capacity absorbed elsewhere.

Those outcomes are not equivalent.

Evidence from Denmark provides a useful reminder that real-world effects can be less dramatic than task-level demonstrations suggest. Humlum and Vestergaard linked surveys of workers in AI-exposed occupations with administrative labour-market data. In the revised study, they found no measurable average effect on earnings or recorded hours two years after ChatGPT's launch, while workers reported that AI had also created new tasks, including oversight and integration work (Humlum and Vestergaard, 2026). Earlier versions of the study reported new AI-related tasks for 8.4% of workers.

The implication for a pilot is to measure the intended business result separately from the time released.

Before rollout, decide what the released capacity should achieve. It might support a one-day rather than two-day response, 20% more volume without additional hiring, a smaller backlog or fewer overtime hours. Assign the capacity to that outcome and measure whether it occurs. "Hours saved" is an input into the result, not the result itself.

Write the pilot decision before running the pilot

A useful decision record is short enough to discuss with the people responsible for delivery. It should state what would justify expanding the application and what would stop it.

Decision to make What to record before testing
Which outcome matters? Separate effort per case, customer turnaround, quality and operating cost; choose the result this pilot must improve
What is being tested? The task, the input sources and the boundaries of the workflow; list the steps that remain manual
Who releases the output? A named owner and the checks required before the result reaches another team or a customer
What counts as acceptable? Error categories and acceptance criteria agreed for this context, with stricter treatment of consequential errors
What would stop or reverse the change? Specific failure conditions, the fallback and the person authorised to suspend use
What must be measured? Production, review, correction, waiting and maintenance effort, including rejected and exceptional cases

Use a representative set of ordinary and difficult cases. Compare the proposed method with the existing one using the same outcome and quality criteria. Record the model, instructions and template version so the test can be repeated after a material change.

At the end, distinguish three decisions: continue experimenting, adopt within the tested scope, or prepare a wider implementation. A successful drafting test may justify the second while leaving the third unresolved. That is a legitimate result, not an incomplete transformation programme.

Scale the process that has actually been tested

A successful test supports a defined use, with a particular set of inputs, controls and responsibilities. Expansion changes that evidence: more volume may overwhelm a reviewer, new document types may introduce errors and another team may depend on a different handover.

Treat expansion as a decision to check those conditions again. The aim is to retain the useful improvement without assuming that every surrounding part of the organisation is ready for it.

Method note. The business examples are hypothetical illustrations. Research findings refer to their stated settings and should not be read as universal estimates or as proof that a particular implementation method causes better performance. The distinction between a technical test and an operating process is a management framework; controls should be proportionate to the actual use, volume and consequences of errors.

← Back to all notes

Continue reading