Technology Adoption

Which tasks should be handed over to AI?

A task-by-task method for SME leaders to find where AI pays, where a simpler tool does better, and when the honest answer is not yet.

In this note

The short answer

Choose the work worth improving before choosing the technology. Start with recurring tasks, compare AI with simpler alternatives and test the few candidates where the benefit survives checking, correction and maintenance.

The decisive question is often whether the output can be checked substantially faster than it can be produced from scratch. A quick draft creates little benefit if its reviewer must redo the work to trust it.

AI is most useful where varied documents, messages or speech need to become a reviewable output. Structured inputs and explicit rules often favour conventional software. The right shortlist may therefore include templates, existing system features and tasks that should simply stop.

Start with the work, not the demonstration

Most first encounters with AI are demonstrations. A vendor, colleague or conference speaker shows a model summarising a contract, answering a customer email or producing a report. The natural reaction is to ask where the company could use the same tool.

That reverses the decision process. A demonstration proves that a tool can perform a chosen example. It does not show that the task occurs often enough to matter, that the output can be checked cheaply, that the necessary data are available, or that AI is better than a simpler form of automation.

The practical question is which pieces of work deserve intervention, and what kind.

AI is one option among several

When a task is slow, repetitive or error-prone, there is a useful order in which to consider the remedies. Moving through it from simple to complex prevents a common mistake: using a language model where something cheaper and more predictable would work better.

Option When it fits Example
Remove the step The step exists from habit or for a reason that no longer applies Stop producing a weekly status report if the project system already contains the same information
Standardise it The task varies more than necessary Use a checklist and template for onboarding a supplier
Use a rule or existing software feature The input is structured and the logic can be specified Match bank transactions with reconciliation rules in accounting software
Use AI as an assistant The input is unstructured and a competent person can review the result as part of normal work Turn a project manager's voice notes into a first draft of a site report
Build AI into the process The input is unstructured, volume is high and checking can be designed Extract order lines from customer purchase orders arriving in different PDF formats

Rule-based software can usually be designed to produce a predictable result from the same structured input. Large language models are probabilistic: they can handle variation that would be awkward to encode as rules, but their outputs can vary and their errors are harder to anticipate.

That trade-off matters. If a task already consists of clean fields and explicit logic, conventional software is often the better tool. One of generative AI's clearest advantages is handling unstructured input: emails, free text, documents, speech and images. That is where earlier automation often became expensive or brittle.

Six questions to ask of every task

Once recurring work has been broken into tasks, six questions are usually enough to rank the first candidates.

Question Why it matters
How often does it occur, and how long does it take? Frequency × duration gives the pool of time potentially available to save. A task performed twice a year rarely justifies much investment.
What form does the input take? Structured data often points to rules or existing software. Unstructured documents, messages, speech and images are where AI is more distinctive.
What does an error cost, and who sees it first? An error in an internal draft can be caught before it matters. An error in an invoice, quotation or regulatory filing may reach someone else first.
Does it depend on tacit knowledge? If the task relies on unwritten knowledge of a customer, supplier or internal exception, that context must be captured or the task should remain with the person who holds it.
Does the result have to enter another system? An extraction that someone must manually retype into the ERP may save much less than the demonstration suggests.
How long does it take to check the result? This is often the deciding question.

Do not spend weeks making the first estimates precise. Rough numbers are enough to identify the few tasks worth testing. Precision should come later, when measuring those candidates before and after a pilot.

And ask the people who actually do the work. Managers often know the broad process but not how often exceptions occur, which steps create rework or how an error is detected in practice.

The task is the unit, not the job

Asking whether AI can replace a role or transform a department is too coarse to be useful. Jobs bundle together tasks with very different characteristics. The relevant unit of analysis is the task.

This is not a new way of thinking about technology and work. Economists have long analysed technological change at task level rather than treating an occupation as one indivisible block (Autor, Levy and Murnane, 2003). Generative AI makes that distinction more important because its capabilities are uneven.

Consider a sales administrator at an industrial distributor. The role might include:

“Can AI do the sales administrator's job?” is not a useful question. Task by task, the picture is clearer. Extracting order lines from varied PDFs may be a strong AI candidate. Order confirmations may already be handled by the ERP. Delivery-date questions depend mainly on accurate ERP data and integration. A dispute with a long-standing customer may depend on context, judgement and relationship history that should stay with a person.

Performance can also change sharply between apparently similar tasks. In an experiment involving 758 consultants, participants using GPT-4 completed 12.2% more tasks and worked 25.1% faster on tasks that fell within the model's capabilities, while producing higher-quality work. On a task deliberately chosen outside that capability frontier, AI users were 19 percentage points less likely to reach the correct answer (Dell'Acqua et al., 2023).

The researchers called this a jagged technological frontier. The implication for a company is simple: do not infer performance on one task from a good demonstration on another. Test the actual work.

Two ways AI enters the work

The distinction between assistance and process automation matters because the economics and controls are different.

1. AI as an assistant

A person chooses when to use the tool, reviews the output and remains accountable for the work. Typical uses include drafting a letter, summarising an email thread, translating a document, preparing for a meeting or challenging an argument before presenting it.

This can create meaningful productivity gains. In a study of 5,172 customer-support agents, access to a generative-AI assistant increased issues resolved per hour by 15% on average. Gains were concentrated among less-experienced and lower-skilled workers; the most experienced workers saw small gains in speed and small declines in quality (Brynjolfsson, Li and Raymond, revised paper, 2024).

Benefits depend on the task, the user and the surrounding organisation.

Assistance is especially useful when the person using AI already knows what good output looks like. A proposal writer, for example, can use AI to get from a blank page to a rough structure, test alternative wording and challenge an argument. The professional judgement does not disappear; it moves toward selecting, checking and improving the draft.

2. AI built into a process

The second mode is more demanding. Here the model performs a recurring step without someone deciding case by case to use it: extracting fields from every incoming invoice, classifying support requests or translating thousands of catalogue entries.

The return can be larger because the volume is larger. So are the design requirements. Someone has to define:

The same underlying task can sit in either mode. Translating an occasional product sheet may be a good assistant use for a bilingual employee. Translating 4,000 catalogue entries into three languages is a process-design problem involving terminology, quality checks, exception handling and ownership.

The important distinction is therefore not “human versus AI”. It is where judgement, checking and accountability sit once AI enters the process.

The economics often depend on checking

Separate the effort saving from the financial return.

Net effort saved = current working time − assisted production, review, correction and transfer time − ongoing maintenance.

For a financial comparison, convert that effort into a monetary value, then subtract setup costs, software costs and the expected cost of errors that escape review. Keep time and money in separate calculations. An hour released is not automatically an hour of cash saved, especially when salaries and staffing remain unchanged.

This prevents two common mistakes: treating review as free and presenting available capacity as a reduction in expenditure.

Compare two outputs that may look equally impressive in a demo.

A draft email to a customer about a delayed delivery may take seconds to generate and less than a minute to check. The facts are known, the tone is familiar and mistakes are visible.

A summary of a 60-page public-tender specification used to decide whether to bid is different. The clause that changes the decision may be a certification requirement, exclusion or penalty regime buried in the document. If a manager must read the full specification to know that the summary is complete, the summary can still help with orientation, but it should not be mistaken for a substitute for the underlying review.

Actual measurement also matters because people are poor judges of their own productivity. In a 2025 randomised study, 16 experienced open-source developers completed 246 real tasks on mature repositories. Before the study they expected AI tools to make them 24% faster; after using them they still believed the tools had made them about 20% faster. Measured completion time was instead 19% longer when AI use was allowed. The useful lesson for this exercise is the mismatch between perceived and measured performance (METR, 2025).

For assistance, one practical test is particularly useful:

Can the person responsible for the work verify the AI output substantially faster than they could produce the result from scratch?

If yes, the task is promising. If verification requires effectively redoing the work, the case is weak unless there is another source of value, such as broader coverage or faster turnaround.

Structure is easier to check than meaning

An extracted document can have the correct format and still contain the wrong information. Software can check whether a required field exists, a date follows the expected pattern or a total balances. It cannot establish from those checks alone that the name, date or amount came from the right place in the source.

Assess these separately. Use automatic validation for structure and consistency, then define how consequential content is checked. For a purchase order, that might mean matching each extracted quantity and product reference with the original before posting it. A clean spreadsheet or valid file is useful evidence of format, not a guarantee of accuracy.

A worked estimate: proposal drafting

Consider a firm preparing 24 proposals a month, using the following assumptions.

Active work per proposal Current method AI-assisted method
Prepare the first draft 45 minutes 10 minutes
Review and correct it 10 minutes 20 minutes
Format and file the final document 5 minutes 5 minutes
Total 60 minutes 35 minutes

The gross saving is 25 minutes per proposal, or 10 hours a month. If maintaining the instructions and examples takes two hours a month, the recurring saving falls to eight hours. Twelve hours of initial setup would leave 84 hours of released capacity in the first year: 8 × 12 − 12.

At an illustrative loaded labour cost of CHF 75 an hour, that capacity is valued at CHF 6,300 before licences and the expected cost of escaped errors. It becomes a cash saving only if expenditure actually changes; otherwise the business must decide what the additional capacity will achieve.

The proposal can still wait several days for approval. That does not erase the effort saving, but it means the pilot has not yet improved customer turnaround. If checking takes longer than assumed, or too many drafts have to be rewritten, the calculation changes. Measure those cases as well as the successful ones.

What the method looks like in practice

Task Input Cost of error Cost of checking Best response
Extract order lines from customer purchase orders (distributor, ~40/day) PDFs in varying formats Moderate; usually visible at confirmation Low: compare with source document AI in the order process, with confirmation before posting
Send appointment reminders (physiotherapy practice) Structured calendar data Low None once configured Existing practice-software feature
Monthly bank reconciliation Structured bank data High Moderate Reconciliation rules in accounting software
Turn site-visit voice notes into reports (engineering office) Speech Moderate Low for the author who attended the visit AI assistant
Answer standard customer questions (online retailer) Free-text email Moderate; customer-facing Low if the draft is grounded in approved answers AI drafts; employee reviews and sends
Translate product sheets into French, German and English Text Moderate; customer-facing Medium; needs a competent language review AI assistant with terminology guidance and review
Write a client proposal Mixed High High; important claims need review AI for structure, first drafts and challenge; author owns final text
Decide whether to bid for a public tender Long unstructured documents High High Human decision; AI may extract requirements for verification against the source

Two patterns matter. First, several tasks that sound like AI opportunities are better served by conventional software. Second, the strongest AI candidates are often not the most visible or strategic tasks; they are the repetitive steps where unstructured information has to be turned into something structured and reviewable.

Where the strongest candidates tend to sit

The most attractive tasks are often routine intake, document extraction, recurring first drafts and translation. They combine repetition with information that is awkward to handle through fixed rules. Their value depends on having a source against which the output can be checked.

Customer-facing decisions need a different assessment. A draft answer reviewed before sending has a checkpoint; an automatic answer reaches the customer before anyone inside the firm sees it. The latter requires stronger testing, approved information sources, escalation and monitoring. High volume alone does not make it a good first project.

Running the exercise in a small firm

This exercise does not require special software. Its main cost is time from the people who know the process.

A practical format is a two-hour session for one process with the two or three people who perform it. Map a typical week or month, then list the recurring tasks. For each task:

  1. estimate frequency and current time spent;
  2. identify the input and output;
  3. work through the options in order: remove, standardise, rule/existing feature, AI assistant, AI in the process;
  4. if AI still looks appropriate, answer the six questions above;
  5. select only the two or three strongest candidates.

Each candidate should leave the session with a named owner, a chosen approach and a baseline measure of how the work is done today. Measure the baseline rather than relying on memory; otherwise there is no reliable way to tell whether the pilot helped.

An empty shortlist is acceptable. It is cheaper than forcing an AI pilot into a process that does not need one.

Data protection belongs in the selection step

Do not wait until implementation to ask what information the task contains and which systems are allowed to receive it.

In Switzerland, the Federal Data Protection and Information Commissioner states that the Federal Act on Data Protection applies directly to AI-supported processing. The FDPIC also states that users of conversational AI have a right to know when they are communicating with a machine and that high-risk processing may require a data-protection impact assessment (FDPIC).

The practical implication is simpler than the legal text: a candidate task is not ready for testing until the company knows what data it contains and whether the proposed tool is permitted to process it.

Be clear about the purpose

A task inventory needs accurate information from the people doing the work. State what the exercise is intended to achieve and whether role or staffing changes are possible. If the aim is to release time for other work, identify that work. Employees should not have to infer the consequences of a technology workshop.

What this method does not answer

A task inventory is useful, but it is not an AI strategy by itself.

It has four important limits.

First, it favours improvements to existing work. It may miss strategic uses that change the product, service or business model. If a competitor can use AI to offer something fundamentally faster, cheaper or different, mapping today's tasks will not answer the whole question.

Second, it assumes the necessary information can be reached. A task can score well on paper but still be a poor first project if the required data are trapped in an old system, inconsistent across files or available only on paper. In that case, the first project may be data access rather than AI.

Third, very small firms may not have enough volume to justify building anything. For them, individual assistance may be the better programme: a small set of approved tools, clear data rules and enough training for employees to learn what those tools do well.

Fourth, the frontier moves. A task that is unattractive today may become viable as models improve or as AI features appear inside software the company already uses. The inventory is therefore worth revisiting periodically; the second pass will be much faster than the first.

The discipline is accepting the answer

The difficult part is not finding possible AI uses. Almost any process can be made to produce some.

The discipline is accepting the result when the best answer is a template, a rule in existing software, a better form, an hour of training on an approved assistant—or nothing for now.

Start with the work. Break it into tasks. Use the simplest tool that reliably improves each one. Measure the result. Then scale only what survives contact with the real process.

Method note. The business examples and proposal calculation are hypothetical, with simplified assumptions. Research results refer to their stated populations, tools and settings; they are not universal estimates of AI performance. Self-reported productivity should be distinguished from measured results. The data-protection discussion provides general information, not legal advice.

← Back to all notes

Continue reading