Skip to main content
AI & Technology

AI Workflow Automation: What Should You Automate First?

Choose a first AI workflow by frequency, variation, risk and review effort. Build a bounded pilot with clear ownership, costs and a defined stop condition.

MattDarm8 min read
Cafe worker beside a tablet, coffee equipment and a counter display
Illustration accompanying MattDarm's guide: AI Workflow Automation: What Should You Automate First?.

Key Takeaways

  • Map the current process before selecting a tool.
  • Frequency alone does not make a task low risk.
  • Compare rules-based automation, AI drafting and tool-using agents separately.
  • Pilot one reversible workflow and measure the result after human review.

The best first automation is not necessarily the task that looks most impressive in a demonstration. It is a process with a clear purpose, known inputs, manageable risk and an owner who can judge whether the result is useful. Without those conditions, automation can move confusion faster rather than remove it.

This guide helps a UK business choose and scope one initial workflow. It does not assume that every repetitive task needs AI or that a pilot will pay for itself within a fixed period. Use the worksheet to make the decision explicit before buying software or connecting customer systems.

Map the task as it happens today

Write down the trigger, inputs, steps, output and receiving person or system. Include the exceptions: missing information, duplicate requests, unclear instructions and cases that need a manager. Ask the person doing the work to review the map rather than relying only on a manager's description.

Record where information is copied between systems and where judgement is actually required. A process may contain several routine steps and one sensitive decision. That suggests automating the routine preparation while keeping the decision under human control.

Identify the current source of truth. If different spreadsheets disagree, an automation needs a rule for resolving the conflict or a way to stop and ask. Giving a model access to all of them does not automatically establish which version is correct.

Create a shortlist based on actual workload

Collect candidate tasks over a normal working period. Record how often each occurs, how long it takes, who does it and what happens when it goes wrong. Include review and follow-up time, not just the visible moment of data entry.

Good candidates might include preparing an internal meeting summary from approved notes, classifying an enquiry for review or creating a draft project checklist. These are examples of possible patterns, not evidence of completed client implementations or guaranteed savings.

Do not dismiss a small task if it creates frequent interruptions, but do not automate it merely because it is annoying. Compare the implementation and maintenance effort with the actual workload. A better form, clearer template or ordinary software rule may remove the problem more simply.

Score frequency and variation separately

Frequency tells you how often an improvement could matter. Variation tells you how difficult it is to specify the correct behaviour. A task performed every day with a fixed input may suit a conventional integration; a task involving ambiguous documents may need interpretation and review.

Use a simple low, medium or high score and write the reason beside it. Avoid a complicated numerical model that gives a false impression of precision. The purpose is to make assumptions discussable, not to prove that a task with the highest score must be automated.

For each candidate, ask what an acceptable output looks like. If the team cannot agree, resolve that before adding AI. An evaluation becomes much easier when the expected result and escalation conditions are explicit.

Treat impact and reversibility as gates

A frequent task can still be high risk. Sending payment reminders, changing account access or issuing refunds may be repetitive, but errors can affect customers, money or confidentiality. Do not label such tasks low risk simply because the steps are familiar.

Ask what the system could change and whether an error can be reversed. A draft in an internal queue is different from a message already sent to a customer. An internal classification is different from a decision that denies someone a service.

Keep consequential actions behind a suitable approval boundary. The agentic AI guide explains why tool permissions matter. A prompt asking the model to be careful is not a substitute for application-level controls.

Choose the simplest suitable approach

Use a rules-based workflow when the trigger and action are predictable. Use AI assistance when interpreting variable language or preparing a draft adds value. Consider a tool-using agent only when its ability to choose actions solves a genuine problem that the simpler approaches cannot handle well.

Write the boundary in ordinary language: “Read approved enquiry fields and suggest a service category; do not send messages or change the customer record.” That is easier to test than “automate our sales process.” Expand the scope only after the first boundary has been demonstrated reliably.

Compare integration ownership as well as model capability. Who maintains credentials, handles provider changes and investigates failed runs? A cheap subscription can still be expensive to operate if the workflow depends on undocumented connections nobody understands.

Prepare the data and access plan

List the exact data needed and exclude information that does not serve the task. Use synthetic examples during early testing where possible. Check whether confidential or personal information is being sent to an external provider and whether the intended use has been approved.

Review retention, access, logging and provider terms. A business account is not blanket permission to upload every document. Use the ICO's AI guidance as a starting point and seek appropriate advice for the actual data and decisions involved.

Store secrets in the approved configuration, not in prompts or shared worksheets. Give the workflow only the permissions it needs and make revocation straightforward. Record the person responsible for approving any later increase in access.

Design a draft-first pilot

Choose a bounded number of representative tasks and a defined period. Have the system prepare an output for human review while the existing process remains available. Do not quietly move from a test queue to live customer actions because the first examples look promising.

An illustrative enquiry-triage pilot could read a permitted set of test messages, propose a category and explain which details are missing. A staff member checks the suggestion before any record changes. Include ambiguous and out-of-scope requests so the system has to demonstrate when it should abstain.

Keep the pilot record separate from real customer reporting. Label synthetic entries and ensure they do not trigger live notifications, marketing or billing. The pilot should reduce uncertainty without creating cleanup work in production systems.

Test failures before judging success

Include missing inputs, conflicting information, unavailable tools and repeated requests. Check whether a retry creates a duplicate entry. Confirm that the workflow stops when approval is refused or when a required condition is not met.

Test untrusted instructions inside the material being processed. An email or document should not be able to grant itself permission to export records or bypass review. The NCSC's secure AI development guidance provides a useful lifecycle framework for these considerations.

Document the expected behaviour for each test. A successful response is not always a completed task; sometimes the correct outcome is to ask for information or decline an action. Rewarding completion at any cost can produce the wrong operating behaviour.

Measure the net effort, not the demonstration speed

Record time spent on preparation, running the workflow, reviewing, correcting and maintaining it. Compare that with the manual baseline for similar tasks. Keep output quality and error severity visible alongside the time measure.

Use a cost worksheet with setup, subscriptions or usage, integration support, review time and expected maintenance. Add an allowance for investigating failures where appropriate. Do not present every saved minute as cash profit; the benefit may be capacity or faster response instead.

Agree what would justify continuation. That could be a reliable reduction in routine preparation with an acceptable review burden. It could also be a decision to stop because the exceptions are too complex. An honest pilot can be successful by preventing a larger unsuitable investment.

Assign ownership and a stop procedure

Name the person responsible for the workflow, its source information and its access permissions. Provide a backup and a way to pause the system without waiting for the original developer. Keep the manual route documented so an outage does not halt the business.

Record which provider, model, prompt and integration versions were tested. Re-evaluate after material changes. A system that worked last month may behave differently after a source update or permission change, even if the visible interface looks the same.

Use the business AI policy template to make staff responsibilities clear. Keep approval for the specific use case separate from a general decision to allow AI tools in the business.

Turn the pilot into a scoped brief

Your brief should contain the task map, baseline, permitted inputs, proposed output, prohibited actions, approval owner, evaluation examples, cost limit and stop condition. Ask suppliers to respond to those specifics rather than a broad request to make the business more automated.

MattDarm can help with AI automation setup, AI workflow automation and AI strategy consulting. Share one current process, including the difficult exceptions. We can then assess whether a simple integration, a draft assistant or a more capable system is justified.

Frequently Asked Questions

What should our first AI automation do?

Choose a task with a clear output, manageable risk and an owner who can evaluate it. Start with a bounded, reversible scope rather than connecting an agent to the whole business.

Does a repetitive task always suit automation?

No. Frequency does not establish low risk or clear rules. Financial, sensitive or customer-facing actions may need careful approval even when the same type of task happens every day.

How long should a pilot run?

Use a period and number of examples sufficient to cover normal work and meaningful exceptions. Set the limit in advance and avoid a universal promise of payback or readiness after a fixed number of days.

What costs should we include?

Include setup, provider usage, integrations, review, correction, maintenance and incident handling. Distinguish improved capacity from a reduction in actual business expenditure.

When should we stop a pilot?

Stop or reduce its scope when it cannot respect permissions, produces unacceptable errors, needs excessive correction or lacks a reliable owner. Keep a manual fallback and document the decision.

AI AutomationWorkflow AutomationAIProductivityUK Business

Share this article

Stay ahead of the curve

Weekly insights on web development, AI, branding & digital marketing. No spam, unsubscribe anytime.

By subscribing you agree to our Privacy Policy. Unsubscribe at any time.

Adam Saez
Alina Stefanovičiūtė
Daniel Ashby
Matt Laybourn
Richard Jones
Paul Campbell

People we've worked with

Real projects, built together

Let’s Grow Your Business Together

Tell us about your project and we’ll show you exactly how we’d grow your business. Book a free 30-minute discovery call, no pressure.