Automation Software and Where AI Actually Fits
Automation software now means two layers since AI arrived. What changed, a worked test for whether a task is ready, and the checks before you trust it.
"Automation software" used to mean one thing: a tool that fires a fixed action when a fixed trigger happens — a new row, a form submission, a scheduled time. That layer still exists and still works the same way it always did. What changed is a second layer sitting on top of it: a model that can read something unstructured — an email, a receipt, a free-text ticket — and decide what the rule-based layer used to need spelled out in advance. Most tools sold as "automation software" today are really both layers stacked together, and knowing which layer is doing which job is the whole trick.
The question before you buy anything
Most searches for automation software jump straight to comparing products. The question that actually decides whether any of them help is narrower: is this specific task a good candidate at all? Find the repetitive part sets out the test — does it happen often, is the input predictable in shape, is a mistake cheap and visible, could you explain the steps to a new hire in five minutes. A task that fails that test will go wrong inside expensive software exactly as fast as it goes wrong inside a free one.
A worked example
Take a concrete task: every Friday, someone collects submitted expense receipts from an inbox, enters the amount, date and vendor into a spreadsheet, and flags anything over the policy limit for a manager to review.
Weak approach: point an automation tool at the whole inbox and ask it to "handle expenses." That hides two different jobs inside one instruction and gives you no way to tell which one broke when something goes wrong.
Better approach: separate the mechanical part from the judgement part before configuring anything. The mechanical part — reading a receipt and pulling out amount, date and vendor into fixed fields — is a narrow extraction task AI genuinely handles well. The judgement part — deciding whether an expense is legitimate — stays with a person. The automation only prepares the case for that person to review faster.
- Trigger — a new receipt lands in a specific inbox folder.
- Extraction fields — amount, date, vendor, read from the receipt image or text.
- Flag rule — a plain comparison against the policy limit, not a model opinion on whether the expense "looks reasonable."
- Output — one row per receipt in a shared sheet, with flagged rows highlighted for a manager to open.
- What stays manual — approving, rejecting, or querying anything flagged.
Business process automation strategy covers scoring a whole queue of candidate tasks this way before picking which one to automate first, rather than starting with whichever one feels most urgent.
Where it goes wrong
The extraction step is where this fails quietly. A model reading a smudged or oddly formatted receipt will produce a confident-looking amount even when it misread a digit — a well-documented property of how these models generate text, not a rare glitch specific to receipts. A wrong amount that is close to the real one is worse than an obviously broken one, because nothing about the output looks wrong until someone checks the original.
The flag rule fails differently: a threshold copied from last year's policy, a currency mismatch between two regions, or a rule that fires on the pre-tax amount when policy means the total — none of these are AI failures at all, just configuration drift that AI-driven extraction makes easier to overlook because the sheet still fills in and still looks complete.
Checks before it runs unattended
- Sample a slice of auto-extracted receipts every week and compare the fields against the original image by eye.
- Confirm the flag threshold matches the current written policy, not a number someone remembers from when the flow was built.
- Check currency and tax handling explicitly if more than one applies — a silent mismatch here is the most common real-world failure.
- Decide in advance what you would look for if the flow silently started misreading receipts — a flagged-rate that goes to zero, or a manager who stops finding anything to review, are both signs worth checking rather than celebrating.
That last habit — deciding up front what a silent failure would look like, rather than waiting to notice one — is the same discipline behind the NIST AI Risk Management Framework, applied to one small flow instead of an enterprise system. IT process automation and CRM workflow automation cover the same pattern applied to a help desk queue and a sales pipeline respectively, if either is closer to your actual task.
Where this sits in the bigger picture
Stanford's AI Index tracks automation and AI adoption climbing faster than measured productivity gains, which is consistent with a lot of deployments automating the wrong half of a task rather than the narrow, checkable half above. A separate Gartner survey of 183 CFOs and senior finance leaders found adoption holding steady rather than accelerating — a sign that even functions under real cost pressure are moving carefully rather than automating everything available.
A Microsoft Research study scoring which real occupational tasks these tools actually apply to keeps landing on the same shape as the expense example: narrow, well-defined sub-tasks, not whole jobs. Indeed's Hiring Lab job-posting data points the same way — demand is shifting toward people who can configure and check these flows, not away from the people who used to do the task by hand. Examples of automation at work and no-code automation are worth reading next if the task you have in mind is not expenses specifically — the separation between mechanical and judgement work holds across most of them, and data automation tools covers connecting a flow like this one directly to a live data source instead of an inbox. Social media automation is a worked example of the same split applied to drafting and scheduling, if the task on your list is a content queue rather than a back-office process.
What to do Monday
- Pick one task you currently do by hand every week and write down which parts are mechanical and which require a judgement call.
- Configure automation software for the mechanical part only — a fixed trigger, fixed fields, a plain comparison rule.
- Keep every judgement call with a person, routed to them faster rather than removed.
- Sample-check the mechanical output against the source for the first few weeks before trusting it unattended.
Coursium teaches this kind of practical judgement — separating what a tool can check from what a person still has to decide. Stay ahead of AI by learning the tools on your phone.