On 4 September 2026 a single sponsored Instagram post set the UK accounting world alight for a week. Amelia Sordell, a British personal branding specialist with 77,000 followers, wrote in a post sponsored by Xero, the cloud accounting software, that she had paid accountants more than 120,000 pounds since 2020. Then she connected Xero to Claude and started building her own management accounts, a job she had been paying 800 pounds a month for somebody else to do.
Accountants revolted on LinkedIn within a day. Laura McKenzie called it the last straw and said she was leaving Xero. Kate Gloudemans said much the same. Damon Anderson, formerly operations director at Xero UK, wrote a blog post hitting back. Xero pulled the post. Managing director Kate Hayward apologised, saying the ad "does not reflect our values … and it should never have run", and Angad Soin added another layer: "The post contradicts everything we stand for", confirming that the company neither wrote nor approved the copy.
Productize works on moving accounting document work across to AI, and keeps what comes out of real engagements as a series of posts. This article lays that task map over the row that just happened, so you can see where the line actually falls.
What was missing from that argument was not information about who was right. It was the unit both sides were arguing in.
Why are both sides arguing in the wrong unit?
Because both sides argued at the level of the profession. One said AI can replace it, the other said it cannot, when the profession is not the unit AI works in. The real unit is the individual task, and every task in an accounting firm comes with its own conditions. No two are alike.
Split the row into two sentences and it gets clearer.
First sentence: Amelia was right. The job she stopped buying was building a management accounts spreadsheet. It is a set of numbers she produces for herself. Nothing is filed, nobody signs it off, and if the numbers drift, she finds out next month at the latest, when the real figures arrive and do not match what she assumed. Work shaped like that genuinely can move. That part is not advertising copy.
Second sentence: Amelia was wrong. She took one task that can move and made it stand in for a whole profession. The 120,000 pounds paid since 2020 was not spent purely on spreadsheets. Bundled inside it were several other things that never got their own line on the invoice. That is where the conclusion breaks.
And the angry side? Arguing at the level of the profession too. The loudest answer on LinkedIn that week was "AI cannot replace accountants", which is true in the aggregate and useless in practice, because nobody in that conversation named a single task and said which ones had already moved and which had not.
There was something real underneath the anger. In July we stood in front of several hundred people from accounting firms, and what we heard most often was not a question about tools. It was four feelings almost nobody says out loud.
- Afraid of being left behind, with no idea where to start
- A sense that AI belongs to technical people and is not an accountant's subject
- Worry about how much there is to learn, on top of a workload that is already heavy
- Clients have started asking about AI, and there is no confident answer to give them
The last one is the hardest to answer when a post like Amelia's lands, because clients bring it up that same week. And "AI cannot replace us" does not make a client any more confident. What makes a client confident is being shown, task by task, which work has moved to the machine, which work is still in human hands, and why.
The map below is that answer. We split accounting work into five groups by the shape of the work, not by the job title of the person doing it.
Group 1Intake and extraction. Already moved, if you bolt a proof layer on the end
This group really has moved, but only where a proof layer sits at the end of it. Without one, all you have bought is speed, and the speed is carrying your errors along faster.
The work here is the pile on the desk every morning. Files come in, get grouped, get named to a system, get checked against what has already arrived so duplicates do not slip through. Then the numbers come off the bills and tax invoices and turn into data a machine can carry on working with, ready for import into the accounting system. AI can do that whole run today.
The problem is how convincingly AI gets it wrong.
We put one bill in front of several models. One read the total as 5,390 baht. Another read the same bill as 4,851 baht. Both answered in exactly the same confident tone. Neither flagged any doubt. If you take the first answer you get and do not check it, you have no way of knowing which one you got that day. The test is written up in verifying OCR accuracy.
This is where the proof layer earns its place. The numbers on a tax document are tied together by equations that check themselves, without anyone having to trust the model's eyesight.
- Net amount multiplied by 0.07 has to equal the VAT printed on the bill
- Net amount plus VAT, less withholding tax, has to equal the net payment
- Services are withheld at 3 percent and rent at 5 percent. If the extracted rate does not match the type of expense, something went wrong back at the reading stage
Those three rules do not make AI read more accurately. They make the misread document announce itself before it reaches a human. Documents that pass every rule flow straight into the accounting system. Documents that trip a rule get bounced out one at a time for a person to look at.
How to lay that whole run out, from file intake through to the gate before import, is written up in detail in extracting invoice data with AI. And if you are still deciding whether to build it or buy it, DIY versus paid invoice OCR tools came out of teaching real accounting teams, not out of a comparison table on a vendor's website.
This group can move, and it should move first. But it moves with its checking rules attached, never on its own.
Group 2Reconciliation and checking. Half moved, deliberately
Half of this group has moved. The half that has not is not stuck because a machine cannot do it. It is stuck because a machine should not.
The first half is matching. Take the books, put them next to the bank statement, and go line by line. What matches gets paired off. What does not match gets pulled into a pile. That is work a person does until their eyes ache, using no judgement whatsoever. A machine does it faster and does not lose concentration halfway through.
The second half is explaining the pile that did not match, and it is an entirely different job. One unmatched line might mean a cheque has not cleared. It might mean something was posted to the wrong date. Or it might mean money left the account and nobody knows about it. All three look identical in the table. What separates them is context, and the context lives in the head of the person who has been looking after that account all year.
So the test we apply to reconciliation work is not "does the AI say it balances". It is "can this report be traced back to every individual item". A report that only says it balances is worth nothing, because there is no way to tell whether it balances because the matching was right or because the matching was invented. How to make a report traceable is in AI bank reconciliation.
Vouching follows the same logic. The job is checking whether a recorded transaction has a real document behind it. AI can pair transactions with documents, and it can surface the ones where no supporting document turns up. But the call on whether a missing document is a control failure or just how this business operates still belongs to a person. That is covered in using AI for vouching.
Hunting for exceptions in the general journal sits in the same group. Ask it to find expenses above a set threshold that were paid in cash, or to run sales tax against reported revenue and see whether the two move together, and a machine sweeps it in minutes. What comes back is a list of things worth asking about. It is not a conclusion about who did something wrong.
Put another way: AI in this group is the one gathering things onto the table. A person is still the one deciding what the things on the table mean.
Group 3Producing documents and reports. Nearly all moved, except anything that binds
This group has moved further than any of the other four, because most of the work is taking data that already exists and shaping it into a format that is already fixed.
A hundred quotations generated from a single Excel file is the cleanest example. The data is all there, the format is settled, and what remains is pure labour. There is no longer a reason for a person to sit and produce those by hand.
One layer up sit the withholding tax certificates, the 50 Tavi that has to be issued to the counterparty every time tax is withheld, and then the monthly returns built on top of them, PND.3 for individual payees and PND.53 for companies. All of it is assembled from the same data that came off the documents back at intake. If the checking rules were laid properly at that stage, there is almost nothing new to do here.
Financial statements can be produced too, shaped up from the trial balance using last year's version as the starting template so the headings and groupings stay consistent. But this needs saying plainly: what comes out is a draft, not a finished set. The difference is that financial statements leave the building, to a bank, a government body or shareholders, and somebody has to sign their name at the bottom.
Financial dashboards sit at the opposite end from financial statements. They are quick to produce and nobody signs them, which makes them just as good a first thing to move as anything at the intake stage. And they happen to be exactly the kind of work Amelia found she could do herself.
The dividing line in this group is sharper than in the others. If the document does not bind you to anyone outside, it can move completely. If it does bind, it moves as far as a draft, and a person checks it before it goes out. Every time.
Group 4Automation across apps. Cost decides this one, not capability
This is the only group where "can it be done" is the wrong question, because the answer is yes. The question worth asking is whether it is worth it.
The work here is an agent chaining steps together across several programs: opening web pages, filling in fields, pulling a file from one place and putting it somewhere else, running on a schedule every morning. When we went and learned how to apply AI to accounting work ourselves, the person teaching it demonstrated the whole thing end to end, an AI filling in a web form to 100 percent completion, and then said flatly in the next breath that it burns through tokens until you are broke.
That was the truest sentence of the day. Driving software from the outside eats several times the tokens that working inside the program does. Every time the agent has to look at the screen, decide where to click, then look at the result again, the meter is running. A job you do once is fine. A job that has to run several times a day, every day, will catch up with the cost of paying a person faster than you expect.
What about routine work where the shape is identical every single time? RPA is better value there, because RPA follows a fixed sequence without having to work anything out afresh on each run. Every click and keystroke was decided in advance, so the cost per run is close to zero. AI is stronger where the work does not look the same twice, where documents arrive in different formats and use different words for the same thing, and something has to be read and interpreted before anything else can happen.
So the workable test in this group is a single question: how variable is this task? Identical every time, use RPA. Highly variable, then pay AI for the flexibility.
And if you are still unsure, the safest approach is to do it by hand for a month, actually timing it, and then compare that against the bill you would run up. Those two numbers answer the question for you.
Group 5Review, and the call on whether it can go out. The group that has not moved
This group has not moved, and the reason is not a technical limit. It is a decision about who carries the blame when something is wrong.
The difference between preparing accounts and reviewing them is a difference of purpose, not of seniority. The person preparing has one goal: get it right and get it in on time. The reviewer's goal is something else entirely. They are weighing up how much the whole set of numbers can be relied on, and then deciding whether it can go outside or has to go back for rework.
A reviewer carries four standing questions, put to every number that looks off.
- Where did this number come from?
- What is the supporting documentation?
- Does this number make sense for this business?
- How does it compare with last year?
And there is a question that comes before all four: what is the client going to do with this set of accounts? Apply for a bank loan, file it with a government body, or hand it to shareholders. The destination decides where you look hardest. A set of accounts going to a lender and a set produced for internal reading get scrutinised at completely different depths, even though the underlying numbers are identical.
That whole skill already has a name, professional skepticism, and we have written it up in full at professional skepticism, the skill AI cannot replace. All this article needs to say is that it is group five, and group five has not moved anywhere.
Where people usually get this wrong is assuming the group is stuck because models are not good enough yet. In fact, put those four questions to a model and it will answer all of them, and answer them quickly. What a model cannot supply is somebody who carries the blame when the answer is wrong. A bank taking a set of accounts into a credit decision does not just want an answer. It wants a person standing behind it.
Should the task in front of you move yet? Four questions
Work through the four questions below one at a time. Once you have answered all four, you will know which group of the map that task belongs to, without waiting for anyone to tell you.
Criteria box · four questions before you hand a task to AI
| Question | Answer like this and it can move | Answer like this and hold off |
|---|---|---|
| 1 · Does this task have a correct answer that a document can prove? | Yes, and it can be re-checked with a rule, such as net amount times the tax rate matching the figure on the bill. It can move, but the checking layer moves with it, always. | The answer depends on interpretation and on the context of the business. No document settles it. |
| 2 · If it goes wrong, when do you find out? | Straight away, or within this month's cycle. | Not until year end, or until somebody outside asks about it. |
| 3 · Who carries the blame when it is wrong? | Nobody has to sign. The consequences stay in house. | The work ends in a person's signature, or somebody outside uses it to make a decision. |
| 4 · Is the cost per run worth it? | High volume, variable work, and a monthly bill lower than the time you get back. | The shape repeats exactly, so something cheaper does the job, or it eats tokens until it stops paying. |
Question four is the one people skip most often, because "it can be done" and "it is worth doing" are two different questions, and the two answers can differ completely for one and the same task.
Try ticking these four off against the work sitting in your week. The tasks that answer yes on all four are usually the ones eating the most of your time already, and the ones nobody on the team wants.
What is left over does not just remain. It gets more expensive
The reassuring line people reach for is "there will always be work left for humans", which makes it sound like leftovers after a division. The reality runs the other way. Once AI speeds up everything around it, what a person holds does not lose value. It gets dearer.
Judgement. Deciding whether to trust a number, in a world with far more numbers on offer than before. It used to take days to assemble one set of figures, and the person assembling them was reviewing them the whole way through as a side effect. Now the same set arrives in minutes, and nobody has reviewed anything along the way.
Taste. Knowing which report actually tells you something and which one is handsome but useless for making a decision. That skill used to be rare because producing a single report cost real effort. Now you can commission several versions before lunch and compare them, which makes the person who can say which one is genuinely usable the deciding voice.
Checking AI's output for what is right and what is wrong. This one only just became a daily job. Nobody had to do it before, and nobody was ever taught it in a classroom. A person who can check a machine's work does it in a minute. A person who cannot takes the whole thing on faith, and that is exactly how one bill with two possible totals flows through into a real number in a real set of accounts.
Understanding the story behind a number. A rising receivables balance means sales are strong, or it means collections have stalled. The number is identical. The story is not, and no tool separates the two without knowing the business.
Now go back to the row. What Amelia could do for herself was produce numbers. What she has been buying from accountants all along, and still needs to buy, is those four things, which have never had a line of their own on any invoice. So when the time came to cut a cost, she could not see them. Nobody had ever written them down where she could.
On stage in July we said one thing to the room: we are not here to tell you to chase AI, we are here to get AI ready for you, with you still being you. And we closed with this: the time you get back, spend it on what AI cannot supply, your expertise, your professional opinion, your decisions.
That closing line holds up just as well this week. You know your clients' businesses better than anyone. All we want is for you to have the time to actually use that knowledge, without having to turn into somebody else.
Who should be answering this question?
People inside the profession. Not software vendors, and not people with large followings.
On the same day the post was pulled, Bill Gates released a 162 second clip explaining why he chose to write about AI now. The reason he gave was "the lack of engagement outside of the industry". Almost everybody speaking loudest on the subject comes from the technology industry, while the people the results will land on have barely any standing in the conversation.
He also said it was the first time he had put it this way: "if we're not careful, the negatives could outweigh the positives", which is a different register from how he has always sounded. And the line closest to the point of this article was this one: "I should only be one of many, many thousands of spokespeople".
Lay that over the row that just went by. The people who got to talk about the future of accounting work that week were one personal branding specialist and a software company that had to issue an apology. The people who do this work every day, and who know better than anyone which tasks have moved and which have not, got to speak as people who were angry, not as people who know.
Nobody can fill that gap except the profession itself. And the way to fill it is not by arguing that AI cannot replace us. It is by pointing, task by task, at what has moved, what is still in our hands, and why.
Take the five group map and the four questions and tick them off against your own work once. You will have an answer ready the next time a client asks.
Frequently asked questions
If clients start producing their own accounts with AI, what is left for the firm?
All of group five, and the second half of group two, which are the parts that already bill higher than data entry. What a client can do for themselves is management accounts that get filed nowhere. Accounts going to a bank for a loan, or to a government body, still need somebody to sign their name at the bottom, and that somebody is not a model.
Where should I start if I only have a few spare hours a week?
Start with group one, because it shows results fastest and it ticks all four criteria. Pick the single document type that arrives in the highest volume, and build one path all the way through, from file intake, through the checking rules, to the import into the accounting system. Then extend to the next document type. Do not start by doing every type at once, because you will never work out where it broke.
Does AI misread bills often, and how would I know which runs went wrong?
It does, and it misreads them confidently. On one bill we tested, one model read 5,390 baht and another read 4,851 baht, and neither signalled any uncertainty. The approach that works is to stop measuring the model's confidence and lay mathematical rules over the output instead, such as net amount times 0.07 having to equal the tax printed on the bill. Anything that fails a rule gets bounced to a person.
Do I need to be able to program to use any of this?
No. What you need is to be able to describe your own work as steps that can be checked right or wrong, which accountants were doing long before AI existed. The person who writes the best checking rules is the person who knows what this type of document normally looks like, not the best coder in the room.
Sources and references
Xero row, 4 September 2026
- SmartCompany, "Xero apologises for Instagram sponsored post" https://www.smartcompany.com.au/marketing/xero-instagram-sponsored-influencer-post-ai-accountants-apology/
Bill Gates, 4 September 2026
- 162 second clip https://www.youtube.com/watch?v=ReogxIL1rBk All quoted lines in this article come from the clip itself, not from the full memo.
Our own work referenced here
- Verifying OCR accuracy https://productize.life/blog/ocr-accuracy-verification/en
- Extracting invoice data with AI https://productize.life/blog/ai-invoice-data-extraction/en
- AI bank reconciliation https://productize.life/blog/ai-bank-reconciliation/en
- Using AI for vouching https://productize.life/blog/ai-vouching/en
- DIY versus paid invoice OCR tools https://productize.life/blog/ai-invoice-ocr-diy-vs-paid/en
- Professional skepticism, the skill AI cannot replace https://productize.life/blog/professional-skepticism/en
Notes
- The five group map was put together from going and learning how to apply AI to accounting work ourselves. Every piece of supporting evidence quoted here comes from work we did.
- Organisation level AI policy and governance is a separate article. This one stays at the level of the individual tasks an accountant does.