A blog or document library that has been running for a few years usually ends up like this. The tags started as a dozen or so. Whoever wrote a post made up new tags on the spot. Posts on the same topic carry ai, AI and artificial-intelligence, each a different tag. A reader clicks a tag and gets 2 posts, when the library holds ten more on that topic. The tags quietly stop doing their job.
Then the thought comes: "Can't we just have AI retag everything? Give it the tag list and done."
The job here is choosing, for each post, which category it belongs in and which tags it gets, from a list you already have. A decision model, the kind of AI we tested on Thai withholding tax in the first post in the series, is built for exactly this kind of pick-from-a-list job. So we tried it for auto-tagging, assigning categories and tags automatically, on the posts of our own blog.
Productize tried a decision model for auto-tagging on all 127 posts of our own blog. Filing accounting documents into categories is a similar kind of problem: you already have a list of categories, and you want a machine to drop each item into the right slot.
Part 1What is a decision model, and why does it fit auto-tagging?
A decision model is an AI that doesn't answer in sentences; it scores every option you send it. You send in one item at a time with a multiple-choice question, and the options are fixed in advance. All that comes back is a score for each option, no other text. When the question allows only one answer, the scores across all options add up to exactly 1. The highest-scoring option is the answer, and its score tells you how sure the model is. When the question is yes or no, each question gets its own "yes" score back.
That shape fits tagging well. The category list and tag list you already have go in as the options. Each one carries a short description of what it covers, and that description is the criterion the model judges by. The list is the options; the descriptions are the criteria. So the model cannot answer outside the list, and no new tag appears out of nowhere.
Here's an example from our blog. For a post about what Claude Code costs, we sent the title, the short description and the section headings, with a question asking it to pick one of 10 categories. The top-scoring category was "SME business" at 0.49. The rest was spread across several categories, such as "AI agent" at 0.16, "Models and cost" at 0.15 and "Claude Code and skills" at 0.15. The top pick didn't even reach half, which means the model was torn. On the yes-or-no questions, one per tag, cost got 0.94, claude-code 0.84 and sme-business 0.73. All 3 tags can score high at once, because they don't share a score. We'll come back to the difference between these two kinds of question several times.
This post is part 2 of our decision model series. The one people are talking about right now is Jev from TypeSafe, which is open only to a first group of early users (reference 3). We did not test Jev. What we actually used is Cloudflare's Clef-flash, which returns scores the same way (reference 2).
Part 2Turning tagging into decision model questions
We sent each post once, and that one request carried two kinds of question. The category was a single question, pick one of our 10 categories, and each category went in with its description. The "Models and cost" category, for example, was described as "choosing, running and paying for AI models, GPU hosting, running models yourself, benchmarks". The tags, short English names such as claude-code or cost, were asked one at a time as yes or no: is this post about this topic? Each question carried that tag's description.
All we sent was the title, the short description and the section headings, in Thai, not the full text. We also changed the category and tag list during the test: it started with 24 tags, one category was renamed partway through, and after we got a person's answer key we revised it again into the list we use now.
We set the confidence line at 0.7 before we started measuring, as the point where the model is fairly sure. When a category reached the line, the model decided on its own. When it didn't, the post went to 2 AI reviewers, OpenAI's Codex and Anthropic's Claude Opus, who read it separately without seeing each other's answers. Tags use the same line: a tag goes on when its "yes" score reaches 0.7. So the Claude Code post above got 3 tags, while its category, at only 0.49, had to be routed.
To know whether the model is right or wrong, you need correct answers to compare against. We ran 3 rounds. Each round sampled about 20 of the 127 posts, and we decided in advance which category and tags were correct. That set is the answer key. Who made the answer key matters more than you'd think, so we say it every time we quote a number. Before each round we wrote down the pass mark first, so we couldn't move it after seeing the results.
The tags on our site came from three places over time:
- When each post was written: a person added its tags.
- Test day, before round 1: we switched to the two-AI method: the 2 AI reviewers tag each post separately, and only the tags both picked go on.
- After the tests: we switched to the top-3 + confirm method, which is what runs now. It's described in the tags section.
That first stage gave us one more comparison for free: the person's old tags, renamed to the new list with a fixed lookup table and no AI involved. We measured them the same way as the model.
Part 3Categories: the decision model can decide when confident, but the answer key still came from AI
In round 1, the category answer key came from the 2 AI reviewers. We sampled 20 posts, had both pick a category, and kept only the posts where they agreed. That left 18 posts. This key came entirely from AI and was never checked against a person. The pass mark we wrote in advance: the answers the model decided on its own had to be at least 9 in 10 right, and it had to decide at least half the posts on its own.
It passed, just barely. The model reached the line on exactly half, 9 posts, and got all 9 right. The 3 posts where it picked the wrong category were all in the half below the line, so in real use they would have reached the reviewers before going live.
After the 18 test posts, we tried the whole blog, all 127 posts. The picture pointed the same way. In the group where the model wasn't confident, its top category matched what the 2 reviewers picked only about half the time. If the reviewers were right, the model really is often wrong when it isn't sure, so routing those posts was worth it. For the confident group in this whole-blog run there was no key to score against, because the key only existed for the test set.
The Claude Code cost post from the start is one of these. Its top category was "SME business", below the line, so the system routed it. Both reviewers picked the same category, "Models and cost". If we had let the model decide everything, this post would have landed in the wrong category.
Everything above was the test. The categories on the live blog were set later, with the list we use now. The table below shows where each post's category came from. No post had its category picked by a person.
When we labeled the live blog, we also had the 2 reviewers pick a category for every post, not just the ones below the line, as an extra check. That's how we saw 3 posts where the model was confident but both reviewers disagreed and agreed on a different category. For those posts we used the reviewers' category. On our blog, the extra check overturned only those 3 confident answers, a small share of the confident group. So the process below checks a sample of the confident group rather than all of it. We have not measured how big that sample must be.
| Where each live category came from | Posts |
|---|---|
| Model reached the line, model's pick used | 73 |
| Model reached the line, but both reviewers picked another category, reviewers' pick used | 3 |
| Model below the line, reviewers agreed | 50 |
| Model below the line, reviewers disagreed, the more confident reviewer's pick used | 1 |
No row in this table has been measured against a person's answer key. So the category side looks workable, but only measured against AIs that agree with each other. Once we got to tags, that mattered far more than we expected.
Part 4Tags: 3 rounds, none passed
Tags need two things measured at once. Take the Claude Code cost post again. The key has 2 tags, claude-code and cost. The model said yes above the line to 3 tags: those 2, plus sme-business, which is a tag, separate from the "SME business" category. It added 3 and 2 were right. That ratio is precision: 2 out of 3. The extra one we count as a wrong tag. Of the 2 tags that should be there, it found both. That ratio is recall: 2 out of 2. A tag that should be there but wasn't found we count as a missed tag. Tag more, and recall tends to go up while precision goes down. Tag less, and it goes the other way.
Round 1 used the same 18 posts. The tag key was the tags the 2 AI reviewers both picked. To pass, it had to find at least as many right tags as the person's old tags did, and at least 7 in 10 of its tags had to be right. The model's precision was fine at 0.75, but its recall was only 0.52, below the person's old tags. Fail. It was too cautious. With the 0.7 line applied to the whole blog, 95 of 127 posts got one tag or none (70 got one, 25 got none).
So in round 2 we tried a combined method: the tags the model said yes to above the line, plus the person's old tags for the same post. We sampled a fresh set of 20 posts, and the key still came from the 2 AI reviewers. The pass mark written before the run was recall 0.8 and precision 0.7, a level where a person wouldn't have much to fix: no more than 1 in 5 of the right tags missing. The combined method found 0.63 of the tags in the key, short of the 0.8 bar. This round also showed something else: the reviewers themselves didn't pick tags very consistently. Count every tag either reviewer put on the same set of posts. If that makes 10 tags, only about 7 were put on by both.
The keys for the first 2 rounds came from AI, so in round 3 we changed who made the key. We sampled another 20 new posts and had the blog's owner tag them by hand, without seeing the machine's answers. The plan was to tag every topic the post really covers, but in practice only the main topics got tagged: 30 tags, 1.5 per post on average. Same pass mark as round 2. Then we compared every method.
Besides the 0.7 line, we also tried the model with its line lowered to 0.5, set up before scoring to see whether a lower line helps. We also measured the two-AI method that was live on the site then. The top-3 + confirm method has the model take its 3 highest-scoring tags, then Anthropic's Claude Sonnet reads the same post and says keep or drop for each, keeping only main topics. We fixed this method, and saved the prompt Sonnet gets, before the person tagged anything.
| Method (against the person's key) | Tags added across 20 posts (person added 30) | Precision | Recall |
|---|---|---|---|
| Model, line at 0.7 | 28 | 0.61 | 0.57 |
| Model, line at 0.5 | 34 | 0.53 | 0.60 |
| Two-AI method (live on the site then) | 45 | 0.47 | 0.70 |
| Top-3 + confirm | 30 | 0.67 | 0.67 |
| Person's old tags (renamed) | 39 | 0.56 | 0.73 |
No method reaches recall 0.8 with precision 0.7. The top-3 + confirm method is the most balanced: it added 30 tags, the same as the person, with 20 right and 10 wrong. If it adds 3 tags, about 1 is wrong, and of 3 tags that should be there, it misses about 1. The person's old tags find more, but add a lot of extras.
Part 5Why might a decision model fit categories better than tags?
Our best explanation is that a choose-one question makes the options compete for the score, while yes-or-no questions, one per tag, have nothing that says how many tags to stop at. We haven't tested this, and the two sides were checked against different keys: categories against AI, tags against a person.
When picking a category, the scores of the 10 categories have to add up to 1. If a post straddles 2 categories, the score splits into 2 chunks, and the winning category ends up below the line on its own. So the 0.7 line can pull the unclear posts out and send them to the reviewers. In a question like this, a low score is itself the signal that the post is ambiguous.
Tags have no pull like that. Each tag is a separate yes or no, so one post can get high yes scores on several tags at once. The number of tags per post then follows the line you set, not how many a person would add. At the 0.7 line, 95 of 127 posts got one tag or none. Lowering the line to 0.5 added more tags, but precision dropped.
A post about renting powerful machines from a service called Modal to run models ourselves is a good example. We tagged it with just one tag, local-llm (running a model yourself instead of calling someone else's service). The model said yes above the line to hosting (renting a place to run systems) and cost, and put local-llm below the line. Scored against the key, that's 2 wrong and 1 missed. But if you read the title and description, the post really is about budget and where to run the model. Neither side is clearly wrong. The only difference is whether secondary topics count as tags.
Even the 2 AI reviewers got stuck here. On categories, they agreed on almost every post, even in the group where the model wasn't confident. On tags, of 10 tags either one added, only about 7 were added by both, as we described in round 2. So part of tagging is taste, whether someone counts secondary topics or not. No method can match any one person's key on every tag.
Part 6An answer key made by the same kind of machine can't measure that machine
At the end of round 2, we mostly trusted the two-AI method the site was running then. The 2 AI reviewers had to agree before a tag went on, which sounds careful. We had made the keys for rounds 1 and 2 this way too.
But we had never measured this method at all. Scored against the keys from the first 2 rounds, it would get full marks, because those keys came from the method itself. It's like letting a student grade their own exam with an answer sheet they wrote.
When the blog owner's key arrived in round 3, this method had added 45 tags across 20 posts. Only 21 matched the person; the other 24 were tags the person didn't add. Its precision was the lowest in the table. Two AIs can agree even on topics a person doesn't count as main topics.
Agreeing doesn't mean being right.
Then it hit us that the category results above, the ones that look workable, were all measured with this same kind of key.
Part 7The tag list a person wrote has holes too
A decision model can only pick from the options on the list. If the list is missing a topic, the model has to reach for the closest option instead. So when round 3 failed, we suspected our own tag list as well, and had Codex trace where the wrong and missed tags came from.
It found real gaps. Our blog has a series of posts we wrote following Andrew Ng's AI skills map, about what ordinary people need to learn; Ng founded DeepLearning.AI. But the list had no tag for people's AI skills at all. The closest-named tag was the one about Claude's skills, so the machine reached for that. Human judgment, and AI and people's work, had nowhere to go either.
The category side had run into this too. In round 2, Claude Opus noticed there was no category for renting machines or places to run systems, so the model squeezed those posts into "Models and cost" with low confidence. So partway through we changed the "Work tools" category into "Infra and tools" (infra meaning the machines and places used to run systems). There are still 10 categories.
But the gaps explain only a small share of the errors. Of the 10 wrong tags the most balanced method added, only about 1 came from a gap in the list. The rest came from a key that only tagged main topics, and from posts sitting on the border between 2 tags. Most missed tags were already on the list; the machine just didn't pick them.
After we got the person's key, we revised the list again into the one we use now. We added tags for people's AI skills, human judgment, AI and work, and filing documents with government agencies, and renamed vague tags so it's clear what they mean.
After the revision, we did not re-score round 3, because that would be grading our own repair. We revised the list after seeing this set of 20 posts' mistakes. A new number needs a fresh sample of posts, with the pass mark written down before a person tags them.
Part 8What should you weigh before you start?
The box above covers the decision model itself. Your own job raises 7 practical questions, worth answering before you build anything.
- How many categories and tags, who owns the list, and what does each description say? The description is the criterion the model reads. We have 10 categories, and a gap like the missing tag for people's AI skills (from Andrew Ng's map) only showed up after a person tagged posts by hand.
- What does a wrong label cost? That sets where you draw the confidence line and how many items a person reviews. A wrong blog tag is cheap. A wrong accounting category throws off the month-end close.
- How many items will fall below the line and go to a second check? That layer has to keep up. In our first round it was about half, and across the whole blog about 4 in 10.
- What will each item give the model to read? We sent only the Thai title, short description and headings, not the full post. Scanned documents need their text pulled out of the image first.
- Do you have a service that returns scores, or a developer to set one up? A regular chat AI doesn't return scores, so you can't use a confidence line. Ours is Clef-flash through Cloudflare Workers AI.
- Who makes the answer key, and when do you measure again? You need a key from a person before you can trust a number, and a new measurement every time the list changes. A key from AI made the 2-AI method look good, but against a person it had the lowest precision. Even our 2 AI reviewers agreed on only about 7 in 10 tags, so no method will match a key on every tag.
- What does each item cost in money and time? We haven't measured this. You'll have to try it on your own items.
Part 9How do you do auto-tagging with a decision model?
From all 3 rounds, we'd suggest 5 steps. The decision model is step 2; the other steps prepare what goes into it and deal with what comes back. Someone who doesn't write code can follow almost every step, except step 2, which needs a service that returns scores, or a developer to set it up.
- A person prepares the options and a key first. Keep the list of categories and tags in one file, each with a short description of what it covers. This list is the decision model's options and the descriptions are its criteria, so write them clearly. Give the list one owner. To add a tag, change the list first. Then sample about 20 items, assign categories and tags yourself before seeing the machine's answers, and write down the pass mark. Pick it from what a wrong item costs you: a wrong blog tag is cheap, a wrong accounting category is not. Ours were 9 in 10 right for categories the model decided alone, and recall 0.8 with precision 0.7 for tags.
- Have the decision model answer one item at a time. Send the title, short description and headings, with one pick-one question for the category and one yes-or-no question per tag. A regular chat AI answers in sentences and doesn't return a score per option, so you can't use a confidence line in the next step. This step needs a service that returns scores (ours is Clef-flash through Cloudflare Workers AI) or a developer to set it up. If you have neither, the honest option is to let AI suggest and have a person decide every item. Scanned receipts need their text pulled out of the image first. We did not measure cost.
- Set the line before you look at results. The owner of the list picks the confidence line before seeing any scores. Ours is 0.7, fixed before scoring. Before round 1, the only thing we tuned was the wording of the questions, on 5 posts kept out of the test. The category and tag list did change between rounds: one category after round 2, and the tags after round 3. Before you use a line, run it on the 20 items you labeled yourself in step 1 and check that the answers above the line are right. In our first round, about half the items fell below the line. Across the whole blog, about 4 in 10 did.
- Choose the route by what a mistake costs. If a mistake is cheap, like a blog category, use the model's answer when it reaches the line, have 2 AIs pick separately for the rest, send disagreements to a person, and have a person spot-check a sample of the items that went through with no person looking. If a mistake is expensive, like an accounting document filed in the wrong category that throws off the month-end close, have a person review every item the model routed, and spot-check the ones the model decided on its own too. Either way, we don't yet know how many spot-checks are enough.
- A person reviews tags on every item. Have the model suggest its top 3 tags, have another AI drop the ones that aren't main topics, then have a person look at every item before publishing, because no method has passed yet. At publish time, have the system check that no tag outside the list slips through. And every time the list changes, measure again on a fresh sample.
Our blog doesn't do all 5 steps yet. For categories, the model decides when it's sure and a layer of 2 AI reviewers checks it, with no person reviewing and no category key from a person yet. For tags, the site uses the top-3 + confirm method, even though it didn't pass the mark, and no person reviews it yet. We chose it because it was the most balanced in the table. For the 9 of 127 posts where Sonnet dropped all 3 tags, we use the model's highest-scoring tag instead, so no post is left without a tag.
What we have done: the tag list lives in one file the blog owner maintains, and every time we update the site, a script checks that no tag outside the list gets through.
Part 10Limits, and what to watch out for
If you use a decision model for auto-tagging this way, there are 3 things we'd ask you to watch. At least one answer key should come from a person, not from the same kind of AI being measured. The category and tag list you wrote yourself needs measuring too. And no tagging method has passed yet, so have a person review tags for now.
The test itself was small: 18 to 20 posts per round. The round 3 key had only 30 tags, so if the machine finds one more or one fewer tag, recall moves by about 0.03. The person's key came from one person, and someone else might tag differently. Everything was tested on one blog, mostly about AI and accounting work. The model saw only the title, description and headings. We used one decision model, Clef-flash, and didn't try Jev or any other model. The current list has no fresh key measuring it yet. A different kind of library may get different results.
If you're starting auto-tagging with a decision model on your own library, sample about 20 items and tag them yourself before choosing a model. Along the way you'll see what your list is missing, and how many tags you actually put on each item. Ours was 1.5, even though we meant to tag every topic each post covered.
FAQ
Q: Can a decision model do all of my auto-tagging on its own?
A: Not yet. For tags, in our test no method reached recall 0.8 and precision 0.7 against the key the blog owner made. The most balanced method scored 0.67 on both. For categories, if a mistake is cheap, let the model decide items that reach the line, have 2 AIs look at the rest, and have a person spot-check. If a mistake is expensive, as with accounting documents, have a person review every item the model routes and spot-check the rest. These category results have not been measured against a person.
Q: Why a decision model? Can't a regular chat AI do the tagging?
A: A chat AI answers in sentences and doesn't return a score per option, so you can't pull out the unsure items with a confidence line. A decision model can only pick from the list you send, and it tells you how sure it is. We did not measure tagging with a chat AI in this test, so we can't say how different the results would be.
Q: If 2 AIs tag together and we keep only the tags they agree on, is that more accurate?
A: In our test, no. Against the person's key, this method had the lowest precision, 0.47. Two AIs can agree even on topics a person doesn't count as main topics.
Q: Can a decision model file accounting documents into categories?
A: It's a similar kind of task. If it means picking one category from a list, routing the unsure items should work. But we only tested blog posts, so make an answer key from your own documents and measure first. And because the model reads text, scanned documents need their text pulled out of the image first. For choosing withholding tax rates, see the withholding tax post.
Q: Did revising the tag list improve the results?
A: We don't know yet. We chose not to re-score the same set of posts, because we revised the list after seeing that set's mistakes. It needs measuring on a fresh set of posts that a person tags first.
- Productize, "Can a Jev-style decision model cut accounting work? We tried it on Thai withholding tax" (part 1): https://productize.life/blog/decision-model-withholding-tax/en
- Cloudflare Workers AI, clef-flash (
@cf/cloudflare/clef-flash): https://developers.cloudflare.com/workers-ai/models/clef-flash/ - Laurence Moroney, "Decision models explained" (2 October 2026), on TypeSafe's Jev: https://laurencemoroney.com/2026/10/02/decision-models-explained.html
- Flavio Copes, "Clef" (Cloudflare Clef and Clef-flash): https://flaviocopes.com/clef/