productize.blog

← Back to home

Written from real work,not a pitch

Notes from real work and from working with AI agents, across product, SaaS, and growth, all the way to building my own products and a quant investing system from zero.

Read by series
SME · AI adoption

Before you buy AI, pick the problem and how you will measure it

If you are bringing AI into your team, the first question is not which one to buy, but which problem to solve and how you will know it is solved.~10 min read · 15 Sep 2026
SME · Service businesses

AI-native services: Y Combinator's playbook is your competitor's plan

The investor playbooks for this space were not written for existing service businesses, but the competitors coming for your work are reading them.~10 min read · 14 Sep 2026
AI Engineering · Part 4

Ng says respect the org's constraints, not where to draw the line

Item four of Andrew Ng's advice says to respect the organisation's priorities and constraints, and that sentence is correct, but it never says where the line actually is. What the map never gives is a way to tell, in the moment, whether you are respecting a constraint or waiting for permission. Part 4 closes the series.~16 min read · 12 Sep 2026
Accounting · AI

AI Has the Intelligence. Accountants Have the Context and Real-Time Accounting Data.

The FlowAccount team walks through three eras of accounting, as seen by the people who build the software: where the cloud-plus-AI era puts the accountant, why real-time accounting data is both something accountants hold and one of their duties, and the one question that shows how ready your firm is.~13 min read · 11 Sep 2026
AI agents · cross-owner trust

A guest file is a prompt your agent obeys

A2A and MCP give you a transport, an agent card, and a task lifecycle. They do not tell you who approves, whether provenance can be forged, or how a revocation actually removes anything. 8 rules from one real cross-house lane.agent to agent protocol · ~11 min read · 2026-09-09
Accounting · AI

AI replaces tasks, not professions. Here is the accounting map.

AI does not replace the accountant, it replaces tasks. Five groups of accounting work, which have moved and which have not, and why the line falls on liability, not capability. It started with a sponsored post: wire your accounting software to AI, drop your accountant.~23 min read · 7 Sep 2026
AI · Verification

After An Alien Mind, put the check on the artifact, not the report

OpenAI's chief scientist writes that reading what a model types out is losing its value as evidence. An Alien Mind read by someone who builds on models, and the one criterion it leaves you with, which you can set tonight.~11 min read · 7 Sep 2026
AI Adoption

The blocker to enterprise AI is not the model, it is that nobody can prove the return

We moderated the "AI Adoption, What Actually Works" panel at vLLM Bangkok Day 2026, wrote all 6 questions, and put 3 people from different ends of the same chain on stage. Here are their answers, rebuilt as criteria you can use. The gates before production, and the 4 sovereignty questions you can actually answer.~22 min read · 7 Sep 2026
AI Engineering · Part 3

Andrew Ng names five skills for using coding agents. The third assumes your tests actually ran.

Andrew Ng opens the third box of his AI engineering skills map into five named skills, from directing the workflow to understanding what a coding agent does inside. Part 3 walks his order one skill at a time, then turns to the question the map never asks: who reviews the reviewer.~14 min read · 5 Sep 2026
AI · Agent Harness

The guardrail that once saved us is now the brake

One day, 21 things to fix, and 8 of them came out of our own earlier fixes. An expired guardrail is not broken: it still fires exactly right, but the problem it was written for has moved. Four signals it is time to retire one, and how to do it at constant capability.~15 min read · 30 Aug 2026
AI Engineering · Part 2

You didn't skip the decision. You made it without looking.

Andrew Ng says software fundamentals still matter because you need them to steer a coding agent. The harder half of his sentence is about knowing which tradeoffs exist at all. A missing answer you can see. A missing question is invisible, and the code compiles either way.~10 min read · 30 Aug 2026
AI Engineering · Part 1

Andrew Ng's AI Engineering Skills Map

The four top-level skills and the six sub-skills expanded one by one, where the map came from, why prompt engineering is on none of the ten boxes, and what it means for engineers, PMs and solo builders.~10 min read · 22 Aug 2026
AI Agents

Good at Exams Is Not the Same as Finishing Real Work

Coding sits around 96 percent, long self driven computer use around 20.6. The 2026 bottleneck moved from building to maintaining. AI Update Bangkok 2026, Part 8~6 min read · 23 Aug 2026
AI Agents

One Agent Is Enough, Until the Evidence Says Otherwise

A framework is the SDK, not the agent. Start with one and add complexity when the evidence demands it. AI Update Bangkok 2026, Part 7~6 min read · 23 Aug 2026
AI Agents

Before an Agent Acts, Someone Has to Approve It First

How smart the model is does not decide how far an agent can reach into the real world. The harness that asks before acting does. AI Update Bangkok 2026, Part 6~6 min read · 23 Aug 2026
Cost & Models

What an AI gateway is, and why the pipe is worth more than what flows through it

A middle layer that owns no GPUs and takes a cut of everything routed. Where the money sits in the AI stack, and the caveat the biggest claim must travel with. AI Update Bangkok 2026, part 5.~9 min · 22 Aug 2026
Cost & Models

When fine-tuning an LLM is worth it: four stages, and the line down the middle

Training has four stages with a line through the middle. The first three create intelligence, the fourth only adjusts behaviour. AI Update Bangkok 2026, part 4.~8 min · 22 Aug 2026
Cost & Models

Quantization in LLMs: the 1% that disappears, and nobody says where

Shrinking a model to fit the machine you own. Vendors publish the damage themselves, about 1% at four bits. Nothing tells you which 1%. AI Update Bangkok 2026, part 3.~12 min · 22 Aug 2026
Reliability

How to read an LLM benchmark without being fooled

Every layer of AI has an announced number and a usable one, and they are never the same. Four questions before believing any figure, plus the map of the whole day. AI Update Bangkok 2026, part 2.~9 min · 22 Aug 2026
Cost & Models

What is Physical AI? When AI leaves the screen and touches real things

Physical AI is the third step after generative and agentic. What changes is not how smart the model gets, but whether you can undo the damage when it gets something wrong. Notes from AI Update Bangkok 2026, part 1.~10 min · 22 Aug 2026
Dev Setup

Terminal browser and markdown tools: what actually works

A browser in the terminal no longer turns the page into text. One runs real Chromium and streams pixels into the pane. How it compares with lynx and w3m, what to edit markdown with, and how to get out.~9 min · 20 Aug 2026
Claude Code · Tooling

The commit lands. Nothing reaches runtime.

A Claude Code marketplace is just a folder with an index file, and you can host your own. We run 9 plugins that way. The layout is small; what costs you is six traps where every command prints green and nothing arrives on the other machine.claude code plugin marketplace · ~8 min read · 2026-08-14
Eval · Reliability

Same task. Different directory. Opposite verdict.

The arm that was supposed to be empty was still holding the rule under test. A flag tells you what it strips, never what it leaves behind. How to prove a baseline is clean, and which old results survive.ai eval baseline · ~8 min read · 2026-08-13
Claude Skills

The skill banned a character. The skill used it 54 times.

The model invoked the skill, opened the file where the rule lives, and still broke that rule 3 runs out of 3. The cause was an example we had labelled good. Four lines flipped it.claude skill examples · ~8 min read · 2026-08-13
Claude Skills

60 skills were never called. We almost deleted them. Then we measured.

A library of 142 skills pays context rent every turn, and 60 showed no trace of ever being called. Measuring first flipped it: 66 of 86 trigger fine. Here are the four layers, and a better test than "has it been used?"claude skill library · ~10 min read · 2026-08-13
AI · Coding Tools

Stop asking which one is better. Ask which one signs off.

Codex and Claude Code really are good at different things. What changed our output quality was not the higher score, it was a rule enforced in code that the model which wrote the work cannot review it.codex vs claude code · ~12 min read · 2026-08-12
AI · Reading and verification

Six AI readers, one 307-page book, and how to trust the result

Reading one book breaks into five jobs and the first two use no model at all. What makes the result trustworthy is not a smarter model but the sequence around it, plus a checker proven able to go red before it is used to measure anything.ai read pdf · ~16 min read · 2026-08-08
AI · Orchestration

Swapping an AI's Brain Without Restarting the Session

The one switch that moves work from the cheap model to the expensive one mid-task cannot be built: the endpoint and the key are read at process start. Two moves work instead, and the day we finished writing them the quota ran out on its own and proved the fallback for free.llm fallback · ~12 min read · 2026-08-07
Databases · Reliability

FalkorDB Rewrote Its Engine in Rust. We Upgraded a Live Graph and Measured It

The Rust engine is real but still a preview, and the tag you would reach for quietly changed meaning. We upgraded a graph holding 3,011 live nodes, measured before and after, and wrote down the three things that make a database upgrade something other than a prayer.falkordb rust · ~11 min read · 2026-08-05
AI · Cost

14 Models, One Budget: LLM Routing in Practice

The full model map, the routing rules that decide which job goes to which model, and three lessons from a real teardown forced by a budget cut.llm routing · ~8 min read · 2026-08-03
AI Agents

system gate > per-agent prompt

Nine agents, three days, zero deliverables. The problem wasn't the prompt, it was the layer the rule lived on: why mandatory behavior belongs in a system gate, not per-agent prompts.AI agent guardrails · ~4 min read · 2026-08-03
Dev Setup

What Is Raycast: A Search Bar That Does More Than Open Apps

Spotlight opens apps and finds files just fine, but the moment something needs pressing over and over every day, it has nothing to offer.spotlight alternative mac · ~10 min read · 2026-08-03
Reliability

Lighthouse score is not a design review

The sweep passed on every axis and a real reader still said it looked like a toy. The problem was the set we chose to measure, not that the thing has no metric.~7 min read · 31 Jul 2026
AI · Framing the ask

The freshest work in your head will crowd out the next answer

Context engineering is deciding what belongs in the context a model sees, not just the sentence you type. Where it differs from prompt engineering.~6 min · Jul 31, 2026
Reliability · Testing

The Passing Test Was Telling Us to Delete Real Customers

One check said the signups table must hold zero rows. Six people had signed up. The quickest way to make it pass again was to delete all six.~5 min · Jul 30, 2026
Accounting · Tax

There Are Two Profit Numbers, and Tax Uses the One Not in Your Books

Same one million baht. Accounting says expense it, tax says add it all back. Both are right, because the two numbers exist for different jobs.~7 min · Jul 30, 2026
Accounting · Review

Preparing Statements ≠ Reviewing Them: The Skill AI Cannot Replace

A preparer asks did I record everything. A reviewer asks does this number make sense. The second question is the part you cannot hand to AI.~8 min · Jul 30, 2026
AI · Tool comparison

Two Ways to Let AI Drive a Browser, Tested for Real

ego-browser vs BrowserClaw, tested on the same task: two architectures of the AI agent browser, and which to pick.~8 min · Jul 29, 2026
Cost & Models

A more accurate model was waiting for $1.45. We chose not to use it.

What it is, the specs that matter, real API cost, the license catch nobody reads, and the one question that tells you when to reach for an open weight model.kimi k3 review · ~7 min read · 2026-07-29
AI · Repo Review

Pool three idle machines into one LLM endpoint you own

A review of mesh-llm, an open-source tool that pools GPUs across machines into one OpenAI-compatible LLM endpoint, from deploying it for real across three machines.~7 min read · Jul 25, 2026
AI · Repo Review

Andrew Ng shipped an AI coworker that works on your machine, and still lets you keep control

A review of openworker, Andrew Ng's local-first AI coworker, from reading the whole codebase. The idea worth borrowing is how it keeps AI controllable at the tool-call boundary.~8 min read · Jul 24, 2026
AI Infrastructure · Cost

The GPU budget doesn't leak under load. It leaks when jobs crash silently

Serverless GPU bills by the second and that part delivers. But a near-6-minute cold start forces an async design, and the budget leaks when jobs break and the CLI never shows it.~7 min read · Jul 24, 2026
AI · Reliability

The daemon that died for two weeks, and nobody noticed

A daemon runs in the background all the time. Its real danger isn't crashing, it's crashing with no signal. Told through one that died silently for two weeks, and how to make yours make a sound.~9 min read · Jul 24, 2026
AI · Design System

AI Built a Convincing Design System, and Got It All Wrong

Asked AI for a design system and it picked a tidy, confident palette that was wrong three times over. The rule that fixed it: reflect what exists, do not invent.~4 min · Jul 16, 2026
Cost & Models

The cheap model that gets to be wrong beats the expensive one that can't

Six hard bugs, measured: a cheaper model allowed to loop 3 times matched the expensive one-shot at 1x cost against 3.76x, because a checker to fire against lets it be right cheaply.~6 min read · Jul 15, 2026
AI Production · Fable 5

Why Claude Fable 5 switches models mid-conversation, and what it costs

Fable 5 hands risky requests to Opus 4.8 automatically. When a normal task gets switched anyway, that is a false alarm on the system side.~4 min · 11 Jul 2026
AI Agents

AI writes the code, but who signs for it

As AI writes more of the code, the missing thing is accountability. The day it became a real file across 5 repos.~3 min · 10 Jul 2026
SME · Inventory

Your cash didn't disappear. It's asleep in the stockroom.

Inventory analysis with a handful of numbers: days of inventory, ABC-XYZ, and a cut score, straight from the spreadsheet you already have.~10 min read · Jul 10, 2026
Accounting · Data Accuracy

The model misread a bill by five hundred baht. What caught it?

One bill, two confident answers. How an arithmetic verification layer catches misreads on its own.~6 min read · Jul 10, 2026
AI Engineering · Reliability

Stop begging the model for JSON. Constrain it with a schema

Three failed runs of a model that promised JSON and rambled instead, until we moved enforcement to the decoding layer.~6 min read · Jul 10, 2026
AI Infrastructure · Cost

Modal GPU on a budget: 7 submerged stumps before the first clean run

Pay-per-second GPU, every stump we hit, and real per-bill costs from our logs.~8 min read · Jul 10, 2026
Accounting · Tax

Certificate in lieu of receipt: when and how to use it for Thai tax deductions

Taxi fares and market purchases come with no receipt. The Revenue Department already has a path for them.~8 min read · Jul 10, 2026
Accounting

Choose an AI Bill Tool by the Job, Not the Feature List

Parnuan, Paypers, FlowAccount, ask ChatGPT/Claude, or DIY each takes your bill somewhere different. How to choose from real accounting work.~7 min · Jul 9, 2026
Product & PM

What is spec-driven development, and why non-coders can do it

It's all over AI coding and sounds like a programmer thing, but the heart of it is thinking clearly about what to build, which non-coders can do.~6 min read · Jul 9, 2026
Product & PM

How to write a PRD your AI can build from (template + example)

You've been told to write a spec first, but the blank page is the hard part. An 8-section template and a real worked example.~8 min read · Jul 9, 2026
Product & PM

A vibe is not a spec: write a PRD your AI can build from

Vibe coding is fast, but without a spec it builds the wrong thing. Here's the minimal PRD that makes AI build right.~7 min read · Jul 9, 2026
Reliability

Build your own social listening tool in a weekend

Wire scrapers to an AI and listen to social yourself in one weekend, plus the five things you must get right.~9 min read · Jul 8, 2026
Dev Setup

Setting up Ghostty + Zsh without the one-command installer

A curl-into-bash script can rewrite your whole shell. How to build it piece by piece, safely.~6 min · 8 Jul 2026
Dev Setup

Ghostty vs iTerm2 vs Terminal.app: the best Mac terminal

All three compared from real use: speed, config, and which fits working with AI agents.~6 min · 8 Jul 2026
Reliability

The safety filter that was silently deleting your data

Gemini's safety filter flagged ordinary content as sensitive and dropped records with no error. The fix: make the extraction step a swappable engine.~9 min · 8 Jul 2026
Knowledge & Memory

AI knowledge management: build a system you can trust, layer by layer

How to assemble an AI knowledge management system layer by layer, each one earning trust before it is believed.~8 min · 8 Jul 2026
Production Guardrails

allowedTools isn't file access: why your agent still gets permission denied

allowedTools allowed Read, yet the agent can't read outside cwd. Glob finds it, Read is denied, because permission and --add-dir are separate layers.~6 min read · Jul 8, 2026
SEO & Discovery

Google search favicon: why it shows a globe instead of your icon

Google shows a globe instead of your favicon when Googlebot can't fetch it. Diagnose with curl, fix at the root.~6 min read · Jul 8, 2026
Funnel & Content

Gated content without a backend: email-unlock on a Cloudflare Worker

Gate content on a fully static site with no backend or HubSpot: crawlable teaser, worker-served payload after unlock, plus the origin trap that nearly left the gate open.~8 min read · Jul 7, 2026
Skills & Plugins

Your AI agents should not all carry the same skills

Six agents, one afternoon, three misses: names that lied, a layer that failed silently, and versions shifting under our feet. Ends with a lint that compares the list against disk.~8 min read · Jul 6, 2026
Claude Skills

“Make it pretty” is the wrong prompt

Every AI-designed site looks the same. A design skill that refuses to code until three strategy questions are answered, told through a real admin-page before/after~8 min read · Jul 6, 2026
Reliability

We Told the Agent to “Verify It Works”. It Picked the Easiest Possible Meaning

A real migration night: every line of the report was green and the system was down. How one runnable acceptance line closes the gap, plus the three enforcement layers that end with proving the checker itself.~5 min read · Jul 6, 2026
AI Agents

We Interviewed the Most Expensive AI Model in the House About Making Cheaper Ones Do Its Job

A real interview with Anthropic’s new top tier: the principles it was trained on, the weaknesses it volunteered, what it is not better at, and six levers that let Opus stand in when costs must come down.~12 min read · Jul 6, 2026
Cost & Models

Claude Opus vs Sonnet vs Fable 5: Which Model for Which Work

The night our priciest model burned its credit mid-task rewrote our fleet rulebook: the criterion is ambiguity, not difficulty, plus six levers that squeeze a mid tier above its weight.~7 min read · Jul 5, 2026
Cost & Model

23 free models on OpenRouter: the quota is real, the queue is not yours

Tested with a real key: the free llama returned 429 nine times in a row with a full quota. A verdict table of which jobs fit which kind of free.~8 min read · Jul 4, 2026
Cost & Model

Your static site can have AI. No backend. No API key.

Add an AI box with Cloudflare Workers AI, with real measured numbers: 124 neurons per call, ~80 free calls a day, and a two-layer cap that keeps the bill at zero by proof.~9 min read · Jul 4, 2026
Cost & Model

Your AI agent doesn't need a bigger machine, it needs a home that never turns off

We moved a full agent fleet to a $14/month VPS and measured for three weeks: 3GB of RAM, not 8GB, and CPU runs out first. With the real invoice and a live three-provider price table.~10 min read · Jul 4, 2026
Cost & Model

A 27B model on a rented GPU with vLLM: the traps nobody writes down

The hands-on half of the GPU-rental story: the busy empty machine, config in two places, what survives a stop, and the tunnel that must reconnect itself.~9 min read · Jul 3, 2026
Cost & Model

Renting a GPU for your own LLM: the expensive part is the idle hours

Per-token APIs vs an hourly GPU vs serverless GPU: live prices from three providers, a real break-even formula, and the data line the price tag never shows.~9 min read · Jul 3, 2026
AI Agents

Claude Fable 5 as the head, everything else as hands

The most expensive model as team lead, four cheaper subagents as hands, one real migration assessment. What worked and what broke.~15 min read · Jul 3, 2026
Claude Skills

Claude Code hooks: a done-voice

Send an agent on a long run and walk away. A hook speaks what finished, and which tab.~8 min read · Jul 3, 2026
Product & PM

AI product management: what's left for PMs when the frameworks become skills

The pillar tying four repo reviews together. When the field's frameworks are a command away, knowledge is cheap and judgment is the scarce part.~8 min read · Jul 2, 2026
Product & PM

The whole GTM playbook in 12 skills, wired to cascade

A review of Maja Voje's GTM Strategist skills. The win is the design, set product context once and each phase feeds the next, not a longer framework list.~7 min read · Jul 2, 2026
Product & PM

The YC president open-sourced his own stack. It's about taste

A review of Garry Tan's gstack. Multiple review lenses encode taste into software, and AI models recommend while users decide.~7 min read · Jul 2, 2026
Product & PM

Product discovery as a pipeline, with two judgment calls baked in

A review of Else van der Berg's product-discovery-skills. Teresa Torres's method, with misfits as signal and importance over prevalence.~6 min read · Jul 2, 2026
Product & PM

The whole PM craft as ~68 skills

A review of Pawel Huryn's pm-skills. Roughly 68 skills across 9 plugins, and the intended-vs-implemented doc-vs-code audit that stands out.~6 min read · Jul 2, 2026
Claude Skills

9arm's skills repo is small, but it has one cost idea worth stealing

Don't burn your expensive model on grunt work. Hand it to a cheap one, keep the expensive model for judgment. The sharp cost idea inside 9arm's small repo.~5 min · Jul 2, 2026
Claude Skills

Addy Osmani agent-skills review: 24 production skills, 72.6k stars

Addy Osmani, a Google web engineer, opened a 24-skill repo covering the full lifecycle. The standout the other three lack: browser testing and performance. The closer of the review series.~6 min · Jul 2, 2026
Claude Skills

mattpocock/skills review: a real engineer's .agents, 216k stars

Matt Pocock opened up his own .agents folder: the v1.2.3 plugin ships 25 skills you fork and remix. The third in the series reviewing trending skill repos, after superpowers and karpathy.~7 min · Jul 2, 2026
Claude Skills

Karpathy Skills review: one CLAUDE.md, 189k stars

A 189k-star repo that distills how AI breaks when it codes into 4 principles in a single CLAUDE.md. More minimal than superpowers, and how to pick between them.~7 min · Jul 2, 2026
Claude Skills

Superpowers review: the AI agent repo blowing up, 243k stars

The AI coding-agent repo everyone's installing. I tried it and read the source: is it worth it, what to take, and the core idea that matched a skill I had already built.~9 min · Jul 2, 2026
AI Agents

What is an AI agent, vs a chatbot?

If you open ChatGPT, ask, and copy the answer out, that is a chatbot, not an agent. Here is the line, and the on-ramp to the 7-layer series.~6 min · Jul 2, 2026
AI Agents

AI Agent Architecture: The 7 Layers a Production Agent Needs

Search "ai agent architecture" and you get textbook diagrams. Here are the 7 real layers that separate a production agent from a demo.~9 min read · 1 Jul 2026
Reliability

The Ported Code Passed Every Test. It Was Still Wrong.

Self-written tests only prove the code matches what you think. Characterization testing uses the original that already runs as the answer key, diffed line by line.~9 min read · 1 Jul 2026
Accounting

From a Ledger Entry to Its Source Document, Checkable

Vouching proves every entry has a real document behind it. Let AI match entries to documents, then surface the ones with none for a person to decide.~7 min · Jul 1, 2026
Accounting

From Statement and Books to a Reconciliation You Can Check

A good reconciliation's output isn't the word "matched". It's every unmatched row with a reason. Let AI match, then classify what didn't.~7 min · Jul 1, 2026
Accounting

From a Photo of an Invoice to Import-Ready Data

The problem isn't that AI can't read your invoices. It's that you get text you still can't import. Here's how to get import-ready data, with the numbers matching the source.~7 min · Jul 1, 2026
Accounting

Before AI helps you file, read the map first

A full walkthrough of DBD e-Filing, then a task analysis of where AI helps and where a person must decide.~5 min · Jul 1, 2026
Claude Skills

A Skill Isn't Just a Longer Prompt

What a Claude Code skill is, how it differs from a prompt, and when to stop re-typing and package one the model reaches for on its own.~5 min · Jun 30, 2026
Claude Skills

Don't Let the LLM Do the Math

A reliable skill isn't the LLM doing everything. Work that must be exact every run is test-locked code; only what needs judgment goes to the model.~8 min · Jun 30, 2026
Reliability

The tool you trust to catch mistakes was the one making them

Sometimes the thing you use to catch mistakes fails silently, then reports "no bugs" while it checked nothing. The way to catch it: make it fail once, on purpose, with mutation testing and a positive control.~7 min · Jun 30, 2026
AI Agents

I ran three AI coders at once. The agents weren't the cost.

Three coders in parallel and the bill barely moved. The expensive part is human time and the defects that slip through. The shape we use: coders on a subscription, a separate reviewer, and a human holding merge.~5 min · Jun 30, 2026
Knowledge & Memory

Your agent has all its rules. That doesn't mean it reaches them in time

When your rules file grows too big to load, what do you cut? The deeper question is why a rule you already stored sometimes doesn't fire when it should.~6 min · Jun 26, 2026
Reliability

When several AIs agree, that isn't proof

Pulling in several AIs to review your work feels more thorough. But the moment they all agree might be the most dangerous one.~4 min · Jun 26, 2026
AI Agents

Hand work to your AI fleet through a board, not a chat

We had AI agents that finish work on their own, but chatting with each one made us the bottleneck. Put the work on your Linear board, tag it for the fleet, and the result comes back on the card. Irreversible work waits for a human.~5 min · Jun 26, 2026
AI Agents

When parallel AI agents corrupt each other's work

Let parallel AI agents write in one shared directory and they trample each other. The fix is structural: one git worktree per task, created by the orchestrator.~5 min · Jun 26, 2026
AI Agents

Don't marry one model: separate the soul, engine, and orchestrator

The day you want a cheaper model, you shouldn't have to rebuild the agent. Separate three layers up front and swapping models becomes a config change.~6 min · Jun 26, 2026
Knowledge & Memory

AGENTS.md and SOUL.md: sharing knowledge across an AI fleet

We run several AI models as a team, but each one's knowledge stays locked in its own format. The fix isn't a big database. It's two open files every model can read.~5 min · Jun 26, 2026
AI Agents

Connect two AI agents across machines without opening a port

The home box is behind NAT and the cloud box has every port closed for safety, yet two AI agents still talk across machines, through an SSH tunnel that already ships on every computer.~7 min · Jun 26, 2026
Privacy & Security

The data that should never flow up

A three-layer defense for data and AI. Cloud knowledge flows down to help; your secrets never flow back up. Plus the three sneakiest leak points.~5 min · Jun 25, 2026
Cost & Models

The best model is often the wrong one

Choose a local LLM by the job, not the benchmark. Qwen vs Gemma from real work, plus QAT, MTP, dense vs MoE.~5 min · Jun 24, 2026
Reliability

Let Codex review the code AI wrote

The writer and the reviewer shouldn't be the same model. Two times a second model caught a bug we'd missed in our own code.~7 min · Jun 23, 2026
Knowledge & Memory

The real memory stack we run: Graphiti + FalkorDB

Behind "give the AI memory" are five wired layers and two bugs that fail silently, plus how to check your agent really remembers.~6 min · Jun 23, 2026
Knowledge & Memory

Design AI memory the way human memory works

Memory that remembers is not about storing, it is about pulling the right thing back. Design it in layers like human memory, with the real tools we run.~9 min · Jun 23, 2026
AI Agents

AI coding agent fleet: the Kanban swarm pattern

One AI coder is easy. The pattern that lets a lead split a goal into cards, run workers in parallel in separate worktrees, and gate the merge with a second-engine review.~8 min · Jun 21, 2026
AI Agents

I Let AI Post to Social for Me, and the 3 Guardrails That Keep It Safe

Auto-posting to Facebook and LinkedIn every day: the step-by-step setup, and the guardrails that keep an irreversible public action from going wrong.~7 min · Jun 21, 2026
Reliability

Why Your AI Agent Lies to You

AI doesn't lie at random. It guesses what's plausible and says it with confidence, and it fools you best when it reports "done."~6 min · Jun 20, 2026
AI Visibility

Make Your Site Agent-Ready: I Did 8, Refused 3

I ran my site through isitagentready. Do the real checks, refuse the ones that advertise capabilities you do not have. The missing checkmarks are deliberate.~7 min · Jun 19, 2026
Privacy & Security

Let Claude cite you, block the trainers

Blocking every AI bot feels safe, but it also shuts the door that lets people find you through AI. How to allow them one at a time on Cloudflare.~7 min read · Jun 18, 2026
Cost & Models

50 MCP servers, six calls

A tool you connect but never call taxes the context window every turn. How to audit and prune.~5 min read · Jun 18, 2026
Claude Skills

Evaluating Claude Code plugins

ponytail and headroom trended together. The test for whether to install.~5 min read · Jun 17, 2026
Cost & Models

You Can Still Run Agents on a Subscription (For Now)

Anthropic announced a billing split then paused it. Right now agents still run on the plan, but the door may close.~8 min read · Jun 16, 2026
AI Visibility

Your Site Is Live. But Who Can See It?

Live does not mean found. 5 steps a one-person business can do in an afternoon so both Google and AI see your site.~6 min read · Jun 16, 2026
AI Agents

Not Every Action Needs a Human

Human in the loop is treated like a light switch. Here's the third option: tier the loop, and decide who decides per action.~8 min · 15 Jun 2026
Knowledge & Memory

Systematize turning speech into a second brain you can trust

Doing it by hand works for one clip, but dozens need a system. Here is the four-stage pipeline that makes the evidence gate run on its own.~8 min · 14 Jun 2026
Knowledge & Memory

Transcribe on your own machine, faster than Whisper

A batch of Thai lecture audio would not finish transcribing all day, until I learned the slow part was the model architecture, not the hardware. Here is the transducer fix.~9 min read · 14 Jun 2026
AI Agents

Bring your AI into Discord, without handing over the keys

Connect AI to Discord so you can chat from anywhere, with a gate that controls who can command the bot, and keeps a chat message from becoming a command.~11 min · 14 Jun 2026
Knowledge & Memory

Turn speech into trustworthy notes, without letting AI make things up

Summarize meetings or lectures into knowledge notes while keeping AI from adding things no one said, then store them so they stay findable.~10 min · 14 Jun 2026