Putting AI to work without the hype tax
Most AI advice tells you how to produce more. Almost none of it tells you whether what you produced was any good. That gap is where the money is lost, and it is the thing worth fixing first.
The methodReading the claims
The AI conversation runs hot in both directions. The useful filter is not optimism or skepticism, it is provenance. Every claim in this document is marked with where it came from, and it is worth applying the same filter to anything else you read this year.
That distinction matters more than usual here, because the loudest voices in AI education are also selling AI education. Their numbers are not necessarily false. They are just unaudited, and they were selected because they were impressive.
Keep theseWhat holds up
Three ideas in circulation right now are genuinely good, cost nothing, and work immediately.
VerifiedMake the tool interview you. The most common mistake is asking for something the tool has no context to deliver. "Help me grow my business" is a question you would not ask a stranger. Adding "ask me fifteen questions before you answer" converts a vague request into a specific one, and it surfaces assumptions you had not articulated yourself.
VerifiedRefine the prompt before you use it. Three steps, two minutes, detailed below. This is the single highest-value technique in circulation.
VerifiedStart from what you resent. List the parts of your week you would happily hand to someone else. That list is your automation candidate list, and it starts from a real constraint rather than from a capability you are trying to find a use for.
Discount theseWhat to treat as marketing
| The claim | How to hold it |
|---|---|
| Sevenfold efficiency gains | Self-reported, unaudited, from a source selling AI certification. Plan for meaningful time savings on specific repetitive tasks. Do not restructure a budget around a multiplier. |
| An "AI team" with names, headshots, and an org chart | The naming genuinely does help people adopt the tool, and that part is real. Underneath the theatre it is a saved instruction file. You need one good saved prompt, not a staffing diagram. |
| Hundreds of ad campaigns generated in minutes | Generating variations was never the bottleneck. Knowing which variation works is the bottleneck, and that still requires running them and measuring. |
| Ninety-three days without writing an email | Impressive as a demo, genuinely risky as a practice. See below. |
Do this nowThe technique worth using today
This takes two minutes and works in any of the major tools. It is the fastest improvement available to you.
Type it messily
Write the request exactly as you would say it out loud, including the parts that feel disorganised. Do not tidy it. Speaking it with dictation works even better, because talking produces fuller requests than typing does.
Ask it to clean the request up
Go to the top of what you wrote and add: Clean up this prompt for clarity and impact: then send. What comes back is usually what you meant but did not know how to phrase.
Ask for the stronger version
Reply with: Turn this into a super prompt. You will get back a version that opens by assigning a role and stating a goal. Save the good ones in a plain document you control. That file becomes genuinely valuable over time.
ImportantYour data, and the tier that does not protect you
The common advice is "do not use the free version, pay for the business version." The instinct is right. The line is drawn in the wrong place, and the difference is the kind of thing that matters if you handle client information.
RiskPaying does not automatically opt you out. The real division is consumer terms versus commercial terms, not free versus paid. On Anthropic's plans, Free, Pro, and Max are all consumer products. Their published update states these plans give users a choice about whether their chats are used to improve models, with retention extending to five years when that setting is on. The same update states it does not apply to services under Commercial Terms, which include Claude for Work, Claude for Government, Claude for Education, and API use.
anthropic.com/news/updates-to-our-consumer-termsSo upgrading a personal plan from free to paid does not, by itself, change your data posture. Two things actually do:
- Open the privacy settings and look. Whatever tool you use, find the training or data-improvement setting and set it deliberately rather than inheriting a default.
- If client confidentiality is involved, move to a business or team plan. That is where the contractual protection lives, and it is usually a modest step up in cost.
Vendor terms change, and they change often. Treat this document as a prompt to check your own settings rather than as a permanent answer.
BoundariesWhat not to hand over
There is real enthusiasm right now for giving AI tools direct access to your inbox so they can read, reply, and send on your behalf. This is the one recommendation worth declining.
RiskAn assistant that can send email is an open door. Language models cannot reliably distinguish instructions written by you from instructions embedded in content they are reading. That means a message arriving in your inbox can contain text the assistant treats as a command. The OWASP Foundation, which publishes the widely used security risk lists for software, ranks this first among risks for AI applications. Their own worked example describes a mailbox assistant with sending capability being tricked by a crafted incoming email into forwarding sensitive information to an attacker.
owasp.org · OWASP Top 10 for LLM Applications 2025 (PDF)The practical version of this is simple and costs you almost nothing:
- Draft with AI. Send with your own hand. Reviewing a draft takes seconds. Retracting a sent message takes a relationship.
- Read-only access is fine. Summarising your inbox is a different risk category from acting on it.
- Anything irreversible gets a human. Sending, publishing, paying, deleting, committing.
The missing pieceHow you will know it is working
This is the part almost no AI advice covers, and it is what separates a tool that saves you money from one that quietly costs you.
Software engineers building AI products use a practice called evaluation. The idea translates cleanly to a small business and takes about an hour to set up.
The principle. You cannot judge an AI tool by how impressive its output feels. You judge it by comparing its output to work you already know is good. Collect examples of your own best work first, then measure against them.
Your version, in four parts
- A reference set. Ten to twenty real examples from your actual work that you were genuinely happy with. Proposals, client updates, whatever you produce most.
- A standard. Written down before you automate anything. What does a good version of this contain? What would make you reject it?
- A check. For the first ten outputs, compare side by side against your standard. Count how many pass.
- A record. Keep the failures. They tell you what to fix, and they are the most useful thing you will collect.
If eight of ten pass, keep the workflow. If four of ten pass, the task is not ready and more prompting will not fix it. That single number is worth more than any productivity claim you will hear this year.
Start hereYour first workflow
One concrete thing, about an hour, difficult to get wrong.
Pick the document you write most often
A proposal, a discovery summary, a client update. Frequency matters more than importance for a first attempt.
Find three past versions you were proud of
These are your reference set. If you cannot find three good ones, that is useful information on its own.
Ask what makes them work
Paste all three in and ask the tool to describe the structure, the tone, and what the strongest parts have in common.
Turn that into a reusable instruction
Ask it to write the description up as an instruction it could follow next time. Save that text in a plain document on your own computer, not inside the chat tool's memory. You want to be able to read it, edit it, and take it with you.
Use it
Next time you write one, paste the saved instruction plus the new facts.
Check the first ten against what you would have written
This is the step everyone skips, and it is the only one that tells you whether any of it worked.
GroundworkBefore you start
Five questions worth answering on paper. They take ten minutes and they prevent most of the expensive mistakes.
- What do you do every week that you would happily pay someone else to do?
- What does a good version of that output look like, specifically? Can you point to one?
- Does any of this touch client information you have promised to keep confidential? This one changes what you are permitted to do, and it is the one most people skip.
- Are you on a personal plan or a business one, and have you actually opened the privacy settings?
- What would you do with the time if it worked?
CalibrationWhat to expect
RealisticMeaningful time back on repetitive writing, first drafts, summarising, reformatting, and research legwork. This is real and it arrives quickly.
RealisticA better starting point on things you find hard to begin. The blank page problem largely goes away.
SlowerAnything requiring judgment about your specific clients, relationships, or market. The tool has no context you have not given it, and giving it that context is itself work.
UnlikelyReplacing a role wholesale, or a step change in revenue that traces to the tool rather than to what you did with the time.
The honest summary: this is a meaningful productivity improvement and a genuine competitive disadvantage to ignore. It is not a business model. The people describing it as one are usually selling something.
Appendix Technical build notes
Implementation detail for a self-hosted multi-model setup. Not required reading for anything above.
Durable Objects, and why they matter for agents
Scott Moss builds on Cloudflare in his AI Engineering course, and Durable Objects are the primitive underneath it. Worth understanding because they solve the problem every stateful agent runs into.
The problem
Serverless functions are stateless by design. Cloudflare's docs put the contrast plainly: Workers are stateless functions where each request may run on a different instance, in a different location, with no shared memory between requests. That is fine for an API endpoint and useless for an agent, which needs to remember a conversation, track in-flight tool calls, and hold a session open.
The usual workaround is to bolt on a database, a cache, and a session store, then coordinate between them. That is three moving parts and a source of race conditions.
What a Durable Object is
Per the same documentation, Durable Objects are stateful compute where each instance has a unique identity, runs in a single location, and maintains state across requests. Each one has its own SQLite database attached, and SQLite runs as a library in the same thread, so query latency is effectively zero.
The guarantee that makes it useful: exactly one instance exists per ID, globally. Ask for the object named conversation-4821 from anywhere in the world and you reach the same single-threaded instance with the same private storage. No locks, no distributed coordination, no double-writes.
Why the Agents SDK is built on them
Cloudflare's Agents SDK makes each agent a Durable Object. That gets you, for free: conversation state that survives crashes and deploys, single-writer correctness so concurrent messages cannot corrupt state, persistent WebSocket connections for streaming, per-agent scheduling via the alarms API so an agent can wake itself up, and hibernation so an idle agent costs nothing.
The design rule from the docs is to model each object around your "atom" of coordination. One per conversation, per user, or per document. The named anti-pattern is a single global object handling everything, which becomes a throughput bottleneck at roughly 500 to 1,000 requests per second.
Worth knowing before committing: the agent-as-Durable-Object model has no drop-in equivalent on AWS or GCP, and the SDK is pre-1.0 and moving fast. Pin versions.
Billing correction
Consumer subscriptions are not API access, and any plan that assumes otherwise will not work.
- Claude Max. Anthropic's help centre states that paid Claude plans and the Console are separate products, and that a subscription does not include API or Console access. A separate metered Console account is required. Reporting also indicates OAuth authentication was restricted to Claude Code and Claude.ai in February 2026, so proxying subscription auth into a gateway is both fragile and against terms.
- Perplexity Pro. Sonar is billed pay-as-you-go against prepaid credits, with a per-token rate plus a per-request search fee. Sources conflict on whether the $5 monthly Pro credit still exists. Verify in the account settings before relying on it.
Cleanest path: keep the subscriptions for interactive use, and open one OpenRouter account with a hard spend cap for everything routed through the gateway.
Hardware roles
Local inference speed is a memory bandwidth problem, not a RAM capacity problem. Capacity decides whether a model fits, bandwidth decides how fast it generates.
| Machine | Role |
|---|---|
| M1 Max, 64GB | Highest bandwidth in the fleet by a wide margin. Heavy local inference and eval runs when awake. |
| M4 Mac Mini, 16GB | Always on. The responsive local endpoint, 8B class via MLX. |
| UGREEN iDX6011 Pro, 64GB | Services, weight storage, embeddings, batch work. TechPowerUp's review found the NPU unused during LLM work with inference falling back to CPU at 20 to 30 percent utilisation, so not the interactive host. The OCuLink port and PCIe Gen4 x8 slot are the real upside: an eGPU there collapses the whole architecture into one box. |
| Synology DS423+ | Backup target. Leave it out of this. |
Bandwidth figures stated from general specification knowledge, not verified against a primary source this session.
Stack
- Inference: MLX or Ollama on the Mac Mini bound to the LAN. Ollama in Docker on the NAS with weights on the NVMe, never the array.
- Gateway: LiteLLM in Docker on the NAS. The only component holding secrets. Per-key spend caps and request logging on from day one. Fallback chains so a sleeping MacBook degrades rather than errors.
- Interface: Open WebUI with exactly one connection, pointed at LiteLLM. SearxNG alongside it for local search.
- Access: Tailscale, not a public tunnel. The gateway holds every key and a spend budget, so it should not be internet-reachable. Cloudflare Tunnel with Access policies is the right tool later for sharing something specific.
- Devin: a dispatched tool behind a Function call, not a model in the dropdown. It has its own session lifecycle and is not a chat-completion endpoint.
Note that /dispatch-style skills are a Claude Code feature and do not carry over to Open WebUI, which uses its own Pipelines and Functions system.
Evals, technical version
Routing across eight endpoints without evals is routing on vibes. Build a golden dataset of 30 to 50 real prompts from actual work, write scorers that grade numerically, run offline against the set before promoting any model in the routing config, and sample production traffic from the LiteLLM logs for online scoring. Braintrust is what the course uses for observability. A TypeScript script looping the golden set against three endpoints and printing a score table is a legitimate first version and takes an evening.
Two adjacent ideas worth taking early: context engineering, meaning curating exactly the tokens needed at inference time rather than stuffing everything in, and structured outputs with a schema so downstream code has something predictable to parse.
Build order
- One evening. Ollama on the Mac Mini, LAN-bound, one model, hit it with curl. Confirm the endpoint before adding orchestration that can hide failures.
- One weekend. LiteLLM plus Open WebUI on the NAS. Two endpoints: Mac Mini and one capped OpenRouter key. Tailscale. PWA on the phone. This alone is the working system.
- Ongoing. Add endpoints one at a time, each with an eval before it becomes a default route. Then knowledge and vector store.
- Optional. Devin behind a Function. The eGPU question.
Do not build
- An agent with send authority on a real inbox.
- Bypass permissions on a machine that matters. Anthropic's docs state it offers no protection against prompt injection and is for isolated containers and VMs only. Auto mode gives fewer prompts with a classifier still checking.
- Local models for anything where correctness matters, until an eval clears them.
- All four layers at once. Phase 2 is useful on its own. Live with it for two weeks first.