Prepared July 2026

Putting AI to work the right way

Most AI advice is about producing more work, faster. Very little of it is about checking whether that work is any good. This document covers both, and it covers the business questions that usually get skipped.

What this covers
  1. Which popular advice holds up and which does not
  2. One technique you can use today
  3. How to tell whether it is actually working
  4. Your data, your client agreements, and who owns the output
  5. A first project to try

How to read AI claims

The AI conversation runs hot in both directions. A simple filter helps: ask where a claim came from before you decide what to do about it.

Every claim in this document is marked. It is worth applying the same marking to anything else you read this year.

CheckedConfirmed against original documentation. The source is listed at the end of the section.
Company claimReported by someone who sells a related product. It may well be true. Do not build a budget on it.
CautionCommon advice that sounds sensible and can cost you money.

This matters more than usual with AI, because many of the loudest voices teaching it also sell AI training. Their numbers are not necessarily false. They are simply unaudited, and they were chosen because they sound impressive.

What is worth doing

Three ideas in wide circulation are genuinely good. They cost nothing and they work right away.

CheckedMake the tool ask you questions first. The most common mistake is asking for something the tool has no way to deliver. "Help me grow my business" is a question you would not ask a stranger on the street. Add "ask me fifteen questions before you answer" and the request becomes specific. It also tends to surface things about your own business you had not put into words.

CheckedImprove the request before you use it. Three steps, about two minutes. This is set out below in full. It is the single most useful technique available right now.

CheckedStart with the work you resent. Write down the parts of your week you would happily hand to someone else. That list is where to begin. It starts from a real problem, rather than from a feature you are trying to find a use for.

What to discount

Four claims come up constantly. Here is how much weight to give each one.

The claimHow to treat it
Sevenfold efficiency gainsSelf-reported by someone selling AI certification. Nobody has audited it. Expect real time savings on repetitive tasks. Do not restructure your finances around a multiplier.
Build an "AI team" with names, photos and job descriptionsThe naming genuinely does help people use the tool, and that part is worth copying. Underneath it, though, it is a saved instruction file. You need one good saved prompt, not a staffing chart.
Hundreds of ad campaigns generated in minutesProducing variations was never the hard part. Knowing which one works is the hard part, and that still means running them and measuring the results.
Ninety-three days without writing an emailImpressive as a demonstration. Genuinely risky as a practice. Explained below.

The prompt technique

This takes about two minutes and works in any of the major tools. It is the fastest improvement available to you.

1

Type it messily

Write the request the way you would say it out loud, including the disorganised parts. Do not tidy it up. Dictating it works even better, because people speak in fuller sentences than they type.

2

Ask the tool to clean it up

Go to the top of what you wrote and add: Clean up this prompt for clarity and impact: then send it. What comes back is usually what you meant but could not phrase.

3

Ask for the stronger version

Reply with: Turn this into a super prompt. The result will open by assigning a role and stating a goal. Save the good ones in a plain document on your own computer. That file becomes genuinely valuable over a few months.

Checking the work

This is the part most AI advice leaves out, and it is what separates a tool that saves you money from one that quietly costs you money.

Software teams building AI products use a practice called evaluation. It sounds technical. It is not. The idea is simply that you cannot judge AI output by how impressive it feels, so you compare it against work you already know is good.

The principle. Collect examples of your own best work first. Then measure the AI output against those examples. Impressive-sounding output is not the same as correct output, and the two are easy to confuse.

Four parts, about an hour to set up

If eight out of ten pass, keep the workflow. If four out of ten pass, the task is not ready and more prompting will not fix it. That single number is worth more than any productivity claim you will hear this year.

The accuracy problem

AI tools produce confident, well-formatted, professional-looking output that is sometimes simply wrong. This is not an occasional bug. It is a known property of how these tools work, and the professional consequences are now well documented.

CautionNever pass on a fact, figure, quote or citation you have not checked yourself. The law firm Norton Rose Fulbright reported in June 2026 that a public database tracking these incidents had documented over 1,148 US court cases involving AI-fabricated citations by lawyers. Penalties have ranged from fines to suspension from practice. The Fifth Circuit observed that the problem shows no sign of slowing down.

Two details from that record matter to you even though you are not a lawyer.

First, the errors are getting harder to spot. Early cases involved obviously invented sources. More recent ones involve real sources with fabricated quotes attached, or real sources cited for things they never said. A quick glance will not catch those.

Second, specialist tools do not solve it. A peer-reviewed study from Stanford's RegLab found error rates of roughly 17 percent to 34 percent even in paid legal research tools built specifically for the job. Using a more expensive tool does not remove your obligation to check.

The practical rule is short. Anything factual that leaves your desk with your name on it gets verified by you at the original source. Statistics, dates, names, quotes, citations, prices.

Sources nortonrosefulbright.com/en-us/knowledge/publications/792d8bf3/ai-in-litigation-update-on-gen-ai-sanctions-in-2026
Stanford RegLab error rates cited via haqq.ai/blog/ai-legal-hallucination-audit. I did not read the underlying study directly, so treat the specific percentages as approximate.

Your data

The common advice is "do not use the free version, pay for the business version." The instinct is right. The line is drawn in the wrong place, and the difference matters if you handle client information.

CautionPaying more does not automatically protect your data. What matters is whether you are on a personal plan or a business plan. Anthropic's Free, Pro and Max plans are all personal plans, even though two of them cost money. Their published update states that these plans give users a choice about whether their conversations are used to improve the models, with data kept for five years when that setting is turned on. The same update states it does not apply to their commercial products, which include Claude for Work, Claude for Government, Claude for Education, and API access.

So upgrading from free to a paid personal plan does not, by itself, change anything about your data. Two things actually do.

These terms change often, and they differ between vendors. Treat this as a reason to check your own account settings, not as a permanent answer.

Source anthropic.com/news/updates-to-our-consumer-terms

Your client agreements

This is the question almost nobody raises, and it is the one most likely to cause you a real problem.

Many consulting agreements contain a confidentiality clause. Those clauses commonly say that client information will not be shared with third parties without permission. Pasting a client's document into an AI tool is, on a plain reading, sharing it with a third party.

Whether that actually breaches your agreement depends on how the clause is written and which plan you are on. The point is that it is a question with a real answer, and most people never ask it.

Three things to do

None of this is a reason to avoid AI. It is a reason to spend twenty minutes on it before a client asks.

Who owns the output

If you deliver work to clients and your agreements transfer ownership of that work, this affects you.

CheckedPurely AI-generated material is not protected by US copyright. The US Copyright Office published a report in January 2025 confirming that human authorship remains required. Work generated entirely by AI is not copyrightable. Writing detailed prompts does not, on its own, create copyright, however much effort went into them. Where a piece contains both human and AI-generated content, only the human contribution is protected.

The Office is equally clear that using AI as an assisting tool is fine. Using it for ideas, for editing, or as part of a larger piece you wrote does not affect protection for the work as a whole. They have registered over a thousand works containing some AI-generated material.

What this means in practice for you is straightforward. If a deliverable is largely AI output with light editing, you may not own what you are transferring to your client. If it is your thinking, your structure and your judgment with AI helping along the way, you are on solid ground.

Two habits keep you there. Make sure the substance and structure are genuinely yours. Keep a rough record of which parts of significant deliverables were AI-assisted, in case ownership is ever questioned.

Sources copyright.gov/ai/Copyright-and-Artificial-Intelligence-Part-2-Copyrightability-Report.pdf
Summary readings: skadden.com/insights/publications/2025/02/copyright-office-publishes-report and congress.gov/crs-product/LSB10922
This is guidance, not settled law. The Copyright Office's position is not binding on courts. It is the clearest guidance currently available. This is not legal advice.

What not to hand over

There is real enthusiasm at the moment for giving AI tools direct access to your inbox so they can read, reply and send on your behalf. This is the one recommendation worth turning down.

CautionAn assistant that can send email is an open door. These tools cannot reliably tell the difference between instructions written by you and instructions hidden inside content they are reading. That means an email arriving in your inbox can contain text the assistant treats as a command from you. The OWASP Foundation, which publishes the security risk lists used across the software industry, ranks this first among risks for AI applications. Their own example describes a mailbox assistant with sending ability being tricked by a crafted incoming email into forwarding sensitive information to an attacker.

The safe version costs you almost nothing.

Source owasp.org — OWASP Top 10 for Large Language Model Applications, 2025 edition (PDF)

What this does to your pricing

This is a business question rather than a technical one, and it catches consultants out.

If you bill by the hour and AI makes you three times faster at a task, you have just cut your own revenue on that task by two thirds. You did better work in less time and you got paid less for it. Hourly billing punishes you for getting more efficient.

There is no single right answer here, but there are three common ones.

Whichever you choose, decide it deliberately. The default outcome is a quiet pay cut.

Why saved prompts stop working

Worth knowing in advance, because it will happen and it is confusing when it does.

AI vendors update their models regularly, often without much announcement. A prompt you refined against one version can produce noticeably different results against the next. Nothing broke and you did nothing wrong. The tool underneath you changed.

Two habits handle this. Keep your saved prompts in a plain document you control, not inside a chat tool's memory, so you can read and edit them. Re-run your reference check every few months, using the same examples from the section above. If quality has slipped, you will see it in the numbers rather than sensing it vaguely.

Your first workflow

One concrete project, about an hour, difficult to get wrong.

1

Pick the document you write most often

A proposal, a discovery summary, a client update. How often you write it matters more than how important it is.

2

Find three past versions you were proud of

These become your reference set. If you cannot find three good ones, that is useful information in itself.

3

Ask what makes them work

Paste all three in and ask the tool to describe the structure, the tone and what the strongest parts have in common.

4

Turn that into a reusable instruction

Ask it to write that description up as an instruction it could follow next time. Save the text in a plain document on your own computer, not in the chat tool's memory. You want to be able to read it, edit it and take it with you if you change tools.

5

Use it

Next time you write one, paste in the saved instruction along with the new facts.

6

Compare the first ten against what you would have written

This is the step everyone skips. It is also the only one that tells you whether any of this worked.

Questions to answer first

Six questions, about fifteen minutes on paper. They prevent most of the expensive mistakes.

What to expect

RealisticReal time back on repetitive writing, first drafts, summarising, reformatting and research legwork. This arrives quickly and it is the main benefit.

RealisticA better starting point on things you find hard to begin. The blank page problem mostly goes away.

SlowerAnything needing judgment about your specific clients, relationships or market. The tool knows nothing you have not told it, and telling it properly is itself work.

UnlikelyReplacing a role outright, or a step change in revenue that traces to the tool rather than to what you did with the time it freed up.

The honest summary is this. AI is a meaningful productivity improvement and a genuine disadvantage to ignore. It is not a business model on its own. The people describing it as one are usually selling something. What you were already good at is still what clients are paying for.

Technical appendix

Implementation notes for a self-hosted multi-model setup. Not needed for anything above.

Durable Objects

Cloudflare's documentation sets out the contrast plainly. Workers are stateless functions, where each request may run on a different instance, in a different location, with no shared memory between requests. That works for an API endpoint. It does not work for an agent, which needs to remember a conversation and hold a session open.

Durable Objects are stateful compute. Each instance has a unique identity, runs in a single location, and keeps its state across requests. Each one has its own SQLite database attached, running in the same thread, so query latency is effectively zero. The key guarantee is that exactly one instance exists per ID, globally. Request the object named conversation-4821 from anywhere and you reach the same single-threaded instance with the same private storage.

The Cloudflare Agents SDK makes each agent a Durable Object. That gives you conversation state surviving deploys, single-writer correctness, persistent WebSockets, per-agent alarms so an agent can wake itself, and hibernation so idle agents cost nothing. The design rule is one object per unit of coordination, meaning one per conversation or per user. The documented anti-pattern is a single global object, which becomes a bottleneck at roughly 500 to 1,000 requests per second.

Two caveats before committing. There is no drop-in equivalent on AWS or GCP. The SDK is pre-1.0 and changes often, so pin versions.

Sourcedevelopers.cloudflare.com/durable-objects/best-practices/rules-of-durable-objects/

Billing

Consumer subscriptions are not API access. Any architecture assuming otherwise will not work.

  • Claude Max. Anthropic's help centre states that paid Claude plans and the Console are separate products, and that a subscription does not include API or Console access. A separate metered Console account is needed. Reporting also indicates OAuth authentication was restricted to Claude Code and Claude.ai in February 2026, so proxying subscription credentials into a gateway is both fragile and against terms.
  • Perplexity Pro. Sonar is billed pay-as-you-go against prepaid credits, with a per-token rate plus a per-request search fee. Sources disagree on whether the $5 monthly Pro credit still exists. Check the account settings before relying on it.

Cleanest approach: keep the subscriptions for interactive use, and open one OpenRouter account with a hard spend cap for anything routed through a gateway.

Hardware

Local inference speed is a memory bandwidth problem, not a RAM capacity problem. Capacity decides whether a model fits. Bandwidth decides how fast it generates.

MachineRole
M1 Max, 64GBHighest bandwidth in the fleet by a wide margin. Heavy local inference and eval runs when it is awake.
M4 Mac Mini, 16GBAlways on. The responsive local endpoint, 8B class via MLX.
UGREEN iDX6011 Pro, 64GBServices, model storage, embeddings, batch work. TechPowerUp's review found the NPU unused during LLM work, with inference falling back to CPU at 20 to 30 percent utilisation. Not the interactive host. The OCuLink port and PCIe Gen4 x8 slot are the real upside, since an eGPU there would collapse this whole architecture into one box.
Synology DS423+Backup target. Leave it out of this.

Bandwidth figures stated from general specification knowledge, not verified against a primary source.

Stack

  • Inference. MLX or Ollama on the Mac Mini, bound to the LAN. Ollama in Docker on the NAS with weights on the NVMe, never the array.
  • Gateway. LiteLLM in Docker on the NAS. The only component holding secrets. Per-key spend caps and request logging enabled from day one. Fallback chains so a sleeping MacBook degrades instead of erroring.
  • Interface. Open WebUI with exactly one connection, pointed at LiteLLM. SearxNG alongside it for local search.
  • Access. Tailscale rather than a public tunnel. The gateway holds every key and a spend budget, so it should not be reachable from the internet. Cloudflare Tunnel with Access policies is the right tool later, for sharing one specific thing.
  • Devin. A dispatched tool behind a Function call, not a model in the dropdown. It has its own session lifecycle and is not a chat-completion endpoint.

Note that /dispatch-style skills are a Claude Code feature. They do not carry over to Open WebUI, which uses its own Pipelines and Functions system.

Evals

Routing across eight endpoints without evals means routing on guesswork. Build a golden dataset of 30 to 50 real prompts from actual work. Write scorers that grade numerically. Run offline against that set before promoting any model in the routing config. Sample production traffic from the LiteLLM logs for ongoing scoring. Braintrust is what Scott Moss uses for observability in the course. A TypeScript script looping the golden set against three endpoints and printing a score table is a legitimate first version and takes an evening.

Two related ideas worth adopting early. Context engineering means curating exactly the tokens needed at inference time rather than stuffing everything in. Structured outputs means constraining responses to a schema so downstream code has something predictable to parse.

Build order

  1. One evening. Ollama on the Mac Mini, LAN-bound, one model, hit it with curl. Confirm the endpoint works before adding orchestration that can hide failures.
  2. One weekend. LiteLLM and Open WebUI on the NAS. Two endpoints: the Mac Mini, and one capped OpenRouter key. Add Tailscale. Install the PWA on the phone. This alone is the working system.
  3. Ongoing. Add endpoints one at a time, each with an eval before it becomes a default route. Then knowledge and vector store.
  4. Optional. Devin behind a Function. The eGPU decision.

Do not build

  • An agent with send authority on a real inbox.
  • Bypass permissions on a machine that matters. Anthropic's documentation states it offers no protection against prompt injection and is for isolated containers and VMs only. Auto mode gives fewer prompts with a classifier still checking.
  • Local models for anything where correctness matters, until an eval clears them.
  • All four layers at once. Phase 2 is useful on its own. Live with it for two weeks first.