The short answer
In 2026, adding AI to a small business's software costs a one-time build — Pythn's proofs of concept start at $5,000 and production AI tools at $20,000 — plus model usage of roughly $2 to $42 a month per use case at small-business volumes on current list prices, and the strongest places to start are document extraction, customer support triage, and internal knowledge search.
The usage figures are our estimates, worked out below from the vendors' published prices and token counts, for volumes like a thousand documents or two thousand emails a month. Ten times the volume is roughly ten times the bill. The three starting points have what makes AI safe to put in front of real work: the task repeats, there's a lot of it, and a person or a rule in code can confirm the result before it matters.
What changed since 2024: cheaper models, longer context, better tool use
Cheaper models
When Anthropic launched its Claude 3 family in March 2024, the top model, Claude 3 Opus, listed at $15 per million input tokens and $75 per million output tokens, and Claude 3 Sonnet at $3 and $15. Today Anthropic's price list puts Claude Opus 5.5 at $4 and $20 and Claude Sonnet 5.5 at $2 and $10. One caveat keeps that comparison honest: Anthropic says its models from Claude Opus 4.7 on use a tokenizer that produces about 30% more tokens for the same text, so the saving per page is smaller than the sticker prices suggest. It's still large.
The other big vendors price in the same range. List prices per million input and output tokens on October 5, 2026:
- Anthropic: Claude Haiku 4.5, $1 and $5; Claude Sonnet 5.5, $2 and $10 (pricing).
- OpenAI: gpt-6-luna, $0.10 and $0.50; gpt-6.1-sol, $2 and $10 (pricing).
- Google: Gemini 3.5 Flash-Lite, $0.30 and $2.50; Gemini 3.8 Flash, $0.75 and $3.75 through December 31, 2026, then $1.50 and $7.50 (pricing).
Two discounts matter as much as the list price. Batch processing — sending work to run later instead of right away — is half price at all three; Anthropic says most batches finish within an hour, with up to 24 hours allowed. And prompt caching bills each repeat of the same instructions at 10% of the normal input price on most Claude models; only the first send costs a little more.
Longer context
Claude 3 launched with a 200,000-token context window — the amount of text a model can consider in one request. Anthropic's current Claude Fable 5.1, Opus 5.5, and Sonnet 5.5 take 1 million tokens, which Anthropic puts at roughly 555,000 words — about a thousand pages of ordinary text in one request.
Capacity isn't free, though. A 1-million-token request to Claude Sonnet 5.5 costs $2 in input alone, every time you send it. For questions your team will ask thousands of times, it's far cheaper to find the few relevant pages first and send only those — which is how the knowledge-search example below works.
Better tool use and structured output
Tool use — a model calling your software's own functions, such as “look up this order” — became generally available across the Claude 3 family on May 30, 2024. Since then, structured outputs have closed the gap that made early integrations brittle: Anthropic now guarantees responses that match the JSON schema you supply. Your code no longer breaks on malformed output. The values inside can still be wrong, which is why every use case below keeps a check in the loop.
Five use cases, with what they cost to run
Every figure here is an estimate at list prices: price per token, times the tokens a task uses, times a monthly volume. Token counts come from Anthropic's documentation — a PDF page is 1,500 to 3,000 text tokens plus an image of the page, which costs up to about 1,600 tokens on Claude Haiku 4.5 and 4,800 on newer models, and a million tokens holds about 750,000 words on Haiku 4.5 and 555,000 on newer models. We took the top of each range and skipped caching, so real bills should come in lower. Each use case is priced on Claude Haiku 4.5 ($1 and $5) and Claude Sonnet 5.5 ($2 and $10).
1. Document extraction
Pull the fields out of invoices, purchase orders, intake forms, or contracts and write them into your database or accounting system. The model returns structured data; your software checks it — totals add up, the vendor exists, the PO number matches — and sends anything that fails, and any high-stakes field, to a person.
- Assumes: 1,000 two-page documents a month, 500 words of instructions, and about 300 words of extracted fields per document.
- Cost: about 1.2¢ to 4.2¢ per document — roughly $12 to $42 a month.
2. Customer support triage
Sort each incoming email or ticket — billing, scheduling, complaint, sales — pull out the order number, route it, and draft a reply for someone to edit and send. Nothing reaches a customer unread.
- Assumes: 2,000 emails a month of about 300 words each, 2,000 words of instructions and policy notes sent with every one, and a 200-word draft back.
- Cost: about 0.4¢ to 1.2¢ per email — roughly $9 to $24 a month. The repeated instructions are more than half of that, so prompt caching would cut the bill roughly in half.
3. Internal knowledge search
Staff ask questions in plain English and get answers drawn from your own manuals, procedures, and past jobs, with the passages it used cited so they can check. The software searches your documents first and sends only the best matches to the model; the search index can live in the PostgreSQL database you may already run, through the open-source pgvector extension.
- Assumes: 2,000 questions a month (about 100 a business day), five retrieved passages totaling 4,000 words per question, and a 250-word answer.
- Cost: about 0.8¢ to 2.1¢ per question — roughly $15 to $42 a month.
4. Drafting from your own records
Quotes, proposals, job summaries, and follow-up emails, drafted from the customer record, your price list, and past examples — then edited and sent by a person who knows the customer.
- Assumes: 300 drafts a month, 3,000 words of context for each, and a 600-word draft.
- Cost: about 0.8¢ to 2.2¢ per draft — roughly $2 to $7 a month.
5. Bulk classification and cleanup
Tag a product catalog, categorize years of transactions, or flag duplicate customer records — work that can run overnight at batch prices.
- Assumes: 50,000 records of about 100 words each, grouped 20 to a request so 1,000 words of instructions go once per group, and a short label back for each record.
- Cost: roughly $7 to $18 for all 50,000 records at batch prices.
Run the first four together at those volumes and the model bill comes to roughly $40 to $115 a month. The software around them is hosted like any other small-business app, typically $20 to $200 a month.
What the build costs
The model bill is the small number. The build is the large one, and it's where the risk gets handled. On our AI & Intelligent Tools page: proofs of concept start at $5,000 and take 2–4 weeks for a focused assistant or document-parsing tool; production AI tools start at $20,000; production retrieval systems with evaluation, guardrails, and monitoring take 6–12 weeks; and high-compliance or high-volume deployments are quoted after discovery. Every one includes an evaluation suite, guardrails, and observability — the parts that turn a demo into something you can leave running.
As with every Pythn project, scope and price are fixed in a written design document before building begins — stage three of the five-stage process. After launch, budget the same 15–20% of the build cost a year that any custom software needs (see maintenance costs); for AI, part of that is re-running the evaluation when a vendor replaces a model. Everything else we build is on the services page.
How to pilot AI without betting the company on it
- Pick one task, not a strategy. High volume, repetitive, and checkable. If nobody can tell whether the output is right, it's the wrong first project.
- Collect real examples with known right answers. Last month's invoices with the correct entries, or past tickets with how they were handled. This becomes the evaluation suite: the same test, re-run every time a prompt or a model changes.
- Run it beside the current process. During the pilot the AI's output is compared, not used. The person still does the work, and you count where the two disagree.
- Keep a person approving anything that leaves the building. Drafts get edited, extracted totals get checked, and answers carry their sources.
- Decide on numbers you wrote down in advance. Error rate, minutes saved per item, cost per item. If the pilot misses them, stop. A $5,000 proof of concept that says “not yet” is a cheap answer.
A Pythn proof of concept is built for exactly this: 2–4 weeks, from $5,000, with the evaluation suite included. If it earns its keep, the production build adds the guardrails, monitoring, and connections to your other systems.
Risks: hallucination, data leakage, and vendor lock-in
Hallucination
NIST's profile of generative-AI risks calls it confabulation: “the production of confidently stated but erroneous or false content”. It's the main reason AI shouldn't make unreviewed decisions. Anthropic's own guidance for reducing it is practical: let the model say it doesn't know, ground answers in direct quotes from your documents, and have it cite a source for each claim. Add checks in code — totals that must add up, IDs that must exist — and a person on anything high-stakes. The goal isn't a model that's never wrong. It's a system where a wrong answer gets caught.
Data leakage
Two questions. First: will the vendor train on your data? On the paid business APIs, the major vendors say no. Anthropic's commercial terms say it “may not train models on Customer Content from Services”; OpenAI says data sent to its API has not been used to train its models since March 1, 2023 unless you opt in; and Google's Gemini API pricing page says paid-tier content is not used to improve its products, while free-tier content is. Free tiers and consumer chat apps are a different deal, so keep company data off them.
Second: can the system itself leak? An assistant that searches your files should only find what the person asking is allowed to see, so permissions belong in the search, not in the prompt. And documents and emails can carry instructions of their own. OWASP ranks prompt injection first in its Top 10 for LLM applications, including indirect injection through the files and websites a model reads. Give the AI no more access than the task needs.
Vendor lock-in
Models retire. Anthropic retired Claude 3 Haiku, from the March 2024 Claude 3 family, on April 20, 2026, and commits to at least 60 days' notice before retiring a publicly released model; even today's Claude Haiku 4.5 is listed as available until “not sooner than October 15, 2026.” The defense is ordinary engineering: keep the model call behind one module in your code, keep the prompts and the evaluation suite in your own repository, and re-run the evaluation before switching models. That's why we build model-agnostic: the model should be a part you can replace, not the foundation.
Frequently asked questions
How much does AI integration cost?
Two parts. The build: at Pythn, proofs of concept start at $5,000 (2–4 weeks) and production AI tools at $20,000. The running cost: model usage of roughly $2 to $42 a month per use case in our worked examples, hosting for the software around it (typically $20 to $200 a month), and maintenance of 15–20% of the build cost a year.
Is AI ready for small business production use in 2026?
For bounded tasks with a check in the loop, yes: extracting fields that get validated, triaging tickets a person approves, answering staff questions with the sources attached. For unsupervised decisions with money or legal consequences, not yet — confabulation is a documented risk, not a rare bug. A simple test: if you can measure whether the output is right, you can deploy it carefully. If you can't, start somewhere else.
Will our data be used to train AI models?
Not on the paid business APIs from Anthropic, OpenAI, or Google, according to their published terms. Google's free Gemini API tier is another matter — its content is used to improve Google's products — and consumer chat apps have their own policies. Use a paid API account that your business owns, and read the terms for whichever tier you're on.
Which AI model should we use?
The cheapest one that passes your evaluation suite. A mid-tier model such as Claude Sonnet 5.5 or OpenAI's gpt-6.1-sol, both $2 and $10 per million tokens, is a sensible place to build; then test a smaller one, such as Claude Haiku 4.5 or gpt-6-luna, against the same examples, and switch if it passes. Your evaluation decides, not a leaderboard — and because models retire, plan on running that comparison again.
Can AI connect to QuickBooks, Stripe, or our CRM?
Yes. The AI part reads and writes through the same APIs any integration uses: extracted invoice fields can become QuickBooks bills, and triaged emails can open CRM tickets. The hard parts are the usual ones — authorization, retries, and data mapping — which our guide to connecting QuickBooks, Stripe, and your CRM covers. Each integration typically adds $2,000 to $5,000 to a project.