A LITTLE CHECK BEFORE YOUR NEXT BIG IDEA

Munch more.
Spend less.

See what AI will cost to process your text, documents or images. Compare models before you spend.

Check my AI costs No sign-up. No API key. Just the math.
SAME INPUT.
VERY DIFFERENT BILLS.
GPT-6 Astra$3,600/mo180K TOKENS × 2,000 RUNS
Claude Sonnet 5.5$720/moSAME WORKLOAD. DIFFERENT RATE.
Gemini 3.1 Flash-Lite$90/mo
97.5%LESS

Example input costs. Different models, different capabilities.

ONE WORKLOAD✳OpenAI✳Anthropic✳Google Gemini✳162 MODELS · 14 PROVIDERS
THE COST PLAYGROUND / 001

Okay, what are
we building?

A prompt. A document. A million requests.
One quick check or an ongoing plan. Your call.

What would you like to work out?

Choose a starting point. You can switch any time.

BUILT FOR BUILDERS, NOT BILL SHOCK.

Less “wait, how much?”
More “look what I built.”

Documents stay on your device.
No AI calls. No accounts. No funny business.

A few loose ends.

Free? What’s the catch?

No catch. The core calculator runs in your browser and makes no AI inference calls, so calculations don’t come with an AI bill. No account or payment required.

What’s included in the number?

Your input tokens, an automatic reply allowance, and published model rates. Choose a one-off check for prices per use, or add a recurring frequency for monthly estimates. If you’re unsure about usage, we keep the comparison per use. Reply costs are included automatically. We check task clues and input length using local rules, with a summary assumption for uploaded documents and a medium-answer assumption for token counts. Include any billed reasoning tokens in that estimate. Optional discounts cover reused input and delayed batch processing. Check the linked provider pricing before committing spend.

Can I throw a document in here?

Yes. One PDF, DOCX, TXT, or MD file up to 10 MB. Extraction happens on your device. Markdown is counted as raw text, including headings and code. For image cost estimates, use the Image cost tab to select a PNG, JPG or WebP up to 5 MB and 8,000 pixels per side. No OCR: scanned image-only PDFs need text extraction elsewhere first. Local token counts are estimates across providers.

So I should just pick the cheapest?

Price is one part of the decision. Test quality and latency on your workload. The size filter checks whether a model has room for your input and estimated reply. It does not measure answer quality or accuracy.

NO BLACK BOXES

A small calculator. Clear assumptions.

How the estimate works

Input cost per run = uncached input × input rate + cached input × cache-read rate. Output cost per run = estimated output tokens × output rate. Divide token-based amounts by 1,000,000 because rates are per million tokens. Total cost per use = input cost + output cost. Monthly total = cost per use × estimated monthly uses. Hourly usage assumes 24 hours × 30 days (720 hours); daily usage assumes 30 days; weekly usage assumes 52 weeks divided by 12 months, rounded to the nearest whole use. A recurring estimate treats every document or request as one model call of the same input and output size. With no usage estimate, costs and savings are shown per use. Reply costs are included automatically. Local rules look for task and length instructions in pasted text or image prompts. Uploaded documents default to a summary; token-only input defaults to a medium answer. You can change the assumed task, shorten or lengthen the reply, or enter an exact count. The displayed planning range is a heuristic scenario range, not a statistical confidence interval. Billed reasoning tokens must be included in that output estimate when applicable. No actual response is generated. The allowance is not a prediction of exact model output.

How we choose the reply allowance

These are our planning rules, not provider predictions. Summaries use 8% of input size, bounded to 200–2,400 tokens. Translation or rewriting uses 115% with a 250-token minimum; extraction uses 70% with a 400-token minimum. Coding uses 80%, bounded to 1,200–12,000 tokens. General answers use a size-based allowance between 450 and 1,800 tokens; image descriptions start at 350. Detected word or sentence instructions take precedence in automatic mode. Shorter halves the allowance; Longer doubles it. We round up to 25 tokens and cap at 1 million. The planning range is 65%–160% of the allowance. These rules cannot infer every task or hidden reasoning requirement.

Token counts and context

Text is counted locally with the o200k_base tokenizer. Treat counts as estimates for the listed models, especially Claude and Gemini, which use different tokenizers. Manual mode applies your supplied count equally across models. Add system prompts, tool schemas, conversation history to your input budget.

Size checks verify the input limit, maximum output length, and combined input plus output for models with a shared context window. For models with separate input limits, input and output are checked independently. Passing these checks is a planning estimate, not a guarantee that a provider will accept a request.

Image cost estimates

One PNG, JPG or WebP per use. Your browser reads its dimensions; no image is sent to a model and no OCR is performed. Image input cost = model-specific image tokens × input price ÷ 1,000,000. Total = image + text prompt + estimated reply. Image tokens also count toward capacity limits. We use direct provider standard rates for the six supported image estimates; cache and batch controls apply only to text and document modes.

OpenAI GPT-4.1 and GPT-4o use 85 base tokens plus 170 per 512 px tile at high detail; GPT-4o-mini uses 2,833 plus 5,667 per tile. Low detail uses only the base allowance. High detail fits within 2,048 px, then reduces a shortest side above 768 px. Claude uses 28 px patches with automatic resizing: Haiku 4.5 has a 1,568 px edge and 1,568-token budget; Sonnet 5.5 and Opus 5.5 have a 2,576 px edge and 4,784-token budget. These are cost estimates, not equivalent image quality. OpenAI image rules · Claude image rules.

Prompt counts use the local text tokenizer and can differ from provider billing. Message framing and other request overhead are excluded. Animated images are outside the estimate; use a still image. Image generation and editing prices are not included. Rates and formulas checked October 5, 2026.

Cache and batch assumptions

The cache percentage is the fraction of input tokens read from an existing cache. Initial cache creation, write premiums, storage duration and minimum cache thresholds are excluded. Batch processing uses separately verified direct-provider rates where listed and assumes the workload can wait. For unverified batch or missing cache pricing, standard rates remain in use and the row is labeled. We never assume a universal discount. Contributor models can have different data-sharing terms; check the source before use.

Your documents stay on your device

PDF, DOCX, TXT and Markdown extraction and token counting happen in your browser. Document content is never sent to an AI provider or our server. Workloads are kept in memory and cleared on reload. Limit: one 10 MB file, 4 million extracted characters, and 1 million locally counted tokens. No OCR; scanned PDFs require text extraction elsewhere first.

What’s outside this estimate

Output or reasoning tokens beyond your estimate, taxes, regional premiums, tools, web search, image generation, unverified image models, audio, retries, cache writes and storage, negotiated discounts, and model quality. The calculator makes no AI inference calls.