An estimate you can explain.
Start with the work, not the tokens.
Choose summarisation, extraction or translation. Paste or upload a representative document and enter your monthly document volume. TokenMunch calculates document size and estimates answer length automatically. Processing happens inside your browser, with no LLM request.
For a summary, the initial answer suggestion is 10% of the document up to 1,000 words. Extraction starts at 20%, up to 1,000 words. Translation starts at the original word count. These are automatic planning assumptions, not measured predictions.
Try a document estimate →What the number includes
The planner estimates input and output text API charges. It adds a 200-token prompt allowance and lets you specify repeated calls and a percentage allowance for retries. Every call is assumed to use the same document and answer length. Different stages of a RAG or agent workflow need separate estimates.
Word counts use 1.4 tokens per word centrally. The planning range uses 1.15–1.8 input tokens per word and 65%–160% of planned output tokens. Local text counts use a shared tokenizer estimate; provider tokenizers can differ. A counted document holds input constant in the range.
The range illustrates uncertainty. It is not a statistical confidence interval or a spending guarantee. The model table tests context and output limits, including the high scenario; it cannot guarantee accuracy or language performance.
What it leaves out
ChatGPT or Claude subscriptions are separate from API charges. The simple planner excludes OCR, embeddings, retrieval, storage, hosting, taxes, external tools, cache writes/storage and reasoning tokens beyond the answer estimate. It assumes no batch or caching discounts.
The advanced calculator supports exact token inputs, a wider endpoint catalog, image estimates and optional caching/batch assumptions. Discount support varies by provider and endpoint; a direct-provider batch rate may not apply through a third-party host. Verify the selected endpoint before budgeting a discount.
Sources, dates and model identity
GPT-6.1 Sol pricing · Claude Sonnet 5.5 pricing · Claude Haiku 4.5 pricing · Gemini 3.1 Flash-Lite pricing
Each result links to its pricing source and shows its endpoint/host. The catalog is a dated snapshot, not a live price guarantee. Public list prices can differ from negotiated rates and invoices. Model aliases can change; record the exact deployed version when evaluating a production switch.
OpenAI pricing ↗ · Anthropic pricing ↗ · Google pricing ↗ · OpenRouter endpoints ↗
All estimates are in USD. There is no implied currency conversion. Saved plans preserve workload assumptions; their displayed estimates use the current catalog and can change after a price update.
Compare forecasts with actual costs
Export metadata from your usage system and reshape it to exactly four CSV columns: model,input_tokens,output_tokens,cost_usd. Each row is one request, with non-negative integer tokens and a billed USD amount. Include failed or retried calls if they incurred charges. Use the same billing month as the forecast.
Quoted fields, extra columns and prompt data are rejected. Files are parsed locally. Saving keeps only aggregate request count, tokens, cost, number of distinct models and import time—not raw rows or model names. A lower bill can reflect less traffic or a changed workload, so it is not automatically a saving.
Download sample CSV ↓Private work, deliberate sharing
Saved projects require sign-in and are tied to your account. Document content and filenames stay in the browser. Project names, numeric assumptions and optional aggregate actuals are stored when you save. Revision history retains previous saved values.
Creating a share link exposes only a snapshot of numeric workload assumptions to anyone holding that link. It excludes your project name, uploaded content, imported totals and history. Replacing or revoking the link invalidates the old link. Someone who has already copied an estimate can retain their own copy.