Qaf AI cost audit
October 9, 2026 · Based on current main, 7 days of production data (Oct 2–8) from Langfuse, Vertex, the Cloudflare AI Gateway and Azure Cost Management. Every finding was checked by two separate reviewers.
Short answer: there is no single runaway bug. The bill is high because each chat turn sends the model far more text than it needs. Three fixes cut about 20–27% of the bill (about $3–4k a week):
- Cap the chat history. The August fix (PR #321) was never merged.
- Send slimmer search results to the model.
- Replace the 36-character citation UUIDs with short ids.
Separately, Langfuse sees only about 45% of real spend, which is why its dashboards look much smaller than the invoice. Gemini 3.7 Flash prices double on Jan 1, 2027, so every saving here doubles too.
Where the money goes
Real weekly spend, Oct 2–8, reconciled from Vertex Monitoring, Cloudflare gateway analytics and Azure Cost Management.
Inside the main chat, input is about 80% of Gemini cost. That input splits into roughly 62% search results, 29–35% chat history and 8–10% system prompt and tool definitions. Terra is mostly the hadith detail sheet (about two thirds) plus chat titles. Rerank and terra appear to be paid from Azure sponsorship credits. The GCP account's cash vs credit split could not be checked.
What to fix, ranked by savings
Savings are real weekly dollars at today's prices: Langfuse-visible figures scaled up for traffic Langfuse misses. All of them double after January 1.
1. Chat history is sent in full on every step
$1.0–1.25k / weekThe chat handler loads the whole conversation with no limit and sends it on every agent step. A turn usually has two steps, so long chats pay for their history twice per message. The August history-window fix (PR #321) is still open, never merged. The two other August cost PRs (#319, #320) were closed unmerged too.
Evidence and fix
packages/api/src/handlers/chat/index.ts:166callsgetMessagesByChatId, which has no LIMIT.pruneMessagesonly removes reasoning and old tool calls; it never shortens history.- Langfuse, 168k turns: step-1 input p50 2.8k tokens, p99 116–133k, max 1.0M. Turns that start at 60k+ tokens are about 2% of turns but 12% of chat cost. Example: a 157-message chat sent 144k and then 159k tokens for one question.
- Prompts over 60k tokens are almost never cached (0–2.5%), so these tokens are paid at full price.
Fix: merge #321 or an equivalent. Keep the last whole user/assistant pairs within about 24–32k tokens, and optionally summarize older turns. Apply the same to the playground handler. Spot-check quality on long chats.
2. Search results carry diacritics, full footnotes and page data the model doesn't need
$1.2–1.7k / weekSearch results are about 62% of everything the main chat sends to Gemini. Each chunk goes to the model with full tashkeel, uncapped footnotes and the pages JSON, about 650 tokens per chunk. Each search delivers 20 chunks.
Evidence and fix
packages/api/src/ai/format-chunk.ts:12-14sendstext_with_tashkeel, footnotes and pages.formatChunkForModelonly drops two id fields.- Measured with Vertex countTokens on 600 production chunks, re-measured independently: stripping tashkeel saves 10–15%, dropping footnotes 14–18% (most chunks have none, but a few are very long), pages 3%. Together that is 28–31% of the payload.
- A stripped text field already exists (the embedding text) but isn't fetched for the model. The August "2.4x from tashkeel" figure used a different tokenizer and overstates it for Gemini.
Fix: give the model stripped text, or keep tashkeel only on the top 3–5 chunks. Cap footnotes at about 200–300 characters and send only the page number. Keep the full chunk for the UI. Gate the change with the existing eval harness, because takhrij lives in footnotes. Then test 20 vs 12–15 chunks per search, which could take savings to about $2.1k a week.
3. Citations are full UUIDs: about 42% of answer tokens
$0.85–1.3k / weekThe model reads two UUIDs per chunk and writes full 36-character UUIDs in every citation. Each one is about 30 tokens at the output price, which is five times the input price. Short per-turn ids (r1, r2…) cut answer tokens by 42–46%. As a side effect, answers stream faster.
Evidence and fix
- countTokens on 26 production answers: 31,975 tokens as written vs 18,446 with short ids.
citation-output-middleware.tsalready rewrites every citation tag in the stream, so it can expand short ids back to UUIDs before anything reaches the client or database.
Fix: assign per-turn aliases in the tool output, keep the alias map in the turn context, and expand aliases in the middleware. Use a short book alias for expand. Run a citation-accuracy eval; this may also reduce mangled ids.
4. Message quotas are skipped for some clients
~$0.5–0.6k / week + revenueThe monetization policy fails open: some requests, including everything from native builds older than 1.3.0, are never metered. No forced update exists, and about 2.5k weekly users are still on pre-1.3 builds. One unmetered account alone ran about 860 turns ($159 traced) in a week. Pro-only features leak through the same path.
Evidence and fix
packages/api/src/lib/monetization-policy.ts:45-65returnsenforce: falsein these cases.MIN_SUPPORTED_CLIENT_VERSIONis still"0.0.0"(TODO).- 86 free users show zero recorded usage despite up to 95 turns a week.
Fix: fail closed by default. Check how 1.0–1.2.x builds render a quota error, then enforce for them or force an update. Stopgap: a per-user daily hard cap that doesn't depend on headers.
5. Some search-count rules in the main prompt may cost more than they're worth
$0.3–0.8k / week (unproven)The production prompt (platform/main-gemini v24) requires "a minimum of two initial searchess for simple questions, and four for complex ones" (typo included). Each extra search adds about 7–11k input tokens to the answer step plus a rerank call. Search counts cluster at exactly 2 and 4.
Fix
Eval v24 with that line removed or narrowed to ruling questions, measuring searches per turn, tokens and judged quality before moving the production label.
6. The hadith detail sheet: $0.19 per tap
$0.2–0.35k / weekEach tap runs a terra agent over 30 un-reranked search results: about 93k input tokens, roughly 13x the cost of a chat turn. Main chat reranks to 20, but hadith, quran and entity skip reranking.
Evidence and fix
packages/api/src/handlers/chat/hadith.ts:64-86runs on terra.agentic/tools.ts:84-110reranks only whentype === 'chat'.- Response caching (4–10% repeat taps) and step caps (93% finish in 3 steps or fewer) are not worth doing.
Fix: rerank to 10–12 (or lower topK) for hadith, quran and entity, and apply the chunk slimming from #2. Moving hadith to Gemini (quran and entity cost about $0.03 per tap) saves more, but needs a tools-then-structured-output split because of the known tool-loop bug.
7. A policy classifier runs on every turn, with thinking on
$0.15–0.25k / weekdetection-type.ts calls the main Gemini route with default thinking on every message. That is about 133 hidden thinking tokens to produce a 2-token answer, and the answer only drives a UI notice (93% of results are "none"). It's small, but almost free to fix: route it to a lite or non-thinking model, or run it only on a first message or new topic. The "minimal" thinking level is rejected by Vertex; "low" saves only a little.
Monitoring and visibility gaps
8. Langfuse sees about half the bill
visibilityVertex, the Cloudflare gateway and Mixpanel all agree: Langfuse captures only 55–61% of production chat calls. Whole turns are missing; partially traced turns are rare. On top of that:
- Thinking tokens stopped being counted when chat moved to the gateway (Sep 12–13), about $560 a week.
- Rerank is logged with no price ($1.19k a week shows as $0).
- Cached tokens are priced at $0.
- Traces stopped entirely from Sep 16 to Sep 26.
Every per-feature or per-user cost number from Langfuse is understated about 2x. The gateway does not duplicate requests: per-request matching shows distinct calls, and the gap predates the gateway.
Where to look
Check LANGFUSE_* and OTEL sampler variables on every Railway replica and service. Check the self-hosted Langfuse ingestion logs for rejects, since some spans carry up to 1M tokens of input. Raise the span processor queue size and log dropped spans. Add hostname and release to traces. Then fix pricing: rerank at $0.0025 per call on one span, cached input at $0.075/M, and reasoning tokens from the gateway route. Until then, treat the Cloudflare gateway's daily cost and Azure Cost Management as the source of truth.
9. No budget alerts on Vertex or Azure AI
guardrailThe GCP billing account has no budgets. The only Azure budget is a $250 placeholder scoped to monitoring services. A spike, an abusive user or the January price change would only surface on the invoice. Fix: a daily check on gateway cost that alerts at more than 1.3x the 7-day average or any user over $10 a day, plus GCP and Azure budgets as a backstop.
Checked and fine — don't chase these
| Area | Finding |
|---|---|
| Step limits / output caps | Turns average about 2 steps (p99 4–5). Only 13 of 168k turns hit the length limit. Neither cap binds. |
| Priority PayGo | 99.99% of Gemini traffic is standard shared tier. The priority-retry code is no longer on the chat path. |
| Cloudflare gateway | No duplicate requests, no retry storms (429s about 0.1%), no markup (BYOK). |
| Rerank setup | Bills 1 unit per call, as expected. Pro with 50→20 was the eval-recommended setup. The only lever is fewer searches. |
| Long-context pricing | 3.7 Flash has no surcharge above 200k tokens. |
| Prompt caching | Prompt ordering is already cache-friendly. The Vertex implicit floor (~6k tokens) and global routing limit hits. Don't pad the prompt to reach it. |
| Abuse | No single abuser dominates (top 10 users are 8% of spend). The Oct 8 spike was organic growth. |
| Embeddings, TTS, transcription, Turbopuffer | A few dollars to about $30 a week in total. No re-embedding is running. |
| Chat titles on terra | About $60 a week. Moving to a lite model is housekeeping, not material. |
| Failed / interrupted turns | Under 1% of spend. Retries are user-initiated only, so there are no retry loops. |
Open questions
- Why does Langfuse lose ~45% of traces? Answering this needs Railway env access and the Langfuse ingestion logs.
- About 6.5k Gemini calls a day bypass the gateway (roughly $1.2–1.4k per 30 days), from other service accounts or local evals in the same GCP project. Who is making them?
- Credits vs cash: how much of the bill is covered by credits, and when do they run out?
- Detailed mode: what share of turns use it? It means high thinking, it isn't traced, and it isn't gated by plan.
Method: 7 parallel investigators (chat loop, side calls, retrieval, Langfuse data, billing, prompts, guardrails) produced 46 findings. Each one was re-checked by two independent skeptics, one on whether it's real and one on dollar impact. A final critic reconciled the conflicting numbers and measured the remaining gaps directly. Savings for findings 1–3 act on mostly separate token pools, so they roughly add up.