Qaf AI cost audit

October 9, 2026 · Based on current main, 7 days of production data (Oct 2–8) from Langfuse, Vertex, the Cloudflare AI Gateway and Azure Cost Management. Every finding was checked by two separate reviewers.

Short answer: there is no single runaway bug. The bill is high because each chat turn sends the model far more text than it needs. Three fixes cut about 20–27% of the bill (about $3–4k a week):

Separately, Langfuse sees only about 45% of real spend, which is why its dashboards look much smaller than the invoice. Gemini 3.7 Flash prices double on Jan 1, 2027, so every saving here doubles too.

~$14.6kReal AI spend per week (~$62–64k per 30 days)
87%Of spend is Gemini 3.7 Flash (the main chat)
~$6.4kWhat Langfuse shows per week (44%)
+$50k/moFrom the Jan 1 price doubling at today's volume

Where the money goes

Real weekly spend, Oct 2–8, reconciled from Vertex Monitoring, Cloudflare gateway analytics and Azure Cost Management.

Gemini 3.7 Flash
$12.5–13k
Cohere rerank
$1.19k
GPT-5.6 terra
$0.66k
Voice, embeddings
~$0.03k

Inside the main chat, input is about 80% of Gemini cost. That input splits into roughly 62% search results, 29–35% chat history and 8–10% system prompt and tool definitions. Terra is mostly the hadith detail sheet (about two thirds) plus chat titles. Rerank and terra appear to be paid from Azure sponsorship credits. The GCP account's cash vs credit split could not be checked.

What to fix, ranked by savings

Savings are real weekly dollars at today's prices: Langfuse-visible figures scaled up for traffic Langfuse misses. All of them double after January 1.

1. Chat history is sent in full on every step

$1.0–1.25k / week
high confidence7–9% of bill

The chat handler loads the whole conversation with no limit and sends it on every agent step. A turn usually has two steps, so long chats pay for their history twice per message. The August history-window fix (PR #321) is still open, never merged. The two other August cost PRs (#319, #320) were closed unmerged too.

Evidence and fix
  • packages/api/src/handlers/chat/index.ts:166 calls getMessagesByChatId, which has no LIMIT. pruneMessages only removes reasoning and old tool calls; it never shortens history.
  • Langfuse, 168k turns: step-1 input p50 2.8k tokens, p99 116–133k, max 1.0M. Turns that start at 60k+ tokens are about 2% of turns but 12% of chat cost. Example: a 157-message chat sent 144k and then 159k tokens for one question.
  • Prompts over 60k tokens are almost never cached (0–2.5%), so these tokens are paid at full price.

Fix: merge #321 or an equivalent. Keep the last whole user/assistant pairs within about 24–32k tokens, and optionally summarize older turns. Apply the same to the playground handler. Spot-check quality on long chats.

2. Search results carry diacritics, full footnotes and page data the model doesn't need

$1.2–1.7k / week
high confidence8–11% of bill

Search results are about 62% of everything the main chat sends to Gemini. Each chunk goes to the model with full tashkeel, uncapped footnotes and the pages JSON, about 650 tokens per chunk. Each search delivers 20 chunks.

Evidence and fix
  • packages/api/src/ai/format-chunk.ts:12-14 sends text_with_tashkeel, footnotes and pages. formatChunkForModel only drops two id fields.
  • Measured with Vertex countTokens on 600 production chunks, re-measured independently: stripping tashkeel saves 10–15%, dropping footnotes 14–18% (most chunks have none, but a few are very long), pages 3%. Together that is 28–31% of the payload.
  • A stripped text field already exists (the embedding text) but isn't fetched for the model. The August "2.4x from tashkeel" figure used a different tokenizer and overstates it for Gemini.

Fix: give the model stripped text, or keep tashkeel only on the top 3–5 chunks. Cap footnotes at about 200–300 characters and send only the page number. Keep the full chunk for the UI. Gate the change with the existing eval harness, because takhrij lives in footnotes. Then test 20 vs 12–15 chunks per search, which could take savings to about $2.1k a week.

3. Citations are full UUIDs: about 42% of answer tokens

$0.85–1.3k / week
medium-high confidence6–9% of bill

The model reads two UUIDs per chunk and writes full 36-character UUIDs in every citation. Each one is about 30 tokens at the output price, which is five times the input price. Short per-turn ids (r1, r2…) cut answer tokens by 42–46%. As a side effect, answers stream faster.

Evidence and fix
  • countTokens on 26 production answers: 31,975 tokens as written vs 18,446 with short ids.
  • citation-output-middleware.ts already rewrites every citation tag in the stream, so it can expand short ids back to UUIDs before anything reaches the client or database.

Fix: assign per-turn aliases in the tool output, keep the alias map in the turn context, and expand aliases in the middleware. Use a short book alias for expand. Run a citation-accuracy eval; this may also reduce mangled ids.

4. Message quotas are skipped for some clients

~$0.5–0.6k / week + revenue
high confidencealso a revenue leak

The monetization policy fails open: some requests, including everything from native builds older than 1.3.0, are never metered. No forced update exists, and about 2.5k weekly users are still on pre-1.3 builds. One unmetered account alone ran about 860 turns ($159 traced) in a week. Pro-only features leak through the same path.

Evidence and fix
  • packages/api/src/lib/monetization-policy.ts:45-65 returns enforce: false in these cases. MIN_SUPPORTED_CLIENT_VERSION is still "0.0.0" (TODO).
  • 86 free users show zero recorded usage despite up to 95 turns a week.

Fix: fail closed by default. Check how 1.0–1.2.x builds render a quota error, then enforce for them or force an update. Stopgap: a per-user daily hard cap that doesn't depend on headers.

5. Some search-count rules in the main prompt may cost more than they're worth

$0.3–0.8k / week (unproven)
medium confidence

The production prompt (platform/main-gemini v24) requires "a minimum of two initial searchess for simple questions, and four for complex ones" (typo included). Each extra search adds about 7–11k input tokens to the answer step plus a rerank call. Search counts cluster at exactly 2 and 4.

Fix

Eval v24 with that line removed or narrowed to ruling questions, measuring searches per turn, tokens and judged quality before moving the production label.

6. The hadith detail sheet: $0.19 per tap

$0.2–0.35k / week
high confidence on cost

Each tap runs a terra agent over 30 un-reranked search results: about 93k input tokens, roughly 13x the cost of a chat turn. Main chat reranks to 20, but hadith, quran and entity skip reranking.

Evidence and fix
  • packages/api/src/handlers/chat/hadith.ts:64-86 runs on terra. agentic/tools.ts:84-110 reranks only when type === 'chat'.
  • Response caching (4–10% repeat taps) and step caps (93% finish in 3 steps or fewer) are not worth doing.

Fix: rerank to 10–12 (or lower topK) for hadith, quran and entity, and apply the chunk slimming from #2. Moving hadith to Gemini (quran and entity cost about $0.03 per tap) saves more, but needs a tools-then-structured-output split because of the known tool-loop bug.

7. A policy classifier runs on every turn, with thinking on

$0.15–0.25k / week
high confidence

detection-type.ts calls the main Gemini route with default thinking on every message. That is about 133 hidden thinking tokens to produce a 2-token answer, and the answer only drives a UI notice (93% of results are "none"). It's small, but almost free to fix: route it to a lite or non-thinking model, or run it only on a first message or new topic. The "minimal" thinking level is rejected by Vertex; "low" saves only a little.

Monitoring and visibility gaps

8. Langfuse sees about half the bill

visibility
high confidence the gap is realroot cause not confirmed

Vertex, the Cloudflare gateway and Mixpanel all agree: Langfuse captures only 55–61% of production chat calls. Whole turns are missing; partially traced turns are rare. On top of that:

Every per-feature or per-user cost number from Langfuse is understated about 2x. The gateway does not duplicate requests: per-request matching shows distinct calls, and the gap predates the gateway.

Where to look

Check LANGFUSE_* and OTEL sampler variables on every Railway replica and service. Check the self-hosted Langfuse ingestion logs for rejects, since some spans carry up to 1M tokens of input. Raise the span processor queue size and log dropped spans. Add hostname and release to traces. Then fix pricing: rerank at $0.0025 per call on one span, cached input at $0.075/M, and reasoning tokens from the gateway route. Until then, treat the Cloudflare gateway's daily cost and Azure Cost Management as the source of truth.

9. No budget alerts on Vertex or Azure AI

guardrail

The GCP billing account has no budgets. The only Azure budget is a $250 placeholder scoped to monitoring services. A spike, an abusive user or the January price change would only surface on the invoice. Fix: a daily check on gateway cost that alerts at more than 1.3x the 7-day average or any user over $10 a day, plus GCP and Azure budgets as a backstop.

Checked and fine — don't chase these

AreaFinding
Step limits / output capsTurns average about 2 steps (p99 4–5). Only 13 of 168k turns hit the length limit. Neither cap binds.
Priority PayGo99.99% of Gemini traffic is standard shared tier. The priority-retry code is no longer on the chat path.
Cloudflare gatewayNo duplicate requests, no retry storms (429s about 0.1%), no markup (BYOK).
Rerank setupBills 1 unit per call, as expected. Pro with 50→20 was the eval-recommended setup. The only lever is fewer searches.
Long-context pricing3.7 Flash has no surcharge above 200k tokens.
Prompt cachingPrompt ordering is already cache-friendly. The Vertex implicit floor (~6k tokens) and global routing limit hits. Don't pad the prompt to reach it.
AbuseNo single abuser dominates (top 10 users are 8% of spend). The Oct 8 spike was organic growth.
Embeddings, TTS, transcription, TurbopufferA few dollars to about $30 a week in total. No re-embedding is running.
Chat titles on terraAbout $60 a week. Moving to a lite model is housekeeping, not material.
Failed / interrupted turnsUnder 1% of spend. Retries are user-initiated only, so there are no retry loops.

Open questions

Method: 7 parallel investigators (chat loop, side calls, retrieval, Langfuse data, billing, prompts, guardrails) produced 46 findings. Each one was re-checked by two independent skeptics, one on whether it's real and one on dollar impact. A final critic reconciled the conflicting numbers and measured the remaining gaps directly. Savings for findings 1–3 act on mostly separate token pools, so they roughly add up.