Sign 1: the bill grows faster than the users
Pull two curves: monthly provider spend and monthly active users of the AI feature. Healthy products track together; a widening gap means cost per user is drifting, and cost per user never drifts for no reason. The usual drivers: conversations lengthening (context compounds, per the chatbot cost math), prompt bloat accumulating from every “small” edit, model mix creeping upward, or the heavy tail widening. An aggregate bill shows none of these until they are expensive; the ratio shows them early.
Sign 2: step changes you cannot explain
Bills should move smoothly with usage. A day or week that jumps without a launch or a traffic event behind it means something structural happened: a runaway loop that ran for hours (the signature in agent loop detection), a prompt edit that silently broke caching (the 25 percent trap), a fallback chain routing to a pricier model, or one new customer whose usage is unlike the rest. If your reaction to a bill jump is opening a spreadsheet rather than a dashboard, that is the sign firing.
Sign 3: you fail the one minute test
The test: name your five most expensive users this month, and their plans, in under a minute. This is the decisive sign because it is not about whether a problem exists, it is about whether you could see one. AI cost problems concentrate in specific accounts, the pattern in one user costing more than they pay, and provider dashboards structurally cannot answer per customer questions, per the OpenAI reporting gap. Failing this test means signs 1 and 2 are invisible to you even when they are happening.
Scoring and next moves
- Zero signs: healthy, keep the ratio from sign 1 on a monthly glance.
- One sign: instrument now, while it is cheap. Attribution added before the incident is two lines; reconstructing after one is archaeology.
- Two or three: you have a margin problem with a bill attached. Instrument first, then let the per user and per feature numbers choose between caps, routing, caching, and repricing, in that order of friction. Optimizing before measuring fixes the wrong feature more often than not.
FAQ
How do I know if my AI feature has a margin problem?
Three signs, in escalating order: your provider bill grows faster than your active users (cost per user is drifting up), your bill has step changes you cannot explain from product changes (a heavy tail or a runaway is forming), and you cannot answer which customers cost the most within one minute (you have no attribution, so problems are invisible until the invoice). Two or more of these and the margin problem is not hypothetical, it is just unmeasured.
Why does the provider bill growing faster than users signal trouble?
Because it means cost per user is rising, and something is driving it: conversations getting longer, prompts accumulating bloat, a heavier model mix, or a widening heavy tail. None of these show up in an aggregate bill until they are large. Bill growth tracking user growth is healthy; bill growth outpacing it is the earliest measurable smell.
What is the heavy tail and why does it decide profitability?
AI usage is extremely skewed: most users are light, a few are 10x to 100x the median. On flat pricing, those few can each cost more than they pay while the average looks healthy, so the product loses money precisely on its most engaged users. Every flat priced AI product has this exposure; the only question is whether it is measured and capped.
What should I do if I recognize these signs?
Instrument before optimizing: attribute cost per user and per feature at call time, then let the data pick the fix, caps on the tail, a cheaper model for the bulk, caching for repeated context, or repricing. Teams that optimize before measuring usually fix the visible feature while the actual leak is a different one; per feature numbers end that guessing.
How fast can I get these numbers if I have none today?
About ten minutes for the first real data point: wrap your existing client in two lines with Weckr, pass your user id and plan, and cost per user against plan price starts accumulating immediately, with runaway detection on by default. The free tier covers 50,000 requests a month, which is plenty to diagnose with.
Keep reading
Pass the one minute test permanently
Weckr exists to make sign 3 impossible to fail: cost and margin per user and per feature, live, with runaway alerts and caps on the tail. The demo shows the exact screens on seeded data, and the integration is ten minutes, free for 50,000 requests a month.