The Closed Books

Detecting and Investigating Anomalous Usage Patterns Before Invoicing

Catch billing errors before invoicing or risk customer disputes and revenue damage.

Contributing Editor · · 9 min read
Cover illustration for “Detecting and Investigating Anomalous Usage Patterns Before Invoicing”
Usage Data Governance · October 5, 2026 · 9 min read · 1,969 words

An invoice is a commitment. Once a metering error survives long enough to appear as a line item on a customer-facing bill, it is no longer an internal data problem but a dispute, a credit, and an explanation that someone at the company has to give. Catching that same error an hour earlier, before the invoice generates, costs nothing: no credit memo, no apology email, no finance team unwinding a finalized record. Post-invoice correction is not the mirror image of pre-invoice detection. Finance has to reverse something that has already been booked, the customer has already seen a number that undermines confidence in every future bill, and engineering has to trace the error backward through an aggregated pipeline instead of reading it off a live event stream. Revenue leakage detection and anomaly detection serve different purposes here: leakage detection catches known errors, like a usage charge that never got applied, while anomaly detection catches the deviations nobody wrote a rule for, like a discount quietly applied outside policy or usage collected but never invoiced. Both matter, and neither recovers its value once the invoice has already gone out. For AI products, usage events fire at millisecond speed, pricing rules interact across multiple dimensions at once, and usage-based pricing is spreading across the industry, so the exposure compounds faster than it did for traditional subscription software, and a single bad billing period can surface as bill shock large enough to push a customer out the door.

The failure modes that pre-invoice detection must catch

AI billing systems break in ways that traditional subscription billing never had to account for, and each failure mode leaves its own signature. Retry logic that protects against dropped events also creates the risk of ingesting the same event twice, and when that duplication surfaces only after the invoice has gone out, the fix (a credit or a discount) often overcorrects and does more damage to revenue than the original double-count did. Usage-based billing for AI companies treats deduplication and idempotency as non-negotiable properties of the system for exactly this reason: a pipeline that doesn't enforce them will eventually count something twice. A second failure mode occurs when a pricing rule changes in one part of the stack but doesn't propagate to every rating layer that needs it, producing a quiet, systematic undercharge that looks like ordinary revenue until someone runs a manual reconciliation and finds the gap; the more pricing dimensions a product has, base subscription, tiered credits, model-tier rates, volume discounts, burst pricing, the more surface area exists for exactly this kind of partial update. A third failure mode involves credit pools, whether from cloud provider promotional programs or prepaid wallets: the pool drains in the background while dashboards keep reporting zero invoiced spend, and the first real signal a customer gets is an invoice priced at full retail, covering days or hours of workload they believed was still covered. A fourth, and increasingly common, failure mode comes from agent loops and tool-call fanout: a single user action can trigger dozens or hundreds of downstream LLM calls, tool invocations, and retrieval steps, and without attribution defined explicitly at the event schema level, none of it can be audited or billed correctly; these anomalies are especially hard to catch before invoicing because the problem isn't one large event but a cluster of small ones whose sum is what matters. A fifth failure mode produces no error and no alert at all: usage gets collected but never invoiced because of a misconfigured threshold, which is among the most common anomaly types precisely because it looks like nothing happened. A sixth failure mode lives in discount and reversal behavior rather than in billing math: renewals processed above policy thresholds, inconsistent regional discount logic, or credits clustering around one product line or region are behavioral deviations, not billing errors, and a rules-based check has no reason to flag them.

The AWS Bedrock billing incident: what happens when detection fails

The AWS Bedrock billing incidents of May 2026 demonstrate what happens when anomaly detection has a coverage gap exactly where AI inference spend concentrates. Customers of Amazon Web Services and Google Cloud received invoices reaching tens of thousands of dollars for AI workloads they never authorized. One developer had set a spending cap in advance and still woke up to a substantial five-figure bill. One AWS customer had anomaly detection enabled and was still charged tens of thousands of dollars for Bedrock inference accumulated over roughly a month of experiments, with no alert ever firing. These were not isolated misconfigurations: coverage gaps of this kind exist across every major provider wherever partner-channel and SaaS-marketplace billing sits outside the main monitoring path. A second cause compounded the first: attackers scraped public GitHub repositories for exposed API keys and converted those keys directly into compute spend. Runaway spend from a stolen credential looks identical to runaway spend from a misconfigured workload unless detection is both real-time and aware of which key is driving the usage. The incidents point to a gap in detection architecture that spans spend-limit interfaces. If a spending cap only applies to one billing surface while inference charges accumulate on another, it was never going to hold, no matter how clearly it was configured.

The event contract fields anomaly detection requires

Anomaly detection can only surface what it can trace, and traceability has to be designed into the event schema before a single invoice runs, not bolted onto the pipeline afterward as a monitoring layer. Usage-based billing for AI companies defines a minimum viable event contract built around a short list of fields, and each one closes a specific gap that would otherwise make detection guesswork. Customer and project identifiers need to be stable and immutable, so that a renamed organization or a restructured account hierarchy doesn't sever the link between past usage and present billing. The action timestamp has to record when the action actually happened, not when the event was emitted, because asynchronous pipelines introduce delays that matter both for which billing period an event belongs to and for how an anomaly's timing gets interpreted. The unit and quantity fields need to record what's being measured, whether that's input tokens, output tokens, or compute seconds, and how much of it occurred, kept atomic unless the pricing model genuinely treats those units as interchangeable. A correlation ID ties the usage event back to the originating request, session, or agent run, and this single field is what lets a team trace a flagged invoice line item back to actual application logs; without it, an investigation starts at an aggregate number with no path backward to the request that produced it. A billable flag and reason code make the billing decision explicit inside the event itself rather than burying it in downstream logic where it becomes much harder to audit later. A schema version field lets a team associate pricing model changes with the specific events they applied to, which is what makes retroactive analysis possible months after the fact. Leaving out the action timestamp distinct from the emit timestamp has a specific cost: late-arriving events land in the wrong billing period and produce apparent anomalies, sudden drops or spikes, that attribution errors cause rather than genuine changes in consumption. For agent and fanout workloads in particular, attribution has to be defined at the schema level, specifying which calls are billable, which tenant or project owns them, and how tool invocations within a session get grouped, because none of that can be reliably reconstructed from an aggregated total after the fact.

Immutable raw event storage as the foundation for retroactive detection

A billing pipeline that aggregates usage events at the moment of ingestion is trading away detection capability for a modest gain in storage efficiency, and the trade is a bad one because the savings are small and the loss is permanent. Once raw events are collapsed into a summary, the pipeline retains the total but not the shape, and a team can no longer tell whether a spike came from one unusually large event, a burst of many small ones, or a duplication artifact that never should have been counted. Immutable raw event storage, meaning original events are retained and any backfill is archived rather than written over the original record, makes several things possible that aggregation forecloses. It allows retroactive anomaly detection, so a pricing rule misapplied months earlier can still be surfaced against the original event record rather than an already-corrupted summary. It allows pricing simulations, where historical usage gets rerated against a proposed new pricing configuration to project revenue impact before the change ever reaches production. It allows a full audit trail, so the lineage from a raw event to a specific invoice line item can be reconstructed whenever a dispute requires it. It allows correction without data loss: when a backfill is needed, whether from a late event batch or a deduplication fix, the original record stays intact and the correction gets applied as a separate record rather than an overwrite that destroys the evidence of what happened. This architecture also separates anomaly detection from billing, because a detection system built on raw events can query that data directly without waiting on the aggregated totals the billing engine produces. If a team cannot retain raw events, it cannot ask, six months later, whether a single unauthorized workload or a thousand small misattributed ones produced a given spike, and that question is often exactly the one a dispute requires answering.

The signals that indicate a real anomaly versus normal consumption variance

If a detection system flags every deviation from the mean, it trains its own operators to ignore it, so the useful work lies in identifying which signals distinguish a genuine pattern change from ordinary variance, and those signals differ by failure mode. Usage velocity per customer is the primary signal for runaway consumption: a sustained rate increase across a defined window, rather than a single noisy spike, is what separates a runaway agent loop from a legitimate burst of real workload. Usage-based billing for AI companies recommends automated pausing paired with a human review queue as the right response to a velocity threshold breach, rather than letting the system adjust the invoice automatically and sort out the judgment call later. Discount and credit pattern drift is a different kind of signal, pointing to misconfiguration or policy violation rather than a metering error: a sudden rise in credit notes tied to one product line, or discount rates climbing in one region without approval, are behavioral deviations that a rules-based check was never built to catch. Contract-to-billing timing gaps point to a process failure rather than a data error: a contract signed in the CRM on one date with billing not activating until weeks later produces lost cash flow and distorted revenue reporting that can look, superficially, like a data discrepancy when the real cause is a broken handoff between systems. Currency and tax variance across systems signals a data propagation failure rather than a configuration mistake: exchange rates that differ across billing, CRM, and finance systems, or local surcharges missing from an invoice, only become visible when the systems are compared against each other, not when any one of them is reviewed in isolation. Credit pool exhaustion deserves its own watch, since a pool approaching zero with no downstream notification produces the same kind of silent failure described earlier, an invoice at full retail with no warning that preceded it. Rules-based checks catch known errors, and anomaly detection catches the deviations nobody wrote a rule for. A detection system built only on thresholds will miss every one of the behavioral drift signals described here, because none of them cross a fixed numeric line, they are visible only as a pattern breaking from what came before.

More in Usage Data Governance