Skip to content

What an auditor needs from an AI usage log

Who asked, when, which model answered, what it cost, and what shape the answer had. The fields that turn an AI log into evidence, the two kinds of row worth keeping apart, and the case for leaving conversation content where it was written.

Updated · 7 min read

Five questions a log answers

An auditor arriving at an AI feature is asking one thing: can this organisation account for what it did. How clever the model is comes second. Five questions carry almost all of the answer.

  1. Who asked. A named person, taken from the signed-in session, not a service account everybody shares.
  2. When. A time on every row, in a sequence a reader can follow in order.
  3. Which model. The model that answered, read from the response, not the setting. A quiet fallback to a different model is exactly the event you want visible.
  4. What it cost. Tokens in, tokens out, and a money figure labelled as the estimate it is.
  5. What shape the answer had. How many rows a query returned, what the assistant read on the way, how many actions it proposed. The shape, without the contents.

A log that answers those five can be handed to a reviewer. A log that answers three starts a conversation about the other two, usually at the worst possible moment.

What one row carries

Identity and place
The signed-in person, and the surface they were on when they asked. Place matters: the same question means something different asked from an audit console and from a migration.
The model
The model id reported by the response.
The tokens
Counts in and counts out. This is the one usage figure that reconciles against a provider invoice.
The read steps
The observable reads the assistant made while composing its answer, kept as labels. This is how a reviewer sees where a number came from.
The proposals
How many approve to run actions the answer put in front of the person. A count. An unusual day is visible without opening anything.

The read steps are the field most logs leave out, and the field a reviewer reaches for first. An answer becomes checkable once you can see what it consulted. Otherwise a reviewer is comparing a paragraph against their own memory of the data. That is the position the log was built to get them out of.

Two kinds of row

Usage from a conversational assistant and usage from a feature inside a tool are different events. They are recorded as different kinds of row.

  • Assistant rows are written where a prompt and an answer exist on the server. They carry the fullest metadata: person, surface, model, tokens, read steps and proposal counts.
  • Tool rows are emitted by each tool as it works. They carry usage alone: which tool, which person, which model, how many tokens. The tools emit no content, so a row has none in it to begin with.

Folding the two together would give one table with half its columns permanently empty. An empty column reads as a missing value rather than a category that never had one. Two honest shapes beat one shape with holes in it.

Content stays with its author

The strongest thing an AI log can do is record everything about a call except its content. By default these rows carry metadata alone. The reasons are about governance, not storage.

A default row records who, where, which model, how many tokens, what was read, and how many actions were proposed.

  • An administrator reading transcripts is reading colleagues’ working notes. That is a different power from reading an audit trail, and it was never the power the console was asked for.
  • A transcript store is a second copy of your data, in a place your classification work never assessed. Whatever a prompt quoted from a data source now lives in a log file as well.
  • Retention becomes a question you have already answered. Usage metadata can sit there for years with nobody losing sleep. Transcripts ask a harder question every year they are kept.

Full capture stays available, and it is switched on deliberately. It is there for a regulated deployment whose own policy requires content audit and says so in writing, and it stays the exception. A team that needs it knows it needs it, and everybody else keeps a log they can retain without a second review.

What a query records

Where the assistant runs a query, the query itself is audited. The statement it ran, how long it took, how many rows came back, and whether it succeeded. The returned cells stay out of the record. The row count is the shape worth keeping, and the data is already where it belongs.

Refusals are recorded with the same weight as successes, and this is the part worth arguing for. A query the guard turned away is the most informative row in the file. It is the one place a reader can watch the boundary do its job. A log that keeps only the calls that worked has quietly deleted its own evidence.

The scope a query ran inside is recorded as a hash, not a name. A reviewer can still tell two scopes apart and follow rows across a session. The log itself stays clear of becoming a directory of who has access to what.

Cost, labelled as an estimate

Money figures in an AI console are computed from token counts and a price table held in the software. That makes them useful for spotting a change in behaviour, and wrong for anything else.

  • They are an operator estimate, and the interface says so. A figure presented as a bill will be treated as one by the first person who screenshots it.
  • Prices move. A table shipped in a release reflects the prices on the day it was written, and your provider invoice is the authority.
  • Token counts are the durable part. They reconcile against your provider, they never go stale, and a cost conversation is better built on them.

The value of the estimate is comparative. One person’s usage against another’s, this month against last, one tool against the rest. Those readings stay sound even when the absolute figure has drifted from the price list.

Bounded, and honest when empty

  • It is bounded. The file is capped and the oldest lines fall away, so an audit log can never fill the disk the running product depends on.
  • It is owner-only on disk, appended one line at a time, and read back newest first.
  • It stays out of the critical path. A failure to write a row is swallowed, not raised. A full disk leaves somebody’s working session intact. The boundary that stops a bad query runs before the log, and is untouched by it.
  • It is honest when empty. With the feature dormant the store holds nothing and the console says so, with no invented figure to fill the card.

The last one sounds like the smallest, and it decides whether anybody believes the other three. A dashboard willing to show zero when there is nothing to show is one you can trust on the day it shows something.