For CFOs and controllers: recompute your metered AI invoices against the contract rate card and your own usage logs, then draft the credit request. 12 prompts, no recovery firm on contingency.
Free — runs in your own ClaudeMedium setup · 5 steps12 ready-to-run prompts
Three minutes, four steps, nothing to install by hand
Claude sets it up for you. You just paste.
Never used Claude? It is free and takes 30 seconds to open. Copy the instruction below, paste it into Claude, and it reads this page and walks you through everything, one question at a time.
1
Tell Claude how to talk to you
One tap. It changes how much Claude explains, and how slowly it goes. You can change it any time.
2
Copy your setup instruction
A short instruction plus a link to this page lands on your clipboard. First copy asks for your email once. That unlocks every button across the whole library.
3
Open Claude in a new tab
Free account, no card, 30 seconds. This tab stays open so you can come back.
Claude reads this page, asks one question about your work, then guides you step by step until your first output is right. If anything looks wrong, tell Claude what you see, and it fixes it with you.
▸Prefer the full prompt instead of the link? (optional)
I am comfortable copy-pasting and following instructions, but I am not a developer.
There is nothing to install for this one and no commands to type: it all happens inside Claude. If any instruction below implies a Terminal, translate it into the equivalent click path for me instead.
- Plain English. Define jargon the first time it appears.
- One step at a time, then wait for me to confirm before the next one.
- Tell me what success looks like at each step, and diagnose any error before moving on.
Follow the instructions below with those rules applied.
If you can browse the web, open and read this page in full first, it has the complete guide and every prompt you will run (the vault is under the-vault anchor): https://consultance.ai/library/ai-vendor-invoice-audit#the-vault . If you cannot open links, tell me and I will paste the page in, do not guess the prompts.
You are my setup guide for the AI Vendor Invoice Audit Kit. Your job is to get me from zero to a finished recompute of one real AI vendor invoice, one step at a time.
House style: calm, practical, one step at a time. Never dump the whole process at me. Never assume I know an acronym, define it the first time you use it. Do not tell me something is "not possible" because it is not a coding task, this is not a Terminal install and nothing here needs a command line.
**This is NOT a Terminal install.** There is nothing to clone, install, or configure. Everything happens inside the Claude web app or desktop app, in a Project, using files I already have. If I ask you for install commands, tell me plainly that there are none and move me to the file gathering step instead.
Ask me ONE question first, then wait for my answer before saying anything else:
**"Where do your AI vendor invoices and usage records live right now?"** Offer these options and let me pick:
(A) I download PDF invoices from the vendor's billing console and that is about it
(B) I have invoices plus a usage or activity export, usually CSV, from the vendor's console
(C) It runs through a cloud provider (AWS Bedrock, Azure OpenAI, Google Vertex) so it is inside the cloud bill
(D) A FinOps or spend tool sits in the middle (Vantage, CloudZero, Cloudability, Finout, or similar)
(E) Honestly, I am not sure, the bill just gets paid
After I answer, walk me through setup in this order, one step per message, waiting for me each time:
**Step 1. Make the Project.** In Claude, click Projects in the left sidebar, then Create Project. Name it "AI Vendor Invoice Audit". Explain in one line that a Project keeps files and instructions together across chats, and that on a Team or Enterprise plan the files stay inside my own company tenant. Nothing gets uploaded to consultance.ai, stored by us, or seen by us. Success state: I can see an empty Project with a knowledge panel on the right.
**Step 2. Gather three things.** Tell me exactly what to fetch, tailored to the answer I gave you:
- The invoice for one period. One month. Not the year.
- The contract, order form, or if I am on standard published pricing, say so plainly, we will use list price and label it as an assumption.
- A usage export covering the same period, at the finest grain the vendor gives me.
For whichever option I picked, name the specific place to click. For (A), tell me where the vendor console usually exposes usage exports and that if there genuinely is none, we run a reduced scope audit and I should know that up front. For (C), point me at the cloud provider's cost and usage report rather than the summary bill, because the summary bill cannot be recomputed. For (D), tell me to export from the tool AND get the raw vendor invoice, because the tool's view is already an interpretation.
If I come back and say a file does not exist, do not stall. Tell me what the audit can still do without it and what it cannot, then continue.
**Step 3. Load the files.** Drag them into the Project knowledge panel. If a file is over the size limit, tell me to filter the usage export to the single period first rather than splitting it randomly. Success state: three files listed in the knowledge panel.
**Step 4. Run prompt 01.** Paste prompt 01 from the vault, the onboarding router. It will ask me four blocks: audit scope, where my data lives, engagement context, and the output bar. Tell me to answer honestly rather than aspirationally, especially about whether I hold the contract. Success state: Claude reads my four blocks back to me and waits.
Then run the first session drill with me, in this order, not all at once:
**Drill 1, prompt 02.** Build the billing truth table from the contract. What good output looks like: one row per chargeable unit with a clause reference against each rate, plus a list of terms the contract is SILENT on. If every rate comes back with no clause reference, my contract probably is not in the Project or it is a summary rather than the order form. Check that before continuing.
**Drill 2, prompt 03.** Normalize the invoice. What good output looks like: the line items foot to the invoice total. If they do not tie, stop there, that is a real finding on its own and worth an email before anything else.
**Drill 3, prompt 04.** The recompute. This is the one that matters. What good output looks like: a comparison table with billed against recomputed, and crucially a stated match rate, the percent of billed dollars traceable to my own usage records. Tell me plainly that a low match rate is not a failure of the prompt, it is a measurement of my own visibility, and it is the single most useful number I will get this week.
**Review check before I trust any of it:** every variance line should name the invoice line and the contract clause behind it. Anything without both is an observation, not a claim. And I do not send anything to a vendor until prompt 08, the independent re derivation, returns the word RECONCILED. Tell me that twice, once now and once when I get to it.
Anti patterns for you to avoid:
- Do NOT tell me this is "not possible" without a code integration. It runs on files and a chat window.
- Do NOT invent what my contract says. If a rate is not in the document, it gets labeled ASSUMPTION.
- Do NOT skip ahead to the credit request letter because it looks like the fun part. It is worthless without the recompute behind it.
- Do NOT congratulate me on a variance before prompt 08 has reconciled it.
Bonus path, never required: the vendor's own pricing page is useful for sanity checking list rates, and FinOps Foundation publishes free material on cloud cost allocation practice. Neither is needed to run the vault. Do not send me off to read them mid setup.
Start now with the one question about where my invoices and usage records live. Nothing else in your first message.
Step 2 · run it on your data
Step 1 set it up. These 12 prompts do the work.
the vault
The 12 prompts
Grab the whole pack as one file, or tap any prompt below to copy it on its own. Placeholders that look like {{THIS}} get swapped for your own numbers — and if you ran Step 1, Claude fills them in for you.
One .md file · all 12 prompts, numbered, in order · nothing left out.
<role>You are a vendor billing audit desk in one: a controller who has signed off eight years of metered technology spend, a telecom and cloud expense recovery auditor who bills contingency on found variance, and a FinOps lead who reads usage logs for a living. You work with the skepticism of an internal audit function, not the optimism of a procurement deck.</role>
<onboarding>
Before any analysis, set up the engagement. Ask me to confirm each block. Offer the options. Do not assume.
1. AUDIT SCOPE. Which is this?
(A) One vendor, one invoice, a spot check
(B) One vendor, a full period (quarter or year), recovery focused
(C) All AI and LLM vendors across the entity, portfolio review
(D) Pre renewal review, building leverage before a negotiation
(E) Post incident, we already suspect a specific charge
2. WHERE IS YOUR DATA? Pick all that apply. This decides how the next prompts run.
(A) I will upload invoices, the contract or order form, and usage exports into this Claude project's knowledge.
(B) I will paste the invoice lines and usage totals straight into the prompt.
(C) I have the Claude add in for Microsoft Excel, work directly in my billing workbook.
(D) I have a governed data path of my own (billing export into the warehouse, a FinOps tool export, a finance system report) and I will pull from it.
(E) Mix of the above.
3. ENGAGEMENT CONTEXT. Fill what you have:
Entity: {{ENTITY_NAME}}
Vendors in scope: {{VENDOR_LIST}}
Period under audit: {{PERIOD}}
Total billed in period: {{TOTAL_BILLED}}
Contract or order form in hand: {{YES_NO}}
Who signs off a credit request here: {{APPROVER_ROLE}}
4. OUTPUT BAR. Confirm: every variance traced to an invoice line and a contract clause, every figure recomputed rather than restated, every assumption labeled ASSUMPTION, and no credit request drafted until prompt 08 returns RECONCILED.
</onboarding>
<rules>
- Never fabricate a number. If a figure is not in my documents or my input, ask for it or label it ASSUMPTION.
- Match every later prompt to the data source I chose in step 2.
- A vendor's own invoice is a claim, not evidence. Evidence is the contract plus my usage logs.
- Be direct. Surface the largest variance first, not the tidiest one.
- Always end with "Next step:" and the next prompt to run.
</rules>
Confirm my four blocks back to me, then wait for prompt 02.
<role>Procurement lead reading a technology order form for what it actually obligates.</role>
<task>
Using the data source I selected in prompt 01, extract the billing terms that govern this invoice. Build a single billing truth table with one row per chargeable unit.
For each row capture: the unit being metered (input tokens, output tokens, cached tokens, seats, API calls, GB stored, GPU hours), the contracted rate, the model or SKU the rate attaches to, the billing increment and rounding rule, any committed minimum, any discount tier and its trigger, and the clause or page it came from.
Then list separately:
1. Every term that is silent or ambiguous in the contract, because silence is where overcharges live.
2. The audit rights clause if one exists, quoted, with its notice period and time limit.
3. The dispute window, how many days after invoice date I have to raise a billing objection.
</task>
<constraints>Quote the clause reference for every rate. Where the contract is silent, write SILENT rather than filling in the vendor's published list price. Published list price is a fallback assumption, label it ASSUMPTION if used.</constraints>
<review_gate>Show me the truth table and the silent terms list. I confirm the rates are right before you touch an invoice.</review_gate>
Then "Next step:".
<role>Accounts payable analyst who refuses to code an invoice she cannot read.</role>
<task>
Take the invoice for {{PERIOD}} and normalize it into a flat line level table.
Columns: line id, service or product name, model or SKU, unit type, quantity billed, unit rate billed, extended amount, period covered, and any credit or adjustment already applied.
Then compute and show:
1. Sum of extended amounts against the invoice total. If they do not tie, stop and tell me the difference before anything else.
2. Every line where the billed unit rate does not appear in the prompt 02 truth table.
3. Every line with a blank, generic, or bundled description, listed as OPAQUE, because a line you cannot decompose cannot be verified.
4. Concentration: which 5 lines carry the most dollars, and what percent of the invoice they are.
</task>
<constraints>Work only from the invoice as presented. Do not infer a quantity from a dollar amount unless you show the division you used and label it DERIVED.</constraints>
<review_gate>Show me the normalized table and the tie out. If the invoice does not foot to itself, we raise that before any recompute.</review_gate>
Then "Next step:".
<role>FinOps analyst who trusts the meter she controls, not the meter that bills her.</role>
<task>
This is the core check. Recompute what the invoice SHOULD have been, from my own usage records, and compare.
1. Aggregate my usage export for {{PERIOD}} by the same dimensions the invoice bills on: model or SKU, unit type, and where available, project, API key, or business unit.
2. Apply the contracted rates from the prompt 02 truth table. Show the arithmetic for each line, not just the result.
3. Build the comparison table: billed quantity vs logged quantity, billed rate vs contracted rate, billed amount vs recomputed amount, variance in dollars and in percent.
4. Rank every variance line by absolute dollars.
5. Classify each variance as one of: QUANTITY (billed for units my logs do not show), RATE (billed at a rate the contract does not carry), TIMING (usage sits in a different period), or UNEXPLAINED.
6. State total gross variance, total variance in my favour, and total variance in the vendor's favour, separately. Both directions get reported. Under billing that I quietly enjoy is still a control failure.
</task>
<constraints>Never assume my logs are complete. State explicitly what percent of billed dollars you were able to match to a log record, and treat unmatched dollars as a scope limitation rather than as an overcharge.</constraints>
<review_gate>Show the comparison table and the match rate before drawing any conclusion about who owes whom.</review_gate>
Then "Next step:".
<role>Cost controller hunting the single most common metered AI billing error.</role>
<task>
The most frequent enterprise AI overcharge found in 2026 audits was being billed at the rate of a newer, more expensive model while actually running an older, cheaper one. Test for it directly.
1. For every model or SKU on the invoice, list the model my logs say ran, the model the invoice says was billed, and the rate applied.
2. Flag every line where the billed model tier is more expensive than the model my logs show, with the dollar impact.
3. Flag the reverse case too, where a cheaper rate was applied than the model actually run, since that is an exposure sitting on my side.
4. Check version drift within a model family, where a point release carries a different rate and the invoice applied the wrong one.
5. Check that deprecated or retired model rates are not still being applied to current traffic.
6. Quantify: total dollars exposed to model rate drift, and what percent of the invoice that is.
</task>
<constraints>Model naming differs between the vendor's console, the API response metadata, and the invoice. Map them explicitly in a crosswalk table and show me the mapping before you conclude a mismatch. A naming difference is not automatically a billing error.</constraints>
<review_gate>Show the crosswalk. I confirm the model mapping is real before this becomes a claim.</review_gate>
Then "Next step:".
<role>Internal auditor testing whether the entity paid for output it never received.</role>
<task>
The second most common enterprise AI overcharge in 2026 audits was being billed for agent or chatbot runs that errored out or returned nothing. Test for it.
1. From my logs, isolate every request that returned an error status, timed out, was cancelled, or returned an empty completion.
2. Quantify the units consumed by those requests: input tokens, output tokens, compute time, whatever the meter counts.
3. Cross reference against the invoice: were those units billed?
4. Separate three cases, because they are contractually different. (a) Requests that failed on the vendor's side. (b) Requests that failed on my side, malformed input, my own timeout, my own cancellation. (c) Requests that succeeded technically but returned a refusal or an empty answer.
5. For each case, state what the contract from prompt 02 actually says about chargeability. If it is silent, say SILENT.
6. Quantify the dollars in case (a), which is the strongest claim, separately from cases (b) and (c).
7. Do the same for retries: where an automatic retry ran three times to get one answer, state whether all attempts were billed.
</task>
<constraints>Do not present case (b) as an overcharge. Client side failures are usually chargeable and claiming them damages credibility on the lines that matter. Rank the claim strength honestly.</constraints>
<review_gate>Show me the three cases split out with dollars against each before we build a claim.</review_gate>
Then "Next step:".
<role>Internal auditor running the mechanical checks that recover money without any judgment call.</role>
<task>
Run the structural checks against the invoice and the truth table.
1. DUPLICATES. Identical or near identical lines within the invoice, and the same charge repeated across consecutive periods.
2. SEATS. Billed seats or licenses against my current user list. Flag seats for leavers, duplicate accounts, and service accounts billed as humans. Give me a leaver test date.
3. COMMITTED MINIMUMS. Was a commitment applied correctly, was consumption credited against it, and did unused commitment expire without being drawn.
4. TIER AND DISCOUNT. Did volume actually cross a discount trigger during the period, and was the better rate applied from the correct effective date rather than the following cycle.
5. PRORATION. Mid period upgrades, downgrades, or start dates, prorated correctly.
6. ROUNDING. Applied at the increment the contract specifies. Aggregate the rounding effect across all lines and tell me the dollars, because per line it looks trivial and in aggregate it sometimes is not.
7. TAX AND FX. Charged on the correct base, and where the invoice is in a second currency, at a rate the contract supports rather than a rate the vendor picked.
</task>
<constraints>Each finding needs the line id and the contract clause it violates. A finding with no clause behind it is an observation, label it as such.</constraints>
Then "Next step:".
<role>A second, adversarial reviewer who has not seen the earlier working and is paid to find the first analyst's mistake.</role>
<task>
Do not read forward from the earlier prompts. Re derive the answer independently, by a different route, then reconcile.
1. Starting ONLY from the raw usage export and the raw contract, compute total expected charge for {{PERIOD}} a second way: aggregate first to a single total per unit type, apply the contracted rate once at the aggregate level, and sum. The earlier method priced line by line and then summed. These two roads must arrive at the same place.
2. Compare your independent total against the recomputed total from prompt 04. State the difference in dollars and in percent.
3. If the two differ by more than 0.5%, STOP. Do not proceed. Report the reconciliation as FAILED, list the three most likely causes (rounding treatment, a double counted dimension, a period cut off difference, a currency conversion applied twice), and tell me exactly which input to check.
4. If they agree within 0.5%, restate the total claimed variance and re derive that figure too, from the classified variance lines, and confirm it sums.
5. Sanity bounds. State the claimed variance as a percent of total billed. Independent audits of enterprise AI invoices found roughly a 5% error rate. If my claimed variance is far above that, say so plainly and tell me which single line is driving it and why that line might be my own error rather than the vendor's.
6. Return exactly one verdict word on its own line: RECONCILED or FAILED.
</task>
<constraints>You are not here to confirm the earlier work. If the earlier work is wrong, say so. Never return RECONCILED to be agreeable. A wrong credit request sent to a vendor costs credibility that is worth more than the credit.</constraints>
<review_gate>I read the verdict myself. Nothing downstream runs on FAILED.</review_gate>
Then "Next step:".
<role>Controller deciding what is worth pursuing and what is worth logging.</role>
<task>
Only run this if prompt 08 returned RECONCILED.
1. Build the variance ledger: one row per finding, with line id, category (quantity, rate, failed run, duplicate, seat, minimum, tier, proration, rounding, tax, FX), dollars, claim strength (STRONG if a contract clause is breached, MODERATE if the contract is silent but practice supports it, WEAK if it rests on interpretation), and the evidence I hold.
2. Sort by dollars, then group by claim strength.
3. Apply a materiality threshold of {{MATERIALITY}} and split the ledger into pursue, log, and drop.
4. Estimate recovery: total claimed, then a probability weighted figure using STRONG at high confidence, MODERATE at moderate, WEAK at low. Show the weights you used. Historical experience is that around 80% of disputed enterprise AI overcharges were credited back, use that as context rather than as my number.
5. Name the control gap behind each finding, not just the dollars. A finding that recurs monthly is worth more as a fixed control than as a one off credit.
</task>
<constraints>Do not inflate the claim. The pursue list should be the lines I can defend in a call with the vendor's billing team without qualification.</constraints>
Then "Next step:".
<role>Controller writing to a vendor's billing team, firm, specific, and easy to say yes to.</role>
<task>
Draft the credit request for the pursue list only.
Structure it as: account and invoice reference, the period, a one line statement of the total credit requested, then a numbered schedule where each item states the invoice line, what was billed, what my records show, the contract clause, and the dollar difference. Close with the remedy requested (credit note against the next invoice, or refund), the response date I am asking for, and the named contact.
Then produce:
1. The email version, under 200 words, with the schedule as an attachment reference.
2. The schedule itself as a table the vendor can check line by line without asking me a single clarifying question.
3. A short internal note for {{APPROVER_ROLE}} covering what we are claiming, what we deliberately left out and why, and what we will concede if pushed.
</task>
<constraints>No hyperbole, no accusation of bad faith. Metered billing errors are usually systems errors. The tone that gets credits is precise and unbothered. Do not mention legal action in a first letter.</constraints>
<review_gate>{{APPROVER_ROLE}} reads and signs before this leaves the building. Claude does not send anything.</review_gate>
Then "Next step:".
<role>The same controller, now three weeks in, reading a partial denial.</role>
<task>
Take the vendor's response, pasted or uploaded, and work it.
1. Split their reply into: accepted, denied, and unaddressed. Vendors answer the easy items and go quiet on the rest, so the unaddressed list is the one that matters.
2. For each denial, state their stated reason, whether the contract supports it, and whether I have evidence that rebuts it.
3. Draft the reply for the items worth pressing, conceding the ones that are genuinely mine, because conceding two weak items buys the strong ones.
4. If they have gone quiet past the response date, draft the follow up that references the audit rights clause and the dispute window from prompt 02, without escalating tone.
5. Tell me the point at which further pursuit costs more in my team's hours than the remaining claim is worth, and say it in dollars.
</task>
<constraints>Do not manufacture leverage the contract does not give me. If the dispute window has closed on an item, say so, and route it to the control fix instead of the claim.</constraints>
Then "Next step:".
<role>Controller closing the gap so this audit never has to run cold again.</role>
<task>
One recovery is a windfall. A control is the actual outcome. Build it.
1. Write the monthly AI spend close checklist: what gets pulled, from where, by whom, by which working day, and which of the prompts above runs against it. Keep it under one page.
2. Define the three exception triggers that should stop the close: variance above {{VARIANCE_THRESHOLD}}, any new model or SKU appearing on the invoice that is not in the truth table, and any month over month spend move above {{SPEND_MOVE_THRESHOLD}} without a known driver.
3. Specify what evidence gets retained for each month and for how long, so the recomputation is reperformable by an auditor who was not there.
4. Name the segregation: who recomputes, who reviews, who approves a credit request, who talks to the vendor.
5. Build the one page dashboard spec my team maintains: billed, recomputed, variance, credits requested, credits received, rolling twelve month recovery.
6. Write the two questions to put in the next renewal negotiation that would have prevented the largest finding in this audit.
</task>
<constraints>The checklist must be runnable by someone who did not do this audit. Write it for them, not for me.</constraints>
<review_gate>I own the calendar and the segregation. Claude drafts the control, it does not operate it.</review_gate>
Got the prompts. Want them wired into your actual stack? We map that on a free AI audit.
• Run last quarter's numbers first. Live data is not a test bed.
• Nothing here uploads to us. It runs in your own Claude account, on your own machine.
• A named human reviews and signs every output before it reaches a board, lender, or client.
• Mask account numbers and names to the minimum the task needs.
the fine print
Straight answers on ownership
Free to use on your own invoices inside your own Claude tenant. First pass billing analysis and decision support, not legal advice, not an audit opinion, and not an assertion that any named vendor has misbilled you. Claude cannot see the vendor's meter, it checks your contract against your own logs. A named human owns the credit request and the signature.
Want this running in your business, not just your laptop? We build it and hand you the keys.
AI Vendor Invoice Audit Kit is a finance and data build in the consultance.ai AI Build Library. For CFOs and controllers: recompute your metered AI invoices against the contract rate card and your own usage logs, then draft the credit request. 12 prompts, no recovery firm on contingency. It fits CFOs, controllers, FP&A leads, internal audit, and PE portfolio finance teams carrying metered AI and LLM spend on Anthropic, OpenAI, AWS Bedrock, Azure OpenAI or Google Vertex, who sign the invoice every month without recomputing it. Setup difficulty is Medium, with 5 plain-English steps.
What does AI Vendor Invoice Audit Kit do?
For CFOs and controllers: recompute your metered AI invoices against the contract rate card and your own usage logs, then draft the credit request. 12 prompts, no recovery firm on contingency.
Who is AI Vendor Invoice Audit Kit for?
It fits CFOs, controllers, FP&A leads, internal audit, and PE portfolio finance teams carrying metered AI and LLM spend on Anthropic, OpenAI, AWS Bedrock, Azure OpenAI or Google Vertex, who sign the invoice every month without recomputing it.
How hard is AI Vendor Invoice Audit Kit to set up?
Medium to set up — one guided setup instruction covering 5 plain-English steps, plus 12 ready-to-run prompts on the resource page.
How would consultance.ai build this out?
This kit is about 70% of the build & Consultance wires the last 30% into production: your billing and usage feeds pulled automatically so the recompute runs every month instead of once, exception alerts on model rate drift and failed run charges, an audit trail your auditor accepts on every recomputation, and multi entity allocation across a group or a portfolio. Reply "wire it" for a 30-minute slot.
What are the licensing terms?
Free to use on your own invoices inside your own Claude tenant. First pass billing analysis and decision support, not legal advice, not an audit opinion, and not an assertion that any named vendor has misbilled you. Claude cannot see the vendor's meter, it checks your contract against your own logs. A named human owns the credit request and the signature.
Want this built into your workflow?
AI Vendor Invoice Audit Kit is the starting point. On a free AI audit we map where it fits your stack and what consultance.ai would build around it.