The AI Accuracy File: What Founders Need Before They Claim Their Product Is Reliable

The FTC's proposed policy statement turns AI accuracy claims into legal exposure. What founders should assemble before marketing an AI product as accurate or reliable.

The FTC has proposed treating misleading AI outputs as a Section 5 problem, and the comment window closes July 31, 2026. For founders, the exposure rarely starts with the model. It starts with the claims made about the model, in decks, demos, onboarding, and contracts, and with whether the company can prove any of it. This article makes the case for an internal AI Accuracy File and walks through what goes in it.


The fastest way for an AI startup to create legal risk is not building an imperfect model. It is selling the model as more accurate, objective, or reliable than the company can prove.

That distinction used to be a marketing hygiene point. As of this month, it is being framed more explicitly as an FTC enforcement theory. In July 2026 the FTC released a proposed policy statement titled Suppression of Accuracy in Artificial Intelligence Systems, with public comments due July 31. Its focus is narrower than a general accuracy rule: undisclosed manipulation or distortion of AI outputs, and systems that behave contrary to what users were reasonably led to expect, framed as potentially unfair or deceptive conduct under Section 5 of the FTC Act, especially where the company's own representations have created expectations about the system's objectivity, accuracy, suitability, or reliability. The agency's AI enforcement docket tells the same story from the other direction, a steady line of actions against companies whose AI marketing, earnings claims, or capability descriptions outran the product.

The uncomfortable part for founders is that the hard question is no longer "did we exaggerate?" It is "can we substantiate?" What is the AI supposed to do, what does it actually do, and does its behavior change after deployment? If the company cannot answer those three questions with evidence it already holds, every confident sentence in the sales deck is a small unsecured liability.

Why isn't accuracy just a model score?

Part of the problem is that "accuracy" does not mean one thing. Inside the product team it might mean a benchmark score on a held-out test set. To a customer it might mean factual correctness of outputs, completion of a task end to end, stable behavior on edge cases, freedom from manipulation, or fitness for one specific business function in one specific industry. A model can score well on the first definition and fail four of the others in the same week.

That ambiguity is exactly where claims risk lives. The product team measures accuracy one way. Marketing describes it another way. Sales implies something broader in a demo. The customer forms an expectation shaped by their own workflow, which nobody at the company has tested. When the FTC talks about representations that create expectations of objectivity, accuracy, suitability, or reliability, it is describing that gap, the space between what was measured and what was reasonably understood.

What did the FTC proposal put on the table?

The proposed policy statement addresses AI systems whose outputs are manipulated, distorted, or otherwise misleading, with particular attention to undisclosed manipulation and distortion, and situates them within the FTC's longstanding Section 5 authority over unfair or deceptive acts and practices. The center of gravity is expectations: where a company has represented, expressly or by implication, that its system is objective, accurate, suitable, or reliable, outputs that behave contrary to those reasonable expectations become an enforcement theory rather than a quality complaint. The important point for founders is that the FTC is not using accuracy only as an engineering metric. It is looking at whether the AI system behaves consistently with what the company led users to expect.

Three things are worth internalizing about the posture. First, it is a proposal, and the July 31, 2026 comment deadline means the final contours may shift; companies with a direct stake have a short window to be heard. Second, it does not create a new statute. Section 5 already reaches deceptive product claims, and the agency has been applying it to AI marketing for several years. The statement consolidates and signals, which in practice is how enforcement priorities get set. Third, a company does not eliminate deception risk merely because the founder sincerely believed the claim. The legal question is whether the claim was truthful, not misleading, and adequately substantiated when made.

Where do founders create accuracy risk without noticing?

Almost never in the obvious place. The website copy usually gets at least one careful read. The risk accumulates in the surfaces around it: sales decks that quote a benchmark from an old model version, demos rehearsed on inputs the model handles well, onboarding flows that quietly suggest outputs can go straight into production, investor decks with capability language nobody reran past the product team, pilots scoped verbally, procurement questionnaires answered under deadline pressure by whoever was free, model cards written once and never updated, and vendor contracts that warrant performance the engineering team was never asked about.

Each surface makes a claim. Very few companies keep an inventory of what all of them are currently saying, which means very few companies could answer a regulator, an enterprise buyer, or a plaintiff who lines three of them up and asks which one is true. Post-deployment drift compounds the problem: a claim that was substantiated at launch can quietly become unsubstantiated two model updates later, while the deck that contains it keeps circulating.

What goes in the AI Accuracy File?

The practical fix is not a policy binder. It is a maintained internal evidence file that does for accuracy claims what a cap table does for equity: one place where the current truth lives. For companies building an AI governance program, this file should become the evidence layer behind the policy language. A working version has eight or nine components.

  • Product claims inventory. Every performance representation currently in circulation, across marketing, sales materials, onboarding, documentation, and contracts, with an owner for each.
  • Benchmark and evaluation summaries. What was tested, on what data, under what conditions, on which model version, and when it was last rerun.
  • Known limitations. Written down, dated, and phrased the way you would defend them, not the way marketing would soften them.
  • Excluded use cases. The workflows the product is not validated for, stated expressly, because silence reads as suitability.
  • Human review requirements. What the product assumes a human checks before an output is relied on, and where the customer is told that.
  • Monitoring logs. Evidence of how behavior holds up in production, because deployment-time drift is where launch-time substantiation goes stale.
  • Incident escalation rules. Who is told what, and how fast, when the system produces a harmful or materially wrong output.
  • Customer-facing disclaimers and contract language. The framing that travels with the claims, kept consistent with the limitations list rather than at war with it.

The test of the file is simple. Pick any external claim the company makes about the product's performance. If you can trace it to something in the file within a few minutes, the claim is an asset. If you cannot, it is exposure with a logo on it.

How do accuracy claims flow into contracts?

Everything in the file eventually touches paper. Performance warranties should describe what the evidence supports, in the language of the evidence, not the language of the homepage. Disclaimers and limitation of liability clauses should match the documented limitations, because a disclaimer that contradicts the sales deck invites an argument about which one the customer reasonably relied on. Service levels and acceptance testing should exercise the conditions the customer will actually run, since an acceptance test passed on curated inputs proves very little about the deployment it precedes. Indemnification and audit rights allocate who pays and who checks when behavior diverges from the claims. The AI vendor contract checklist in this cluster walks the clause-level detail, and the indemnification article covers the allocation fights specifically.

Founders tend to think of the contract as the place to disclaim their way out of accuracy exposure. It is better understood as the place where the company's claims and its evidence either line up or visibly fail to. A signed disclaimer helps. A signed disclaimer sitting next to a deck that promised the opposite helps the other side.

What will diligence ask?

Enterprise buyers and investors are converging on the same four questions the FTC's framework implies. What did you say the system does? What evidence supports that claim? What happens when it fails? Who is responsible when it does? Security reviews and procurement questionnaires already probe monitoring and evaluation practices, and diligence on AI companies increasingly prices the gap between marketed capability and documented capability, because whoever writes the check inherits that gap. A maintained accuracy file turns those questions into a same-day document production. Its absence turns them into escrows, special indemnities, and price adjustments. The AI diligence package article covers the buyer-side view of the same file.

What is the founder takeaway?

The risk is not that the product is imperfect. Every AI product is imperfect, and regulators, buyers, and courts all know it. The risk is making claims that outrun the company's evidence, and the new FTC posture converts that gap from a credibility problem into a legal one. The companies that will move fastest through the next two years of AI procurement and enforcement are the ones that can hand over the file. Building it now, while it is short, costs a fraction of reconstructing it later under a civil investigative demand or a diligence deadline.


This article is for informational purposes only and does not constitute legal advice. Every company's situation is different, and you should consult with qualified legal counsel before making compliance decisions based on the developments discussed here.

Consilium Law advises growth-stage companies on AI product claims, vendor contracting, and AI governance programs calibrated to U.S. and EU regulatory obligations. If your sales materials make performance claims your evaluation records cannot back, that is worth fixing before someone else finds it.

Contact

If this touches the work in front of you, start a conversation.

Send a short note about what changed, what you are building, and where legal judgment needs to sit closer to the work.

Disclaimer. This article is provided for informational purposes only and does not constitute legal advice. Readers should consult independent counsel before acting on any analysis. The views expressed are solely those of the author and do not necessarily reflect the views of Consilium Law LLC.