Skip to content
← Back to blog
Research·August 26, 2026·8 min read

Foundation model transparency scores fell from 58 to 41 in a year. Here's what that means for the model your product runs on.

Stanford's Foundation Model Transparency Index fell from 58 to 41 in 2025, and the five biggest labs all clustered in the same untested middle.

The organization that grades AI companies on how much they actually disclose about their own models just posted its worst average score since it started grading. Stanford's Foundation Model Transparency Index (FMTI) fell from 58 out of 100 in the 2024 edition to 40.69 in the December 2025 edition, undoing a year of real improvement and landing the industry almost exactly back where it started in 2023. The five labs a funded product company is most likely building its roadmap on, Amazon, Anthropic, Google, Meta and OpenAI, all scored within fifteen points of each other, clustered in the same untested middle, not because they compete on transparency but because none of them currently has to.

What the transparency index actually scores

The FMTI is now in its third year, run by researchers at Stanford's Center for Research on Foundation Models along with MIT and Princeton. It grades a company's flagship model against 100 public indicators split across three domains: upstream (how the model was built, meaning training data, compute, and the labor behind it), model (the model's own properties, evaluations and release process), and downstream (what happens after the model ships, meaning usage monitoring, incident handling, and how customer data gets treated). A company earns a point only when the information is actually public, not when the team believes the company probably handles things responsibly behind closed doors. The December 2025 edition scored 13 companies, including Alibaba and DeepSeek for the first time, the widest the index has ever been even as the average score fell.

The number that reversed itself

Context matters, because the 2024 jump wasn't really about companies changing their behavior. It came from the index switching to a proactive-submission format that let companies newly disclose things they had been quietly doing all along, like the exact FLOP count and hardware used to train a flagship model. Fourteen companies submitted reports in 2024 and the average leapt from 37 to 58. In 2025, only seven did: AI21 Labs, Amazon, Google, IBM, Meta, OpenAI and Writer. The other six companies scored, Alibaba, Anthropic, DeepSeek, Midjourney, Mistral and xAI, were assessed entirely from public information the FMTI team could find on its own, because those companies declined to submit anything. Several more of the 23 companies the team approached, including Apple, Microsoft, Nvidia, Cohere and Baidu, declined to participate or never responded. Some of the drop is a stricter ruler: the team says it raised the bar on several indicators this year. Most of it isn't. Meta's score was cut roughly in half year over year, Mistral's by more than two-thirds, and OpenAI, one of the companies that did submit a report, still dropped 14 points.

Frontier labs didn't compete to be transparent. They agreed, without agreeing, not to finish last

Five of the scored companies belong to the Frontier Model Forum, the industry group Amazon, Anthropic, Google, Meta and OpenAI formed to coordinate on frontier AI safety. All five landed in the middle of the index: Anthropic scored 46, Google 41, Amazon 39, OpenAI 35, and Meta 31, the lowest of the group and the only one whose disclosure pattern doesn't closely track the other four. The FMTI researchers found Anthropic and OpenAI's practices strikingly similar (matching on 85 of the 100 indicators) even though Anthropic scored 11 points higher: Anthropic discloses a near-superset of what OpenAI does, with two narrow exceptions. What's notable is that Anthropic didn't submit a report this year. It was scored from public information alone and still beat the company that did the paperwork. Eight of the thirteen companies, including OpenAI, Anthropic, Google and Amazon, have signed the EU's voluntary AI Act Code of Practice, and signatories score only marginally higher than non-signatories, almost entirely on the downstream domain (things like published usage policies). The domain where a buyer's actual risk sits, training data and compute, doesn't move for having signed anything.

Where the opacity actually concentrates

Break the score down by domain and the pattern gets specific. Upstream, meaning training data and compute, is both the least transparent domain overall (a mean of just 9.2 of 34 possible points) and one of the most unevenly scored. Two companies, Midjourney and Mistral, disclose nothing at all in that domain. Three others, OpenAI, xAI and Anthropic, score at most 3 of the available 34 points, which puts the two highest-scoring frontier labs in the same neighborhood as the industry's weakest performers on exactly the question a buyer most needs answered: what did this model actually learn from, and could that create legal exposure that lands on the client using it? That's not hypothetical. Anthropic settled the Bartz v. Anthropic class action for using roughly 500,000 pirated books to train its models, a settlement that received final court approval on July 20, 2026, according to the Authors Guild. No FMTI indicator would have flagged that risk in advance, because no company discloses enough about its training data for an outside evaluator to check.

A worked example

Picture a 40-person Series B company whose core product runs a single reasoning step through one vendor's API, wired in eighteen months ago when the model was the obvious best choice and nobody was thinking about a fallback. The VP of Engineering is now negotiating a 12-month enterprise contract at a steep discount for committing to volume. The vendor's FMTI score dropped double digits this year and nobody on the product team noticed, because nothing about the product broke and no announcement mentioned it. The discount for locking in a year is real money. So is the fact that a transparency score falling this fast, on a vendor whose training-data practices were already close to unscored, is exactly the kind of signal that shows up months before a deprecation notice, a licensing dispute, or a quiet model swap that changes behavior underneath a product nobody re-tested. Signing for a year doesn't remove that risk. It just removes the option to react to it.

What to check before a vendor relationship becomes the roadmap

  • Build an eval harness against your own examples, so a vendor's transparency score, or lack of one, isn't the only signal you have when something changes.
  • Read a vendor's FMTI history, not just this year's score. A five-point drop is noise. A twenty-point drop, like Meta's or Mistral's, is the vendor telling you something changed even if it never says what.
  • Ask specifically about training-data litigation exposure and indemnification terms before signing a volume contract, not after a suit is filed.
  • Keep the model layer swappable. Pinning a stack to a model ID nobody revisits is the same mistake whether the trigger is a deprecation notice or a transparency score that quietly collapsed, and it's the same discipline behind why agents that look fine in a demo keep failing in production.
  • Here's the honest exception: if your product doesn't touch regulated or sensitive data and switching vendors costs an afternoon, not a quarter, this index is a nice-to-know, not a blocker. Building an abstraction layer around a single well-behaved API call is over-engineering for a team that could just switch if it had to.

None of this means the frontier labs are hiding something disqualifying. Most of what they're opaque about, they're opaque about together, which is its own kind of information. It means the transparency score attached to whatever model runs your product's core loop is worth checking before you sign a longer contract than your fallback plan can undo. If you want a second read on whether your current model dependency is a real risk or just an annoyance, our engineering team can walk through it, and our Silicon Valley team is a short call away. Tell us what you're running and we'll give you a straight answer.

Sources

Frequently asked questions.

The Foundation Model Transparency Index (FMTI) is an annual scorecard, now in its third year, run by Stanford's Center for Research on Foundation Models with researchers from MIT and Princeton. It grades AI companies against 100 public indicators covering training data and compute, model properties and evaluation, and post-deployment practices. The December 2025 edition scored 13 companies.

The average FMTI score fell from 58 out of 100 in the 2024 edition to 40.69 in the December 2025 edition, reversing a year of gains and landing close to the index's 2023 starting average of 37. Meta's score was roughly cut in half over that year and Mistral's fell by more than two-thirds, while OpenAI, one of only seven companies that submitted a self-reported disclosure, still dropped 14 points.

In the December 2025 FMTI, Anthropic scored 46, Google 41, Amazon 39, OpenAI 35 and Meta 31, with all five Frontier Model Forum members clustered within 15 points of each other in the middle of the index. IBM led the field at 95 and Writer scored 72, while Midjourney and xAI tied for the lowest score at 14.

Training data and compute make up the FMTI's upstream domain, and it's the least transparent of the three domains the index measures, averaging just 9.2 of 34 possible points in the December 2025 edition. Two companies disclosed nothing in that domain at all, and three others, including Anthropic and OpenAI, scored 3 points or fewer, meaning no FMTI indicator can tell a buyer in advance what a model was actually trained on.

Not by itself. A single-point drop is noise, but a double-digit swing, like the ones Meta and Mistral posted in the index's December 2025 edition, signals a real change worth investigating. The more durable fix is keeping the model layer swappable and testing against your own eval set, so a vendor's disclosure, or its absence, isn't the only warning a team gets before something changes underneath a shipped product.