First Try methodology

A score you can inspect.

First Try measures whether leading AI systems can accurately understand a product from its public experience. The benchmark is designed to make ambiguity visible—not to manufacture a perfect answer.

What we test

We evaluate a preserved snapshot of the submitted public product website. The benchmark looks for consistent understanding of the product, intended audience, core value, and next action.

  • Product identity and primary job
  • Target customer and use case
  • Differentiation and value
  • Pricing and conversion path, when public

Three independent AI families

The same preserved material and evaluation schema are submitted independently to GPT-5.6 Sol, Claude Opus 5, and Grok 4.6 through OpenRouter. Each model produces its own interpretation and evidence without seeing the other answers. The public profile shows the breakdown rather than hiding disagreement behind one number.

Machine Clarity Score

Each model scores product identity, audience, value, and action from 0–100. Its model score is the mean of those four dimensions. The final score weights the three-model mean at 85% and cross-model agreement at 15%. Evidence quotes are reviewed against the preserved material before publication; reviewers do not rewrite model answers or improve a score.

The result measures clarity and consistency in this benchmark. It is not a certification or a promise of commercial performance.

Ranking contract

You cannot buy a better position. Rankings use the latest valid result; fees, traffic, and company size do not affect rank.

Snapshots and evidence

Every published result identifies the tested date, benchmark version, exact model IDs, reasoning levels, and preserved website snapshot. Historical results remain attached to the profile when a company re-tests.

Re-tests

A re-test is a new independent benchmark after a meaningful benchmark-relevant change. If the submitted experience appears unchanged, Mlola may decline or defer the run. The latest valid benchmark—not the highest-ever score—determines current rank.