Analytics and marketing cookies

    Optional analytics and advertising cookies help us improve your experience and measure marketing performance. Learn more

    UserApproved.ai
    Benchmark v1 · August 2026

    Fewer findings, more that matter

    UserApproved surfaced 5 findings, 3 rated critical by independent CRO experts. Claude and GPT surfaced 12 combined, with just 1 rated critical.

    • 3 independent CRO reviewers
    • All 3 results rated blind
    • One store, one run per system
    Critical-finding densityPer system

    How much of the output actually mattered?

    60% of UserApproved findings were rated critical, versus 14% for Claude and 0% for GPT.

    UserApproved

    60%

    3 of 5 findings rated critical

    GPT Sol 5.6

    0%

    0 of 5 findings rated critical

    Claude Opus 5

    14%

    1 of 7 findings rated critical

    Rated criticalRated as not changing a decision

    Overall quality score

    UserApproved scored higher than both leading general agents on all four dimensions, with the widest gaps on tracking real evidence across sources and on defining actionable experiments.

    Rated by three external CRO reviewers, blind to which system produced each result.

    What the scores mean

    Four ways an AI report wastes your team's time

    Can I trust it, or does someone have to go check?

    When a claim isn't tied to a screenshot, a session recording, or a number you can pull yourself, somebody on your team has to go verify it before anyone will act on it. And it only takes one invented claim to make a merchant distrust the entire report. Every finding we publish carries the artifact it came from, and the arithmetic is printed so you can recompute it.

    Evidence grounded

    UserApproved3.20
    GPT2.57
    Claude2.23

    Did it do the thinking, or hand me raw observations?

    A report can tell you mobile bounce is high, and separately that shipping cost appears late in checkout, and leave you to connect them. That connection is the actual analysis. Most systems hand it back to you and call it a finding.

    Strategic synthesis

    UserApproved3.50
    GPT2.93
    Claude2.85

    Is this worth a meeting and a sprint slot?

    Every finding costs you attention: someone reads it, someone argues about it, someone decides whether it goes in the backlog. Most AI findings are not worth that, and you only find out after you've spent the meeting.

    Impact and usefulness

    UserApproved3.20
    GPT2.63
    Claude2.35

    Do I know what to do on Monday?

    "Improve your mobile experience" is the problem restated as advice. A finding should name what to change, where, and what to watch afterward, so the next step is doing the work rather than another week deciding what the work is.

    Action specificity

    UserApproved3.70
    GPT2.76
    Claude3.00
    Method

    How this was run

    How the benchmark worked

    The store

    A live US ecommerce storefront, not a demo site. Audited on desktop and mobile. Named to prospects under NDA.

    Same brief

    Act as an ecommerce growth expert and identify actionable revenue opportunities.

    Same available access

    The same store data, APIs, MCP integrations, and browser tools were available.

    External review

    Three external CRO professionals rated the results from all three systems.

    Blinded scoring

    Reviewers did not know which system produced each result.

    How a finding is scored

    Every finding is rated 1 to 4 on each of the four dimensions, for a maximum of 16 points, then normalized to 100. Every score on this page is the mean across the three reviewers, from the August 4, 2026 run.

    4
    One minor flaw a reader would not act on differently
    3
    Passes, with a real and citable flaw
    2
    Several real flaws, or one that materially weakens the finding
    1
    Barely passes; a reader would distrust the pillar
    Coverage, runtime, and cost
    Coverage, elapsed time, and quality-per-dollar comparison
    SystemCoverageElapsedQuality / $
    UserApproved AgentsDesktop + mobile · 4 persona paths46:523.3
    GPT Sol 5.6, HighDesktop · 1 primary path40:003.7
    Claude Opus 5, X-highDesktop · 1 primary path39:002.0

    GPT had the highest reported quality-per-dollar index. UserApproved had the highest quality score and the longest elapsed time while choosing broader coverage. Provider accounting can differ, so this is a run-specific index.

    The supported conclusion

    In this pilot, external blinded reviewers rated UserApproved higher on all four quality dimensions and rated a higher concentration of its findings critical. This is directional evidence for the system design, not proof of universal superiority.

    The model is not the product. The product is the system that decides what deserves to ship.

    Run the same investigation on your store

    Start with a free baseline scan, then connect your commerce context to see which findings deserve action.

    Free baseline audit · No credit card · Evidence stays attached to every finding