Methodology v1
Score what AI can actually operate.
Works in Chat separates documented availability from independently reproduced capability. Vendors can buy distribution and analytics. They cannot buy a better score.
Read / query10%Can the AI reliably retrieve live data and context?
Create15%Can it create meaningful records, content or objects?
Update15%Can it modify existing state accurately?
Execute20%Can it trigger consequential workflows and actions?
End-to-end25%Can it complete a real job rather than isolated commands?
UI independence15%How much can be done without reopening the original SaaS UI?
Score states
Provisional — derived from official docs, vendor materials and integration metadata.
Verified — at least one representative workflow has reproducible evidence and independent confirmation.
Stale — evidence is old, the integration changed, or recent tests conflict.
Commercial independence
Sponsored launches are visibly labeled. Paid plans can unlock analytics, benchmarking, verification operations and launch distribution.
Organic score and rank are never sold. Evidence changes scores; payments do not.