Changelog
Every grade that has ever changed, newest first. Corrections, rubric changes, product changes and scope changes are labeled separately, because “we were wrong” and “we changed how we measure” mean different things. We publish our own corrections first.
Scope. Amazon Q Developer is retired from the board. AWS closed it to new signups on 15 May 2026 and it reaches end of support on 30 April 2027, with Kiro named as the migration destination, and Kiro is already a row here, so the board was grading a product alongside its own replacement. Three rows are renamed to the product a buyer can find: Azure AI Foundry is Microsoft Foundry, AWS Bedrock AgentCore is Amazon Bedrock AgentCore (AWS brands it Amazon, never AWS), and Databricks Mosaic AI is Databricks Unity AI Gateway. The word “Mosaic” appears in none of the source URLs behind this board, and Databricks titles the docset those sources come from “AI governance with Unity AI Gateway”.
Corrections. Fourteen cells were grading against one band while their own text argued another. Six grades move: Agentforce Log Quality and Copilot Studio Containment to Yellow, Kiro RT Logging and Kiro Guardrails to Green, Atlassian Rovo Data Boundary to Green, and the Gemini app’s Model Routing from a security gap to Red. A per-user model picker is a weak lever, but it is not the absence of one. Five more keep their grade and get text that argues for it.
Four of those rulings were wrong, and are reversed here. Re-reading all eleven decisions against each cell’s own primary source, rather than against the sentence the cell had quoted from it, refutes four. Three fail the same way, and it is worth naming: the quote stopped one sentence before the answer. ServiceNow Data Boundary quoted “BYOK supports the same AI models that the ServiceNow AI Platform supports” and concluded that nothing said whose account the inference runs in; the next paragraph says requests go “through your cloud AI provider account instead of ServiceNow AI Platform infrastructure”. Glean Model Routing concluded no customer base URL was documented, from a page reading “Optional: Enter a custom endpoint in Base URL”. Notion AI Log Quality called cost attribution missing by any documented means, which is the Red criterion, from a page that says admins “can see exact usage in the Notion credits dashboard” and an API that makes per-agent credits a required field. All three return to their prior grades. v1.1 learned not to treat a page’s existence as evidence of its contents; this is the same error one level in, treating one quoted sentence as evidence of the page.
The fourth was published, so it is corrected in public. This entry previously said Kiro’s Tool Governance stayed Yellow “because the vendor’s own page says MCP governance fails open”. That phrase appears nowhere on kiro.dev. The page says the opposite, twice: MCP “fails closed if the client cannot reach the governance API”. We attributed to a vendor the inverse of what it wrote and made it the sole stated reason for refusing a grade change. The grade is unaffected and the cell’s own text was always right about why: enforcement is client-side, Kiro warns it can be circumvented by a user with administrative access to their machine, and governance reaches IAM Identity Center and API-key auth but not Builder ID or social login. Kiro Tool Governance stays Yellow for those reasons.
Corrections published after v1.2 shipped. Four absence scopes named a plan or a brand that no longer exists, which makes a scope unreadable: Replit “Teams” (sunset March 2026), Snowflake “Intelligence” (now CoWork), ServiceNow “Enterprise tier” (replaced by Prime in April 2026, and the row label already said Prime, so the board disagreed with itself about which SKU it graded), and Databricks still carrying “Mosaic AI” two renames on. Microsoft Agent 365 was named as a mechanism in five justifications and quoted in none, so the licence gates behind the Foundry Containment and Agent Identity greens now carry the sentence that states them. ChatGPT Enterprise Data Boundary moves to Green. It read Yellow because every lever sat in OpenAI’s plane, argued from a page that documents Enterprise Key Management for eligible Enterprise and Edu workspaces, while the Codex row was Green on that same customer-held key. One row cannot inherit a workspace’s settings and out-grade the workspace on them.
A grade that survived its challenge. Snowflake Cortex Agents was flagged as a probable wrong grade on Agent Identity: Snowflake now documents a
SERVICE_AGENT user type, which is this column’s whole subject. On reading the documentation it is an available user type rather than the principal Cortex Agents executes as, so the security gap stands, and the cell now says that outright instead of leaving a reader to wonder whether we had missed it.Scope. Agent Identity joins as a ninth control surface, graded on the mode each product runs in rather than on a capability its platform has somewhere. In 21 of 29 products the agent runs as the person who started it. Agent identity is removed from the Containment layer map at the same time. It was averaged into one grade while being cited as the differentiator in another, which moved that published figure to 5 layers. Two domains are then added back to the containment map, from receipts already on the cells: the cloud control plane (IAM or policy in a cloud account the organisation owns and administers: five products, led by AgentCore, where StopRuntimeSession is an IAM-permissioned call in the org’s own account. This entry said six until 2026-09-02; Amazon Q Developer was one of them and was retired from the board in this same release, so six was never what v1.2 shipped) and the data platform (a grant or role in the warehouse holding the data: Snowflake’s CORTEX_AGENT_USER, Databricks service principals). Both were being absorbed into the product admin plane, which asserts the cut runs through the vendor’s console, the opposite of what those receipts say. Rubric change, no grade moves; the figure derives to 7. Seven products are added: Atlassian Rovo, Slack AI, Notion AI, Perplexity Enterprise, ServiceNow AI Agents, Databricks Agent Bricks and Snowflake Cortex Agents.
Rubric. Guardrails is re-graded against the redefined floor. A guardrail must be able to block, across all three classes. Four bands moved. The tags moved on almost everything: action coverage from roughly 4 of 22 to 22 of 23, content coverage from 0 of 22 to 13 of 23. The column had been grading data-loss prevention and calling it guardrails.
Corrections. Two columns turned out to be grading on rules they never published. Real-Time Logging was applying an unwritten “is it real time?” test while its own Green band names an org-owned bucket; three products documented exactly that and were marked down for it. Data Boundary’s Red band read “one or more not enforceable”, which collapsed partially coveredand absent into one grade and, read literally, would have put half that column in Red. Both band texts now say what the columns do.
Evidence. Every cell that had never been individually re-read was audited: 22 held, 18 justifications corrected, 6 grades moved. Every receipt that no script could read was replaced with one that can, 33 cells across two passes, except two Agentforce cells, where the vendor’s documentation renders entirely in the browser and even archive captures store only the loading state. That limit is published rather than hidden, because a vendor whose docs a machine cannot read should not be quietly graded on weaker evidence than everyone else.
Our own errors. A demotion was held back because the only evidence for it sat behind a wall; readable sources then showed the opposite and the grade stayed. Two quotes were found to be near-verbatim rather than verbatim, a paraphrase and an added word, neither of which changed a claim, which is why the check is mechanical. And the verifier itself was wrong four times, each time by quietly not checking something: an exclusion list naming readable hosts, a timeout that reclassified readable pages, a retry that skipped the failure it was built for, and a request header that made one vendor’s entire site return 404.
Rubric. Receipts now carry the document date and the sentence the cell rests on, not just a link. The evidence ladder splits third-party sources into disinterested (T3a) and commercially interested (T3b, which includes us), and defines T4 as user-generated content that may never carry a grade. An absence claim now requires an anchor page we read, a search date, and the terms we searched. Grades name ownership; coverage is tagged, never folded into a band. Containment publishes its mapping: a Green needs one cut that is org-owned, targeted, and reaches the product’s default execution path.
Corrections. Two receipts described GitHub issues as open when both were closed. A Microsoft model list named a Google model. One cell rendered Green while its own text read “Graded Yellow”. The only graded cell on the board with no receipt at all now has five. Root URLs standing in for specific claims were replaced with the pages that document them, and a receipt citing SDK documentation was removed, since SDK-class entries are excluded from grading.
Grades that moved, and why they moved. Nine cells asserted that a control did not exist. On six of them the vendor documents the control, and those cells were wrong: Replit and Lovable on connector governance, ChatGPT Enterprise, Replit, Devin and Kiro on model routing, Amazon Q on the difference between a user preference and no control at all. Three absences were checked and upheld. The statistic for products with no model-path insertion point fell from 9 to 4 as a result. That figure is the one v1.1 shipped and is frozen here: it read from the live dataset until 2026-09-02, so later releases kept rewriting a finished release’s result. It stands at 3 today, for reasons belonging to v1.2.
Scope. The A2A column becomes Agent Interop Governance, scoped to the question rather than to one specification, because five products implement ACP and under the old scope every one of them read “not applicable, this is not a gap” over a real ungoverned surface. Windsurf is renamed Devin Desktop rather than merged into Devin: the two have opposite postures on that column. Claude (Team & Enterprise) is added, closing a gap where every competing chat product was graded and the vendor flagged in our conflicts note was not.
Evidence, re-read. Every receipt was checked against the cell it sat under. Thirty did not support the claim above them and were removed, which left fifteen cells with no defensible grade; those rendered as a dot rather than a grade until all fifteen were re-sourced from vendor documentation. Every quote on this page has been re-fetched and matched against the page it cites. So has every number: two figures on the first version of this board appeared in no source at all, one of them under a Red grade, and both are gone.
Grades that moved on re-reading. Microsoft 365 Copilot goes Red to Yellow on Agent Interop. Microsoft documents a live A2A endpoint publishing a conformant agent card, so the claim that it does not participate was false. Copilot Cowork goes Red to not applicable on the same column: Red asserts a control surface exists and is vendor-enforced, while the cell asserted the surface does not exist, and ten other products saying exactly that are graded not applicable. Replit goes Red to Yellow on Data Boundary, where a residency control enforceable org-wide had been missed. Amazon Q Developer goes Green to Yellow on Log Quality, because AWS documents prompt logging and then excludes users subscribed in member accounts of an AWS Organization, the ordinary enterprise arrangement.
Density. No Green on this board now rests on a single page. Requiring a second source stopped being a formality when seven of the twenty-two cells it touched turned out to carry a claim their second page qualified or contradicted: an eligibility gate limiting a logging path to one enterprise type, a preview feature excluding a whole agent class, a zero-retention posture that must be requested rather than configured. A single source is not merely thin. The page that announces a capability is rarely the page that scopes it.
Our own errors. A documentation-index sweep behind two of our absence records covered a quarter of the corpus and was recorded as complete; redone, the conclusions held and two further wrong grades surfaced. One cell asserted a default that a page it cited denies. Anthropic’s own documentation disagrees with itself about whether Cowork exports prompt content by default, and we had picked a side and then kept citing the other one. That cell now states the conflict. Every headline number on this page is computed from the dataset rather than typed into it.
Where we stand. Zaun sells the org-owned control fabric this map grades toward, so the direction of our interest is on the record: we are arguing that organizations should own these enforcement points. Everything else about our position is neutral. Zaun deploys on any hyperscaler or on-premises, routes inference through any provider, and holds no commercial relationship with any vendor graded here. No vendor pays Zaun, sponsors placement, or sees a grade before publication. Check us without trusting us: every graded cell carries tiered vendor documentation, the methodology is published, and the dispute channel below is public. Our own published analysis is tiered T3b, commercially interested, wherever it appears as evidence, on the same terms as any other vendor selling into this category. What would change these grades: vendor documentation showing an org-owned enforcement point we missed, or showing that one we credited does not exist or does not reach the default execution path.
Are you a vendor on this board? If we have you wrong, show us the doc and we fix it in public: [email protected]. Every dispute is investigated against the documentation and resolved in this public changelog, whichever way it goes. Disputes may target grades, activation states, or receipts, and every changed cell gets a dated entry.