ZaunZaun
Zaun Research

Methodology
How this board is graded.

This document is the rubric the index is graded against. It is published because the board’s central claim is that you can check our work without trusting us, and most of the ways it offers to do that depend on this page existing. Sections are numbered and stable; justifications on the board cite them by number.

1. What the index measures

Whose infrastructure enforces: yours, or the vendor’s?

One question, applied 261 times. Not how good the console is. Not how much of the product a control reaches. Not how mature it is. Those are real, and they are carried in tags, footnotes and activation states. The grade itself names ownership of the enforcement point, because that is the property that survives a vendor’s roadmap, and the one an enterprise can build a runbook against.

1.1 Public sources only

Every grade derives from publicly available material. No product testing, no customer telemetry, no briefings, no NDA documents. That constraint has a consequence we state rather than hide: a grade describes documented controllability. Where a vendor ships a control but does not document its operational parameters, we grade the documentation gap and say so.

A control an organization cannot operate from published material is a weaker control. A vendor can fix that grade by writing documentation, which is the right incentive. It also means the board is falsifiable in both directions by anyone with a browser.

2. Eligibility and scope

2.1 The eligibility rule

A product earns a row when it is agentic (takes actions, not just returns text), enterprise-adoptable, has an administrative surface someone can configure on behalf of others, and is publicly documented.

2.2 The generally-available-tier rule

We grade the highest generally available enterprise tier. Controls gated behind a higher SKU carry the tier activation state. Waitlisted and invite-only capabilities are not GA and do not carry a grade. The buyer this board is written for has already paid; grading the free tier describes a product they are not running.

2.3 The possibility gate

A capability is graded when the published documentation is sufficient for an organization to operate it: turn it on, configure it, and know what it covers. A capability documented as existing but whose operational parameters are undocumented is graded as the documentation gap it is.

2.4 Exclusions

SDK-class entries, code libraries such as agent SDKs, are excluded. A library has no administrative surface; its governance is whatever the application built on it implements. It follows that SDK documentation may not be used to grade a product either.

Products that pass the eligibility rule and are not yet graded are named rather than left silent. At v1.1 that list was Atlassian Rovo, Slack AI, Notion AI, Perplexity Enterprise, ServiceNow AI Agents, Databricks and Snowflake Cortex; all seven were graded in v1.2 and the list was not updated with them, so this section spent a release naming graded rows as ungraded. The current queue is GitLab Duo Agent Platform, Amazon Quick, IBM watsonx Orchestrate and UiPath. That is a scope limit, not a judgement.

2.5 The agentgateway baseline

An organization-operated gateway sitting between its users and the model provider, terminating credentials it issues, applying policy it authors, and emitting logs to a destination it controls.

A product scores Green on a gateway-dependent surface when it can be fronted by that gateway without loss of function, and the vendor documents how. The baseline is a measuring stick, not a product recommendation. Zaun sells into this category, which is why the baseline is stated in vendor-neutral terms and why our own published analysis is tiered T3b (§3.1).

3. Evidence

3.1 The evidence tiers

TierClassLoad-bearing?
T1Vendor reference documentation: product docs, admin guides, API references, security and trust pages, termsYes
T2Vendor announcements: blogs, changelogs, release notes, message-center entriesYes, with a document date
T3aDisinterested third party: CVE and GHSA records, non-vendor advisories, academic work, regulatory filings, standards bodies, mainstream technical pressYes
T3bCommercially interested third party: analysis from companies selling into this category, including Zaun, and including any vendor graded here writing about anotherCorroborating only
T4User-generated: issue trackers, community forums, discussion boardsNo

T4 may evidence exactly one thing: the existence and recorded state of a community request. Never the state of a product. A closed feature request does not establish that a capability shipped; only vendor documentation does.

A Green requires T1. Full controllability is a claim about a documented, operable mechanism.

3.2 What a receipt must be

Every receipt carries a deep link, a tier, the date we checked it, the date the document was published, and the sentence the cell rests on, quoted.

The quote is the discipline. If you cannot quote the sentence, you have not established the claim, and if the quote does not say what the justification says, the receipt is wrong. It also turns revalidation into a diff: does this sentence still appear on this page?

Every quote we can fetch is re-fetched and string-matched against its page, weekly. About one in seven sits on a page a script cannot read: a PDF, an app that renders in the browser, a host that refuses automated requests. Those are checked by eye instead. We state the difference rather than rounding it up to “all verified”, because an assertion whose scope is wider than its evidence is the exact failure this index exists to expose.

A receipt must document the mechanism named, cover the plan tier graded, cover the protocol graded, and be about the graded product. Documentation roots and index pages are not receipts. A dated claim needs a dated source. Minimum two receipts per graded cell.

Every figure this board states must appear in a quote on that cell. Not that a page exists. The number itself, in a sentence we quoted from a page that cell cites. A statistic is the most quotable thing on a comparison board and the least checkable by eye, which is exactly why it needs a machine to check it. Two numbers on the first draft of this board were invented, and one of them sat under a Red grade.

A justification usually makes several claims, so a receipt may carry more than one quoted sentence from the same page. Every one is re-fetched and matched against that page like the first.

Current state, computed from the dataset: 1070 receipts, of which 1070 carry the quote and 624 do not yet have an established document date. Both gaps are visible on every affected receipt rather than hidden behind a uniform validation stamp.

3.3 Evidence density, per row

How many distinct pages stand behind each product’s row, beside how many Greens that row earned. A vendor should be able to see whether their grade rests on more or fewer sources than a competitor’s without taking our word for the comparison.

Counted by unique URL, not by receipt: one page cited under four labels is one source. At v1.0 this ran the wrong way, with the most-credited row resting on the fewest pages. Publishing it is what makes that hard to repeat.

ProductGreensUnique sourcesReceipts
Amazon Bedrock AgentCore83946
Gemini Enterprise Agent Platform82633
Claude Code72940
Claude Cowork72639
Microsoft Foundry63842
OpenAI Codex CLI62237
Databricks Unity AI Gateway43136
Glean42937
Copilot Studio33237
Atlassian Rovo23237
GitHub Copilot22829
Kiro22834
Perplexity Enterprise22639
Notion AI22335
ServiceNow AI Agents17082
Salesforce Agentforce14653
Replit12733
Devin AI12735
Gemini Code Assist12436
ChatGPT Enterprise12334
Cursor12331
Google Gemini (chat app)12029
Snowflake Cortex Agents12033
Slack AI11928
Microsoft 365 Copilot Cowork11430
Devin Desktop02938
Claude (Team & Enterprise)02428
Microsoft 365 Copilot02429
Lovable02030

4. Grading

4.1 Bands

BandMeans
GFull controllability. Enforcement lives in the organization’s own control fabric: its gateway, IdP tenant, cloud account, MDM-authored policy, network, SIEM. The org can change or revoke it without the vendor
YLimited controllability. The vendor’s plane, configured by the org. The org authors policy; the vendor evaluates it
RNo controllability. Vendor-authored and vendor-enforced
Security gap. No control surface exists. Requires a documented absence record (§5)
n/aThe surface does not apply to this product’s architecture. This is not a gap. Requires a positive receipt, never a silence
·Held. The evidence does not yet meet this standard, so no grade is published

4.2 Grade ownership, tag coverage

Partial coverage does not move a grade. It is recorded as a coverage footnote and, where the column defines them, a coverage tag. A column may make coverage a band criterion only if its spec says so in advance, and it then applies to every row identically.

There is one definitional limit: a control that does not reach the default execution path is not full controllability. That is not a coverage downgrade, it is the Green definition applied. A mechanism that excludes the existing fleet does not let the organization enforce.

4.3 Ceiling-grading, and what it does not mean

The grade is a ceiling. It answers: if your organization turns this on, whose plane enforces it? That is a property of the product.

The activation marker is potential energy. It answers: is it on? That is a property of your deployment, and the index cannot observe your deployment, which is precisely why the marker exists rather than a lower grade.

Every marked cell names which state applies (opt-in, tier-gated, deployment-gated, build-required, preview, or gateway-dependent) and the panel states what turning it on requires. Read together: the richest logging on this board is not flowing until someone turns it on. The grade tells you what you would own. The marker tells you that you probably do not own it yet. Neither discounts the other.

4.4 Preview

Vendor-shipped previews grade as shipped, carrying the preview state. The alternative penalises vendors for documenting early and rewards silence. A preview cannot carry a Green where it is scoped too narrowly to reach the default execution path, and a preview date needs a dated receipt.

4.5 Justifications

A justification names the mechanism, says where enforcement happens, and stops. Every mechanism it names has a receipt documenting that mechanism. A grade is never delivered in prose. The reader is not asked to re-grade a cell based on a sentence inside it.

5. Absence

Absence grading is the most error-prone method on this board. It is also unavoidable: “there is no control here” is frequently the finding. So it carries the most structure.

Not documented in the vendor’s published materials as of DATE, anchored at URL, searched with these terms.

That is what an absence claim asserts, and all it asserts. Never “the vendor does not support X”. The first is a claim about the documentary record that a vendor can refute with one link. The second is a claim about an implementation we have not tested.

Required: anchor URLs where the control would live, each one resolved and read, never a documentation root, ideally adjacent to a capability the vendor does document: the search date, the verbatim search terms, which documentation index was enumerated, the plan tier and deployment mode covered, and who ran it.

What does not establish absence: a 404 on a guessed URL, zero search-engine results, silence on a marketing page, an unread sibling page, or a forum thread saying the feature is missing.

“Not documented” is never “not applicable.” It resolves to a gap if the record is complete, and to held if it is not. Absence claims expire at their column’s revalidation interval and render stale rather than silently continuing to assert a negative.

No cell is currently held: all 261 either carry a grade the evidence standard supports or are marked not applicable. The hold mechanism is the release valve, not a permanent state. Every cell held at v1.1 was resolved or re-sourced.

6. The columns

Each column has its own spec stating the question it asks, its bands, what evidence satisfies each, and what is explicitly out of scope.

ColumnAsks
Model RoutingCan the org insert itself in the model path: gateway, base URL, BYOK, model pinning?
Tool & Connector GovernanceCan the org decide which tools and connectors an agent may reach?
GuardrailsCan the org apply its own enforcement policy to the prompt, response and tool path, and where does the block happen?
Agent Interop GovernanceCan the org govern this product’s participation in cross-boundary agent-to-agent protocols: A2A, ACP, or any successor?
RT LoggingMechanism and latency: does telemetry arrive at a destination the org controls?
Log QualityDo the logs carry prompt and response content, agent actions, cost attribution, and identity?
ContainmentAt which layers can the org cut this agent off?
Data BoundaryCan the org enforce retention control, training-use exclusion, and residency?
Agent IdentityDoes the agent have a principal the org issues, scopes, revokes and sees in logs, in the mode the product actually runs in?

Two rules cut across columns:

  • Guardrails has a floor: a guardrail must be able to block. System prompts, custom instructions, convention files, agent skills, documentation and post-hoc detection are excluded. Without this floor every product qualifies through custom instructions.
  • Admission is Tool Governance; runtime blocking is Guardrails. “This agent may use that connector” is admission. “This agent may not pass a customer record to it” is a guardrail.
  • Agent Identity grades the default execution mode, not the platform’s capability. When a developer runs a coding agent in their IDE, the agent acts as the developer: their token, their permissions, their name in the audit log. That a service-account path exists elsewhere in the same vendor’s platform does not change what happened on that laptop, so it is carried as a marker and never as a grade. There is no “not applicable” on this column: every agentic product has an execution principal, and a product with no separate one is a security gap, not an inapplicable surface.
  • Agent Interop Governance is scoped to the question, not to one specification. Any protocol by which a product exchanges work with an agent outside its own trust boundary is in scope: A2A and ACP today. Scoping to a single protocol name would let a real ungoverned inter-agent surface be graded “not applicable, this is not a gap.”

7. Adjudication and disputes

Contested grades are argued from both sides before they land: a vendor case from public sources, a buyer case from an enterprise on the top GA SKU who has already paid, then an adjudication against what the column actually measures. Ties go to the buyer. Both cases must quote documentation; a case with no quotes is not heard.

Rulings publish alongside the grades, so a vendor writing to us finds the argument already made and answered, or finds that it has not been, in which case they may be right.

What changes a grade: vendor documentation showing an org-owned enforcement point we missed, or showing that one we credited does not exist or does not reach the default execution path, or a published date, mechanism or scope contradicting a receipt we cite.

What does not: a briefing, a roadmap, a private assurance, an NDA document. We will not move a grade on evidence a reader cannot check.

8. Revalidation

Receipts carry both the check date and the document date, so staleness is visible rather than inferred. Revalidation asks whether the quoted sentence still appears on the page; a cell whose quote has gone is held and read by a person, not silently re-graded.

Intervals are set per column: shorter where preview density is high, where the column’s Green share is high (a stale Green over-credits and nobody complains; a stale Red under-credits and the vendor tells us within a week), and where the surface depends on a protocol under active standardisation. Floor 30 days, ceiling 120. Cells past their interval render as stale. The current values are drafts, not measurements. Computing them from observed documentation-change rates needs a document date on each receipt, and 624 of 1070 do not yet carry one. A cell aging past a rough interval is a better signal than a cell that never ages, which is what this board had, so the drafts ship and the computation waits for the history to get deep enough to measure.

9. Corrections and versioning

Four entry classes, never mixed: CORRECTION (we were wrong, the product did not change), RUBRIC (the question or bands changed, grades may move with no product change), CHANGE (the vendor shipped something), SCOPE (a row or column was added, merged, renamed, retired). A rubric change lists every cell it moves.

We publish our own corrections first, and with the same standing as anyone else’s. The v1.1 changelog carries a list of errors this research made about its own work, including a mislabeled receipt of exactly the kind it flagged on the board and a double-counted advisory.

The root cause is published too, because it is the most useful thing we learned: every error, on the board and in the audit of the board, came from treating a page’s existence as evidence of its contents. Open the page. Quote the sentence. Record the date.

Each released version’s dataset is frozen, and citations carry the version.

10. Conflicts

Zaun sells the org-owned control fabric this index grades toward. That is the direction of our interest and it belongs on the record: we are arguing that organizations should own these enforcement points, and this board is built to make that argument checkable rather than to make it persuasive.

Everything else about our position is neutral. Zaun deploys on any hyperscaler or on-premises and routes inference through any provider, so no vendor on this board is one we depend on. No graded vendor pays Zaun, sponsors placement, or reviewed a grade before publication.

Our own published analysis is tiered T3b, commercially interested, wherever it appears as evidence, on the same terms as any other vendor selling into this category.

Two structural checks against our own bias, both visible on the board: evidence density per row is published, so a vendor can see whether their grade rests on more or fewer sources than a competitor’s; and advisories are footnoted under one rule for everyone, which is why an advisory against a vendor we use sits beside one against a vendor we do not, in the same release.

If we have a vendor wrong, show us the document: [email protected]. Every dispute is investigated against the documentation and resolved in the public changelog, whichever way it goes.

← Back to the map