.reveal,.reveal.in{opacity:1!important;transform:none!important}

The state of agent-operability in media tech

AXI measures whether autonomous software can actually operate media-and-entertainment platforms — not whether vendors talk about AI, but whether agents can discover capabilities, authenticate, orchestrate work, and economically transact through them.

Why this exists

Autonomous-agent workflows are arriving in media operations faster than vendor documentation suggests. A platform that cannot be discovered, authenticated, or driven by software at scale becomes a bottleneck — not because its features are weak, but because its operability is. AXI measures the second.

The goal is a public, methodology-transparent benchmark that gives buyers, integrators, and vendors a shared yardstick: where each platform sits today on a 0–100 scale of agent-operability, which pillars carry the most weight, and which concrete gaps would move the score most.

How a score is built

Each product is assessed against the rubric version attached to its release using multiple independent evaluations and a consistent evidence standard. A structured scoring process produces pillar results and a composite score. Published releases and operator review keep the process consistent and reproducible.

See Methodology for the full framework, scoring approach, and classification bands, or Rankings for the current leaderboard.

How to read the results

Start with Summary for cohort-level results, then explore individual products in Rankings. Select a release to keep scores, classifications, and methodology in the same period. Compare products assessed under the same rubric or a documented scoring-equivalent patch; a score change can reflect new evidence or evaluation differences as well as a changed product.

Explore the release summary →

What this isn't

  • Not a quality or feature ranking. A platform can have outstanding domain features and still score poorly on agent-operability.
  • Not paid placement or a vendor scorecard. Vendors do not pay to be included.
  • Not a one-shot benchmark. Releases are versioned and products are reassessed as the methodology and vendor surfaces evolve.

1. Executive Summary

The Agent Experience Index (AXI) is a benchmarking framework for evaluating the degree to which media and entertainment technology platforms are operable by autonomous software systems.

AXI does not measure whether a company has AI features.

AXI measures whether autonomous coding agents and operational agents can discover, authenticate, configure, orchestrate, monitor, govern, and economically interact with a platform through machine-operable interfaces under realistic enterprise governance constraints.

AXI is therefore an operational infrastructure benchmark, not an AI feature benchmark.

The primary audience includes:

  • media technology buyers
  • vendors
  • product leaders
  • infrastructure teams
  • ecosystem partners
  • analysts
  • investors
  • agentic workflow builders

2. Core Thesis

The next generation of enterprise software competition may increasingly depend on which platforms are most operable by autonomous systems.

Traditional enterprise software was designed around:

  • human operators
  • graphical user interfaces
  • manual configuration
  • implementation specialists
  • professional services teams
  • seat-based usage patterns

AXI assumes that future-ready infrastructure increasingly supports:

  • machine-operable workflows
  • autonomous provisioning
  • programmatic governance
  • software-to-software execution
  • agent-mediated orchestration
  • governed machine access
  • usage-based or transactional economic models

The framework measures:

autonomous operability under realistic enterprise governance conditions.


3. What AXI Measures

AXI measures:

  • agent-accessible documentation
  • machine execution surfaces
  • CLI/MCP/equivalent agent surfaces
  • execution surface completeness
  • autonomous provisioning capability
  • governed operational accessibility
  • workflow composability
  • operational observability
  • machine-operable governance
  • commercial alignment with autonomous consumption
  • public agentic-strategy commitments and roadmap signals

AXI does not measure:

  • overall product quality
  • customer satisfaction
  • market share
  • feature breadth
  • generic AI sophistication
  • AI marketing
  • embedded copilots
  • chatbot functionality

A company may have sophisticated AI features and still score poorly if its platform is not operable by autonomous systems.

A company may have little AI branding and score highly if it exposes strong machine-operable infrastructure.


4. Foundational Assessment Question

Every AXI assessment is governed by one question:

Could a modern autonomous coding agent independently implement and operate this platform using available machine interfaces and operational documentation with minimal human UI dependency?

This question governs scoring, evidence collection, and recommendations.


5. Public Discoverability vs Governed Operability

AXI does not assume that all operational interfaces should be fully public.

Enterprise systems, especially in media and entertainment, may have valid reasons to restrict access:

  • rights complexity
  • security concerns
  • premium content protection
  • financial workflows
  • compliance requirements
  • partner ecosystem boundaries

AXI distinguishes between:

Concept Meaning
Public Agent Discoverability Can autonomous systems discover and understand operational capabilities?
Governed Agent Operability Can autonomous systems operate the platform within controlled enterprise boundaries?

A platform may score highly without unrestricted public exposure if it supports governed autonomous operation.

However, systems that depend on undocumented workflows, human-only onboarding, UI-only administration, or professional-services mediation should score lower.


6. Evidence Accessibility Tiers

AXI recognizes four evidence/access tiers.

Tier Description
E1 Fully public anonymous access
E2 Authenticated developer access
E3 Partner or ecosystem access
E4 Brokered or governed operational access

Governed access is not inherently penalized.

Opaque access is penalized when it prevents reproducible assessment or autonomous operation.

Consistent Evidence Discovery

Before each assessment, AXI gathers a bounded set of relevant public and first-party evidence for the product being evaluated.

Every evaluator receives the same evidence set and the same scoring instructions. Claims must be supported by captured source material, and incomplete or conflicting results are held for review rather than published.

This consistent evidence standard makes published results comparable across products and releases while preserving an auditable record of how each assessment was produced.


7. Multi-Model Assessment Model

AXI assessments should be performed by multiple independent LLM evaluators.

Recommended evaluators:

  • ChatGPT
  • Gemini
  • Claude

During a native scoring attempt, each evaluator receives:

  • the same rubric
  • the same deterministic first-party evidence set
  • the same scoring instructions
  • the same assessment template

Each evaluator returns:

  • entity-match decision and confidence
  • fired signal keys and applied penalty keys by category
  • lifecycle coverage estimates where applicable
  • evidence notes
  • assertion-level evidence bindings for every scored signal, penalty, and positive lifecycle-coverage estimate
  • recommendations

AXI centrally derives numeric category subtotals and the final score from those three structured determinations; evaluators do not supply numeric scores.

Final AXI score:

A product is assessed by all three evaluators in parallel. A score is complete and publishable only when ChatGPT, Claude, and Gemini each return exactly one valid determination for every rubric category. A 1/3 or 2/3 result is retained only as a draft partial so a later healing run can fill the missing cells without repeating successful calls.

For audit and disagreement analysis, AXI also stores a standalone score derived from each evaluator's own determinations. These three audit values are not the published consensus score. Historical evaluator rows created before standalone derivation was introduced remain labeled as legacy consolidated values rather than being retroactively reinterpreted.

Account-level provider failures are not assessment outcomes. Authentication, billing, or exhausted-quota failures halt the batch and cancel its remaining provider calls; they are recorded as infrastructure failures rather than model refusals or partial evidence. Ordinary evaluator refusals and transient errors remain eligible for the draft-partial and healing path above.

For ordinary categories, each signal and penalty is multiplied by the fraction of all three evaluators that detected it, after which the category floor and cap are applied. For lifecycle-coverage categories, the three coverage estimates are averaged and the matching rubric bucket is selected. The final AXI score is the sum of the resulting category subtotals. This consensus aggregation rewards capabilities the models agree on without silently changing the denominator when a provider fails.

The assessment, all three evaluator audit rows, category components, and consolidated recommendations are published as one transaction through the canonical publication RPC. Direct draft-to-published updates are rejected. A failure in any part leaves no partially published scorecard. Once published, evaluator rows and rubric content are immutable; a rubric change requires a new version. Provider raw_output retains the evaluator's complete structured response; the validated parsed determinations are numeric scoring authority. Public signal/penalty breakdowns and the flattened assertion-evidence index are verified from those determinations at the publication boundary. Every scored assertion must bind to an integrity-checked source in its persisted evidence packet. Recommendation source members must match recommendations in the corresponding persisted evaluator output; agreement, representative text, median impact, and ordering are then derived from the exact member multiset rather than accepted as caller labels.

Human role:

  • define rubric
  • define weights
  • review anomalies
  • improve methodology
  • manage evidence quality

Humans are prevented from overriding scores directly.


8. Scores and Classifications

AXI should publish both numeric scores and classifications.

Numeric scores preserve precision and enable trend tracking.

Classifications improve legibility and sharing. The score-range → classification band table is published on the methodology page, driven by the active rubric.

Score rangeClassification
90100Agent Native
8089Agent Operable
6579Agent Capable
5064Emerging Agentic
3549Human Mediated
2034Human Dependent
019Human Operated

9. AXI Scoring Framework

Total score: 100

Category 1 — Agent-Accessible Documentation

Objective: Can an autonomous agent learn how to operate the platform from available documentation?

Signal Rationale
API documentation Establishes machine-accessible operational surface
OpenAPI / Swagger / machine-readable schemas Enables machine parsing and tooling
Executable implementation examples Helps agents implement correctly
CLI documentation Indicates non-UI operational design
MCP or equivalent agent protocol documentation Indicates explicit agent-oriented access
Authentication flow documentation Enables autonomous identity and access setup
Infrastructure-as-code examples Enables repeatable deployment
Workflow orchestration examples Demonstrates composability
Troubleshooting and recovery guidance Helps agents recover from failure

Penalties:

Condition
Sales-gated technical documentation
Static/manual-style docs without executable guidance
Missing auth documentation
Documentation materially incomplete or stale

Category 2 — Machine Execution Surfaces

Objective: Can autonomous systems execute operational tasks programmatically?

Signal Rationale
Mature operational APIs Core execution surface
Official SDKs Supports implementation
Webhooks/eventing Supports event-driven workflows
Async job support Supports scalable operations
Retry/idempotency semantics Supports safe autonomous execution
Workflow APIs Enables orchestration
Declarative orchestration Enables machine-composable workflows
CLI execution capability Indicates machine-first operation
MCP or equivalent execution interface Enables structured agent interaction
Operational health/observability surfaces Supports monitoring and remediation

Penalties:

Condition
No CLI or equivalent machine execution surface
No MCP or equivalent agent execution interface
UI-only execution of core workflows
Read-only machine interfaces only

Category 3 — Execution Surface Completeness

Objective: How completely do CLI/MCP/API/agent surfaces cover the operational lifecycle?

This category addresses sincerity of implementation. A vendor should not receive high agentic-readiness credit for exposing a superficial CLI, toy MCP server, read-only adapter, or narrow wrapper that covers only a small portion of platform functionality.

Lifecycle areas:

Lifecycle Area
Discover
Provision
Operate
Monitor
Recover
Govern
Deprovision

Scoring:

Coverage
0–10% lifecycle coverage
10–25% lifecycle coverage
25–50% lifecycle coverage
50–70% lifecycle coverage
70–90% lifecycle coverage
90–100% lifecycle coverage

Rules:

  • Do not count endpoints mechanically.
  • Assess lifecycle coverage, not command count.
  • Read-only coverage is materially weaker than execution authority.
  • Admin/governance coverage matters.
  • Recovery and observability coverage matter.
  • Thin wrappers should receive low scores.

Penalties:

Condition
CLI exists but covers only trivial/read-only functionality
MCP exists but is demo-only or non-operational
Machine surfaces omit provisioning entirely
Machine surfaces omit recovery/remediation entirely
Machine surfaces omit governance/admin entirely

Category 4 — Autonomous Provisioning and Configuration

Objective: Can autonomous systems provision, configure, and administer the platform without significant human mediation?

Signal Rationale
Self-service provisioning Enables autonomous onboarding
Usage-based onboarding Aligns with elastic operation
API-key or token-based auth Enables machine identity
OAuth/service principals Enables governed enterprise automation
Infrastructure-as-code support Enables reproducible deployment
Configuration-as-code support Enables versioned configuration
Machine-operable administration Reduces UI dependency

Penalties:

Condition
Human onboarding required for operational access
UI-required provisioning
Mandatory professional services for implementation

Category 5 — Commercial Agentic Alignment

Objective: Can autonomous systems economically transact with the platform?

Signal Rationale
Usage-based pricing Aligns with machine consumption
API monetization model Indicates machine-oriented business model
Transactional economics Supports autonomous operations
Elastic scaling Supports variable agent workloads
Public pricing transparency Improves discoverability
Autonomous payment/provisioning Enables self-operating workflows

Penalties:

Condition
Pure seat-based licensing
Enterprise-sales-gated operational access
Mandatory professional services dependency

Category 6 — Governance and Operational Trust

Objective: Can enterprises safely delegate operational authority to autonomous systems?

Signal Rationale
Service account support Enables controlled machine identity
Auditability Enables traceability
RBAC controls Enables scoped permissions
Policy enforcement Enables governed operation
Rights/security boundaries Critical in media environments
Approval layers Supports controlled autonomy
Operational observability Supports governance and monitoring

Category 7 — Agentic Strategy Signals

Objective: Has the company made public commitments toward agentic operability, even if the operational surface isn't fully built yet? This pillar captures intent — meaningful for trajectory but not for current operability. A company that publicly states an MCP roadmap is materially further along than one that's silent on it. Maximum contribution is 10/100; marketing alone caps a company at F-tier on the overall score.

Signal Rationale
Published blog post, whitepaper, or press release about agentic/autonomous operability strategy Stated direction signals strategic intent
Public commitment to MCP server availability (roadmap, coming-soon, or GA announcement) MCP-specific intent is a directly relevant signal
Public commitment to expanding CLI or SDK coverage API/CLI commitment is foundational to agent operability
Leadership (CEO, CTO, or DevRel) has spoken publicly about agentic strategy Leadership commitment correlates with delivery
Available via at least one third-party orchestration platform or marketplace Third-party operability counts toward agent reachability
Developer-relations function actively focused on agent / automation use cases Investment in agent DevRel predicts faster operational progress

Penalties: none. This pillar scores stated intent only — the absence of public commitments yields no credit rather than a deduction, and its 10-point cap prevents public messaging alone from producing a high overall score.

Documented assertion corrections

An independently reviewed discrepancy may be corrected from a retained response by the same evaluator using the same evidence packet. Only the explicitly reviewed assertions change; the remaining decisions are preserved. Such a result is labeled as derived and retains both complete structured evaluator outcomes, the exact correction mask and source provenance. It does not claim that one provider invocation authored the combined result. Scores are recalculated using the unchanged rubric and aggregation policy, without a minimum score or a requirement to agree with a previous month.

Targeted assertion corrections may request a partial response from the same evaluator against the unchanged packet. An independently reviewed exact mask can be applied to the immutable original through deterministic replay, preserving all unrelated decisions. The resulting complete assessment is explicitly derived; no full-provider re-endorsement is implied. Its explicit per-field invocation metadata preserves the original cohort rubric authority while attributing new prompt/model/usage fields to the accepted targeted invocation. Full original and partial-response provenance remain distinguishable in the exported record. Separately attributed independent corroboration may support an unchanged scored assertion without rewriting the original provider evidence.

When an accepted correction disproves a dependent recommendation's premise or estimated impact, an independent reviewer may withdraw that exact recommendation. The deletion, original content hash, assertion dependency and reason are retained as reviewer provenance; it is not attributed to the evaluator. Withdrawal supplies no replacement advice and cannot change scores. Unrelated or still-useful compound recommendations remain intact.