1. Executive Summary
The Agent Experience Index (AXI) is a benchmarking framework for evaluating the degree to which media and entertainment technology platforms are operable by autonomous software systems.
AXI does not measure whether a company has AI features.
AXI measures whether autonomous coding agents and operational agents can discover, authenticate, configure, orchestrate, monitor, govern, and economically interact with a platform through machine-operable interfaces under realistic enterprise governance constraints.
AXI is therefore an operational infrastructure benchmark, not an AI feature benchmark.
The primary audience includes:
- media technology buyers
- vendors
- product leaders
- infrastructure teams
- ecosystem partners
- analysts
- investors
- agentic workflow builders
2. Core Thesis
The next generation of enterprise software competition may increasingly depend on which platforms are most operable by autonomous systems.
Traditional enterprise software was designed around:
- human operators
- graphical user interfaces
- manual configuration
- implementation specialists
- professional services teams
- seat-based usage patterns
AXI assumes that future-ready infrastructure increasingly supports:
- machine-operable workflows
- autonomous provisioning
- programmatic governance
- software-to-software execution
- agent-mediated orchestration
- governed machine access
- usage-based or transactional economic models
The framework measures:
autonomous operability under realistic enterprise governance conditions.
3. What AXI Measures
AXI measures:
- agent-accessible documentation
- machine execution surfaces
- CLI/MCP/equivalent agent surfaces
- execution surface completeness
- autonomous provisioning capability
- governed operational accessibility
- workflow composability
- operational observability
- machine-operable governance
- commercial alignment with autonomous consumption
- public agentic-strategy commitments and roadmap signals
AXI does not measure:
- overall product quality
- customer satisfaction
- market share
- feature breadth
- generic AI sophistication
- AI marketing
- embedded copilots
- chatbot functionality
A company may have sophisticated AI features and still score poorly if its platform is not operable by autonomous systems.
A company may have little AI branding and score highly if it exposes strong machine-operable infrastructure.
4. Foundational Assessment Question
Every AXI assessment is governed by one question:
Could a modern autonomous coding agent independently implement and operate this platform using available machine interfaces and operational documentation with minimal human UI dependency?
This question governs scoring, evidence collection, and recommendations.
5. Public Discoverability vs Governed Operability
AXI does not assume that all operational interfaces should be fully public.
Enterprise systems, especially in media and entertainment, may have valid reasons to restrict access:
- rights complexity
- security concerns
- premium content protection
- financial workflows
- compliance requirements
- partner ecosystem boundaries
AXI distinguishes between:
| Concept |
Meaning |
| Public Agent Discoverability |
Can autonomous systems discover and understand operational capabilities? |
| Governed Agent Operability |
Can autonomous systems operate the platform within controlled enterprise boundaries? |
A platform may score highly without unrestricted public exposure if it supports governed autonomous operation.
However, systems that depend on undocumented workflows, human-only onboarding, UI-only administration, or professional-services mediation should score lower.
6. Evidence Accessibility Tiers
AXI recognizes four evidence/access tiers.
| Tier |
Description |
| E1 |
Fully public anonymous access |
| E2 |
Authenticated developer access |
| E3 |
Partner or ecosystem access |
| E4 |
Brokered or governed operational access |
Governed access is not inherently penalized.
Opaque access is penalized when it prevents reproducible assessment or autonomous operation.
Consistent Evidence Discovery
Before each assessment, AXI gathers a bounded set of relevant public and first-party evidence for the product being evaluated.
Every evaluator receives the same evidence set and the same scoring instructions. Claims must be supported by captured source material, and incomplete or conflicting results are held for review rather than published.
This consistent evidence standard makes published results comparable across products and releases while preserving an auditable record of how each assessment was produced.
7. Multi-Model Assessment Model
AXI assessments should be performed by multiple independent LLM evaluators.
Recommended evaluators:
During a native scoring attempt, each evaluator receives:
- the same rubric
- the same deterministic first-party evidence set
- the same scoring instructions
- the same assessment template
Each evaluator returns:
- entity-match decision and confidence
- fired signal keys and applied penalty keys by category
- lifecycle coverage estimates where applicable
- evidence notes
- assertion-level evidence bindings for every scored signal, penalty, and
positive lifecycle-coverage estimate
- recommendations
AXI centrally derives numeric category subtotals and the final score from those
three structured determinations; evaluators do not supply numeric scores.
Final AXI score:
A product is assessed by all three evaluators in parallel. A score is complete
and publishable only when ChatGPT, Claude, and Gemini each return exactly one
valid determination for every rubric category. A 1/3 or 2/3 result is retained
only as a draft partial so a later healing run can fill the missing cells
without repeating successful calls.
For audit and disagreement analysis, AXI also stores a standalone score derived
from each evaluator's own determinations. These three audit values are not the
published consensus score. Historical evaluator rows created before standalone
derivation was introduced remain labeled as legacy consolidated values rather
than being retroactively reinterpreted.
Account-level provider failures are not assessment outcomes. Authentication,
billing, or exhausted-quota failures halt the batch and cancel its remaining
provider calls; they are recorded as infrastructure failures rather than model
refusals or partial evidence. Ordinary evaluator refusals and transient errors
remain eligible for the draft-partial and healing path above.
For ordinary categories, each signal and penalty is multiplied by the fraction
of all three evaluators that detected it, after which the category floor and
cap are applied. For lifecycle-coverage categories, the three coverage
estimates are averaged and the matching rubric bucket is selected. The final
AXI score is the sum of the resulting category subtotals. This consensus
aggregation rewards capabilities the models agree on without silently changing
the denominator when a provider fails.
The assessment, all three evaluator audit rows, category components, and
consolidated recommendations are published as one transaction through the
canonical publication RPC. Direct draft-to-published updates are rejected. A
failure in any part leaves no partially published scorecard. Once published,
evaluator rows and rubric content are immutable; a rubric change requires a new
version. Provider raw_output retains the evaluator's complete structured
response; the validated parsed determinations are numeric scoring authority.
Public signal/penalty breakdowns and the flattened assertion-evidence index are
verified from those determinations at the publication boundary. Every scored
assertion must bind to an integrity-checked source in its persisted evidence
packet. Recommendation source members must match recommendations in the
corresponding persisted evaluator output; agreement, representative text,
median impact, and ordering are then derived from the exact member multiset
rather than accepted as caller labels.
Human role:
- define rubric
- define weights
- review anomalies
- improve methodology
- manage evidence quality
Humans are prevented from overriding scores directly.
8. Scores and Classifications
AXI should publish both numeric scores and classifications.
Numeric scores preserve precision and enable trend tracking.
Classifications improve legibility and sharing. The score-range → classification
band table is published on the methodology page, driven by the active rubric.