1. Executive Summary
The Agentic Readiness Index (ARI) is a benchmarking framework for evaluating the degree to which media and entertainment technology platforms are operable by autonomous software systems.
ARI does not measure whether a company has AI features.
ARI measures whether autonomous coding agents and operational agents can discover, authenticate, configure, orchestrate, monitor, govern, and economically interact with a platform through machine-operable interfaces under realistic enterprise governance constraints.
ARI is therefore an operational infrastructure benchmark, not an AI feature benchmark.
The primary audience includes:
- media technology buyers
- vendors
- product leaders
- infrastructure teams
- ecosystem partners
- analysts
- investors
- agentic workflow builders
2. Core Thesis
The next generation of enterprise software competition may increasingly depend on which platforms are most operable by autonomous systems.
Traditional enterprise software was designed around:
- human operators
- graphical user interfaces
- manual configuration
- implementation specialists
- professional services teams
- seat-based usage patterns
ARI assumes that future-ready infrastructure increasingly supports:
- machine-operable workflows
- autonomous provisioning
- programmatic governance
- software-to-software execution
- agent-mediated orchestration
- governed machine access
- usage-based or transactional economic models
The framework measures:
autonomous operability under realistic enterprise governance conditions.
3. What ARI Measures
ARI measures:
- agent-accessible documentation
- machine execution surfaces
- CLI/MCP/equivalent agent surfaces
- execution surface completeness
- autonomous provisioning capability
- governed operational accessibility
- workflow composability
- operational observability
- machine-operable governance
- commercial alignment with autonomous consumption
- public agentic-strategy commitments and roadmap signals
ARI does not measure:
- overall product quality
- customer satisfaction
- market share
- feature breadth
- generic AI sophistication
- AI marketing
- embedded copilots
- chatbot functionality
A company may have sophisticated AI features and still score poorly if its platform is not operable by autonomous systems.
A company may have little AI branding and score highly if it exposes strong machine-operable infrastructure.
4. Foundational Assessment Question
Every ARI assessment is governed by one question:
Could a modern autonomous coding agent independently implement and operate this platform using available machine interfaces and operational documentation with minimal human UI dependency?
This question governs scoring, evidence collection, and recommendations.
5. Public Discoverability vs Governed Operability
ARI does not assume that all operational interfaces should be fully public.
Enterprise systems, especially in media and entertainment, may have valid reasons to restrict access:
- rights complexity
- security concerns
- premium content protection
- financial workflows
- compliance requirements
- partner ecosystem boundaries
ARI distinguishes between:
| Concept | Meaning |
|---|---|
| Public Agent Discoverability | Can autonomous systems discover and understand operational capabilities? |
| Governed Agent Operability | Can autonomous systems operate the platform within controlled enterprise boundaries? |
A platform may score highly without unrestricted public exposure if it supports governed autonomous operation.
However, systems that depend on undocumented workflows, human-only onboarding, UI-only administration, or professional-services mediation should score lower.
6. Evidence Accessibility Tiers
ARI recognizes four evidence/access tiers.
| Tier | Description |
|---|---|
| E1 | Fully public anonymous access |
| E2 | Authenticated developer access |
| E3 | Partner or ecosystem access |
| E4 | Brokered or governed operational access |
Governed access is not inherently penalized.
Opaque access is penalized when it prevents reproducible assessment or autonomous operation.
Deterministic First-Party Discovery
Before model evaluation, ARI resolves the product's website and retrieves a shared first-party evidence bundle. The collector:
- retrieves the resolved product website;
- probes
/llms.txtat the resolved origin; - parses its links; and
- retrieves a bounded set of product-matching documents on first-party domains.
During a native scoring attempt, ARI runs one evidence-harvest stage per product. The three provider-native search surfaces may contribute URLs to that stage, but they do not run inside the evaluator scoring calls. ARI unions and canonicalizes their discoveries, captures the resulting public pages through the same bounded retrieval path, and produces one immutable evidence packet. All three structured evaluators receive the same packet hash and the same tool-free prompt. Search-surface variance is therefore packet-construction provenance rather than model-judgment variance.
ARI pins the provider controls that the selected model and pinned SDK expose: OpenAI reasoning effort, Gemini temperature and thinking budget, Claude manual-thinking budget, or Anthropic effort for adaptive-thinking models. The requested harvest models and resolved harvest settings are hash-bound in the shared evidence packet, while the independently resolved structured-scoring settings are persisted on each evaluator run. Unsupported controls are explicit JSON nulls. One residual remains: the pinned Anthropic SDK cannot express adaptive-thinking mode itself, so adaptive Claude runs pin effort but record thinking mode and budget as null; this is not a claim that thinking was disabled. Historical SQL null settings are unmeasured and are never reconstructed from current defaults. Evidence harvest and structured scoring also record stage duration, provider finish reason, token counts when exposed, and normalized failure messages. These diagnostics use the existing hash-pinned evidence, model-run, assessment, and run-note surfaces; they do not intentionally add provider response text. On the MediaClaw path, each successful cell retains its diagnostics with its model-run row, and a complete three-cell fan-in aggregates the map into the assessment snapshot. Until then, the model-run rows—not a partial assessment's top-level diagnostic map—are the complete per-cell authority.
For structured scoring, providers select compact packet-local excerpt IDs and
use the exact signal, penalty, or coverage assertion-type literal.
Providers are instructed to preserve every leading zero exactly as printed.
Mediafier derives each excerpt's canonical source and hydrates its URL, kind,
hash, and exact captured text before validation. A shorter numeric ordinal is
canonicalized only when its zero-padded candidate occurs exactly once in the
current immutable packet; the assertion identity and received-to-canonical
mapping are retained in the attempt diagnostic. Unknown, over-padded, or
ambiguous selectors remain rejected. Every canonicalized binding still passes
the unchanged evidence validators. A normally completed response that fails
the local response/evidence contract may receive
up to two correction attempts against the same immutable packet, with the
literal and selector rules repeated. Each correction requires the immediately
preceding response to complete normally and then fail local validation.
Evidence harvest is not repeated, usage from all scoring responses is counted,
and evaluator refusal, truncation/non-stop completion, or provider-call
failure is not retried by this correction path.
For an otherwise valid response, cross-assertion excerpt reuse is checked across
the complete response and every detected conflict is returned in one correction
diagnostic. This does not weaken the assertion-specific evidence rule or add
correction attempts; it prevents the bounded correction budget from revealing
one already-present reuse conflict at a time.
Once a full response has entity_match=true and its determination fields pass
the local schema and rubric-key checks, Mediafier freezes those decision fields.
If only the evidence contract remains invalid, subsequent eligible corrections
switch to a narrow evidence-rebinding schema: the provider may return only the
frozen category/assertion identity and a replacement packet-local excerpt ID.
Mediafier merges those bindings into the frozen local response; the provider
cannot change signals, penalties, coverage, notes, confidence, match rationale,
or recommendations during rebinding. Per-attempt diagnostics record the
response contract and frozen-decision hash so the merge boundary is auditable.
A response whose decision fields are not yet valid continues through the
full-response correction contract instead.
Both modes share the existing maximum of three structured-scoring responses,
the same packet, and the same usage accounting and no-correction boundaries.
Any provider-derived rejection diagnostic echoed into a correction is placed in
an entropy-suffixed untrusted-data delimiter and cannot alter system authority.
The public-content path is HTTPS-only, validates public DNS answers and every manual redirect hop, rejects unsupported content types, and enforces its byte limit while streaming. The persisted shared packet includes deterministic first-party and registry bundles, provider discovery provenance, resolved captures, errors, the assertion-source catalog, harvest usage, and one SHA-256 packet hash. Model-authored harvest summaries are discarded. Subject metadata and every retrieved source are zero-width-normalized and placed in source-specific, entropy-suffixed data delimiters; trusted rubric and scoring instructions stay in the model's system role. Provider-native search/tool results are explicitly classified as untrusted even where the provider owns their internal framing. With that boundary installed, the assertion catalog may include pinned primary sources plus registry and provider captures. Unfetchable URLs remain discovery provenance and cannot support assertions. Provider discoveries remain attributed to the provider that found them, while top-level capture and retrieval-error projections are deduplicated across evaluator outcomes. One product-level deadline spans first-party and registry collection, provider discovery, public capture, and all structured scoring calls; cancellation reaches the pinned retrieval transport. Every fired signal, applied penalty, and positive lifecycle-coverage determination records its own evidence binding: assertion type/key, source identity and kind, exact URL, supporting excerpt, and captured-content hash. New v2 assertion-source content and its hash use the exact normalized prompt view; the immutable packet separately retains the raw retrieval capture. Replay and write validation reject unknown packet versions, require a v2 evaluator catalog to match the hash-pinned packet catalog exactly, and retain v1's exact raw-content semantics plus its historical primary-source subset for older packets. At publication the database independently recomputes the packet's canonical SHA-256, verifies its product binding, and treats whitespace-only or generically reused excerpts as missing evidence. The packet, rather than the duplicated snapshot projection, is the v2 replay authority. ARI rejects an assertion when its binding is missing, the URL/source is absent from the shared packet's admissible assertion-source catalog, or the excerpt/hash does not match captured content. Penalties require at least one primary source. A model- authored URL is never retrieved after scoring to widen the admissible set. A claim that documentation exists is not sufficient evidence when the documentation itself was not retrieved during evidence gathering.
Healing may combine previously stored evaluator cells with newly collected gap cells. Newly collected gap cells share one packet for that heal attempt, while the assessment-level shared-packet fields remain null when historical cells do not carry the same packet hash. Mixed outcomes retain each distinct packet once by hash and map every evaluator cell to its packet hash, or to explicit null for legacy cells. Replay verifies available hashes and restores the original assertion catalog before the shared write boundary runs again. The evidence snapshots and per-cell prompt hashes preserve that provenance; a healed result does not claim that historical cells were produced from the newest evidence bundle. Consolidation is allowed only when each source run contains an explicit expected-product cohort snapshot, so products that failed all three evaluators cannot disappear from completion checks.
7. Multi-Model Assessment Model
ARI assessments should be performed by multiple independent LLM evaluators.
Recommended evaluators:
- ChatGPT
- Gemini
- Claude
During a native scoring attempt, each evaluator receives:
- the same rubric
- the same deterministic first-party evidence set
- the same scoring instructions
- the same assessment template
Each evaluator returns:
- entity-match decision and confidence
- fired signal keys and applied penalty keys by category
- lifecycle coverage estimates where applicable
- evidence notes
- assertion-level evidence bindings for every scored signal, penalty, and positive lifecycle-coverage estimate
- recommendations
ARI centrally derives numeric category subtotals and the final score from those three structured determinations; evaluators do not supply numeric scores.
Final ARI score:
A product is assessed by all three evaluators in parallel. A score is complete and publishable only when ChatGPT, Claude, and Gemini each return exactly one valid determination for every rubric category. A 1/3 or 2/3 result is retained only as a draft partial so a later healing run can fill the missing cells without repeating successful calls.
For audit and disagreement analysis, ARI also stores a standalone score derived from each evaluator's own determinations. These three audit values are not the published consensus score. Historical evaluator rows created before standalone derivation was introduced remain labeled as legacy consolidated values rather than being retroactively reinterpreted.
Account-level provider failures are not assessment outcomes. Authentication, billing, or exhausted-quota failures halt the batch and cancel its remaining provider calls; they are recorded as infrastructure failures rather than model refusals or partial evidence. Ordinary evaluator refusals and transient errors remain eligible for the draft-partial and healing path above.
For ordinary categories, each signal and penalty is multiplied by the fraction of all three evaluators that detected it, after which the category floor and cap are applied. For lifecycle-coverage categories, the three coverage estimates are averaged and the matching rubric bucket is selected. The final ARI score is the sum of the resulting category subtotals. This consensus aggregation rewards capabilities the models agree on without silently changing the denominator when a provider fails.
The assessment, all three evaluator audit rows, category components, and
consolidated recommendations are published as one transaction through the
canonical publication RPC. Direct draft-to-published updates are rejected. A
failure in any part leaves no partially published scorecard. Once published,
evaluator rows and rubric content are immutable; a rubric change requires a new
version. Provider raw_output retains the evaluator's complete structured
response; the validated parsed determinations are numeric scoring authority.
Public signal/penalty breakdowns and the flattened assertion-evidence index are
verified from those determinations at the publication boundary. Every scored
assertion must bind to an integrity-checked source in its persisted evidence
packet. Recommendation source members must match recommendations in the
corresponding persisted evaluator output; agreement, representative text,
median impact, and ordering are then derived from the exact member multiset
rather than accepted as caller labels.
Human role:
- define rubric
- define weights
- review anomalies
- improve methodology
- manage evidence quality
Humans are prevented from overriding scores directly.
8. Scores and Classifications
ARI should publish both numeric scores and classifications.
Numeric scores preserve precision and enable trend tracking.
Classifications improve legibility and sharing. The score-range → classification band table is published on the methodology page, driven by the active rubric.
| Score range | Classification |
|---|---|
| 90–100 | Agent Native |
| 80–89 | Agent Operable |
| 65–79 | Agent Capable |
| 50–64 | Emerging Agentic |
| 35–49 | Human Mediated |
| 20–34 | Human Dependent |
| 0–19 | Human Operated |
9. ARI Scoring Framework
Total score: 100
Category 1 — Agent-Accessible Documentation
Objective: Can an autonomous agent learn how to operate the platform from available documentation?
| Signal | Rationale |
|---|---|
| API documentation | Establishes machine-accessible operational surface |
| OpenAPI / Swagger / machine-readable schemas | Enables machine parsing and tooling |
| Executable implementation examples | Helps agents implement correctly |
| CLI documentation | Indicates non-UI operational design |
| MCP or equivalent agent protocol documentation | Indicates explicit agent-oriented access |
| Authentication flow documentation | Enables autonomous identity and access setup |
| Infrastructure-as-code examples | Enables repeatable deployment |
| Workflow orchestration examples | Demonstrates composability |
| Troubleshooting and recovery guidance | Helps agents recover from failure |
Penalties:
| Condition |
|---|
| Sales-gated technical documentation |
| Static/manual-style docs without executable guidance |
| Missing auth documentation |
| Documentation materially incomplete or stale |
Category 2 — Machine Execution Surfaces
Objective: Can autonomous systems execute operational tasks programmatically?
| Signal | Rationale |
|---|---|
| Mature operational APIs | Core execution surface |
| Official SDKs | Supports implementation |
| Webhooks/eventing | Supports event-driven workflows |
| Async job support | Supports scalable operations |
| Retry/idempotency semantics | Supports safe autonomous execution |
| Workflow APIs | Enables orchestration |
| Declarative orchestration | Enables machine-composable workflows |
| CLI execution capability | Indicates machine-first operation |
| MCP or equivalent execution interface | Enables structured agent interaction |
| Operational health/observability surfaces | Supports monitoring and remediation |
Penalties:
| Condition |
|---|
| No CLI or equivalent machine execution surface |
| No MCP or equivalent agent execution interface |
| UI-only execution of core workflows |
| Read-only machine interfaces only |
Category 3 — Execution Surface Completeness
Objective: How completely do CLI/MCP/API/agent surfaces cover the operational lifecycle?
This category addresses sincerity of implementation. A vendor should not receive high agentic-readiness credit for exposing a superficial CLI, toy MCP server, read-only adapter, or narrow wrapper that covers only a small portion of platform functionality.
Lifecycle areas:
| Lifecycle Area |
|---|
| Discover |
| Provision |
| Operate |
| Monitor |
| Recover |
| Govern |
| Deprovision |
Scoring:
| Coverage |
|---|
| 0–10% lifecycle coverage |
| 10–25% lifecycle coverage |
| 25–50% lifecycle coverage |
| 50–70% lifecycle coverage |
| 70–90% lifecycle coverage |
| 90–100% lifecycle coverage |
Rules:
- Do not count endpoints mechanically.
- Assess lifecycle coverage, not command count.
- Read-only coverage is materially weaker than execution authority.
- Admin/governance coverage matters.
- Recovery and observability coverage matter.
- Thin wrappers should receive low scores.
Penalties:
| Condition |
|---|
| CLI exists but covers only trivial/read-only functionality |
| MCP exists but is demo-only or non-operational |
| Machine surfaces omit provisioning entirely |
| Machine surfaces omit recovery/remediation entirely |
| Machine surfaces omit governance/admin entirely |
Category 4 — Autonomous Provisioning and Configuration
Objective: Can autonomous systems provision, configure, and administer the platform without significant human mediation?
| Signal | Rationale |
|---|---|
| Self-service provisioning | Enables autonomous onboarding |
| Usage-based onboarding | Aligns with elastic operation |
| API-key or token-based auth | Enables machine identity |
| OAuth/service principals | Enables governed enterprise automation |
| Infrastructure-as-code support | Enables reproducible deployment |
| Configuration-as-code support | Enables versioned configuration |
| Machine-operable administration | Reduces UI dependency |
Penalties:
| Condition |
|---|
| Human onboarding required for operational access |
| UI-required provisioning |
| Mandatory professional services for implementation |
Category 5 — Commercial Agentic Alignment
Objective: Can autonomous systems economically transact with the platform?
| Signal | Rationale |
|---|---|
| Usage-based pricing | Aligns with machine consumption |
| API monetization model | Indicates machine-oriented business model |
| Transactional economics | Supports autonomous operations |
| Elastic scaling | Supports variable agent workloads |
| Public pricing transparency | Improves discoverability |
| Autonomous payment/provisioning | Enables self-operating workflows |
Penalties:
| Condition |
|---|
| Pure seat-based licensing |
| Enterprise-sales-gated operational access |
| Mandatory professional services dependency |
Category 6 — Governance and Operational Trust
Objective: Can enterprises safely delegate operational authority to autonomous systems?
| Signal | Rationale |
|---|---|
| Service account support | Enables controlled machine identity |
| Auditability | Enables traceability |
| RBAC controls | Enables scoped permissions |
| Policy enforcement | Enables governed operation |
| Rights/security boundaries | Critical in media environments |
| Approval layers | Supports controlled autonomy |
| Operational observability | Supports governance and monitoring |
Category 7 — Agentic Strategy Signals
Objective: Has the company made public commitments toward agentic operability, even if the operational surface isn't fully built yet? This pillar captures intent — meaningful for trajectory but not for current operability. A company that publicly states an MCP roadmap is materially further along than one that's silent on it. Maximum contribution is 10/100; marketing alone caps a company at F-tier on the overall score.
| Signal | Rationale |
|---|---|
| Published blog post, whitepaper, or press release about agentic/autonomous operability strategy | Stated direction signals strategic intent |
| Public commitment to MCP server availability (roadmap, coming-soon, or GA announcement) | MCP-specific intent is a directly relevant signal |
| Public commitment to expanding CLI or SDK coverage | API/CLI commitment is foundational to agent operability |
| Leadership (CEO, CTO, or DevRel) has spoken publicly about agentic strategy | Leadership commitment correlates with delivery |
| Available via at least one third-party orchestration platform or marketplace | Third-party operability counts toward agent reachability |
| Developer-relations function actively focused on agent / automation use cases | Investment in agent DevRel predicts faster operational progress |
Penalties: none. This pillar scores stated intent only — the absence of public commitments yields no credit rather than a deduction, and its 10-point cap prevents public messaging alone from producing a high overall score.