1. Executive Summary

The Agentic Readiness Index (ARI) is a benchmarking framework for evaluating the degree to which media and entertainment technology platforms are operable by autonomous software systems.

ARI does not measure whether a company has AI features.

ARI measures whether autonomous coding agents and operational agents can discover, authenticate, configure, orchestrate, monitor, govern, and economically interact with a platform through machine-operable interfaces under realistic enterprise governance constraints.

ARI is therefore an operational infrastructure benchmark, not an AI feature benchmark.

The primary audience includes:

  • media technology buyers
  • vendors
  • product leaders
  • infrastructure teams
  • ecosystem partners
  • analysts
  • investors
  • agentic workflow builders

2. Core Thesis

The next generation of enterprise software competition may increasingly depend on which platforms are most operable by autonomous systems.

Traditional enterprise software was designed around:

  • human operators
  • graphical user interfaces
  • manual configuration
  • implementation specialists
  • professional services teams
  • seat-based usage patterns

ARI assumes that future-ready infrastructure increasingly supports:

  • machine-operable workflows
  • autonomous provisioning
  • programmatic governance
  • software-to-software execution
  • agent-mediated orchestration
  • governed machine access
  • usage-based or transactional economic models

The framework measures:

autonomous operability under realistic enterprise governance conditions.


3. What ARI Measures

ARI measures:

  • agent-accessible documentation
  • machine execution surfaces
  • CLI/MCP/equivalent agent surfaces
  • execution surface completeness
  • autonomous provisioning capability
  • governed operational accessibility
  • workflow composability
  • operational observability
  • machine-operable governance
  • commercial alignment with autonomous consumption
  • public agentic-strategy commitments and roadmap signals

ARI does not measure:

  • overall product quality
  • customer satisfaction
  • market share
  • feature breadth
  • generic AI sophistication
  • AI marketing
  • embedded copilots
  • chatbot functionality

A company may have sophisticated AI features and still score poorly if its platform is not operable by autonomous systems.

A company may have little AI branding and score highly if it exposes strong machine-operable infrastructure.


4. Foundational Assessment Question

Every ARI assessment is governed by one question:

Could a modern autonomous coding agent independently implement and operate this platform using available machine interfaces and operational documentation with minimal human UI dependency?

This question governs scoring, evidence collection, and recommendations.


5. Public Discoverability vs Governed Operability

ARI does not assume that all operational interfaces should be fully public.

Enterprise systems, especially in media and entertainment, may have valid reasons to restrict access:

  • rights complexity
  • security concerns
  • premium content protection
  • financial workflows
  • compliance requirements
  • partner ecosystem boundaries

ARI distinguishes between:

Concept Meaning
Public Agent Discoverability Can autonomous systems discover and understand operational capabilities?
Governed Agent Operability Can autonomous systems operate the platform within controlled enterprise boundaries?

A platform may score highly without unrestricted public exposure if it supports governed autonomous operation.

However, systems that depend on undocumented workflows, human-only onboarding, UI-only administration, or professional-services mediation should score lower.


6. Evidence Accessibility Tiers

ARI recognizes four evidence/access tiers.

Tier Description
E1 Fully public anonymous access
E2 Authenticated developer access
E3 Partner or ecosystem access
E4 Brokered or governed operational access

Governed access is not inherently penalized.

Opaque access is penalized when it prevents reproducible assessment or autonomous operation.

Deterministic First-Party Discovery

Before model evaluation, ARI resolves the product's website and retrieves a shared first-party evidence bundle. The collector:

  1. retrieves the resolved product website;
  2. probes /llms.txt at the resolved origin;
  3. parses its links; and
  4. retrieves a bounded set of product-matching documents on first-party domains.

During a native scoring attempt, ARI runs one evidence-harvest stage per product. The three provider-native search surfaces may contribute URLs to that stage, but they do not run inside the evaluator scoring calls. ARI unions and canonicalizes their discoveries, captures the resulting public pages through the same bounded retrieval path, and produces one immutable evidence packet. All three structured evaluators receive the same packet hash and the same tool-free prompt. Search-surface variance is therefore packet-construction provenance rather than model-judgment variance.

ARI pins the provider controls that the selected model and pinned SDK expose: OpenAI reasoning effort, Gemini temperature and thinking budget, Claude manual-thinking budget, or Anthropic effort for adaptive-thinking models. The requested harvest models and resolved harvest settings are hash-bound in the shared evidence packet, while the independently resolved structured-scoring settings are persisted on each evaluator run. Unsupported controls are explicit JSON nulls. One residual remains: the pinned Anthropic SDK cannot express adaptive-thinking mode itself, so adaptive Claude runs pin effort but record thinking mode and budget as null; this is not a claim that thinking was disabled. Historical SQL null settings are unmeasured and are never reconstructed from current defaults. Evidence harvest and structured scoring also record stage duration, provider finish reason, token counts when exposed, and normalized failure messages. These diagnostics use the existing hash-pinned evidence, model-run, assessment, and run-note surfaces; they do not intentionally add provider response text. On the MediaClaw path, each successful cell retains its diagnostics with its model-run row, and a complete three-cell fan-in aggregates the map into the assessment snapshot. Until then, the model-run rows—not a partial assessment's top-level diagnostic map—are the complete per-cell authority.

For structured scoring, providers select compact packet-local excerpt IDs and use the exact signal, penalty, or coverage assertion-type literal. Providers are instructed to preserve every leading zero exactly as printed. Mediafier derives each excerpt's canonical source and hydrates its URL, kind, hash, and exact captured text before validation. A shorter numeric ordinal is canonicalized only when its zero-padded candidate occurs exactly once in the current immutable packet; the assertion identity and received-to-canonical mapping are retained in the attempt diagnostic. Unknown, over-padded, or ambiguous selectors remain rejected. Every canonicalized binding still passes the unchanged evidence validators. A normally completed response that fails the local response/evidence contract may receive up to two correction attempts against the same immutable packet, with the literal and selector rules repeated. Each correction requires the immediately preceding response to complete normally and then fail local validation. Evidence harvest is not repeated, usage from all scoring responses is counted, and evaluator refusal, truncation/non-stop completion, or provider-call failure is not retried by this correction path. For an otherwise valid response, cross-assertion excerpt reuse is checked across the complete response and every detected conflict is returned in one correction diagnostic. This does not weaken the assertion-specific evidence rule or add correction attempts; it prevents the bounded correction budget from revealing one already-present reuse conflict at a time. Once a full response has entity_match=true and its determination fields pass the local schema and rubric-key checks, Mediafier freezes those decision fields. If only the evidence contract remains invalid, subsequent eligible corrections switch to a narrow evidence-rebinding schema: the provider may return only the frozen category/assertion identity and a replacement packet-local excerpt ID. Mediafier merges those bindings into the frozen local response; the provider cannot change signals, penalties, coverage, notes, confidence, match rationale, or recommendations during rebinding. Per-attempt diagnostics record the response contract and frozen-decision hash so the merge boundary is auditable. A response whose decision fields are not yet valid continues through the full-response correction contract instead. Both modes share the existing maximum of three structured-scoring responses, the same packet, and the same usage accounting and no-correction boundaries. Any provider-derived rejection diagnostic echoed into a correction is placed in an entropy-suffixed untrusted-data delimiter and cannot alter system authority.

The public-content path is HTTPS-only, validates public DNS answers and every manual redirect hop, rejects unsupported content types, and enforces its byte limit while streaming. The persisted shared packet includes deterministic first-party and registry bundles, provider discovery provenance, resolved captures, errors, the assertion-source catalog, harvest usage, and one SHA-256 packet hash. Model-authored harvest summaries are discarded. Subject metadata and every retrieved source are zero-width-normalized and placed in source-specific, entropy-suffixed data delimiters; trusted rubric and scoring instructions stay in the model's system role. Provider-native search/tool results are explicitly classified as untrusted even where the provider owns their internal framing. With that boundary installed, the assertion catalog may include pinned primary sources plus registry and provider captures. Unfetchable URLs remain discovery provenance and cannot support assertions. Provider discoveries remain attributed to the provider that found them, while top-level capture and retrieval-error projections are deduplicated across evaluator outcomes. One product-level deadline spans first-party and registry collection, provider discovery, public capture, and all structured scoring calls; cancellation reaches the pinned retrieval transport. Every fired signal, applied penalty, and positive lifecycle-coverage determination records its own evidence binding: assertion type/key, source identity and kind, exact URL, supporting excerpt, and captured-content hash. New v2 assertion-source content and its hash use the exact normalized prompt view; the immutable packet separately retains the raw retrieval capture. Replay and write validation reject unknown packet versions, require a v2 evaluator catalog to match the hash-pinned packet catalog exactly, and retain v1's exact raw-content semantics plus its historical primary-source subset for older packets. At publication the database independently recomputes the packet's canonical SHA-256, verifies its product binding, and treats whitespace-only or generically reused excerpts as missing evidence. The packet, rather than the duplicated snapshot projection, is the v2 replay authority. ARI rejects an assertion when its binding is missing, the URL/source is absent from the shared packet's admissible assertion-source catalog, or the excerpt/hash does not match captured content. Penalties require at least one primary source. A model- authored URL is never retrieved after scoring to widen the admissible set. A claim that documentation exists is not sufficient evidence when the documentation itself was not retrieved during evidence gathering.

Healing may combine previously stored evaluator cells with newly collected gap cells. Newly collected gap cells share one packet for that heal attempt, while the assessment-level shared-packet fields remain null when historical cells do not carry the same packet hash. Mixed outcomes retain each distinct packet once by hash and map every evaluator cell to its packet hash, or to explicit null for legacy cells. Replay verifies available hashes and restores the original assertion catalog before the shared write boundary runs again. The evidence snapshots and per-cell prompt hashes preserve that provenance; a healed result does not claim that historical cells were produced from the newest evidence bundle. Consolidation is allowed only when each source run contains an explicit expected-product cohort snapshot, so products that failed all three evaluators cannot disappear from completion checks.


7. Multi-Model Assessment Model

ARI assessments should be performed by multiple independent LLM evaluators.

Recommended evaluators:

  • ChatGPT
  • Gemini
  • Claude

During a native scoring attempt, each evaluator receives:

  • the same rubric
  • the same deterministic first-party evidence set
  • the same scoring instructions
  • the same assessment template

Each evaluator returns:

  • entity-match decision and confidence
  • fired signal keys and applied penalty keys by category
  • lifecycle coverage estimates where applicable
  • evidence notes
  • assertion-level evidence bindings for every scored signal, penalty, and positive lifecycle-coverage estimate
  • recommendations

ARI centrally derives numeric category subtotals and the final score from those three structured determinations; evaluators do not supply numeric scores.

Final ARI score:

A product is assessed by all three evaluators in parallel. A score is complete and publishable only when ChatGPT, Claude, and Gemini each return exactly one valid determination for every rubric category. A 1/3 or 2/3 result is retained only as a draft partial so a later healing run can fill the missing cells without repeating successful calls.

For audit and disagreement analysis, ARI also stores a standalone score derived from each evaluator's own determinations. These three audit values are not the published consensus score. Historical evaluator rows created before standalone derivation was introduced remain labeled as legacy consolidated values rather than being retroactively reinterpreted.

Account-level provider failures are not assessment outcomes. Authentication, billing, or exhausted-quota failures halt the batch and cancel its remaining provider calls; they are recorded as infrastructure failures rather than model refusals or partial evidence. Ordinary evaluator refusals and transient errors remain eligible for the draft-partial and healing path above.

For ordinary categories, each signal and penalty is multiplied by the fraction of all three evaluators that detected it, after which the category floor and cap are applied. For lifecycle-coverage categories, the three coverage estimates are averaged and the matching rubric bucket is selected. The final ARI score is the sum of the resulting category subtotals. This consensus aggregation rewards capabilities the models agree on without silently changing the denominator when a provider fails.

The assessment, all three evaluator audit rows, category components, and consolidated recommendations are published as one transaction through the canonical publication RPC. Direct draft-to-published updates are rejected. A failure in any part leaves no partially published scorecard. Once published, evaluator rows and rubric content are immutable; a rubric change requires a new version. Provider raw_output retains the evaluator's complete structured response; the validated parsed determinations are numeric scoring authority. Public signal/penalty breakdowns and the flattened assertion-evidence index are verified from those determinations at the publication boundary. Every scored assertion must bind to an integrity-checked source in its persisted evidence packet. Recommendation source members must match recommendations in the corresponding persisted evaluator output; agreement, representative text, median impact, and ordering are then derived from the exact member multiset rather than accepted as caller labels.

Human role:

  • define rubric
  • define weights
  • review anomalies
  • improve methodology
  • manage evidence quality

Humans are prevented from overriding scores directly.


8. Scores and Classifications

ARI should publish both numeric scores and classifications.

Numeric scores preserve precision and enable trend tracking.

Classifications improve legibility and sharing. The score-range → classification band table is published on the methodology page, driven by the active rubric.

Score rangeClassification
90100Agent Native
8089Agent Operable
6579Agent Capable
5064Emerging Agentic
3549Human Mediated
2034Human Dependent
019Human Operated

9. ARI Scoring Framework

Total score: 100

Category 1 — Agent-Accessible Documentation

Objective: Can an autonomous agent learn how to operate the platform from available documentation?

Signal Rationale
API documentation Establishes machine-accessible operational surface
OpenAPI / Swagger / machine-readable schemas Enables machine parsing and tooling
Executable implementation examples Helps agents implement correctly
CLI documentation Indicates non-UI operational design
MCP or equivalent agent protocol documentation Indicates explicit agent-oriented access
Authentication flow documentation Enables autonomous identity and access setup
Infrastructure-as-code examples Enables repeatable deployment
Workflow orchestration examples Demonstrates composability
Troubleshooting and recovery guidance Helps agents recover from failure

Penalties:

Condition
Sales-gated technical documentation
Static/manual-style docs without executable guidance
Missing auth documentation
Documentation materially incomplete or stale

Category 2 — Machine Execution Surfaces

Objective: Can autonomous systems execute operational tasks programmatically?

Signal Rationale
Mature operational APIs Core execution surface
Official SDKs Supports implementation
Webhooks/eventing Supports event-driven workflows
Async job support Supports scalable operations
Retry/idempotency semantics Supports safe autonomous execution
Workflow APIs Enables orchestration
Declarative orchestration Enables machine-composable workflows
CLI execution capability Indicates machine-first operation
MCP or equivalent execution interface Enables structured agent interaction
Operational health/observability surfaces Supports monitoring and remediation

Penalties:

Condition
No CLI or equivalent machine execution surface
No MCP or equivalent agent execution interface
UI-only execution of core workflows
Read-only machine interfaces only

Category 3 — Execution Surface Completeness

Objective: How completely do CLI/MCP/API/agent surfaces cover the operational lifecycle?

This category addresses sincerity of implementation. A vendor should not receive high agentic-readiness credit for exposing a superficial CLI, toy MCP server, read-only adapter, or narrow wrapper that covers only a small portion of platform functionality.

Lifecycle areas:

Lifecycle Area
Discover
Provision
Operate
Monitor
Recover
Govern
Deprovision

Scoring:

Coverage
0–10% lifecycle coverage
10–25% lifecycle coverage
25–50% lifecycle coverage
50–70% lifecycle coverage
70–90% lifecycle coverage
90–100% lifecycle coverage

Rules:

  • Do not count endpoints mechanically.
  • Assess lifecycle coverage, not command count.
  • Read-only coverage is materially weaker than execution authority.
  • Admin/governance coverage matters.
  • Recovery and observability coverage matter.
  • Thin wrappers should receive low scores.

Penalties:

Condition
CLI exists but covers only trivial/read-only functionality
MCP exists but is demo-only or non-operational
Machine surfaces omit provisioning entirely
Machine surfaces omit recovery/remediation entirely
Machine surfaces omit governance/admin entirely

Category 4 — Autonomous Provisioning and Configuration

Objective: Can autonomous systems provision, configure, and administer the platform without significant human mediation?

Signal Rationale
Self-service provisioning Enables autonomous onboarding
Usage-based onboarding Aligns with elastic operation
API-key or token-based auth Enables machine identity
OAuth/service principals Enables governed enterprise automation
Infrastructure-as-code support Enables reproducible deployment
Configuration-as-code support Enables versioned configuration
Machine-operable administration Reduces UI dependency

Penalties:

Condition
Human onboarding required for operational access
UI-required provisioning
Mandatory professional services for implementation

Category 5 — Commercial Agentic Alignment

Objective: Can autonomous systems economically transact with the platform?

Signal Rationale
Usage-based pricing Aligns with machine consumption
API monetization model Indicates machine-oriented business model
Transactional economics Supports autonomous operations
Elastic scaling Supports variable agent workloads
Public pricing transparency Improves discoverability
Autonomous payment/provisioning Enables self-operating workflows

Penalties:

Condition
Pure seat-based licensing
Enterprise-sales-gated operational access
Mandatory professional services dependency

Category 6 — Governance and Operational Trust

Objective: Can enterprises safely delegate operational authority to autonomous systems?

Signal Rationale
Service account support Enables controlled machine identity
Auditability Enables traceability
RBAC controls Enables scoped permissions
Policy enforcement Enables governed operation
Rights/security boundaries Critical in media environments
Approval layers Supports controlled autonomy
Operational observability Supports governance and monitoring

Category 7 — Agentic Strategy Signals

Objective: Has the company made public commitments toward agentic operability, even if the operational surface isn't fully built yet? This pillar captures intent — meaningful for trajectory but not for current operability. A company that publicly states an MCP roadmap is materially further along than one that's silent on it. Maximum contribution is 10/100; marketing alone caps a company at F-tier on the overall score.

Signal Rationale
Published blog post, whitepaper, or press release about agentic/autonomous operability strategy Stated direction signals strategic intent
Public commitment to MCP server availability (roadmap, coming-soon, or GA announcement) MCP-specific intent is a directly relevant signal
Public commitment to expanding CLI or SDK coverage API/CLI commitment is foundational to agent operability
Leadership (CEO, CTO, or DevRel) has spoken publicly about agentic strategy Leadership commitment correlates with delivery
Available via at least one third-party orchestration platform or marketplace Third-party operability counts toward agent reachability
Developer-relations function actively focused on agent / automation use cases Investment in agent DevRel predicts faster operational progress

Penalties: none. This pillar scores stated intent only — the absence of public commitments yields no credit rather than a deduction, and its 10-point cap prevents public messaging alone from producing a high overall score.

The Agentic Readiness Index is an experimental benchmark from Mediafier.