Boxbox.com
Box's API program shows strong foundations in documentation discovery and machine-readable contract quality, but two areas are putting partner integrations at risk: agents and partners have no safe place to experiment before hitting production, and mutating operations lack the idempotency guidance needed to prevent duplicate writes during retries. Closing the sandbox gap and adding idempotency documentation should be the immediate priorities, followed by surfacing code samples and an agent-readable instruction file to accelerate partner onboarding.
API DesignA clean, typed, well-governed API contract agents can reason about3 pass2 warn0 fail88A
| Signal | Points | Findings | Rationale | |
|---|---|---|---|---|
| pass | Auth declared & discoverablevia docs | 25/25 | Authentication is required across the API (declared globally or on every operation), so an agent knows credentials are needed. Investigated: docs 100%, wellknown 100%, spec 75%. | Agents can only authenticate when auth is declared, scoped, and discoverable. |
| warn | Machine-readable, versioned contractvia spec | 20.8/25 | The API sets a version ("2024.0") but exposes no versioning scheme in the URL, a header, or the media type, so an agent can't pin to a specific version. Investigated: spec 83%, docs 67%. | A current OpenAPI version with a declared versioning scheme lets agents reason about the contract. |
| warn | Schema coverage & depthvia spec | 15.3/25 | Only 83% of the 297 operations document both request and response schemas (target 95%+), so an agent can't reliably call the rest. Investigated: spec 61%. | Typed, complete request/response schemas are what make agent function-calling possible. |
| pass | Security & governance hygienevia spec | 15/15 | No credential-shaped strings detected in spec. Investigated: spec 100%, wellknown 0%. | No leaked secrets, no critical lint violations, no OWASP API Top-10 spec smells, and a published vulnerability-disclosure channel. |
| pass | Example coveragevia spec | 10/10 | Only 38% of parameters and responses (586 of 1559) include example values, so an agent has little grounding in real payload shapes. Investigated: spec 100%, docs 0%. | Examples carry shape semantics schemas under-specify — for humans and agents alike. |
Developer ExperienceThe context both developers and agents need to integrate fast — onboarding, code samples, complete descriptions and worked examples4 pass0 warn1 fail85A
| Signal | Points | Findings | Rationale | |
|---|---|---|---|---|
| pass | Self-service developer portalvia docs | 29/29 | Self-service signup available at https://account.box.com/signup/developer, with a free tier or sandbox documented — an agent can onboard without contacting sales. Investigated: docs 100%. | Self-serve key/account creation is the fast first call for partners, with no sales gate. |
| pass | Quickstart presentvia docs | 25/25 | Quickstart at https://developer.box.com/guides/box-ai/ai-tutorials/prerequisites includes a runnable first-call code sample. Investigated: docs 100%. | A quickstart is the fastest path from landing page to first successful call. |
| pass | Description completenessvia spec | 15/15 | Only 88% of operations and parameters have substantive descriptions (6 of 8 ops, 2 of 2 params; target 90%+). Investigated: spec 100%. | Complete descriptions are the context humans and agents need to use endpoints. |
| pass | Changelog publishedvia spec | 13/13 | changelog is documented on the docs site at https://developer.box.com/changelog, even though it isn't declared in the API spec. Investigated: spec 100%, docs 100%. | A published changelog lets partners track changes without surprise. |
| fail | Code samples in docsvia docs | 0/18 | No code samples detected across 17 sampled docs pages, so a coding agent gets no ready-to-use examples. Investigated: docs 0%. | Multi-language samples shorten time-to-first-call. |
Agent DiscoveryPartners and their agents can find your APIs — llms.txt, registries, crawlable and reachable docs3 pass1 warn0 fail95A+
| Signal | Points | Findings | Rationale | |
|---|---|---|---|---|
| pass | Docs reachable, not hard auth-gatedvia docs | 30/30 | All 50 sampled pages are publicly accessible. Investigated: docs 100%. | Agents can only index and fetch docs they can reach — past auth gates and over correct HTTP semantics. |
| pass | Registry & SDK presencevia docs | 28/28 | Indexed on Context7 (websites/ascii_dev, 748 snippets). Investigated: docs 100%, sdk 100%, cli 0%, mcp 0%, wellknown 0%. | Listing in MCP registries and publishing SDKs puts the API where agents and their tooling look. |
| warn | llms.txt present, valid & comprehensivevia docs | 24.4/30 | llms.txt covers 151/516 sitemap doc pages (29%); 365 missing; 31 llms.txt links not in sitemap (may indicate stale links or incomplete sitemap). Investigated: docs 81%. | A valid, comprehensive llms.txt is the machine-readable entry point for agents. |
| pass | Crawlable / AEOvia wellknown | 12/12 | robots.txt lets all monitored AI crawlers reach the docs paths. Investigated: wellknown 100%. | Bots allowed plus a fresh sitemap make docs findable by agent crawlers. |
Agent UnderstandingAgents can correctly interpret your APIs — machine-readable errors, consistent descriptions, structured data, parseable docs3 pass2 warn1 fail80A
| Signal | Points | Findings | Rationale | |
|---|---|---|---|---|
| pass | Machine-readable errors (RFC 9457)via docs | 28/28 | Only 2 distinct 4xx/5xx error codes are documented (plus 0 catch-all "default" responses), so an agent has limited failure branching. Investigated: docs 100%, spec 50%. | RFC 9457 problem details and a documented error-code inventory let agents parse failures without burning tokens. |
| pass | Operation purpose clarityvia spec | 25/25 | 99% of operations (293 of 297) have both a clear summary and a descriptive name, so an agent can pick the right endpoint. Investigated: spec 100%. | Agents select the right endpoint from its summary + operationId; clear, named operations make tool-selection reliable — the strongest driver of correct tool choice. |
| warn | Agent-navigable, token-efficient docsvia docs | 14.7/22 | 1 of 47 pages have substantive content differences between markdown and HTML (avg 7% missing); 2 failed to fetch. Investigated: docs 67%. | Server-rendered, clean, small-footprint docs are what an agent can cheaply fetch and parse correctly. |
| pass | Docs structured datavia docs | 8/8 | Machine-readable JSON-LD Article markup on 14 of 14 assessed pages (100%), with dateModified present. Investigated: docs 100%. | Structured data (JSON-LD/schema.org) on docs pages gives agents an unambiguous parse target and is what answer engines cite. Detected on the JS-rendered head (Firecrawl) for a bounded page budget, so JS-injected JSON-LD is now caught; pages we can't render are excluded rather than failed. |
| warn | Description consistency across surfaces | 0.5/7 | Mean pairwise description similarity across 2 surfaces (spec, docs) is 3% (threshold 35% for full credit). | Every surface tells the same story about what the product is. |
| fail | Agent instructions file (AGENTS.md)via wellknown | 0/10 | No AGENTS.md at the site root or /.well-known/, so coding agents have no ready-made setup and usage instructions. Investigated: wellknown 0%. | An AGENTS.md gives coding agents explicit setup, auth, and usage instructions to interpret and operate the API — beyond llms.txt's link index. |
Agent UsabilityAgents have the context to use your APIs reliably, not just find them0 pass3 warn2 fail48F
| Signal | Points | Findings | Rationale | |
|---|---|---|---|---|
| warn | Rate-limit signalingvia spec | 11/22 | No rate-limit response headers are documented, so an agent can't tell when it is approaching a limit and will get throttled. Investigated: spec 50%, docs 50%. | Machine-readable rate-limit headers let agents throttle adaptively. |
| warn | Pagination documented & consistentvia spec | 11/22 | The 1 list endpoint(s) expose no pagination parameters, forcing an agent into all-or-nothing reads. Investigated: spec 50%, docs 50%. | Consistent, documented pagination lets agents traverse collections. |
| warn | Runnable collection with test scriptsvia platform | 4.5/9 | No test scripts found in the workspace's public collections, so an agent can't use them to verify API behavior. Investigated: platform 50%. | A public, maintained collection with assertions is runnable truth agents validate against. |
| fail | Idempotency documentedvia spec | 0/27 | Only 0% of mutating operations (0 of 166) document idempotency, so an agent's retries can create duplicate writes. Investigated: spec 0%, docs 0%. | Documented idempotency lets agents retry safely. |
| fail | Sandbox separationvia spec | 0/20 | None of the 16 declared server(s) is a sandbox/test host, and no test-key prefixes are documented, so agents can only hit production. Investigated: spec 0%, docs 0%. | An isolated environment lets agents exercise destructive operations safely. |
Resources Discovered
The public resources we found for Box — the evidence behind the score. All discovered from public sources; nothing here requires access to your systems.
| Agent hints | Context7 (748) |
|---|---|
| APIs analyzed | 3 — Box Platform API, Main API, Vendors API |