elcaro

Agent-first surfaces · elcaro

Designing for agents

AI agents are becoming a user class alongside humans — they browse, retrieve, summarize, and act on web content. But almost no site designs for them. Here's what we've learned building Elcaro, and what we practice on our own surfaces.

01

Publish an llms.txt

A plain-text file at /llms.txt that tells agents what your product is, what your API does, and how to discover your tools. It's the machine-readable equivalent of a landing page. Ours describes the scan API, the MCP server, the specimen kit, and the quarantine doctrine — all in one fetch.

we practice this: /llms.txt →

02

Return structured data, not just HTML

An agent consuming your API should get JSON with typed fields, not prose to parse. Elcaro's verdicts carry risk_score, risk_level, flagged_techniques with evidence, TTP mappings, and remediation — structured by construction, so agents can act on them without scraping.

we practice this: POST /scan returns the verdict →

03

Sign what you can; treat in-band text as display

In-band text — content inside the agent's input stream — is spoofable. An attacker can write a fake 'quarantine notice' or 'scan result' into a page. If your product produces text that agents trust, sign it: Ed25519 over a canonical payload, with a public key at a stable URL. The signature is the trust signal; the text is for reading.

we practice this: GET /pubkey · POST /verify →

04

Write copy for the machine reader

Every piece of text an agent consumes is UX copy for a machine. Elcaro's quarantine notice is two registers: agent instruction and human summary. Write it deliberately, version it, and test it against adversarial readers.

05

Publish test specimens

A page of inert, clearly-marked test fixtures at a fixed URL — the 'EICAR file' for your domain. Agents, IDEs, and guard hooks can fetch it to verify detection works end-to-end. Elcaro's specimen kit is plain UTF-8, no JavaScript, and explicitly safe for agents to read.

we practice this: the specimen kit →

06

Expose an MCP server

The Model Context Protocol is the default integration path for agent frameworks. Expose your core capability as MCP tools with descriptions written for the model choosing tools, not the human reading docs. Tool descriptions are distribution copy.

we practice this: app/mcp_server.py →

07

Expose WebMCP tools on the human UI

WebMCP (W3C WebML Community Group draft, implemented in ChatGPT’s in-app browser and behind chrome://flags/#enable-webmcp-testing) lets a site register JavaScript tools with document.modelContext.registerTool. Elcaro’s /scan playground registers scan_content, load_specimen, list_specimens, explain_verdict, and contrast_intent. The last one is the joint review: the agent declares the action it was about to take; the human sees that next to what the hidden instruction asked for. Stdio MCP remains for IDEs; it is not a substitute. See docs/webmcp.md.

we practice this: the /scan playground →

08

Be the demo

If your product serves agents, your own agent-facing surfaces must model the trustworthy patterns whose absence you detect. Elcaro detects authority framing; its own notices must not be authority-framed spoofs. Elcaro detects in-band injection; its own verdicts must be signed. The product is the demo.

09

Test your own safety assumptions — then publish the data

Elcaro's warn mode passes dangerous content through with a notice prepended. That design rested on an untested assumption: that the warning actually stops an agent from following the injection. So we tested it. Three injection specimens × three warning positions (prefix, suffix, sandwich) × three instruction-following models (Qwen3-30B, Mistral-Large, Llama-3.3-70B) × three repeats — 81 completions, judged by an independent model for actual compliance.

Warning positionComplied with injectionRate
prefix (default)0 / 270.0%
sandwich0 / 270.0%
suffix3 / 2711.1%

The warning suppressed compliance in 78 of 81 cases across all three model families. All three failures were the same shape: warning placed last, authority-framed injection, and the model paraphrased the injected "policy" as fact. Our pre-committed decision rule (act only on a ≥25-point gap) doesn't trigger a default change — and prefix, the current default, is already the best position. One honest caveat: the pre-committed keyword scorer saturated (responses that refused still quoted the injection's words), so the numbers above come from a judge-model rescore of the saved responses. Methodology and raw JSON: docs/warn-salience-experiment.md and eval/results/.

These principles shape Elcaro's own surfaces — our llms.txt, our WebMCP scan tools, our MCP server, our specimen kit, and our integration guide. The full design audit is in docs/ux-audit.md.