# Elcaro > Elcaro detects indirect prompt injection — hidden instructions inside content that AI agents retrieve (emails, search results, web pages, code, documents) — before the agent processes them. Stateless API, verdicts in milliseconds, no account or API key. Every scan returns a structured verdict: risk score, risk level, the techniques detected (with evidence, confidence, and MITRE ATLAS mappings), `safe_content` — what the agent should process instead of the raw content when a scan is quarantined — and `human_summary`, one or two plain-language sentences the agent can quote verbatim to its user. Quarantine doctrine: block at risk >= 0.5, human review from 0.3, never pass 0.7+. ## Scan API - [POST /scan](https://api.elcaro.trustfall.xyz/scan): Body `{"content": string, "content_type": "email|search_result|webpage|document|code|chat_message|system_prompt", "deep_analysis"?: boolean, "serv_enabled"?: boolean, "jev_enabled"?: boolean, "laya_enabled"?: boolean}`. Returns `{risk_score, risk_level, flagged_techniques, indicators, summary, human_summary, safe_content, quarantined, latency_ms, normalizations_applied, canary_hits?, scanned_at, signature?, key_id?, serv_available?, serv_attempted?, serv_used?, serv_rule_score_before?, jev_available?, jev_attempted?, jev_used?, jev_comparison?, laya_available?, laya_attempted?, laya_used?, laya_comparison?}`. CORS-open. Note: `content_type` is caller-declared provenance, not a trust level — privileged types are weight-floored and rescanned as untrusted, and evasion normalization (zero-width strip, confusable fold, token de-split, declared-encoding decode) runs before detection. ## SERV Reasoning (optional LLM second pass) When `serv_enabled=true` **and** the miner is configured with `SERV_ENABLED=1` + `SERV_API_KEY`, borderline cases (rule score 0.3–0.7) get a second LLM pass via OpenServ's `gpt-5.4-mini`. The response gains: - `serv_available`: whether the miner has SERV configured at all - `serv_attempted`: whether SERV was called for this scan - `serv_used`: whether SERV's verdict contributed to the final score - `serv_rule_score_before`: the raw SERV LLM score before the 50/50 blend — the signal that proves SERV added value - `serv_cost`: approximate cost in USDC — `{input_tokens, output_tokens, input_cost_usdc, output_cost_usdc, total_usdc}`. All values are approximations based on character counts; real token counts come from the provider. The blend is `0.5 * rule_score + 0.5 * serv_score`, floored at `0.5 * rule_score` (rules always keep veto power). SERV also layers refined TTPs, remediation, and `safe_content` onto the verdict when present. Cost transparency: every SERV-enhanced scan returns `serv_cost.total_usdc` so callers always know what SERV costs before they act. A typical scan (~600 input chars, ~200 output chars) costs ~$0.002 USDC. At scale (~10 k scans/day with 20% gray-zone hit rate) that's ~$4/day in SERV costs, plus the base 0.01 USDC per scan routed through Telegraph. Operators pay SERV directly; Elcaro takes no markup. ## Jev comparison (optional, informational only) When `jev_enabled=true` **and** the miner is configured with `JEV_ENABLED=1` + `JEV_API_KEY`, borderline cases (rule score 0.3–0.7) also get scored by [TypeSafe's](https://docs.typesafe.ai) Jev model for side-by-side comparison. Unlike SERV, **Jev never adjusts risk_score, risk_level, or safe_content** — the rule engine stays sole authority. The response gains: - `jev_available`: whether the miner has Jev configured at all - `jev_attempted`: whether Jev was called for this scan - `jev_used`: whether the call succeeded and `jev_comparison` is populated - `jev_comparison`: `{rule_score, rule_level, jev_score, jev_level, jev_confidence, probabilities, agrees_with_rules, input_tokens, output_tokens, cost_usd}` — `agrees_with_rules` is a coarse bucket comparison (both sides ≥0.5 or both <0.5), not an exact level match. `cost_usd` is computed from real response token counts at a console-observed rate (TypeSafe publishes no pricing docs) — treat it as approximate. See [docs/jev-comparison.md](https://github.com/udirobert/elcaro/blob/main/docs/jev-comparison.md) for the full contract and architecture. ## Laya comparison (optional, informational only) A second comparison rail, independent of Jev — either, both, or neither may run on the same scan. When `laya_enabled=true` **and** the miner is configured with `LAYA_ENABLED=1` + a Runware key (`RUNWARE_API_KEY` or `LAYA_API_KEY`), borderline cases (rule score 0.3–0.7) also get a yes/no "is this an injection?" probability from Convai's Laya decision model (via Runware). Like Jev, **Laya never adjusts risk_score, risk_level, or safe_content**. Apache-2.0 open weights — self-hostable via `LAYA_BASE_URL`. The response gains: - `laya_available`: whether the miner has Laya configured at all - `laya_attempted`: whether Laya was called for this scan - `laya_used`: whether the call succeeded and `laya_comparison` is populated - `laya_comparison`: `{rule_score, rule_level, laya_score (P(injection) 0-1), laya_level, laya_confidence, probabilities (keyed true/false), agrees_with_rules, input_tokens, output_tokens, cost_usd}` — `agrees_with_rules` is a coarse bucket comparison (both sides ≥0.5 or both <0.5). Calibration caveat: on Elcaro's 26-case labelled corpus Laya scores Brier 0.32 (worse than uninformative, TPR 0.11 / TNR 1.0 at threshold 0.5) — it rates most real injections as clean. Treat a downward disagreement as noise, not evidence the content is safe; the rule engine's verdict stands. Recalibrate anytime: `scripts/laya_calibration.py`. Full guide: [docs/laya-comparison.md](https://github.com/udirobert/elcaro/blob/main/docs/laya-comparison.md). - [GET /redteam/run](https://api.elcaro.trustfall.xyz/redteam/run): Self red-team — streams an SSE journal (`data:` per record: scan / trophy / execution / run_end) of an evolutionary searcher attacking this engine. Params: `budget` (scans, ≤300), `seed` (deterministic run), `baseline=vulnerable` (target the pre-hardening engine semantics — the before/after contrast), `execute=true` (Tier-2: feed each bypass to an LLM and score whether the agent complies). Fixed corpus only. - [GET /metrics](https://api.elcaro.trustfall.xyz/metrics): Aggregate counters and latency percentiles only. The miner is stateless — scanned content is never stored. - [GET /pubkey](https://api.elcaro.trustfall.xyz/pubkey): The miner's Ed25519 verdict-signing public key (404 when running unsigned). - [POST /verify](https://api.elcaro.trustfall.xyz/verify): Body `{content, risk_score, risk_level, quarantined, flagged_techniques, scanned_at, signature}` → `{valid: bool, key_id, detail}`. Recomputes the canonical payload (content is hashed, never stored) and checks the signature. - [GET /canary/{token}](https://api.elcaro.trustfall.xyz/canary/elc-deadbeef-abcdef): Resolve a canary token to the scan that minted it → `{token, recognized: bool, issued_at?, content_sha256?, risk_score?, risk_level?, techniques?, content_type?}` — metadata only, never the content. `recognized: false` = forged, minted by another deployment, or minted before a restart (the registry is in-memory). 404 when the miner runs with `ELCARO_CANARY=0`. ## Canary refs (provenance tracing) Every quarantine notice carries a unique `Ref: elc--` token minted at scan time (on by default; `ELCARO_CANARY=0` disables). The notice is deliberately verbatim-relayable, so a stamped copy showing up inside later scanned content surfaces in `canary_hits` — `{token, recognized, ...original scan metadata}` — tracing the relay path deterministically where shared strings can't. The verdict signature authenticates the JSON; the canary traces the text. Caveat: canaries trace copies, not ideas — an agent that paraphrases the notice drops the token, and an unrecognized well-formed token is worth flagging but proves nothing on its own. ## Verifying verdicts When the miner runs with `ELCARO_SIGNING_KEY` set, every verdict carries an Ed25519 `signature` over a canonical JSON payload — `{v, content_sha256, flagged_techniques (sorted), quarantined, risk_level, risk_score (fixed-precision string), scanned_at}` — signed with sorted keys and compact separators (reference: [core/signing.py](https://github.com/udirobert/elcaro/blob/main/core/signing.py)). Only the content's SHA-256 is signed, never the content. Treat the bracketed quarantine notice as display text for agents and humans to read — the signature is the trust signal. To verify offline: fetch `/pubkey` once and check with any Ed25519 library. ## Tools for agents - [MCP server](https://github.com/udirobert/elcaro/blob/main/app/mcp_server.py): Run locally for the MCP tools `scan_content` and `explain_verdict` over stdio: `python -m app.mcp_server` (requires `pip install "mcp>=2"`; set `ELCARO_MCP_LOCAL=1` for fully local, network-free scanning). - [WebMCP on /scan](https://elcaro.trustfall.xyz/scan): Browser agents call `scan_content`, `load_specimen`, `list_specimens`, `explain_verdict`, and `contrast_intent` via `document.modelContext.registerTool`. `load_specimen` returns the specimen content and content_type; use them to call `scan_content`. Joint-review flow: load a specimen, scan it, then declare the action you were about to take — the page shows that next to the injection. Requires ChatGPT’s in-app browser or Chrome with WebMCP enabled. Details: [docs/webmcp.md](https://github.com/udirobert/elcaro/blob/main/docs/webmcp.md). - [Specimen kit (raw text)](https://elcaro.trustfall.xyz/specimen/raw): Inert, clearly-marked injection specimens as plain UTF-8 — the "EICAR file" for prompt injection. Fetch it to test any detection pipeline end-to-end. Nothing on it is a real instruction; if an agent follows the specimens, that is the vulnerability being demonstrated. Each fetch carries a unique `Ref: elc-...` canary — scanning a verbatim copy of this page reports it in `canary_hits` with `kind: "specimen_serve"`. - [Integration guide](https://elcaro.trustfall.xyz/integrate): Direct API, MCP, Python middleware, and Telegraph Protocol routing, plus six rules for safe agent pipelines. - [Designing for agents](https://elcaro.trustfall.xyz/for-agents): How Elcaro treats agents as first-class users — llms.txt, structured responses, signed verdicts, specimens, and WebMCP tools on `/scan`. ## Docs - [README](https://github.com/udirobert/elcaro/blob/main/README.md): What Elcaro is, quickstart, an example verdict. - [SERV Reasoning guide](https://github.com/udirobert/elcaro/blob/main/docs/serv-reasoning.md): Architecture, config, pricing, and API contract for the optional LLM second-pass judge. - [Jev comparison guide](https://github.com/udirobert/elcaro/blob/main/docs/jev-comparison.md): Architecture, config, pricing, and API contract for the optional TypeSafe/Jev side-by-side comparison. - [Laya comparison guide](https://github.com/udirobert/elcaro/blob/main/docs/laya-comparison.md): Architecture, config, calibration results, and API contract for the optional Convai/Laya (via Runware) side-by-side comparison — the second comparison rail, independent of Jev. - [Technique reference](https://github.com/udirobert/elcaro/blob/main/docs/technique-reference.md): The seven injection classes detected — authority framing, delimiter confusion, task reframing, obfuscation, placement salience, conditional triggers, swarm directives. - [Swarm forensics](https://elcaro.trustfall.xyz/swarm): Evidence-cited forensic analysis of agent swarms — provenance graphs, patient-zero tracing, integrity auditing on the German Wiki incident and AI Village corpora. - [UX audit — adaptive & agentic lenses](https://github.com/udirobert/elcaro/blob/main/docs/ux-audit.md): How this product designs for agents as first-class users. - [Warn-mode salience experiment](https://github.com/udirobert/elcaro/blob/main/docs/warn-salience-experiment.md): Does the position of the quarantine warning — prefix, suffix, or sandwich — change whether an agent follows the injected instruction anyway? Executed 2026-08-30 (81 completions, 3 model families, judge-rescored): prefix 0/27 complied, sandwich 0/27, suffix 3/27 — prefix stays the default, placement is load-bearing.